Padmi

AI Agent Evaluation Engineer

BangalorePosted 2 months ago
Software QualitySeniorFull Time
Apply at cloudfulcrum

Opens the source posting on foundit.in

Source description

About the role

View original

Job Description : Job Title: AI Agent Evaluation Engineer Experience Level: 6 - 8 Years (Minimum 6 years in Software QA required) Work Distribution: 70% Automation Testing / 30% Manual Testing Primary Focus Areas: Responsible AI, Safety Evaluations, and Google ADK Role Summary: We are seeking a seasoned QA professional to lead evaluation and testing efforts for AI and LLM-based systems. The ideal candidate will have deep experience in validating conversational agents, conducting safety and adversarial assessments, and working with Google's Agent Development Kit (ADK) and Vertex AI ecosystem. Mandatory Requirements: AI / LLM Testing Expertise At least 2 years dedicated experience testing or evaluating AI systems, conversational agents, or large language models (LLMs). Proven track record of designing, executing, and reviewing AI evaluation frameworks and test suites, including both automated and exploratory evaluations. Safety & Red Teaming Hands-on experience with safety evaluations is mandatory. Demonstrated practice in red teaming, adversarial testing, jailbreaking, toxicity/bias measurement, and other responsible AI safety assessments. Google ADK Knowledge Must have direct experience with or strong conceptual understanding of the Google Agent Development Kit (ADK) and its role in building and evaluating AI agents. ADK is an open-source framework for developing, testing, and deploying agentic systems and integrates with tools such as Vertex AI. Familiarity with Vertex AI (including Agent Builder and related evaluation services) is required. Technical Requirements Core Programming & Automation Strong proficiency in Python for test scripting, automation frameworks, and data manipulation. Practical experience with PyTest for creating robust automated test suites. AI Safety & Prompt Challenges Experience identifying and implementing tests for prompt injections, adversarial inputs, jailbreak scenarios, and robustness checks. Evaluation & Tooling Familiarity Candidates should be familiar with at least some of the following AI evaluation tools, libraries, or frameworks: Langsmith DeepEval Ragas Giskard Hugging Face evaluation tools These tools support structured testing and performance evaluation for LLMs and agent systems. Skills : AI Agent Evaluation Engineer Education : Bachelors Specialization : Any Specialization

One address, no account. We’ll tell you when matching roles go live.

More at cloudfulcrum

Related open roles

View all roles