Source description
About the role
We are seeking an innovative and detail-oriented Agentic AI Test Engineer to join our AI Engineering team. The ideal candidate will have strong expertise in AI/ML testing, LLM validation, autonomous AI agents, and software quality engineering. You will be responsible for validating the performance, reliability, safety, and accuracy of AI agents, multi-agent systems, and GenAI-powered applications across mobile, web, and enterprise environments. This role requires a strong blend of AI testing (70%) and software automation/quality engineering (30%) to ensure robust and scalable AI-driven products. The candidate will have responsibilities across the following functions: Agentic AI and GenAI Testing (70%): Design and execute testing strategies for Agentic AI systems, autonomous workflows, and LLM-powered applications.Validate AI agents' reasoning, planning, memory, tool usage, and decision-making capabilities.Perform functional, regression, performance, and behavioural testing for AI agents.Develop test scenarios for prompt engineering, response validation, hallucination detection, and output consistency.Evaluate AI model performance using metrics such as accuracy, reliability, latency, relevance, and safety.Conduct multi-agent orchestration testing and agent-to-agent interaction validation.Test integrations with LLMs (GPT, Claude, Gemini, Llama, etc. ), APIs, vector databases, and retrieval systems (RAG).Ensure compliance with AI governance, ethical AI, bias testing, security, and privacy standards.Identify risks related to hallucination, prompt injection, data leakage, adversarial attacks, and model drift.Create automated validation frameworks for AI-generated outputs. QA and Automation Testing (30%): Develop and maintain automation test scripts using frameworks such as Selenium, Playwright, Appium, Cypress, or API automation tools.Perform functional, API, integration, and end-to-end testing.Work with CI/CD pipelines for automated quality validation.Conduct defect tracking, root cause analysis, and collaborate with engineering teams for issue resolution.Create and maintain test plans, test cases, and quality documentation. Requirements: Bachelor's/Master's degree in Computer Science, Engineering, AI, or related field.4+ years of QA/Test Automation experience with exposure to AI/GenAI testing.Hands-on experience in LLM testing, Prompt Testing, RAG validation, and Agentic AI systems.Strong understanding of AI evaluation frameworks, model behaviour analysis, and autonomous systems testing.Experience with Python, Java, or JavaScript for automation and testing.Knowledge of Prompt Engineering, LangChain, LangGraph, CrewAI, AutoGen, MCP, OpenAI APIs, or AI orchestration frameworks.Experience with API testing tools such as Postman or Rest Assured.Familiarity with cloud platforms (AWS, Azure, GCP) is a plus.Strong analytical, debugging, and problem-solving skills. Good to Have: Experience in AI Safety Testing / Responsible AI.Exposure to Mobile Testing (Android/iOS) using Appium.Understanding of MLOps, vector databases, embeddings, and RAG pipelines.Experience with performance testing tools such as JMeter or Locust. We are seeking an innovative and detail-oriented Agentic AI Test Engineer to join our AI Engineering team. The ideal candidate will have strong expertise in AI/ML testing, LLM validation, autonomous AI agents, and software quality engineering. You will be responsible for validating the performance, reliability, safety, and accuracy of AI agents, multi-agent systems, and GenAI-powered applications across mobile, web, and enterprise environments. This role requires a strong blend of AI testing (70%) and software automation/quality engineering (30%) to ensure robust and scalable AI-driven products. The candidate will have responsibilities across the following functions: Agentic AI and GenAI Testing (70%): Design and execute testing strategies for Agentic AI systems, autonomous workflows, and LLM-powered applications.Validate AI agents' reasoning, planning, memory, tool usage, and decision-making capabilities.Perform functional, regression, performance, and behavioural testing for AI agents.Develop test scenarios for prompt engineering, response validation, hallucination detection, and output consistency.Evaluate AI model performance using metrics such as accuracy, reliability, latency, relevance, and safety.Conduct multi-agent orchestration testing and agent-to-agent interaction validation.Test integrations with LLMs (GPT, Claude, Gemini, Llama, etc. ), APIs, vector databases, and retrieval systems (RAG).Ensure compliance with AI governance, ethical AI, bias testing, security, and privacy standards.Identify risks related to hallucination, prompt injection, data leakage, adversarial attacks, and model drift.Create automated validation frameworks for AI-generated outputs. QA and Automation Testing (30%): Develop and maintain automation test scripts using frameworks such as Se
More at Deutsche Telekom Digital Labs