Padmi

AI QA Engineer LLM, RAG & Prompt Testing

Delhi NCRPosted 1 month ago
Software QualitySeniorFull Time; Regular
Apply at STATUSNEO INC

Opens the source posting on shine.com

Source description

About the role

View original

Work location - Hyderabad About the Role We are looking for an AI QA Engineer with hands-on experience in testing Generative AI, LLM-based applications, AI chatbots, and RAG systems. This role is not suitable for candidates with only manual testing or generic Selenium automation experience. The ideal candidate should have practical exposure to AI response validation, hallucination detection, prompt testing, grounding verification, RAG retrieval validation, and LLM evaluation frameworks. Key Responsibilities Test AI chatbots, LLM-powered applications, and GenAI workflows across functional, regression, API, and end-to-end scenarios.Validate RAG retrieval results, including context relevance, retrieval accuracy, source grounding, and response completeness.Evaluate AI-generated responses for hallucinations, factual accuracy, answer relevancy, consistency, and faithfulness.Verify grounding and citations by checking whether model responses are supported by retrieved source documents.Perform prompt validation, prompt robustness testing, jailbreak testing, and prompt injection attack testing.Validate model responses against expected outcomes, golden datasets, business rules, and acceptance criteria.Build or execute AI evaluation suites using tools/frameworks such as RAGAS, DeepEval, LangChain evaluation, LangSmith, prompt evaluation frameworks, or custom Python-based metrics.Test multi-turn chatbot conversations for context retention, fallback handling, safety behaviour, and response consistency.Collaborate with product, engineering, data science, and QA teams to define AI testing strategy and quality metrics.Log AI defects clearly with prompt, context, source data, actual response, expected response, screenshots/logs, and evaluation reason.Support automation using Python, Playwright, Selenium, REST API testing, Postman, GitHub Actions, Jenkins, or similar tools.Must-Have Skills 5+ years of QA / SDET / automation testing experienceHands-on experience in AI testing, GenAI testing, LLM testing, chatbot testing, or RAG testingStrong understanding of:Hallucination detectionGrounding validationPrompt testingPrompt injection / jailbreak testingRAG retrieval validationAI model response evaluationExperience validating model outputs against expected outcomes, source documents, golden datasets, or evaluation metricsWorking knowledge of Python, JavaScript, or TypeScriptExperience with automation tools such as Playwright, Selenium, PyTest, or REST AssuredAPI testing experience using Postman / REST APIsFamiliarity with CI/CD pipelines such as GitHub Actions, Jenkins, GitLab CI, or Azure DevOpsExperience with Jira, Agile/Scrum, test planning, defect reporting, and regression testingGood-to-Have Skills Experience with RAGAS, DeepEval, LangSmith, Langfuse, LlamaIndex, LangChain, or LLM-as-a-JudgeExposure to OpenAI, Azure OpenAI, Claude, Gemini, or other LLM APIsKnowledge of vector databases such as FAISS, Pinecone, ChromaDB, or WeaviateUnderstanding of embeddings, chunking, retrieval, context precision/recall, and faithfulness metricsExperience testing agentic AI workflows, AI agents, MCP-based workflows, or multi-agent systemsPerformance testing exposure for AI systems, including latency, TTFC, streaming consistency, and response quality. .

One address, no account. We’ll tell you when matching roles go live.

More at STATUSNEO INC

Related open roles

View all roles