Source description
About the role
Exp-6-8 Years Notice-Immediate-serving 10 Days Location-Pune (Hinjewadi-Phase 3) Mode-2 Days from Office Key Responsibilities : Design and execute test cases fo r LLM agents, RAG pipelines, agentic workflo ws, and AI-assisted decision too lsValidate AI outputs agains t ground tru th using structured accuracy scoring (NAICS, risk exposure flags, hazard group mappin g)Detec t hallucinations, reasoning gaps, source fabrication, and misattributio ns in model-generated conte ntRu n multi-model comparative testi ng across GPT, Claude, Gemini, and Perplexity — evaluating accuracy, latency, and output completene ssTes t prompt versions iterative ly and track accuracy changes across prompt cycl esValidat e citation accuracy, document ingestion pipelin es, and cross-document context handli ngDesig n edge case and negative tes ts for AI-specific failure modes — content filter triggers, tool call limits, missing documents, and incomplete synthes isPerfor m regression testi ng after model upgrades, prompt changes, and backend fixes, and maintain structured QA sign-off in JI RA What Makes This Role Different from Traditional QAYou evalua te whether an AI is reasoning correc tly — not just whether the UI behaves as expec tedYou bui ld evaluation rubrics for non-deterministic outp uts and app ly LLM-as-a-Judge techniq ues to score quality at sc aleYou treat eve ry model or prompt change as a potential accuracy regress ion, not just a functional oneYou understand that in live AI system s, a passing test today does not guarantee a passing test tomor row Required Skill sets7–8 years of QA experience w ith minimum 2 years in Generative AI / LLM-based proj ectsHands-on experience test ing chatbots, RAG systems, or agentic AI pipel inesProven ability to perf orm ground truth valida tion and detect hallucinations and reasoning fail uresFamiliarity w ith multi-model evalua tion, prompt-aware testing, and JIRA-based defect repor tingPreferred Skill setsBackground in insurance or regulated indust ries; exposure to underwriting or risk classification conc eptsFamiliarity w ith Azure OpenAI, AWS Bed rock, or SharePoint-integrated AI environm entsKnowledge of AI governance, content filtering, and PII redac tion valida tion
More at Fulcrum Digital