Source description
About the role
HIRING ALERT! AI QA ENGINEER ! Role: AI QA Engineer Work Location: Navi Mumbai/ Bangalore Experience: 3-6 Years Preferred Immediate Joiner. Applicants please share your updated resumes on [HIDDEN TEXT] Role Overview We are looking for an AI QA Engineer with hands-on experience working on GenAI platforms and workflows, particularly AI observability, evaluation, testing, or safety platforms similar in nature to Arize, Cekura, Fiddler, Giskard, WhyLabs, Phoenix, or equivalent systems. This role requires direct exposure to GenAI bots, RAG pipelines, and AI workflows in real platforms, not theoretical knowledge or generic QA experience. You will evaluate AI behavior, analyze failures, and contribute to improving the quality, reliability, and safety of GenAI systems. This is not a traditional QA role and not suitable for candidates who have only tested UI, APIs, or deterministic systems. Key Responsibilities Execute hands-on AI evaluation and QA on: GenAI bots (chatbots, copilots, agents) RAG-based workflows AI evaluation and observability platforms Design and execute: Prompt-based test scenarios Multi-turn conversation tests Edge-case and adversarial evaluations Evaluate AI behavior for: Hallucinations Grounding failures Inconsistency and drift Safety and reliability issues Work within existing AI platforms (evaluation, observability, monitoring) to: Analyze traces, metrics, and outputs Interpret model behavior across runs Document findings clearly in: Evaluation reports Platform dashboards Structured test summaries Collaborate with: AI QA Lead on methodology and review ML and Product teams to provide actionable feedback Must-Have Skills & Experience 3–5+ years of experience in AI QA, AI evaluation, or GenAI testing Hands-on experience working on GenAI platforms, such as: AI evaluation platforms AI observability or monitoring tools GenAI safety or reliability platforms Direct experience testing: GenAI chatbots or agents RAG pipelines and workflows Multi-step AI interactions Strong understanding of: Prompt-based testing Non-deterministic AI behavior Quality vs correctness in GenAI systems Ability to clearly explain why an AI response is good or bad Good-to-Have Experience with platforms like Arize, Cekura, Fiddler, Giskard, Phoenix, WhyLabs, or similar Familiarity with evaluation metrics and scoring approaches Exposure to AI safety, adversarial testing, or red-teaming Basic scripting or data analysis skills
More at QualityKiosk Technologies Private Limited