Padmi

Data Scientist

IndiaPosted 3 months ago
Data Science And StatisticsSeniorFull Time; Regular
Apply at Giggso

Opens the source posting on shine.com

Source description

About the role

View original

As a Data Scientist, you will be responsible for designing and implementing end-to-end evaluation frameworks to assess the performance, reliability, and safety of multi-agent AI systems. Your role will involve leading experimentation and A/B testing efforts to systematically test hypotheses, validate model improvements, and track performance across agent iterations. Additionally, you will curate and maintain high-quality ground truth datasets to enable accurate and reproducible evaluation of multi-agent outputs. It will be your responsibility to identify and address reliability and accuracy gaps across agent workflows, failure modes, and edge cases in production-like environments. To excel in this role, you are expected to stay current on emerging research in agentic AI, LLM evaluation, and multi-agent coordination to continuously improve framework design. Key Responsibilities: - Design and implement end-to-end evaluation frameworks for multi-agent AI systems - Lead experimentation and A/B testing efforts for validating model improvements - Curate and maintain high-quality ground truth datasets for accurate evaluation - Identify and address reliability and accuracy gaps in production-like environments - Stay updated on emerging research to improve framework design Qualifications Required: - Proficiency in Python and ML frameworks - Hands-on experience with LLM APIs and agentic frameworks (LangChain, LlamaIndex, Semetic KernalI) - Familiarity with evaluation tooling such as Ragas, DeepEval, LangSmith, or similar - Experience with data pipelines, experiment tracking (MLflow, W&B), and CI/CD for ML workflows - Strong foundation in statistics, NLP, prompt engineering, experimental design, and A/B testing methodology - Proficiency in Azure ML, Azure OpenAI Service, and Azure AI Foundry for model deployment, evaluation, and orchestration - Familiarity with Azure Monitor and Application Insights for tracking reliability and performance of deployed agent systems As a Data Scientist, you will be responsible for designing and implementing end-to-end evaluation frameworks to assess the performance, reliability, and safety of multi-agent AI systems. Your role will involve leading experimentation and A/B testing efforts to systematically test hypotheses, validate model improvements, and track performance across agent iterations. Additionally, you will curate and maintain high-quality ground truth datasets to enable accurate and reproducible evaluation of multi-agent outputs. It will be your responsibility to identify and address reliability and accuracy gaps across agent workflows, failure modes, and edge cases in production-like environments. To excel in this role, you are expected to stay current on emerging research in agentic AI, LLM evaluation, and multi-agent coordination to continuously improve framework design. Key Responsibilities: - Design and implement end-to-end evaluation frameworks for multi-agent AI systems - Lead experimentation and A/B testing efforts for validating model improvements - Curate and maintain high-quality ground truth datasets for accurate evaluation - Identify and address reliability and accuracy gaps in production-like environments - Stay updated on emerging research to improve framework design Qualifications Required: - Proficiency in Python and ML frameworks - Hands-on experience with LLM APIs and agentic frameworks (LangChain, LlamaIndex, Semetic KernalI) - Familiarity with evaluation tooling such as Ragas, DeepEval, LangSmith, or similar - Experience with data pipelines, experiment tracking (MLflow, W&B), and CI/CD for ML workflows - Strong foundation in statistics, NLP, prompt engineering, experimental design, and A/B testing methodology - Proficiency in Azure ML, Azure OpenAI Service, and Azure AI Foundry for model deployment, evaluation, and orchestration - Familiarity with Azure Monitor and Application Insights for tracking reliability and performance of deployed agent systems

One address, no account. We’ll tell you when matching roles go live.

More at Giggso

Related open roles

View all roles