Full Description As an experienced Data Scientist and Agentic AI Developer at our company, your role will involve designing, developing, and evaluating intelligent AI agents to deliver high-quality, reliable, and compliant solutions. This position requires a blend of data science expertise, AI/ML engineering capabilities, and hands-on experience building agentic AI systems. Key Responsibilities:* - Design anddevelop agentic AI systems powered by Large Language Models (LLMs) with tool-calling capabilities - Build and optimize multi-step reasoning workflows and agent orchestration frameworks - Implement retrieval-augmented generation (RAG) pipelines for knowledge-intensive applications - Integrate external tools, APIs, and databases into agent workflows - Deploy and monitor production-grade AI agents at scale - Develop comprehensive evaluation frameworks using LLM-as-a-judge methodologies - Implement automated scoring systems for output quality metrics (correctness, helpfulness, coherence, relevance) - Design and execute robustness testing including adversarial attack scenarios - Monitor and reduce hallucination rates and ensure factual accuracy - Track performance metrics including latency, throughput, and cost-per-interaction - Analyze agent performance data to identify improvement opportunities - Build custom evaluation pipelines and scoring rubrics - Conduct A/B testing and statistical analysis for model optimization - Create dashboards and visualization tools for stakeholder reporting - Implement RAGAs (Retrieval Augmented Generation Assessment) frameworks - Ensure AI systems meet ethical standards including bias detection and fairness - Implement safety guardrails to prevent harmful content generation - Develop compliance monitoring systems for regulatory frameworks (EU AI Act, GDPR, HIPAA, DPDP) - Document transparency and explainability measures - Establish human oversight protocols Required Skills & Qualifications:* - Strong proficiency in Python; experience with AI/ML frameworks (LangChain, LangSmith, Phoenix, or similar) - Hands-on experience with GPT, Claude, or other frontier models; prompt engineering and fine-tuning - Deep understanding of NLP, deep learning architectures, and model evaluation - Experience with MLOps tools, vector databases, and observability platforms - Proficiency in SQL, data pipelines, and ETL processes - Understanding of agentic AI architectures and autonomous systems - Knowledge of RAG systems and information retrieval techniques - Familiarity with LLM evaluation methodologies and benchmarks - Experience with conversational AI and dialogue systems - Understanding of AI safety, alignment, and interpretability - Experience designing evaluation rubrics and scoring systems - Proficiency with automated evaluation frameworks (RAGAs, custom evaluators) - Understanding of quality metrics: coherence, fluency, factual accuracy, hallucination detection - Knowledge of performance metrics: latency optimization, token usage, throughput analysis - Experience with user experience metrics (CSAT, NPS, turn count analysis) What you'll Deliver:* - Production-ready agentic AI systems with measurable quality improvements - Comprehensive evaluation frameworks with automated scoring - Performance dashboards and reporting systems - Documentation for technical specifications and compliance standards - Continuous improvement strategies based on data-driven insights, ## Full Description As an experienced Data Scientist and Agentic AI Developer at our company, your role will involve designing, developing, and evaluating intelligent AI agents to deliver high-quality, reliable, and compliant solutions. This position requires a blend of data science expertise, AI/ML engineering capabilities, and hands-on experience building agentic AI systems. Key Responsibilities:* - Design anddevelop agentic AI systems powered by Large Language Models (LLMs) with tool-calling capabilities - Build and optimize multi-step reasoning workflows and agent orchestration frameworks - Implement retrieval-augmented generation (RAG) pipelines for knowledge-intensive applications - Integrate external tools, APIs, and databases into agent workflows - Deploy and monitor production-grade AI agents at scale - Develop comprehensive evaluation frameworks using LLM-as-a-judge methodologies - Implement automated scoring systems for output quality metrics (correctness, helpfulness, coherence, relevance) - Design and execute robustness testing including adversarial attack scenarios - Monitor and reduce hallucination rates and ensure factual accuracy - Track performance metrics including latency, throughput, and cost-per-interaction - Analyze agent performance data to identify improvement opportunities - Build custom evaluation pipelines and scoring rubrics - Conduct A/B testing and statistical analysis for model optimization - Create dashboards and visualization t