Source description
About the role
Location & Work Mode: Hyderabad (On-Site) Open to Relocation Employment Type: Full-time Experience Level: 35 years About the Role: You will architect and manage resilient, production-grade AI systems at scale within the Statistical Experiments and Machine Learning Lab at the Department of Data Science and AI, IIT Madras, while working for the National Institute of Rural Development and Panchayati Raj (NIRDPR). This role focuses on building operational stability for research platforms, optimizing infrastructure costs, and ensuring that ML models remain reliable and performant in high-demand academic and experimental environments. Key Responsibilities System Architecture: Design robust, scalable architectures for AI services, prioritizing uptime and seamless integration with existing core infrastructure.Operational Resilience: Implement failover mechanisms and health monitoring to maintain the stability of production models under varying loads.Cost Efficiency: Optimize GPU/CPU utilization and cloud resource allocation to reduce the total cost of ownership for large-scale AI deployments.Infrastructure Management: Leverage IaC and container orchestration to provide a stable, reproducible environment for the entire ML lifecycle.Performance Tuning: Conduct deep-level optimization of model inference and data retrieval to meet strict latency and throughput requirements. Must-Have Requirements Strong fundamentals in Probability, Statistics and Mathematics.Solid grasp of Deep Learning mechanics (backpropagation, attention mechanisms) and fluency in PyTorch or TensorFlow.Cloud Infrastructure management (e.g., AWS SageMaker) with proficiency in cloud architecture.Advanced specialization in GenAI/LLMOps with demonstrable experience in scaling production-grade AI systems and optimizing infrastructure for cost and performance.Strong problem-solving mindset with the ability to translate complex technical constraints into clear, non-technical business outcomes for stakeholders.Proficiency with AI coding assistants (e.g., Cursor, Cloud Code, GitHub Copilot) for rapid development and demo creation.Experience with ML/AI frameworks and deployment tools (e.g., VLLM, MLflow).Proven experience in Solution Architecture.Preferred / Nice-to-Have Hands-on GenAI/LLM experience (Fine-tuning, RAG, Vector DBs)Experiment tracking (MLflow / Weights & Biases); LLMOps exposure.Ability to maintain clean Technical Documentation.Education Degree in CS, ML, math, or related field; M.Tech/MS a plus for depth. Location & Work Mode: Hyderabad (On-Site) Open to Relocation Employment Type: Full-time Experience Level: 35 years About the Role: You will architect and manage resilient, production-grade AI systems at scale within the Statistical Experiments and Machine Learning Lab at the Department of Data Science and AI, IIT Madras, while working for the National Institute of Rural Development and Panchayati Raj (NIRDPR). This role focuses on building operational stability for research platforms, optimizing infrastructure costs, and ensuring that ML models remain reliable and performant in high-demand academic and experimental environments. Key Responsibilities System Architecture: Design robust, scalable architectures for AI services, prioritizing uptime and seamless integration with existing core infrastructure.Operational Resilience: Implement failover mechanisms and health monitoring to maintain the stability of production models under varying loads.Cost Efficiency: Optimize GPU/CPU utilization and cloud resource allocation to reduce the total cost of ownership for large-scale AI deployments.Infrastructure Management: Leverage IaC and container orchestration to provide a stable, reproducible environment for the entire ML lifecycle.Performance Tuning: Conduct deep-level optimization of model inference and data retrieval to meet strict latency and throughput requirements. Must-Have Requirements Strong fundamentals in Probability, Statistics and Mathematics.Solid grasp of Deep Learning mechanics (backpropagation, attention mechanisms) and fluency in PyTorch or TensorFlow.Cloud Infrastructure management (e.g., AWS SageMaker) with proficiency in cloud architecture.Advanced specialization in GenAI/LLMOps with demonstrable experience in scaling production-grade AI systems and optimizing infrastructure for cost and performance.Strong problem-solving mindset with the ability to translate complex technical constraints into clear, non-technical business outcomes for stakeholders.Proficiency with AI coding assistants (e.g., Cursor, Cloud Code, GitHub Copilot) for rapid development and demo creation.Experience with ML/AI frameworks and deployment tools (e.g., VLLM, MLflow).Proven experience in Solution Architecture.Preferred / Nice-to-Have Hands-on GenAI/LLM experience (Fine-tuning, RAG, Vector DBs)Experiment tracking (MLflow / Weights & Biases); LLMOps exposure.Ability to maintain clean Technical Documentation.Education Degree in
More at Department of Data Science & AI, IIT Madras