Source description
About the role
Work Location:Bangalore/Chennai/Hyd NP: Immediate joiners Job Description: Build and manage AI/ML platforms and pipelines (data, training, deployment, monitoring) Ensure high availability, scalability, and performance of RL models in production Implement CI/CD for ML (MLOps) including model versioning, rollback, and automation Set up observability frameworks (model performance, drift, system health) Enable real-time decisioning systems for yield optimization use cases Ensure security, compliance, and cost optimization of AI workloads Support disaster recovery and resilience engineering for AI systems. Preferred Skills AWS (EKS, EC2, S3, IAM, CloudWatch, Cost Explorer)MLflow, DVC, Airflow/MWAAPython, Kubernetes, DockerGrafana, Prometheus, ELK/OpenSearch, OpenTelemetryCI/CD platforms (GitHub Actions, GitLab CI, Jenkins)Reinforcement Learning (RL) infrastructure and model operationsMLOps, FinOps, Site Reliability Engineering (SRE), AIOps Ideal Candidate 815 years of experience in Cloud, Platform Engineering, DevOps, MLOps, or AI Infrastructure.Proven experience building AI/ML platforms and production-grade machine learning systems.Strong expertise in AWS-native services and cloud cost optimization.Ability to architect end-to-end AI platforms that simplify the experience for Data Scientists while providing enterprise-grade governance and observability. Work Location:Bangalore/Chennai/Hyd NP: Immediate joiners Job Description: Build and manage AI/ML platforms and pipelines (data, training, deployment, monitoring) Ensure high availability, scalability, and performance of RL models in production Implement CI/CD for ML (MLOps) including model versioning, rollback, and automation Set up observability frameworks (model performance, drift, system health) Enable real-time decisioning systems for yield optimization use cases Ensure security, compliance, and cost optimization of AI workloads Support disaster recovery and resilience engineering for AI systems. Preferred Skills AWS (EKS, EC2, S3, IAM, CloudWatch, Cost Explorer)MLflow, DVC, Airflow/MWAAPython, Kubernetes, DockerGrafana, Prometheus, ELK/OpenSearch, OpenTelemetryCI/CD platforms (GitHub Actions, GitLab CI, Jenkins)Reinforcement Learning (RL) infrastructure and model operationsMLOps, FinOps, Site Reliability Engineering (SRE), AIOps Ideal Candidate 815 years of experience in Cloud, Platform Engineering, DevOps, MLOps, or AI Infrastructure.Proven experience building AI/ML platforms and production-grade machine learning systems.Strong expertise in AWS-native services and cloud cost optimization.Ability to architect end-to-end AI platforms that simplify the experience for Data Scientists while providing enterprise-grade governance and observability.
More at Sonata Software