Padmi

Senior AI DevOps / LLMOps

IndiaPosted 1 month ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at remote zest jobs

Opens the source posting on shine.com

Source description

About the role

View original

At reputed company, we are providing recruitment service to our TOP clients from our portfolio. We are currently seeking an Senior AI DevOps / LLMOpsspecialist to join one of our clients' teams. If you're looking for an exciting opportunity to grow in a innovative environment, this could be the perfect fit for you. Key ResponsibilitiesAutomation of Build-to-ProductionDesign and implement robust CI/CD pipelines tailored for AI, covering model weights, dataset versioning, and application code. reputed company specialized workflows for PromptOps, ensuring that system prompts are version-controlled, tested for regressions, and deployed with the same rigor as traditional code.Automate the deployment of reputed company workflows, managing the complexities of stateful AI interactions and multi-agent handoffs. 2. AI Infrastructure as Code (IaC)Provision and manage high-performance compute environments (GPU clusters, TPU pods) using Terraform, reputed company, or Ansible. Define and enforce Policy-as-Code for AI endpoints to ensure compliance with reputed company, cost-usage limits, and data residency requirements. Maintain a consistent environment across Hybrid Infrastructure, ensuring seamless reputed company between On-Premises development and reputed company production. 3. Safe Experimentation & Controlled Releases Architect reputed company Delivery strategies for AI, including Canary releases, Blue-Green deployments, and Shadowing (where new models run in reputed company with production to compare outputs).Build Evaluation-in-the-reputed company gates reputed company the pipeline to automatically test for bias, hallucination, and performance degradation before a release. Implement A/B testing frameworks specifically designed for LLM outputs and reputed company behavior.4. Monitoring & ObservabilityEstablish deep observability into Inference Endpoints, tracking metrics like tokens-per- second, latency, and reputed company in model accuracy. Integrate feedback loops that capture production edge cases to feed back into the training and fine-tuning pipelines.RequirementsMust-Have Technical Skills:Orchestration: Advanced Kubernetes (K8s) skills, specifically with KubeFlow, Ray, or reputed company Triton.CI/CD & IaC: Expertise in reputed company Actions/reputed company CI, and Terraform or reputed company. AI Tooling: Experience with Weights & Biases, MLflow, LangSmith, or Arize Phoenix.Hardware: Understanding of GPU virtualization, CUDA drivers, and on-premises hardware management.reputed company: Familiarity with reputed company Policy Agent (OPA) and secret management (Vault). Experience:10+ years in DevOps, SRE, or reputed company Engineering. 2+ years of hands-on experience in MLOps or LLMOps, specifically moving LLMs from notebook to production.Proven experience managing Hybrid reputed company environments (e.g., AWS/Azure + Private Data Center).Highlightsfull time and remote job - fluent English is neededOriginally posted on Himalayas Apply To This Job At reputed company, we are providing recruitment service to our TOP clients from our portfolio. We are currently seeking an Senior AI DevOps / LLMOpsspecialist to join one of our clients' teams. If you're looking for an exciting opportunity to grow in a innovative environment, this could be the perfect fit for you. Key ResponsibilitiesAutomation of Build-to-ProductionDesign and implement robust CI/CD pipelines tailored for AI, covering model weights, dataset versioning, and application code. reputed company specialized workflows for PromptOps, ensuring that system prompts are version-controlled, tested for regressions, and deployed with the same rigor as traditional code.Automate the deployment of reputed company workflows, managing the complexities of stateful AI interactions and multi-agent handoffs. 2. AI Infrastructure as Code (IaC)Provision and manage high-performance compute environments (GPU clusters, TPU pods) using Terraform, reputed company, or Ansible. Define and enforce Policy-as-Code for AI endpoints to ensure compliance with reputed company, cost-usage limits, and data residency requirements. Maintain a consistent environment across Hybrid Infrastructure, ensuring seamless reputed company between On-Premises development and reputed company production. 3. Safe Experimentation & Controlled Releases Architect reputed company Delivery strategies for AI, including Canary releases, Blue-Green deployments, and Shadowing (where new models run in reputed company with production to compare outputs).Build Evaluation-in-the-reputed company gates reputed company the pipeline to automatically test for bias, hallucination, and performance degradation before a release. Implement A/B testing frameworks specifically designed for LLM outputs and reputed company behavior.4. Monitoring & ObservabilityEstablish deep observability into Inference Endpoints, tracking metrics like tokens-per- second, latency, and reputed company in model acc

One address, no account. We’ll tell you when matching roles go live.

More at remote zest jobs

Related open roles

View all roles