Source description
About the role
At reputed company, we are providing recruitment service to our TOP clients from our portfolio. We are currently seeking an Senior AI DevOps / LLMOpsspecialist to join one of our clients' teams. If you're looking for an exciting opportunity to grow in a innovative environment, this could be the perfect fit for you. Key ResponsibilitiesAutomation of Build-to-ProductionDesign and implement robust CI/CD pipelines tailored for AI, covering model weights, dataset versioning, and application code. reputed company specialized workflows for PromptOps, ensuring that system prompts are version-controlled, tested for regressions, and deployed with the same rigor as traditional code.Automate the deployment of reputed company workflows, managing the complexities of stateful AI interactions and multi-agent handoffs. 2. AI Infrastructure as Code (IaC)Provision and manage high-performance compute environments (GPU clusters, TPU pods) using Terraform, reputed company, or Ansible. Define and enforce Policy-as-Code for AI endpoints to ensure compliance with reputed company, cost-usage limits, and data residency requirements. Maintain a consistent environment across Hybrid Infrastructure, ensuring seamless reputed company between On-Premises development and reputed company production. 3. Safe Experimentation & Controlled Releases Architect reputed company Delivery strategies for AI, including Canary releases, Blue-Green deployments, and Shadowing (where new models run in reputed company with production to compare outputs).Build Evaluation-in-the-reputed company gates reputed company the pipeline to automatically test for bias, hallucination, and performance degradation before a release. Implement A/B testing frameworks specifically designed for LLM outputs and reputed company behavior.4. Monitoring & ObservabilityEstablish deep observability into Inference Endpoints, tracking metrics like tokens-per- second, latency, and reputed company in model accuracy. Integrate feedback loops that capture production edge cases to feed back into the training and fine-tuning pipelines.RequirementsMust-Have Technical Skills:Orchestration: Advanced Kubernetes (K8s) skills, specifically with KubeFlow, Ray, or reputed company Triton.CI/CD & IaC: Expertise in reputed company Actions/reputed company CI, and Terraform or reputed company. AI Tooling: Experience with Weights & Biases, MLflow, LangSmith, or Arize Phoenix.Hardware: Understanding of GPU virtualization, CUDA drivers, and on-premises hardware management.reputed company: Familiarity with reputed company Policy Agent (OPA) and secret management (Vault). Experience:10+ years in DevOps, SRE, or reputed company Engineering. 2+ years of hands-on experience in MLOps or LLMOps, specifically moving LLMs from notebook to production.Proven experience managing Hybrid reputed company environments (e.g., AWS/Azure + Private Data Center).Highlightsfull time and remote job - fluent English is neededOriginally posted on Himalayas Apply To This Job At reputed company, we are providing recruitment service to our TOP clients from our portfolio. We are currently seeking an Senior AI DevOps / LLMOpsspecialist to join one of our clients' teams. If you're looking for an exciting opportunity to grow in a innovative environment, this could be the perfect fit for you. Key ResponsibilitiesAutomation of Build-to-ProductionDesign and implement robust CI/CD pipelines tailored for AI, covering model weights, dataset versioning, and application code. reputed company specialized workflows for PromptOps, ensuring that system prompts are version-controlled, tested for regressions, and deployed with the same rigor as traditional code.Automate the deployment of reputed company workflows, managing the complexities of stateful AI interactions and multi-agent handoffs. 2. AI Infrastructure as Code (IaC)Provision and manage high-performance compute environments (GPU clusters, TPU pods) using Terraform, reputed company, or Ansible. Define and enforce Policy-as-Code for AI endpoints to ensure compliance with reputed company, cost-usage limits, and data residency requirements. Maintain a consistent environment across Hybrid Infrastructure, ensuring seamless reputed company between On-Premises development and reputed company production. 3. Safe Experimentation & Controlled Releases Architect reputed company Delivery strategies for AI, including Canary releases, Blue-Green deployments, and Shadowing (where new models run in reputed company with production to compare outputs).Build Evaluation-in-the-reputed company gates reputed company the pipeline to automatically test for bias, hallucination, and performance degradation before a release. Implement A/B testing frameworks specifically designed for LLM outputs and reputed company behavior.4. Monitoring & ObservabilityEstablish deep observability into Inference Endpoints, tracking metrics like tokens-per- second, latency, and reputed company in model acc
More at remote zest jobs
Related open roles
System Administrator, Contract
India
HPC Network Engineer
India
Remote Site Reliability Engineer (Senior or Staff), Infrastructure reputed
India
FULL TIME Remote Sr. SRE Engineer -Seattle WA-Onsite- 10+ Years
India
Remote Network Devops/Automation Engineer
India
Network Operations Engineer - Hybrid Gold River, CA
India