Source description
About the role
We are looking for an experienced MLOps Engineer to design, build, deploy, and manage enterprise-scale Machine Learning and Generative AI solutions. The ideal candidate should have strong expertise in Python, CI/CD, Docker, Kubernetes (AKS), Azure Cloud, and MLOps platforms to operationalize AI models in production environments. Key Responsibilities MLOps & AI Platform Engineering Design, develop, and maintain end-to-end MLOps pipelines for model training, validation, deployment, monitoring, retraining, and retirement.Build and operate scalable AI/ML platforms across development, testing, and production environments.Deploy and manage production-grade Machine Learning and Generative AI applications.Support both batch and real-time inference workloads.Develop production-ready applications and automation using Python.Build and maintain Git-based CI/CD pipelines for automated deployments.Implement Infrastructure as Code (IaC) using Terraform, ARM Templates, Bicep, or CloudFormation. Containerization & Cloud Containerize AI applications using Docker.Deploy and manage workloads on Kubernetes, preferably Azure Kubernetes Service (AKS).Design cloud-native solutions on Azure (preferred), AWS, or GCP.Implement deployment strategies such as Blue-Green, Canary, and Rolling Deployments. Generative AI & LLMs Develop and deploy LLM-based applications.Implement Retrieval-Augmented Generation (RAG) architectures.Work with Prompt Engineering, Embeddings, and Vector Databases.Integrate AI solutions with Azure OpenAI, OpenAI, Anthropic Claude, or AWS Bedrock.Build Agentic AI solutions using frameworks like LangChain, LangGraph, or similar orchestration tools. Monitoring & Operations Implement monitoring, logging, metrics, tracing, and alerting for AI services.Monitor model performance, latency, availability, and model drift.Provide production support and resolve deployment or infrastructure issues. Security & Governance Implement IAM, RBAC, secrets management, encryption, and network security.Ensure enterprise compliance, governance, and audit readiness for AI platforms. Required Skills 513 years of experience in MLOps, AI Engineering, or Platform Engineering.Strong programming skills in Python.Hands-on experience with Linux.Experience deploying Machine Learning models into production.Expertise with Docker and Kubernetes (AKS preferred).Experience building CI/CD pipelines using Azure DevOps, GitHub Actions, Jenkins, or GitLab CI.Experience with Infrastructure as Code (Terraform, ARM, Bicep, or CloudFormation).Experience with Azure, AWS, or GCP cloud platforms.Strong understanding of MLOps lifecycle and model governance. Preferred Skills Azure Cloud (Preferred)Azure Machine LearningMLflow, Kubeflow, SageMaker, or Vertex AIAzure OpenAI/OpenAI/Anthropic/AWS BedrockLangChain, LangGraph, Semantic Kernel, CrewAIRAG, Embeddings, Vector Databases (Pinecone, FAISS, ChromaDB, Weaviate)Prometheus, Grafana, Azure Monitor, ELK StackGit, REST APIs, FastAPI, FlaskHigh-availability AI platform designITIL or Enterprise Service Management knowledge Mandatory Skills PythonMLOpsModel Development & DeploymentCI/CDDockerKubernetes / AKSAzure CloudGitInfrastructure as Code (Terraform/ARM/Bicep)Linux Skills: azure,ci/cd,mlops,kubernetes We are looking for an experienced MLOps Engineer to design, build, deploy, and manage enterprise-scale Machine Learning and Generative AI solutions. The ideal candidate should have strong expertise in Python, CI/CD, Docker, Kubernetes (AKS), Azure Cloud, and MLOps platforms to operationalize AI models in production environments. Key Responsibilities MLOps & AI Platform Engineering Design, develop, and maintain end-to-end MLOps pipelines for model training, validation, deployment, monitoring, retraining, and retirement.Build and operate scalable AI/ML platforms across development, testing, and production environments.Deploy and manage production-grade Machine Learning and Generative AI applications.Support both batch and real-time inference workloads.Develop production-ready applications and automation using Python.Build and maintain Git-based CI/CD pipelines for automated deployments.Implement Infrastructure as Code (IaC) using Terraform, ARM Templates, Bicep, or CloudFormation. Containerization & Cloud Containerize AI applications using Docker.Deploy and manage workloads on Kubernetes, preferably Azure Kubernetes Service (AKS).Design cloud-native solutions on Azure (preferred), AWS, or GCP.Implement deployment strategies such as Blue-Green, Canary, and Rolling Deployments. Generative AI & LLMs Develop and deploy LLM-based applications.Implement Retrieval-Augmented Generation (RAG) architectures.Work with Prompt Engineering, Embeddings, and Vector Databases.Integrate AI solutions with Azure OpenAI, OpenAI, Anthropic Claude, or AWS Bedrock.Build Agentic AI solutions using frameworks like LangChain, LangGraph, or similar orchestration tools. Monitoring & Operations Implement monitoring, log
More at Zorba AI