Source description
About the role
Qualifications
-
Bachelor’s degree in Computer Science, Engineering, Information Technology, Data Science, or a related field, or equivalent practical experience.
-
Three or more years of experience in machine learning engineering, DevOps, site reliability engineering, cloud engineering, software production support, or AI system operations.
-
Experience independently supporting business-critical applications or services in a production environment.
-
Strong working knowledge of at least one major cloud platform, such as AWS, Google Cloud Platform, or Microsoft Azure.
-
Hands-on experience with container technologies and orchestration platforms such as Docker and Kubernetes.
-
Experience with machine learning platforms or MLOps tools such as MLflow, Kubeflow, Vertex AI, SageMaker, or Azure Machine Learning.
-
Proficiency with Python and SQL for troubleshooting, scripting, data investigation, and automation.
-
Experience supporting CI/CD pipelines and infrastructure-as-code tools such as Terraform, CloudFormation, or equivalent technologies.
-
Ability to investigate application logs, infrastructure metrics, distributed-system failures, data-quality issues, and service dependencies.
-
Experience with incident management, problem management, root cause analysis, and production change-management practices.
-
Strong analytical, troubleshooting, and communication skills.
-
Ability to manage multiple priorities and make sound technical decisions in a fast-paced, 24/7 support environment.
-
Ability to participate in an on-call rotation, including occasional support outside standard business hours.
-
Strong English communication skills, including listening, speaking, and note-taking, with the ability to collaborate effectively in cross-functional teams.
Preferred Qualifications
-
Experience with monitoring, observability, and alerting tools such as Prometheus, Grafana, Datadog, Splunk, Cloud Monitoring, or CloudWatch.
-
Experience developing production-quality automation using Python, Bash, or another scripting language.
-
Understanding of the complete machine learning lifecycle, including data preparation, model training, validation, deployment, monitoring, retraining, and retirement.
-
Experience supporting real-time inference services, batch-scoring pipelines, generative AI applications, or AI agents.
-
Familiarity with service-level indicators, service-level objectives, error budgets, and site reliability engineering practices.
-
Experience with model monitoring, data drift, model drift, performance degradation, and AI-specific operational risks.
-
Familiarity with secure software-development practices, identity and access management, secrets management, and cloud security controls.
-
Experience mentoring junior engineers or leading technical incident investigations.
-
Experience supporting globally distributed systems, teams, or customers.
Why You'll Love Working Here 100% salary during probation period Annual Leave: 18 days/ year Five “Recharge Days” – Extra days, in addition to company holidays. 2 days WFH/ weef Flexible Friday afternoon Full salary insurance 13th-month bonus 1 day off for birthday Advanced health insurance (Generali) Regular engagement activities: sport clubs, internal event… Support Macbook and Monitor
Working location: Waseco building, 10 Pho Quang, Tan Son Hoa ward, Ho Chi Minh city, Vietnam.
More at Yum!
Related open roles
KFC Experience - Senior Frontend Engineer (ReactJS focus)
Ho Chi Minh, Dong Nam Bo, Viet Nam · Hybrid
Site Reliability Engineer III
Vietnam · Hybrid
Sr Machine Learning Engineer
Los Angeles · Hybrid
Sr. Frontend Software Engineer
Ho Chi Minh, Dong Nam Bo, Viet Nam · Hybrid
Sr. Software Engineer I
Vietnam · Hybrid
Machine Learning Engineer II
Dallas–Fort Worth · Hybrid