Source description
About the role
Technical
Proven experience with AWS managed services: SageMaker, Bedrock, EKS, EC2, S3, Lambda, API Gateway, RDS, CloudTrail, and CloudWatch.
Proven experience with Kubernetes, both on cloud and on-premises.
Proven experience designing and operating multi-stage ETL pipelines for ML training and inference.
Proven experience setting up observability for ML models (training and inference), such as with Weights & Biases (W&B).
Solid understanding of platform engineering best practices and patterns.
Hands-on experience with Infrastructure-as-Code tools (Terraform, Helm, or similar).
Leadership & Collaboration
Ability to work closely with ML engineers, backend engineers, and platform stakeholders on shared, cross-functional systems.
Comfortable establishing and enforcing platform patterns and best practices across teams.
Clear communicator, able to align infrastructure decisions with the needs of model development and deployment workflows.
Mindset
Reliability-minded: builds systems that are observable, scalable, and built to last, not just functional.
Curious about the full ML lifecycle, from data ingestion and training through to production deployment, rather than infrastructure in isolation.
Pragmatic and standards-driven, with a bias toward reusable platform patterns over one-off solutions.
Nice to Have
Experience securing managed and self-hosted AI platforms, including ChatUI integrations, MCP servers, and backend services.
Familiarity with Apache Kafka and event-driven architectures.
Hands-on experience with Snowflake integrations.
Experience with ETL-as-Code frameworks such as dbt.
Hands-on experience with workflow orchestrators such as Prefect or equivalent (e.g., Airflow, Dagster).
Proven experience with AWS IAM and account management.
Familiarity with ML frameworks such as PyTorch, Hugging Face, or scikit-learn.
What This Role Is Not:
This is not a pure DevOps or SRE role. You will work directly with ML systems, training pipelines, and model deployment - not just maintain cloud infrastructure.
This is not a data engineering role. While you will build and operate data pipelines, the focus is enabling ML training and inference workflows, not analytics or business reporting.
More at Protolabs
Related open roles
Mechanical Engineer – CNC Manufacturing & CAD/CAM Automation
Hyderabad · Onsite
Ml Platform & Full Stack Ai Engineer Hyderabad (India)
Hyderabad
Sr. Quality Engineer - Corp Systems (Hyderabad)
Hyderabad
Product Data Analyst (India)
India
Sr. SQE Performance Testing (Hyderabad)
Hyderabad
IC1 - ML Platform & Full stack AI Engineer
Hyderabad
