Source description
About the role
As an experienced ML Platform Architect, you will be responsible for architecting and evolving a scalable, secure ML platform that supports the end-to-end ML lifecycle. Your key responsibilities will include: - Designing and implementing core ML infrastructure for model training, hyperparameter tuning, experiment tracking, and model registry using cloud native and open source technologies. - Orchestrating ML workflows using tools such as Kubeflow, SageMaker, MLflow, Vertex AI, or Databricks to ensure reproducibility and robust automation. - Partnering with Data Engineers to build reliable, high-quality data pipelines and feature pipelines that provide model-ready datasets at scale. - Implementing and improving CI/CD for ML, including automated testing, validation, and safe rollout/rollback of models and data pipelines. - Optimizing ML workload performance and cost across compute, storage, and networking layers on public cloud platforms. - Embedding observability and governance into the platform, including logging, tracing, model performance monitoring, and drift detection. - Collaborating with security, compliance, and data governance teams to ensure adherence to standards for security, privacy, and auditability. - Providing technical leadership and mentorship to other engineers, influencing architectural decisions, coding standards, and best practices for ML platform and MLOps. - Continuously exploring and applying AI/automation to improve developer and data scientist productivity, platform reliability, and operational efficiency. Qualifications Required: - Demonstrated ability to embrace AI and use data-driven insights for innovation and continuous improvement. - Bachelors or Masters degree in Computer Science, Engineering, or a related field. - 10+ years of software engineering experience, including 5+ years working on ML platforms or infrastructure. - Expertise in building large-scale distributed systems and microservices, with a solid understanding of system design and architecture. - Strong programming skills in Python, Go, or Java, with an emphasis on writing clean, testable, maintainable code. - Hands-on experience with containerization and orchestration, for example Docker and Kubernetes. - Familiarity with MLOps tools such as MLflow, Kubeflow, SageMaker, and their integration into an end-to-end ML platform. - Cloud platform experience on AWS, GCP, or Azure, including core services for compute, storage, networking, and identity. - Experience with statistical learning algorithms and deep learning approaches, along with knowledge of training and deployment processes. - Strong communication, leadership, and problem-solving skills with the ability to work effectively in a global, distributed environment. Preferred Qualifications: - Experience with real-time model inference, streaming ML pipelines, and low-latency serving. - Deep knowledge of model governance, reproducibility, and monitoring, including experiment lineage and versioning. - Understanding of model performance metrics, drift detection, and monitoring for data drift and concept drift. - Exposure to feature stores and workflow orchestration tools in production environments. - Familiarity with regulatory and compliance considerations for ML systems, including model auditability and data privacy laws. - Experience with real-time data pipelines and streaming technologies. - Experience using infrastructure-as-code tools for CI/CD of platform components. - Domain experience in insurance or related industries or a demonstrated ability to ramp quickly in highly regulated domains. As an experienced ML Platform Architect, you will be responsible for architecting and evolving a scalable, secure ML platform that supports the end-to-end ML lifecycle. Your key responsibilities will include: - Designing and implementing core ML infrastructure for model training, hyperparameter tuning, experiment tracking, and model registry using cloud native and open source technologies. - Orchestrating ML workflows using tools such as Kubeflow, SageMaker, MLflow, Vertex AI, or Databricks to ensure reproducibility and robust automation. - Partnering with Data Engineers to build reliable, high-quality data pipelines and feature pipelines that provide model-ready datasets at scale. - Implementing and improving CI/CD for ML, including automated testing, validation, and safe rollout/rollback of models and data pipelines. - Optimizing ML workload performance and cost across compute, storage, and networking layers on public cloud platforms. - Embedding observability and governance into the platform, including logging, tracing, model performance monitoring, and drift detection. - Collaborating with security, compliance, and data governance teams to ensure adherence to standards for security, privacy, and auditability. - Providing technical leadership and mentorship to other engineers, influencing architectural decisions, coding standards, and best practices for ML platform