Source description
About the role
Provider of platform for machine learning model training and deployment. It offers tools for fine-tuning, deploying, and observing machine learning models. Design, develop, and maintain backend and platform services in Python Build and evolve REST and gRPC APIs Develop deployment, evaluation, and orchestration workflows Write production-ready, scalable, and maintainable code Implement robust testing, logging, and error-handling practices Support containerized deployments using Docker and Kubernetes Build systems with strong observability through metrics, logs, and tracing Investigate and resolve production issues Improve platform reliability, scalability, and performance Participate in architecture and technical design discussions Collaborate closely with machine learning and infrastructure teams Own platform components throughout their lifecycle from development to production support Build and scale core platform and backend services that power AI model deployment, orchestration, monitoring, and operational workflows. Own production-grade systems, APIs, internal tooling, and infrastructure-adjacent services with a focus on reliability, scalability, observability, and developer experience. Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
More at Coderound Ai