Source description
About the role
As an AI Platform Engineer, your role involves building, managing, and enhancing AI/ML infrastructure, workflows, and automation pipelines to create scalable platforms for training and deploying machine learning models. You will work closely with data scientists and platform engineering teams to enable efficient resource management and scalable operations across cloud and hybrid ecosystems. Key Responsibilities: - Design, deploy, and manage Kubernetes environments optimized for AI/ML applications and workloads to ensure scalability, reliability, and performance of containerized AI platforms. - Implement and manage GPU orchestration solutions such as Run:ai for workload scheduling and resource optimization, enabling efficient GPU allocation and utilization for AI model training and inference. - Build and maintain Python-based automation tools and machine learning pipelines for infrastructure deployment using Terraform and configuration management through Ansible. - Develop and maintain Jupyter Notebook environments to support experimentation, research, and collaborative model development. - Configure and optimize NVIDIA Enterprise Suite technologies including CUDA, NeMo Framework, Triton, TensorRT, and GPU drivers to support accelerated AI computing. - Implement MLOps standards and practices covering model lifecycle management, CI/CD pipelines, monitoring, and governance using tools such as MLflow and Kubeflow. - Partner with data scientists, ML engineers, and platform teams to enhance scalability, operational efficiency, and resource utilization across cloud and hybrid infrastructures. Required Skills & Experience: - Strong programming expertise in Python with hands-on experience using ML frameworks like TensorFlow and PyTorch. - Practical experience with Kubernetes and container orchestration technologies. - Familiarity with GPU workload scheduling platforms like Run:ai. - Strong experience in infrastructure automation using Terraform and configuration management with Ansible. - Experience working with Jupyter Notebooks in AI/ML development environments. - Good understanding of NVIDIA Enterprise Suite technologies including CUDA, NeMo Framework, Triton, and GPU drivers. - Knowledge of MLOps concepts, workflows, and tools such as MLflow and Kubeflow. - Experience deploying, managing, and scaling AI/ML workloads within cloud or hybrid infrastructure environments. As an AI Platform Engineer, your role involves building, managing, and enhancing AI/ML infrastructure, workflows, and automation pipelines to create scalable platforms for training and deploying machine learning models. You will work closely with data scientists and platform engineering teams to enable efficient resource management and scalable operations across cloud and hybrid ecosystems. Key Responsibilities: - Design, deploy, and manage Kubernetes environments optimized for AI/ML applications and workloads to ensure scalability, reliability, and performance of containerized AI platforms. - Implement and manage GPU orchestration solutions such as Run:ai for workload scheduling and resource optimization, enabling efficient GPU allocation and utilization for AI model training and inference. - Build and maintain Python-based automation tools and machine learning pipelines for infrastructure deployment using Terraform and configuration management through Ansible. - Develop and maintain Jupyter Notebook environments to support experimentation, research, and collaborative model development. - Configure and optimize NVIDIA Enterprise Suite technologies including CUDA, NeMo Framework, Triton, TensorRT, and GPU drivers to support accelerated AI computing. - Implement MLOps standards and practices covering model lifecycle management, CI/CD pipelines, monitoring, and governance using tools such as MLflow and Kubeflow. - Partner with data scientists, ML engineers, and platform teams to enhance scalability, operational efficiency, and resource utilization across cloud and hybrid infrastructures. Required Skills & Experience: - Strong programming expertise in Python with hands-on experience using ML frameworks like TensorFlow and PyTorch. - Practical experience with Kubernetes and container orchestration technologies. - Familiarity with GPU workload scheduling platforms like Run:ai. - Strong experience in infrastructure automation using Terraform and configuration management with Ansible. - Experience working with Jupyter Notebooks in AI/ML development environments. - Good understanding of NVIDIA Enterprise Suite technologies including CUDA, NeMo Framework, Triton, and GPU drivers. - Knowledge of MLOps concepts, workflows, and tools such as MLflow and Kubeflow. - Experience deploying, managing, and scaling AI/ML workloads within cloud or hybrid infrastructure environments.
More at LinkedIn
Related open roles
Data Engineer - Remote, Full-Time (Kochi)
India
Senior Enterprise Engineer - DNS, DHCP, IPAM
Bangalore
Associate Engineer, Data Center
United States
Senior Enterprise Engineer - Enterprise Infrastructure
Bangalore
Staff Network Engineer
San Francisco Bay Area · Hybrid
Principal Data Architect
India