Padmi
LinkedIn logo
LinkedIn

professional networking · talent acquisition

AI Platform Engineer

IndiaPosted 3 months ago
Infrastructure And DatabasesMid-levelFull Time; Regular
Apply at LinkedIn

Opens the source posting on shine.com

Source description

About the role

View original

As an AI Platform Engineer, your role involves building, managing, and enhancing AI/ML infrastructure, workflows, and automation pipelines to create scalable platforms for training and deploying machine learning models. You will work closely with data scientists and platform engineering teams to enable efficient resource management and scalable operations across cloud and hybrid ecosystems. Key Responsibilities: - Design, deploy, and manage Kubernetes environments optimized for AI/ML applications and workloads to ensure scalability, reliability, and performance of containerized AI platforms. - Implement and manage GPU orchestration solutions such as Run:ai for workload scheduling and resource optimization, enabling efficient GPU allocation and utilization for AI model training and inference. - Build and maintain Python-based automation tools and machine learning pipelines for infrastructure deployment using Terraform and configuration management through Ansible. - Develop and maintain Jupyter Notebook environments to support experimentation, research, and collaborative model development. - Configure and optimize NVIDIA Enterprise Suite technologies including CUDA, NeMo Framework, Triton, TensorRT, and GPU drivers to support accelerated AI computing. - Implement MLOps standards and practices covering model lifecycle management, CI/CD pipelines, monitoring, and governance using tools such as MLflow and Kubeflow. - Partner with data scientists, ML engineers, and platform teams to enhance scalability, operational efficiency, and resource utilization across cloud and hybrid infrastructures. Required Skills & Experience: - Strong programming expertise in Python with hands-on experience using ML frameworks like TensorFlow and PyTorch. - Practical experience with Kubernetes and container orchestration technologies. - Familiarity with GPU workload scheduling platforms like Run:ai. - Strong experience in infrastructure automation using Terraform and configuration management with Ansible. - Experience working with Jupyter Notebooks in AI/ML development environments. - Good understanding of NVIDIA Enterprise Suite technologies including CUDA, NeMo Framework, Triton, and GPU drivers. - Knowledge of MLOps concepts, workflows, and tools such as MLflow and Kubeflow. - Experience deploying, managing, and scaling AI/ML workloads within cloud or hybrid infrastructure environments. As an AI Platform Engineer, your role involves building, managing, and enhancing AI/ML infrastructure, workflows, and automation pipelines to create scalable platforms for training and deploying machine learning models. You will work closely with data scientists and platform engineering teams to enable efficient resource management and scalable operations across cloud and hybrid ecosystems. Key Responsibilities: - Design, deploy, and manage Kubernetes environments optimized for AI/ML applications and workloads to ensure scalability, reliability, and performance of containerized AI platforms. - Implement and manage GPU orchestration solutions such as Run:ai for workload scheduling and resource optimization, enabling efficient GPU allocation and utilization for AI model training and inference. - Build and maintain Python-based automation tools and machine learning pipelines for infrastructure deployment using Terraform and configuration management through Ansible. - Develop and maintain Jupyter Notebook environments to support experimentation, research, and collaborative model development. - Configure and optimize NVIDIA Enterprise Suite technologies including CUDA, NeMo Framework, Triton, TensorRT, and GPU drivers to support accelerated AI computing. - Implement MLOps standards and practices covering model lifecycle management, CI/CD pipelines, monitoring, and governance using tools such as MLflow and Kubeflow. - Partner with data scientists, ML engineers, and platform teams to enhance scalability, operational efficiency, and resource utilization across cloud and hybrid infrastructures. Required Skills & Experience: - Strong programming expertise in Python with hands-on experience using ML frameworks like TensorFlow and PyTorch. - Practical experience with Kubernetes and container orchestration technologies. - Familiarity with GPU workload scheduling platforms like Run:ai. - Strong experience in infrastructure automation using Terraform and configuration management with Ansible. - Experience working with Jupyter Notebooks in AI/ML development environments. - Good understanding of NVIDIA Enterprise Suite technologies including CUDA, NeMo Framework, Triton, and GPU drivers. - Knowledge of MLOps concepts, workflows, and tools such as MLflow and Kubeflow. - Experience deploying, managing, and scaling AI/ML workloads within cloud or hybrid infrastructure environments.

One address, no account. We’ll tell you when matching roles go live.

More at LinkedIn

Related open roles

View all roles