Source description
About the role
Job Summary: We are seeking a highly skilled Senior AI Infrastructure Management Engineer with expertise in Azure, AWS, and AI/ML deployment environments. This role demands deep technical knowledge of Linux, DevOps practices, cloud architecture, and AI/ML operations (MLOps/AIOps). The ideal candidate will be responsible for architecting, deploying, and maintaining scalable and secure infrastructure for enterprise AI applications. Key Responsibilities: Linux System Expertise - Manage and optimize Linux systems (CentOS, Ubuntu, Red Hat) - Perform kernel tuning, file system configuration, and network optimization - Develop shell scripts for automation and system management. Cloud Infrastructure (AWS & Azure) - Design and implement secure, scalable cloud architectures on AWS and Azure - Use services like EC2, S3, Lambda, Azure VMs, Blob Storage, and Functions - Manage hybrid and multi-cloud environments and ensure seamless integration Infrastructure as Code (IaC) - Automate infrastructure provisioning using Terraform, CloudFormation, or similar tools - Maintain infrastructure versioning and ensure traceability of changes - Enforce DevSecOps best practices and secure configurations AI/ML Infrastructure Management - Deploy and manage cloud infrastructure for AI/ML workloads including GPUs - Scale resources (GPU/CPU) for training and inference workloads - Deploy AI/ML apps using Docker and Kubernetes - Ensure high availability, performance, and reliability of AI applications - Work on MLOps/AIOps pipelines for model deployment and monitoring Qualifications: - Bachelor's degree in Computer Science, Engineering, or related field - 6+ years of experience in Infrastructure/Cloud/DevOps roles - Strong experience with AWS, Azure, and Linux systems - Experience in AI/ML infrastructure setup and management - Proficient in scripting: Python, Bash, PowerShell - Hands-on with Kubernetes, Docker, and cloud-native services - Experience with DevSecOps principles and CI/CD tools - Certifications (preferred): - AWS Solution Architect Associate / Cloud Practitioner - Azure DevOps Engineer / Administrator - Certified Kubernetes Administrator (CKA) Preferred Skills: - Experience with GPU cluster management for AI workloads - Strong knowledge of cloud security and compliance - Familiarity with real-time monitoring and logging tools - Exposure to up-to-date data stacks and AI lifecycle management .
More at Hungry Bird It Consulting Services