Source description
About the role
Platform Site Reliability Engineer (AWS | Kubernetes | AI Platform) C2H Company NetConnect Global Employment Type Contract to Hire (12 Months) Location Bangalore (Hybrid) Experience 6–9 Years Budget Up to 2.20 Lakhs/Month CTC Job Description NetConnect Global is looking for an experienced Platform Site Reliability Engineer (Platform SRE) to manage production-scale AWS and Kubernetes infrastructure. The ideal candidate should have strong expertise in cloud platform operations, infrastructure automation, performance tuning, and production reliability. Experience supporting AI/LLM platforms is highly preferred. Key Responsibilities Manage and optimize AWS cloud infrastructure and Kubernetes (EKS) clusters. Ensure platform availability, scalability, and production reliability. Perform memory, container, and application performance tuning. Build and enhance CI/CD pipelines and infrastructure automation. Monitor production environments using Splunk and Grafana. Manage incidents, changes, problems, and CAPA activities following ITIL processes. Support AI/LLM platform workloads and cloud-native services. Mandatory Skills AWS (EC2, ECS, EKS, Lambda, IAM, S3, RDS, API Gateway, VPC, EFS, SNS, SQS, EventBridge, CloudFormation) Kubernetes (EKS), Helm & Production Cluster Management Terraform and/or CloudFormation GitHub Actions / Jenkins / GitLab Python or Bash Scripting YAML Splunk & Grafana ServiceNow & JIRA ITIL Framework (Incident, Change, Problem & CAPA Management) Memory & Performance Optimization Production Platform Support Preferred Skills AI/LLM Platform Operations MCP or Agentic AI environments Multi-cloud exposure (AWS, Azure, GCP) Database performance tuning (RDS) Strong troubleshooting and stakeholder management skills Notice Period: Immediate to 30 Days Preferred Employment Type: Contract to Hire (C2H) Company: NetConnect Global (Deployment with a Leading Enterprise Client)
More at Net Connect