Source description
About the role
Design, implement, and maintain scalable, highly available infrastructure on AWS/GCP . Automate deployments, monitoring, and incident response using Terraform, Kubernetes, and CI/CD pipelines . Optimize system performance, troubleshoot incidents, and implement blameless postmortems . Enhance observability with Prometheus, Grafana, and distributed tracing tools . Collaborate with development teams to implement best practices in reliability engineering . Requirements: 3+ years of experience in SRE, DevOps, or infrastructure engineering. Strong programming skills in Python, Go, or Bash for automation. Deep understanding of Kubernetes, Docker, Terraform, and cloud-native architectures . Expertise in monitoring, logging, and alerting systems. Experience with incident management and performance tuning.
More at Unacademy Group