Padmi
PwC logo
PwC

M&A transaction services · financial due diligence

Senior Associate - Azure SRE DevOps Kubernetes

ChennaiPosted 2 months ago
Software engineeringSeniorFull Time; Regular
Apply at PwC

Opens the source posting on shine.com

Source description

About the role

View original

As a Site Reliability Engineer (SRE) at the Global Capability Center (GCC) team of PwC, you will play a crucial role in supporting highly scalable, global platforms. Your primary responsibility will be to ensure system reliability, availability, performance, and operational excellence across production environments. This hands-on role demands strong expertise in cloud infrastructure, automation, production support, and incident management, as you collaborate closely with global engineering and product teams. Responsibilities: - Ensure high availability and reliability of large-scale, business-critical systems - Monitor, troubleshoot, and resolve production incidents; participate in on-call rotations - Define and track SLIs, SLOs, and error budgets - Perform root cause analysis (RCA) and drive preventive actions - Build and maintain automation, scripts, and runbooks to reduce operational toil - Support and optimize cloud environments (Azure/AWS/GCP) - Work with containers and orchestration platforms (Docker, Kubernetes) - Improve system resilience through capacity planning, performance tuning, and failover testing - Collaborate with global engineering, platform, and security teams Required Skills & Qualifications: - 5+ years of experience in SRE, DevOps, Production Engineering, or similar roles - Solid experience with Linux/Unix systems - Proficiency in at least one programming/scripting language: Python / Go / Java / Bash - Hands-on experience with: - Cloud platforms (Azure / AWS / GCP) - Kubernetes & Docker - CI/CD pipelines - Monitoring & observability tools (Prometheus, Grafana, Datadog, ELK, Azure Monitor, etc.) - Infrastructure as Code (Terraform, ARM, CloudFormation, Pulumi) Good to Have: - Experience working in a Global Capability Center (GCC) or global delivery model - Exposure to 24x7 production support environments - Knowledge of SRE best practices (SLIs, SLOs, error budgets) - Cloud or Kubernetes certifications - Experience with FinOps or cost optimization What We Offer: - Opportunity to work on global-scale platforms - Strong engineering and reliability-focused culture - Exposure to global stakeholders and teams - Career growth across SRE, Platform Engineering, and Leadership tracks - Competitive compensation and learning opportunities Preferred Skill Sets: AWS, GCP Years of Experience Required: 5-8 years Education Qualification: Btech BE At PwC, we believe in providing equal employment opportunities without any discrimination, fostering an environment where each individual can contribute to personal and firm growth while bringing their true selves to work. As a Site Reliability Engineer (SRE) at the Global Capability Center (GCC) team of PwC, you will play a crucial role in supporting highly scalable, global platforms. Your primary responsibility will be to ensure system reliability, availability, performance, and operational excellence across production environments. This hands-on role demands strong expertise in cloud infrastructure, automation, production support, and incident management, as you collaborate closely with global engineering and product teams. Responsibilities: - Ensure high availability and reliability of large-scale, business-critical systems - Monitor, troubleshoot, and resolve production incidents; participate in on-call rotations - Define and track SLIs, SLOs, and error budgets - Perform root cause analysis (RCA) and drive preventive actions - Build and maintain automation, scripts, and runbooks to reduce operational toil - Support and optimize cloud environments (Azure/AWS/GCP) - Work with containers and orchestration platforms (Docker, Kubernetes) - Improve system resilience through capacity planning, performance tuning, and failover testing - Collaborate with global engineering, platform, and security teams Required Skills & Qualifications: - 5+ years of experience in SRE, DevOps, Production Engineering, or similar roles - Solid experience with Linux/Unix systems - Proficiency in at least one programming/scripting language: Python / Go / Java / Bash - Hands-on experience with: - Cloud platforms (Azure / AWS / GCP) - Kubernetes & Docker - CI/CD pipelines - Monitoring & observability tools (Prometheus, Grafana, Datadog, ELK, Azure Monitor, etc.) - Infrastructure as Code (Terraform, ARM, CloudFormation, Pulumi) Good to Have: - Experience working in a Global Capability Center (GCC) or global delivery model - Exposure to 24x7 production support environments - Knowledge of SRE best practices (SLIs, SLOs, error budgets) - Cloud or Kubernetes certifications - Experience with FinOps or cost optimization What We Offer: - Opportunity to work on global-scale platforms - Strong engineering and reliability-focused culture - Exposure to global stakeholders and teams - Career growth across SRE, Platform Engineering, and Leadership tracks - Competitive compensation and learning opportunities **Prefer

One address, no account. We’ll tell you when matching roles go live.

More at PwC

Related open roles

View all roles