Source description
About the role
As a Site Reliability Engineer (SRE) at the Global Capability Center (GCC) team of PwC, you will play a crucial role in supporting highly scalable, global platforms. Your primary responsibility will be to ensure system reliability, availability, performance, and operational excellence across production environments. This hands-on role demands strong expertise in cloud infrastructure, automation, production support, and incident management, as you collaborate closely with global engineering and product teams. Responsibilities: - Ensure high availability and reliability of large-scale, business-critical systems - Monitor, troubleshoot, and resolve production incidents; participate in on-call rotations - Define and track SLIs, SLOs, and error budgets - Perform root cause analysis (RCA) and drive preventive actions - Build and maintain automation, scripts, and runbooks to reduce operational toil - Support and optimize cloud environments (Azure/AWS/GCP) - Work with containers and orchestration platforms (Docker, Kubernetes) - Improve system resilience through capacity planning, performance tuning, and failover testing - Collaborate with global engineering, platform, and security teams Required Skills & Qualifications: - 5+ years of experience in SRE, DevOps, Production Engineering, or similar roles - Solid experience with Linux/Unix systems - Proficiency in at least one programming/scripting language: Python / Go / Java / Bash - Hands-on experience with: - Cloud platforms (Azure / AWS / GCP) - Kubernetes & Docker - CI/CD pipelines - Monitoring & observability tools (Prometheus, Grafana, Datadog, ELK, Azure Monitor, etc.) - Infrastructure as Code (Terraform, ARM, CloudFormation, Pulumi) Good to Have: - Experience working in a Global Capability Center (GCC) or global delivery model - Exposure to 24x7 production support environments - Knowledge of SRE best practices (SLIs, SLOs, error budgets) - Cloud or Kubernetes certifications - Experience with FinOps or cost optimization What We Offer: - Opportunity to work on global-scale platforms - Strong engineering and reliability-focused culture - Exposure to global stakeholders and teams - Career growth across SRE, Platform Engineering, and Leadership tracks - Competitive compensation and learning opportunities Preferred Skill Sets: AWS, GCP Years of Experience Required: 5-8 years Education Qualification: Btech BE At PwC, we believe in providing equal employment opportunities without any discrimination, fostering an environment where each individual can contribute to personal and firm growth while bringing their true selves to work. As a Site Reliability Engineer (SRE) at the Global Capability Center (GCC) team of PwC, you will play a crucial role in supporting highly scalable, global platforms. Your primary responsibility will be to ensure system reliability, availability, performance, and operational excellence across production environments. This hands-on role demands strong expertise in cloud infrastructure, automation, production support, and incident management, as you collaborate closely with global engineering and product teams. Responsibilities: - Ensure high availability and reliability of large-scale, business-critical systems - Monitor, troubleshoot, and resolve production incidents; participate in on-call rotations - Define and track SLIs, SLOs, and error budgets - Perform root cause analysis (RCA) and drive preventive actions - Build and maintain automation, scripts, and runbooks to reduce operational toil - Support and optimize cloud environments (Azure/AWS/GCP) - Work with containers and orchestration platforms (Docker, Kubernetes) - Improve system resilience through capacity planning, performance tuning, and failover testing - Collaborate with global engineering, platform, and security teams Required Skills & Qualifications: - 5+ years of experience in SRE, DevOps, Production Engineering, or similar roles - Solid experience with Linux/Unix systems - Proficiency in at least one programming/scripting language: Python / Go / Java / Bash - Hands-on experience with: - Cloud platforms (Azure / AWS / GCP) - Kubernetes & Docker - CI/CD pipelines - Monitoring & observability tools (Prometheus, Grafana, Datadog, ELK, Azure Monitor, etc.) - Infrastructure as Code (Terraform, ARM, CloudFormation, Pulumi) Good to Have: - Experience working in a Global Capability Center (GCC) or global delivery model - Exposure to 24x7 production support environments - Knowledge of SRE best practices (SLIs, SLOs, error budgets) - Cloud or Kubernetes certifications - Experience with FinOps or cost optimization What We Offer: - Opportunity to work on global-scale platforms - Strong engineering and reliability-focused culture - Exposure to global stakeholders and teams - Career growth across SRE, Platform Engineering, and Leadership tracks - Competitive compensation and learning opportunities **Prefer
More at PwC
Related open roles
IN_Associate_GenAI and Agentic AI_GCC_Advisory_Bangalore
Bangalore
Generative AI- Data Scientist-Senior Associate (Telangana)
Hyderabad
In Senior Associate Sap Abap Sap Aith Advisory Bhubaneswar Bhubaneswar
India
Senior Associate_Dell Boomi Integration Developer_NetSuite Finance Solutions
India
Python Software Developer ( GenAI/FastAPI/Django)
Hyderabad
Salesforce LSC/Platform Developer-SA (Pune)
Mumbai