Source description
About the role
Site Reliability Engineer (SRE) We are seeking a highly skilled Site Reliability Engineer (SRE) with expertise in Windows, Linux, Databases, Networking, and Voice/Unified Communications to ensure the reliability, availability, and performance of enterprise infrastructure and services. This role bridges the gap between operations and engineering, applying software engineering principles to infrastructure problems and driving a culture of automation, observability, and continuous improvement across heterogeneous environments. Responsibilities Windows Infrastructure Administration - Design and plan the upgrade for Windows Server environments (2016/2019/2022), ensuring high availability, reliability, and performance - Automate routine administration tasks using PowerShell scripting and configuration management tools (Ansible, SCCM) - Ensure OS-level security hardening, compliance, and access control enforcement Linux Infrastructure Administration - Design and plan the upgrade for Linux environments (RHEL, CentOS, Ubuntu) across on-premises and cloud platforms - Automate infrastructure tasks using Bash, Python scripting, and tools such as Ansible or Puppet - Enforce security baselines, SELinux/AppArmor policies, and patch compliance Database Administration & Reliability - Monitor database performance, optimize queries, and manage indexing strategies to meet SLOs - Collaborate with DBAs and application teams to resolve database-related incidents and capacity concerns - Ensure database security, access controls, and audit compliance across all platforms Network Operations & Reliability - Monitor and maintain network infrastructure including routers, switches, firewalls, and load balancers - Troubleshoot network incidents affecting availability, latency, and throughput across LAN/WAN/SD-WAN environments - Collaborate with network engineering teams on capacity planning, topology changes, and cloud network integration Voice & Unified Communications Administration - Manage and maintain Voice and Unified Communications platforms including Cisco CUCM, CUBE, Unity Connection, or equivalent - Monitor call quality, diagnose VoIP issues, and ensure SLA compliance for voice services - Perform patching, upgrades, and configuration management of voice infrastructure components - Coordinate with telecom vendors and carriers for circuit management and issue resolution Cross-Functional Responsibilities - Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets across infrastructure domains - Lead incident response, root cause analysis, and post-mortem reviews to drive systemic improvements - Build and maintain monitoring, alerting, and observability frameworks using tools such as Grafana, Prometheus, Datadog, or Splunk - Collaborate with application, cloud, and security teams to support platform reliability and modernization programs - Drive automation initiatives to eliminate toil and improve operational efficiency across all managed domains - Provide Tier 2/3 support for critical infrastructure incidents Qualifications - Bachelors degree in Computer Science, Information Technology, or related technology field preferred - Minimum of 5 years of hands-on experience across infrastructure domains (Windows, Linux, Network, Database, or Voice) - Proven experience with SRE principles including SLOs, error budgets, incident management, and blameless post-mortems - Strong scripting and automation skills in Python, Bash, or PowerShell - Experience with monitoring and observability platforms (Grafana, Prometheus, Datadog, Splunk, or equivalent) - Solid working knowledge of ITIL principles and ITSM practices - Familiarity with cloud platforms (Azure) and hybrid infrastructure environments - Current understanding of industry trends, reliability engineering methodologies, and DevOps practices - Outstanding verbal and written communication skills - Excellent attention to detail with strong analytical and problem-solving capabilities - Strong interpersonal skills and ability to collaborate across technical and non-technical teams Site Reliability Engineer (SRE) We are seeking a highly skilled Site Reliability Engineer (SRE) with expertise in Windows, Linux, Databases, Networking, and Voice/Unified Communications to ensure the reliability, availability, and performance of enterprise infrastructure and services. This role bridges the gap between operations and engineering, applying software engineering principles to infrastructure problems and driving a culture of automation, observability, and continuous improvement across heterogeneous environments. Responsibilities Windows Infrastructure Administration - Design and plan the upgrade for Windows Server environments (2016/2019/2022), ensuring high availability, reliability, and performance - Automate routine administration tasks using PowerShell scripting and configuration management tools (Ansible, SCCM)
More at PwC Service Delivery Center
Related open roles
Senior Azure Advanced Analytics Engineer(Azure Data engineer) (Telangana)
Hyderabad
IN_Senior Associate_Devops Engineer_GCC_One Consulting_Bangalore
Bangalore
Senior Azure Advanced Analytics Engineer(Azure Data engineer) (Hyderabad)
Hyderabad
IN_Senior Associate_DevOps Engineer_GCC_Advisory_Bangalore
Bangalore
In Senior Associate Sre Gcc Telangana (India)
Hyderabad
IN_Manager_ Data Engineer_Data Analytics_Advisory_Pune (Pune)
Mumbai