Padmi

Site Reliability Engineer-Windows,Linux,observability (Bengaluru)

BangalorePosted 2 months ago
Infrastructure And DatabasesMid-levelFull Time; Regular
Apply at PwC Service Delivery Center

Opens the source posting on shine.com

Source description

About the role

View original

Site Reliability Engineer (SRE) We are seeking a highly skilled Site Reliability Engineer (SRE) with expertise in Windows, Linux, Databases, Networking, and Voice/Unified Communications to ensure the reliability, availability, and performance of enterprise infrastructure and services. This role bridges the gap between operations and engineering, applying software engineering principles to infrastructure problems and driving a culture of automation, observability, and continuous improvement across heterogeneous environments. Responsibilities Windows Infrastructure Administration - Design and plan the upgrade for Windows Server environments (2016/2019/2022), ensuring high availability, reliability, and performance - Automate routine administration tasks using PowerShell scripting and configuration management tools (Ansible, SCCM) - Ensure OS-level security hardening, compliance, and access control enforcement Linux Infrastructure Administration - Design and plan the upgrade for Linux environments (RHEL, CentOS, Ubuntu) across on-premises and cloud platforms - Automate infrastructure tasks using Bash, Python scripting, and tools such as Ansible or Puppet - Enforce security baselines, SELinux/AppArmor policies, and patch compliance Database Administration & Reliability - Monitor database performance, optimize queries, and manage indexing strategies to meet SLOs - Collaborate with DBAs and application teams to resolve database-related incidents and capacity concerns - Ensure database security, access controls, and audit compliance across all platforms Network Operations & Reliability - Monitor and maintain network infrastructure including routers, switches, firewalls, and load balancers - Troubleshoot network incidents affecting availability, latency, and throughput across LAN/WAN/SD-WAN environments - Collaborate with network engineering teams on capacity planning, topology changes, and cloud network integration Voice & Unified Communications Administration - Manage and maintain Voice and Unified Communications platforms including Cisco CUCM, CUBE, Unity Connection, or equivalent - Monitor call quality, diagnose VoIP issues, and ensure SLA compliance for voice services - Perform patching, upgrades, and configuration management of voice infrastructure components - Coordinate with telecom vendors and carriers for circuit management and issue resolution Cross-Functional Responsibilities - Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets across infrastructure domains - Lead incident response, root cause analysis, and post-mortem reviews to drive systemic improvements - Build and maintain monitoring, alerting, and observability frameworks using tools such as Grafana, Prometheus, Datadog, or Splunk - Collaborate with application, cloud, and security teams to support platform reliability and modernization programs - Drive automation initiatives to eliminate toil and improve operational efficiency across all managed domains - Provide Tier 2/3 support for critical infrastructure incidents Qualifications - Bachelors degree in Computer Science, Information Technology, or related technology field preferred - Minimum of 5 years of hands-on experience across infrastructure domains (Windows, Linux, Network, Database, or Voice) - Proven experience with SRE principles including SLOs, error budgets, incident management, and blameless post-mortems - Strong scripting and automation skills in Python, Bash, or PowerShell - Experience with monitoring and observability platforms (Grafana, Prometheus, Datadog, Splunk, or equivalent) - Solid working knowledge of ITIL principles and ITSM practices - Familiarity with cloud platforms (Azure) and hybrid infrastructure environments - Current understanding of industry trends, reliability engineering methodologies, and DevOps practices - Outstanding verbal and written communication skills - Excellent attention to detail with strong analytical and problem-solving capabilities - Strong interpersonal skills and ability to collaborate across technical and non-technical teams Site Reliability Engineer (SRE) We are seeking a highly skilled Site Reliability Engineer (SRE) with expertise in Windows, Linux, Databases, Networking, and Voice/Unified Communications to ensure the reliability, availability, and performance of enterprise infrastructure and services. This role bridges the gap between operations and engineering, applying software engineering principles to infrastructure problems and driving a culture of automation, observability, and continuous improvement across heterogeneous environments. Responsibilities Windows Infrastructure Administration - Design and plan the upgrade for Windows Server environments (2016/2019/2022), ensuring high availability, reliability, and performance - Automate routine administration tasks using PowerShell scripting and configuration management tools (Ansible, SCCM)

One address, no account. We’ll tell you when matching roles go live.

More at PwC Service Delivery Center

Related open roles

View all roles