Source description
About the role
Description We are seeking a Site Reliability Engineer (SRE) with 7+ years of experience to support and enhance the reliability, availability, and performance of critical banking systems at Truist. The role requires strong hands on expertise in cloud native platforms, observability, automation, and incident management, with a focus on reliability engineering and operational excellence Required Technical Skills Cloud & Infrastructure - Microsoft Azure - Kubernetes - OpenShift Observability & Monitoring - Datadog - Dynatrace / AppDynamics - Splunk - Jenkins - Ansible Automation & CI/CD - Python - Kafka - RabbitMQ - Exposure to Java and Node.js Production Support & Incident Management - Strong experience handling major incidents - Production support in high availability, mission critical environments - Root cause analysis and reliability improvement - Engineer and enhance observability across systems and platforms - Define, implement, and track SLIs and SLOs - Design and build automation for recovery and self healing - Apply cloud native resiliency and failure isolation patterns - Lead major incident response with an engineering driven approach - Drive system level root cause fixes - Reduce long term incident volume through reliability engineering initiatives Desired Skills - Analyze, optimize, and enable CI/CD pipelines to improve reliability outcomes Supplementary Skills (Positive to Have) Advanced use of AIOps for predictive reliability insights ", Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying. .
More at Diverse Lynx