Source description
About the role
Site Reliability Engineer (SRE) Purpose & Overall Relevance for the Organization: Maintain and enhance monitoring framework (data collection, alert aggregation, dashboarding) and implement and enhance alerting logic (framework). Enable proactive incident alert and resolution leveraging knowledge scripts. Identify and detect repetitive incidents (stability, reliability) and develop solutions to fix problems. Work on technical resolution for incidents and identify technical root cause. Ensure tool standards, exploit tool capability to fine-tune product reliability. Integrate incident, release, monitoring, and alerting tools into the overall ecosystem. Measure and report SLI, MTTx in periodic reviews, analyze deviations, and take actions to closure. Update runbooks with changes to process/tools. Drive postmortems to arrive at remedial actions. Participate in On-Call Incident Technical Support. Ensure production release guidelines (entry/exit) and implementation are adhered to for changes to Production. Support CI/CD pipeline implementation and integration to quality, security. Scale systems sustainably through mechanisms like automation; evolve systems by pushing for changes that improve reliability and velocity. Key Responsibilities: Maintain and enhance monitoring framework (data collection, alert aggregation, dashboarding). Implement and enhance alerting logic (framework). Enable proactive incident alert and resolution leveraging knowledge scripts. Identify and detect repetitive incidents (stability, reliability) and develop solutions to fix problems. Work on technical resolution for incidents and identify technical root cause. Ensure tool standards, exploit tool capability to fine-tune product reliability. Integrate incident, release, monitoring, alerting tools into the overall ecosystem. Measure and report SLI, MTTx in periodic reviews, analyze deviations, and take actions to closure. Update runbooks with changes to process/tools. Drive postmortems to arrive at remedial actions. Participate in On-Call Incident Technical Support. Ensure production release guidelines (entry/exit) and implementation are adhered to for changes to Production. Support CI/CD pipeline implementation and integration to quality, security. Scale systems sustainably through mechanisms like automation; evolve systems by pushing for changes that improve reliability and velocity. Required Skills: At least 5 years overall IT experience with 3 years in relevant area (DevOps / SRE). Strong awareness and experience of working with Site Reliability Engineering principles. Good understanding of public cloud offerings such as AWS components like EC2, IAM, RDS, Cloudwatch, Database (Redis, RDS, Dynamo DB). Hands-on experience on enterprise toolsets such as Grafana, Instana, Prometheus, ELK Stack, etc. Exposure to networking concepts (SSH, FTP, TCP/IP, DNS, Load balancing, CDN, etc.). Experience in any scripting language (bash / python / perl). Good experience with CI/CD pipelines including BitBucket, Jenkins. Experience operating high-availability, fault-tolerant, scalable, distributed software in production: building monitoring into your code, tweaking dashboards, defining alerts. Knowledge of Agile software development principles, including using JIRA. Experience in a 24/7 high availability production environment. Excellent organizational, verbal, and written communication skills. Aptitude to be a good team player and the desire to learn and implement new technologies. Knowledge of ITIL processes. Nice to Have: Experience with building Rest APIs, API Integration, and Web Services is preferred. Knowledge in Messaging and Streaming frameworks like RabbitMQ / Kafka. Exposure to languages such as Typescript, Nodejs. ITIL V4 Foundation certified.
More at Adidas