Source description
About the role
Azure Cloud Sr. Site Reliability Engineer (SRE) Key Responsibilities: Experience in technical role-using Cloud Native services IaaS, and PaaS platforms Responsible for ensuring the reliability, availability, observability, and operational excellence of cloud-native platforms running on Azure and Kubernetes Hands-on experience in cloud-native architecture design, implementation of distributed, fault-tolerant enterprise applications for Cloud Hand on Azure Kubernetes Services (AKS) and manage it using IaC principles (Terraform and GitHub) Hands-on multi-tier architecting skills Sound knowledge of Infrastructure design on Compute, Storage, Network Hands-on cloud on App Gateway, API gateway, Front Door, Azure WAF and Azure Monitor and certificates Strong knowledge & Hands-on scripting language: YAML, Python, Shell, Bash, PowerShell, etc. Expertise in CI/CD using Jenkins or GitHub Actions to build re-usable release pipelines Experience in at least one configuration management tool (Ansible, Chef) Strong debug and troubleshoot skills on service bottlenecks throughout the whole software stack Measure and monitor availability, latency, and overall system health Experience investigating incidents impacting connected IoT devices is preferred Incident Investigation & Root Cause Analysis on alerts from IoT devices, cloud, and applications Strong Observability with SREs principles of building and enhance dashboards, alerts, KPIs, and SLIs/SLOs Responsible for production support IPC, RCA, and Operations runbooks to excel stability and reliability Experience: Overall 8+ Years of relevant technical experience Min of 2 years experience in managing Kubernetes and observability implementation Excellent communication skills, self-motivated and self-starter, strong investigative mindset Ability to indecently troubleshoot incidents, and possess ownership mentality Ability to collaborate with developers and engineering teams Preferred skills: Participate and lead working directly with technical teams from enterprise customers Collaborate and build automations that enable fellow SREs to operate at high speed and wide scale Perform periodic service reviews to proactively improve application performance, reliability and stability Create and update runbooks to automate the resolution quicker and learn Experience in managing immutable infrastructure-as-code and microservices model on Kubernetes Experience in SREs 7 practices to implement SLO/SLI/SLA and error budgets on Grafana Cloud Source Code Management, Continuous Integration and Continuous Delivery using tools BS degree in Computer Science, Engineering, or other highly technical, scientific discipline Expertise in leveraging open-source tooling such as Prometheus, Grafana, or Loki Certifications Required: Azure solution architect certification (AZ 303 & 304) AZ-400: Designing, Implementing Microsoft DevOps Solutions