Source description
About the role
Job Overview: We are seeking a skilled and experienced Service Reliability Analyst to join our diverse team as part of newly created Service Reliability Centre (SRC). In this role, you will help improve the availability and performance of Arm infrastructure by utilising Arms AI Operations (AIOPS) and observability platforms. You will collaborate closely with development and platform teams to build and maintain robust observability and response processes. Responsibilities: - Lead the analysis and resolution of infrastructure incidents across physical and virtual servers, storage, identity, and engineering platforms. - Drive proactive monitoring, tuning, and optimization of systems using Dynatrace and other observability tools. - Look for opportunities to adapt automation to support the AIOps platform - Conduct root cause analysis of incidents and implement preventive measures. - Management of incidents to suppliers and Arms technical on-call rotas as appropriate - To log all issues in the Service Management Tool and manage them to completion within EIT service levels and quality criteria matrix - Work on a shift pattern, on a 24/7/365 operating model, while being able to work independently and flexibly in response to emergencies or critical issues Required Skills and Experience: - 3-6 years of hands-on experience in Platform Operations, or Infrastructure Support roles. - Solid experience with observability tools managing and optimising an enterprise observability (e.g., Dynatrace, Datadog, Splunk) for real-time monitoring, alerting, and diagnostics. - Proficiency in one or more scripting or programming languages (e.g., Python, Java,.NET, Node.js, Ansible or JavaScript). - Practical knowledge of infrastructure automation using Ansible, including writing and managing playbooks. - Understanding of UAM and IAM across on Premise, OUD/LDAP and Azure AD, including fault finding and access issues. - Experience supporting Windows and Linux operating systems - Experience wit .
More at ARM