Source description
About the role
This position is within the Site Reliability Engineering (SRE) team, which is part of the broader Cloud Platforms Team. Our comprehensive suite of SaaS solutions, distributed systems, and product integrations supports our internal stakeholders in running critical business operations. As a Cloud Engineer, you will be responsible for maintaining and enhancing our cloud infrastructure across supported cloud platforms. This includes orchestrating deployments and supporting our industry-leading SaaS solution. As an integral part of the SRE team, you will ensure our platforms are highly available and resilient through continuous monitoring and proactive improvement suggestions. You will collaborate closely with Development and Delivery engineering teams to uphold contracted Service Level Objectives (SLOs), ensuring both internal and external systems maintain the reliability and uptime required for user needs. Key Responsibilities and Requirements ● Provide third-line support for infrastructure and applications. ● Assume ownership of fault resolution, coordinating the effort and collaborating with engineering teams to address faults raised against supported elements, networks, or applications. ● Manage and maintain the daily shift coverage schedule for the team. ● Conduct one-on-one reviews for team members within the Cloud Platform Team. ● Own the deployment process, ensuring regular service improvements are delivered by the engineering team. ● Lead the investigation into Root Cause Analysis (RCA) for service-impacting incidents, produce necessary reports, and coordinate the delivery of fixes to mitigate future occurrences. ● Utilize and maintain our monitoring platforms. Essential Experience ● Experience working in a distributed, multi-cloud environment (Azure, AWS, or GCP). ● Experience managing and leading technical teams. ● Excellent fault-finding and troubleshooting abilities. ● Proficiency in Linux (Debian/Ubuntu) and Windows Server. ● Knowledge of SQL. ● Expertise in Software Defined Networking. ● Strong understanding of Cloud and Platform Security. ● Experience with Monitoring Solutions and Incident Management. ● Experience with Docker/Containerization.
More at Kroll