Source description
About the role
WHAT YOU WILL BE DOING Identify, debug and troubleshoot break fix issues and take it to a resolution Automate the proactive detection and prevention of SLA breaches, outages, and issues, before it impacts customers Configure observability and monitoring platforms to send actionable alerts, and create automated SOPs for such alerts Provide Level 4 support for the system within agreed service levels Implement and manage the effectiveness of Incident, Service Request, Change and Problem management processes for the services area Provide daily/weekly report for overall health of the cloud platform and effectiveness of automation, monitoring for incident prevention and resolution Responsible for the maintenance of system configurations and process documentation, operating procedures, and infrastructure support documentation Work closely with business teams and DevSecOps teams on for activities related to supporting the IAM service offerings WHAT YOU BRING Strong understanding of Cloud offerings - AWS or Azure preferred Hands-on experience in internals of cloud components (ex. Public Cloud, Containers, Kubernetes, Applications, APIs) Database : Extensive experience in database operations SQL Strong hands-on experience with Scripting (ex. Shell, Javascript, Python, Groovy) Experience working with global Customers and strong customer focus Excellent written and verbal communication skills Knowledge of Application Performance Monitoring (APM) and Site Reliability Engineering(SRE) Identity and Access Management domain knowledge is a preference Hands-on implementation/maintenance experience of the IGA platform is a preference IGA: Saviynt or Sailpoint or relevant Experience with SSL Certificates and SSO technologies is a preference Bachelor s degree or an equivalent experience 10+ years of industry experience in design, development, customization, configuration, deployment of cloud platforms, automation frameworks, database systems. Passionate about operating and scaling cloud services per the expected SLA and constantly looking into data for operational insights Knowledge of Java/J2EE and SQL Experience with ticketing systems, monitoring and automation platforms ex: Freshdesk, Datadog, Prometheus, Grafana, RunDeck Understanding of SLAs and the importance of being within SLAs Disclaimer : This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
More at Saviynt
