Source description
About the role
The SRE team supports vendor applications such as Client, Databricks, AWS Athena, and Redshift. While the engineering team builds a layer on top of these platforms, SREs have limited control over the underlying vendor products. This means the team must focus on defining SLI and SLO that are defensible and ensure that these are monitored and alerted on. Service-level objective (SLO) is a key element of a service-level agreement (SLA) between a service provider and a customer. SLO's are agreed upon as a means of measuring the performance of the Service Provider and are outlined as a way of avoiding disputes between the two parties based on misunderstanding. Service Level Indicator (SLI) is a carefully defined quantitative measure of some aspect of the level of service of a system. SLI's are closely aligned to metrics and are often a simple derivation from a raw metric (e.g. a SLI for response time can be the 'response time' metric represented as a percentage). Observability Solution: Grafana / Dynatrace based SLI/SLO enablement with integration to Service Status, Client Cortex, Databricks Metadata Lake Golden Signal alerting for (Availability, Response Times, Throughput, Error Rates) using Client for identified set of metrics available in Databricks metadata lake and work towards defining the Golden Signals and how they will be calculated. Basic dashboard will be developed that includes Golden Signals and SLI/SLO metrics that are aligned with stakeholder. Alerting based on the Golden Signals will be integrated with Client so that we do not have to monitor mailboxes for email notifications. We should have at least a basic Availability alerting to complete this milestone. The data will be visualized using Grafana/Dynatrace. Alerting based on Client tickets will be created, and appropriate runbooks to handle these alerts will be put in place. Monitoring solutions will help detect any failures in the target system(Databricks or Client), issues with AWS Lambda that cause monitoring to fail because fresh data is not available if AWS Lambda is down. Mechanism to be designed to monitor the freshness of data as well as monitor the monitoring infrastructure using jobs as well as key components of our monitoring toolset. This includes monitoring the SLI/SLO jobs that fetch data from system tables into our Metadata Lake, monitoring availability of key AWS services like Lambda, monitoring data freshness in our dashboards, etc.
More at Thoughtwave Software and Solutions