Padmi

Site Reliability Engineer

New YorkPosted 1 month ago
Infrastructure And DatabasesUnspecified
Apply at Thoughtwave Software and Solutions

Opens the source posting on atsapp.swarmhr.com

Source description

About the role

View original

Position: Site Reliability Engineer Location: Onsite- Jersey City, NJ Duration: Long Term on W2 only Background: • The SRE team supports vendor applications such as Client, Databricks, AWS Athena, and Redshift. • While the engineering team builds a layer on top of these platforms, SREs have limited control over the underlying vendor products. • This means the team must focus on defining SLI and SLO that are defensible and ensure that these are monitored and alerted on. • Service-level objective (SLO) is a key element of a service-level agreement (SLA) between a service provider and a customer. • SLO's are agreed upon as a means of measuring the performance of the Service Provider and are outlined as a way of avoiding disputes between the two parties based on misunderstanding. • Service Level Indicator (SLI) is a carefully defined quantitative measure of some aspect of the level of service of a system. • SLI's are closely aligned to metrics and are often a simple derivation from a raw metric (e.g. a SLI for response time can be the 'response time' metric represented as a percentage). Observability Solution: • Grafana / Dynatrace based SLI/SLO enablement with integration to Service Status, Client Cortex, Databricks Metadata Lake • Golden Signal alerting for (Availability, Response Times, Throughput, Error Rates) using Client for identified set of metrics available in Databricks metadata lake and work towards defining the Golden Signals and how they will be calculated. Basic dashboard will be developed that includes Golden Signals and SLI/SLO metrics that are aligned with stakeholder. • Alerting based on the Golden Signals will be integrated with Client so that we do not have to monitor mailboxes for email notifications. • We should have at least a basic Availability alerting to complete this milestone. The data will be visualized using Grafana/Dynatrace. • Alerting based on Client tickets will be created, and appropriate runbooks to handle these alerts will be put in place. • Monitoring solutions will help detect any failures in the target system(Databricks or Client), issues with AWS Lambda that cause monitoring to fail because fresh data is not available if AWS Lambda is down. • Mechanism to be designed to monitor the freshness of data as well as monitor the monitoring infrastructure using jobs as well as key components of our monitoring toolset. • This includes monitoring the SLI/SLO jobs that fetch data from system tables into our Metadata Lake, monitoring availability of key AWS services like Lambda, monitoring data freshness in our dashboards, etc. Additional context: • Establish a clear, well-documented definition of Service Health, encompassing Availability, Response Times, Job Health, and AWS Infrastructure Health. • This should leverage clearly defined Service Level Indicators (SLIs) and Service Level Objectives (SLOs), along with the specific metrics that will be used for monitoring. • Develop simplified, high-level dashboards that provide an aggregated overview of system health. More detailed, low-level metrics for in-depth analysis should be separated into their own dedicated dashboards. • Ensure alignment among all stakeholders, including Line of Business (LOB) partners and the Engineering team, regarding what constitutes a defensible SLO for the Databricks and Client services. • Create highly accurate alerts in Client that promptly capture incidents and events. • The closure codes of Client tickets will be used to assess whether each incident was significant or a false alarm. • This will be measured and tracked as part of ongoing Incident Management enhancements, with Observability work serving as a foundational step for these improvements. • Provide clear visualization of SLOs for Databricks and Client, ensuring accurate incident reporting. • Error Budgets should be transparently calculated and displayed in a dashboard to inform and prioritize future service improvements. Thanks & Regards Raju Phone: 630 491 8638 Email: raju@qwavetech.com

One address, no account. We’ll tell you when matching roles go live.

More at Thoughtwave Software and Solutions

Related open roles

View all roles