Padmi

SENIOR ENGINEER - Google cloud Monitoring and Logging

BangalorePosted 1 month ago
Infrastructure And DatabasesSenior
Apply at Happiest Minds Technologies

Opens the source posting on jobs.happiestminds.com

Source description

About the role

View original

Job Description: Vendor Partner: Happiest Minds Technologies Client: DoubleVerify Team: NOC Tier 1 SRE Operations Support Designation (DV): Support Engineer 1 Work Schedule: 24x7 Shift-Based (Rotational) Role overview: The Support Engineer I (Tier 1) is the first line of defense for DoubleVerify's production environment. Operating as part of the Happiest Minds NOC team embedded within DV's TechOps organization, this role is responsible for monitoring, triaging, and resolving production alerts across multiple platforms and products 24 hours a day, 7 days a week. The engineer is expected to work strictly within defined Standard Operating Procedures (SOPs), escalate intelligently to Tier 2 / Tier 3 teams when needed, and contribute to reducing overall escalation rates over time. Responsibilities: Monitor production alert channels (PagerDuty, Grafana, Nagios, Site24x7, Slack alert channels) continuously during assigned shift Triage incoming alerts, assess severity, and respond within defined SLA targets Follow SOPs to investigate, acknowledge, silence, or resolve alerts across products including Pinnacle, Social Integrations, Programmatic, Publisher, Measurement, and TechOps Raise Freshservice tickets for all actionable alerts and maintain accurate ticket status throughout the lifecycle Execute alert silencing via AlertManager using correct durations; avoid short silencing windows that cause ticket re-inflation Escalate unresolved or out-of-scope issues to Tier 2 and Tier 3 teams via PagerDuty with sufficient context and findings documented Validate root cause before escalating , avoid passing incomplete or unclear information upstream. Coordinate with DBA, Analytics, SI, DataOps, and SRE teams through Slack alert threads, following proper tagging protocols Avoid directly pinging on-site DV team members; route escalations via defined channels and the @support_group alias Execute Kubernetes pod scaling procedures (e.g., incremental scaling via Grafana/Lens) as directed by Tier 2 leads Monitor Salesforce loader jobs and reprocessing pipelines, following documented runbooks Run Snowflake queries using the alert conditions Trigger pipeline reruns (DAG tasks, Rivery rivers, Rundeck jobs) per SOP when applicable Verify certificate status on Site24x7 and confirm renewals post-update Category Tools Monitoring & Alerting Grafana, Nagios, PagerDuty, Site24x7, Prometheus/AlertManager Kubernetes OpenLens, AWX, GCP Console Data Platforms Snowflake, Rivery, Vertica, Splunk Automation & CI/CD Rundeck, GitLab Pipelines Ticketing Freshservice, Jira Communication Slack (multiple channels), Google Meet Reqiured Skills: 2 years of experience in a NOC, production support, or IT operations role Familiarity with Linux command-line basics and reading system logs Basic understanding of cloud infrastructure concepts (GCP, Kubernetes preferred) Ability to read and execute SQL queries (Snowflake experience is a plus) Exposure to monitoring and alerting tools (Grafana, Nagios, or similar) Strong attention to detail and ability to follow documented procedures accurately Good written communication skills in English for Slack coordination and ticket documentation Ability to work in a 24x7 rotational shift environment, including nights and weekends Job Description: Vendor Partner: Happiest Minds Technologies Client: DoubleVerify Team: NOC Tier 1 SRE Operations Support Designation (DV): Support Engineer 1 Work Schedule: 24x7 Shift-Based (Rotational) Role overview: The Support Engineer I (Tier 1) is the first line of defense for DoubleVerify's production environment. Operating as part of the Happiest Minds NOC team embedded within DV's TechOps organization, this role is responsible for monitoring, triaging, and resolving production alerts across multiple platforms and products 24 hours a day, 7 days a week. The engineer is expected to work strictly within defined Standard Operating Procedures (SOPs), escalate intelligently to Tier 2 / Tier 3 teams when needed, and contribute to reducing overall escalation rates over time. Responsibilities: Monitor production alert channels (PagerDuty, Grafana, Nagios, Site24x7, Slack alert channels) continuously during assigned shift Triage incoming alerts, assess severity, and respond within defined SLA targets Follow SOPs to investigate, acknowledge, silence, or resolve alerts across products including Pinnacle, Social Integrations, Programmatic, Publisher, Measurement, and TechOps Raise Freshservice tickets for all actionable alerts and maintain accurate ticket status throughout the lifecycle Execute alert silencing via AlertManager using correct durations; avoid short silencing windows that cause ticket re-inflation Escalate unresolved or out-of-scope issues to Tier 2 and Tier 3 teams via PagerDuty with sufficient context and findings documented Validate root cause before escalating , avoid passing incomplete or unclear information upstream. Coordinate with DBA, Analytics, SI, DataOps, and SRE teams through Slack alert threads, following proper tagging protocols Avoid directly pinging on-site DV team members; route escalations via defined channels and the @support_group alias Execute Kubernetes pod scaling procedures (e.g., incremental scaling via Grafana/Lens) as directed by Tier 2 leads Monitor Salesforce loader jobs and reprocessing pipelines, following documented runbooks Run Snowflake queries using the alert conditions Trigger pipeline reruns (DAG tasks, Rivery rivers, Rundeck jobs) per SOP when applicable Verify certificate status on Site24x7 and confirm renewals post-update Category Tools Monitoring & Alerting Grafana, Nagios, PagerDuty, Site24x7, Prometheus/AlertManager Kubernetes OpenLens, AWX, GCP Console Data Platforms Snowflake, Rivery, Vertica, Splunk Automation & CI/CD Rundeck, GitLab Pipelines Ticketing Freshservice, Jira Communication Slack (multiple channels), Google Meet Reqiured Skills: 2 years of experience in a NOC, production support, or IT operations role Familiarity with Linux command-line basics and reading system logs Basic understanding of cloud infrastructure concepts (GCP, Kubernetes preferred) Ability to read and execute SQL queries (Snowflake experience is a plus) Exposure to monitoring and alerting tools (Grafana, Nagios, or similar) Strong attention to detail and ability to follow documented procedures accurately Good written communication skills in English for Slack coordination and ticket documentation Ability to work in a 24x7 rotational shift environment, including nights and weekends DVFNOC01-1-1 Job Description: Vendor Partner: Happiest Minds Technologies Client: DoubleVerify Team: NOC Tier 1 SRE Operations Support Designation (DV): Support Engineer 1 Work Schedule: 24x7 Shift-Based (Rotational) Role overview: The Support Engineer I (Tier 1) is the first line of defense for DoubleVerify's p... Job Description: Vendor Partner: Happiest Minds Technologies Client: DoubleVerify Team: ... Job Description: Vendor Partner: Happiest Minds Technologies Client: DoubleVerify Team: NOC Tier 1 SRE Operations Support Designation (DV): Support En... Grafana,Linux OS,Kubernetes,GCP,NAGIOS,Site24x7,Sre

One address, no account. We’ll tell you when matching roles go live.

More at Happiest Minds Technologies

Related open roles

View all roles