Source description
About the role
Shift / On-call: 24x7 operations with rotational shifts including weekend & on-call Role Summary: Provide L1.5/L2-style triage for production issues from monitoring alerts and end-user reports. Perform initial application troubleshooting using Logic Monitor, AppDynamics, Azure Application Insights, logs, and basic Azure checks, manage Major Incidents and escalations with strong technical context, maintain high-quality documentation/runbooks, and uplift reliability through SRE-aligned practices (SLIs/SLOs, alert quality). Key Responsibilities: Own ticket lifecycle: create, investigate, update, resolve incidents/service requests using runbooks, ensure SLA compliance. Identify potential Major Incidents, trigger escalation, join bridges, and provide structured updates (impact, scope, findings, next actions). Perform initial application/API troubleshooting beyond basic checks: error/latency analysis, dependency triage, log-based diagnosis. Use AppDynamics and Azure App Insights to analyze performance and availability Maintain detailed work logs, update/create runbooks/KB articles, propose alert tuning and monitoring improvements. Support change/maintenance windows (alert suppression/reactivation) and validate pre/post health. Must-Have Experience: 6-7+ years in NOC/Production Support/Application Support/Operations. Azure Application Insights hands-on (failures/performance/dependencies, basic query capability preferred). Strong SRE fundamentals: SLIs/SLOs/SLAs, alert noise reduction, runbook-driven operations, error budget, RCA, automation etc . Incident & Escalation Management: strong ownership, prioritization, SLA discipline, high-quality ticket notes. Major Incident Handling: bridge participation, stakeholder communication, structured incident updates. Application Troubleshooting: HTTP basics, identifying app vs dependency vs infra symptoms. Good understanding of Kubernetes monitoring and alerting configuration. Log Analysis: time correlation, error pattern analysis, evidence collection for escalations. APM/Observability: hands-on with AppDynamics for triage (transactions, errors, latency, dashboards). ITIL Awareness: Understanding of Incident/Major Incident/Change processes, Problem Management understanding Communication & Documentation: runbooks/KB creation and clear written/verbal communication. Good-to-Have Azure fundamentals for triage (Azure Monitor/Log Analytics, resource/service health signals). Exposure to APM tool Certifications (Good-to-Have) ITIL Foundation (preferred) Cloud: Azure Fundamentals (AZ-900) or higher (AZ-104 a plus) APM/Observability: AppDynamics and/or Dynatrace certifications (or equivalent observability certs)
More at R1 RCM HOLDCO INC