Source description
About the role
Job Description Site Reliability Engineer (SRE) L2 Position: Site Reliability Engineer (SRE) L2 Experience: 56 Years Location: Noida Employment Type: Full-Time Role Summary We are seeking an experienced Site Reliability Engineer (SRE) L2 with strong expertise in Microsoft Azure, Dynatrace, and Application Monitoring to support business-critical cloud applications. The ideal candidate will possess hands-on experience in Azure PaaS services, observability platforms, incident response, root cause analysis, and production support. The role requires proactive monitoring, troubleshooting, and ensuring high availability and reliability of cloud-hosted applications. Key Responsibilities Monitor, troubleshoot, and support production applications hosted on Microsoft Azure. Perform end-to-end troubleshooting of application and infrastructure incidents. Investigate alerts and production issues using Azure Monitor, Application Insights, Log Analytics, and Dynatrace. Analyze application logs and telemetry using Kusto Query Language (KQL) and Dynatrace Query Language (DQL). Monitor and troubleshoot Azure API Management (APIM), Azure Functions, Service Bus, and Azure-native services. Identify the root cause of application failures by tracing requests across APIs, Azure Functions, messaging services, and backend components. Configure and manage alerts, dashboards, and monitoring rules to ensure proactive incident detection. Utilize Dynatrace features such as Smartscape, Problems & Events, Distributed Tracing, Synthetic Monitoring, and Davis AI for performance analysis. Participate in production on-call rotations and provide support for P1/P2 incidents. Lead technical troubleshooting during major incidents and collaborate with development and infrastructure teams for timely resolution. Prepare detailed Root Cause Analysis (RCA) reports and recommend preventive measures. Work closely with DevOps, Application Development, and Cloud Infrastructure teams to improve platform reliability and observability. Support continuous service improvements through automation and operational excellence initiatives. Required Technical Skills Microsoft Azure Azure Monitor Application Insights Log Analytics Kusto Query Language (KQL) Azure API Management (APIM) Azure Functions Azure Service Bus Azure Alerts & Action Groups Azure Portal Monitoring & Observability Dynatrace Problems & Events Feed Smartscape Distributed Tracing Synthetic Monitoring Dynatrace Query Language (DQL) Alternative tools (acceptable): New Relic Datadog Incident Management Production Support P1/P2 Incident Handling Major Incident Management Root Cause Analysis (RCA) Problem Management SLA Management On-call Support More Info : komal.kumari@nlbtech.com
More at NLB Services