Source description
About the role
You are an experienced L1/L2 Monitoring & Incident Management Specialist responsible for supporting centralized monitoring and incident command operations for business-critical, client-facing applications across mainframe and distributed environments. Your role involves proactive monitoring, rapid incident response, effective stakeholder communication, and coordination across technical teams to ensure service availability and operational excellence. Key Responsibilities: - Provide 24x7 monitoring and operational support for client-facing applications and services. - Monitor real-time system alerts and respond promptly to incidents impacting business operations. - Act as the primary point of contact during incidents, ensuring timely communication with clients and stakeholders. - Perform initial incident triage, impact assessment, and preliminary root cause analysis. - Coordinate with Infrastructure, Application, and Vendor teams to facilitate incident resolution and service restoration. - Manage and track incidents through their complete lifecycle, ensuring adherence to defined SLAs. - Handle production issues including batch failures, system alerts, service degradation, and application outages. - Escalate critical issues appropriately through ITSM platforms such as ServiceNow. - Maintain accurate incident records, status updates, and post-incident documentation. Required Skills & Experience: - Solid experience with enterprise monitoring and alert management tools. - Hands-on experience with IT Service Management (ITSM) platforms, preferably ServiceNow. - Solid understanding of incident management processes and incident lifecycle management. - Experience working in production support environments with 24x7 operational coverage. - Exposure to mainframe systems, distributed applications, APIs, and integrated technology environments. - Knowledge of service restoration processes, escalation management, and operational support best practices. You should possess the ability to work effectively in high-pressure, mission-critical environments, strong analytical and problem-solving skills, excellent communication and stakeholder management skills, ability to coordinate across multiple teams, a strong sense of ownership and accountability, and the ability to prioritize and manage multiple incidents simultaneously while maintaining service quality. Preferred Qualifications: - Experience supporting enterprise-scale production environments. - Familiarity with SLA-driven support models and operational governance. - Understanding of ITIL principles and service management frameworks. - Experience working with both mainframe and distributed technology stacks. You are an experienced L1/L2 Monitoring & Incident Management Specialist responsible for supporting centralized monitoring and incident command operations for business-critical, client-facing applications across mainframe and distributed environments. Your role involves proactive monitoring, rapid incident response, effective stakeholder communication, and coordination across technical teams to ensure service availability and operational excellence. Key Responsibilities: - Provide 24x7 monitoring and operational support for client-facing applications and services. - Monitor real-time system alerts and respond promptly to incidents impacting business operations. - Act as the primary point of contact during incidents, ensuring timely communication with clients and stakeholders. - Perform initial incident triage, impact assessment, and preliminary root cause analysis. - Coordinate with Infrastructure, Application, and Vendor teams to facilitate incident resolution and service restoration. - Manage and track incidents through their complete lifecycle, ensuring adherence to defined SLAs. - Handle production issues including batch failures, system alerts, service degradation, and application outages. - Escalate critical issues appropriately through ITSM platforms such as ServiceNow. - Maintain accurate incident records, status updates, and post-incident documentation. Required Skills & Experience: - Solid experience with enterprise monitoring and alert management tools. - Hands-on experience with IT Service Management (ITSM) platforms, preferably ServiceNow. - Solid understanding of incident management processes and incident lifecycle management. - Experience working in production support environments with 24x7 operational coverage. - Exposure to mainframe systems, distributed applications, APIs, and integrated technology environments. - Knowledge of service restoration processes, escalation management, and operational support best practices. You should possess the ability to work effectively in high-pressure, mission-critical environments, strong analytical and problem-solving skills, excellent communication and stakeholder management skills, ability to coordinate across multiple teams, a strong sense of ownership and accountability, and the abilit
More at IntraEdge
Related open roles
Fircosoft Application Specialist / Engineer (Secunderabad)
Hyderabad
Fircosoft Application Specialist / Engineer
India
Oracle CPQ Support & Administration Specialist
Chennai
Production Support Analyst (Command Center) (Pune)
Mumbai
Sr. Software Engineer - SRE Support
Mumbai
Bank Argo Teller Technical Analyst
United States