Padmi

Production Support Analyst (Command Center)

MumbaiPosted 2 months ago
IT supportMid-levelFull Time; Regular
Apply at IntraEdge

Opens the source posting on shine.com

Source description

About the role

View original

You are an experienced L1/L2 Monitoring & Incident Management Specialist responsible for supporting centralized monitoring and incident command operations for business-critical, client-facing applications across mainframe and distributed environments. Your role involves proactive monitoring, rapid incident response, effective stakeholder communication, and coordination across technical teams to ensure service availability and operational excellence. Key Responsibilities: - Provide 24x7 monitoring and operational support for client-facing applications and services. - Monitor real-time system alerts and respond promptly to incidents impacting business operations. - Act as the primary point of contact during incidents, ensuring timely communication with clients and stakeholders. - Perform initial incident triage, impact assessment, and preliminary root cause analysis. - Coordinate with Infrastructure, Application, and Vendor teams to facilitate incident resolution and service restoration. - Manage and track incidents through their complete lifecycle, ensuring adherence to defined SLAs. - Handle production issues including batch failures, system alerts, service degradation, and application outages. - Escalate critical issues appropriately through ITSM platforms such as ServiceNow. - Maintain accurate incident records, status updates, and post-incident documentation. Required Skills & Experience: - Solid experience with enterprise monitoring and alert management tools. - Hands-on experience with IT Service Management (ITSM) platforms, preferably ServiceNow. - Solid understanding of incident management processes and incident lifecycle management. - Experience working in production support environments with 24x7 operational coverage. - Exposure to mainframe systems, distributed applications, APIs, and integrated technology environments. - Knowledge of service restoration processes, escalation management, and operational support best practices. You should possess the ability to work effectively in high-pressure, mission-critical environments, strong analytical and problem-solving skills, excellent communication and stakeholder management skills, ability to coordinate across multiple teams, a strong sense of ownership and accountability, and the ability to prioritize and manage multiple incidents simultaneously while maintaining service quality. Preferred Qualifications: - Experience supporting enterprise-scale production environments. - Familiarity with SLA-driven support models and operational governance. - Understanding of ITIL principles and service management frameworks. - Experience working with both mainframe and distributed technology stacks. You are an experienced L1/L2 Monitoring & Incident Management Specialist responsible for supporting centralized monitoring and incident command operations for business-critical, client-facing applications across mainframe and distributed environments. Your role involves proactive monitoring, rapid incident response, effective stakeholder communication, and coordination across technical teams to ensure service availability and operational excellence. Key Responsibilities: - Provide 24x7 monitoring and operational support for client-facing applications and services. - Monitor real-time system alerts and respond promptly to incidents impacting business operations. - Act as the primary point of contact during incidents, ensuring timely communication with clients and stakeholders. - Perform initial incident triage, impact assessment, and preliminary root cause analysis. - Coordinate with Infrastructure, Application, and Vendor teams to facilitate incident resolution and service restoration. - Manage and track incidents through their complete lifecycle, ensuring adherence to defined SLAs. - Handle production issues including batch failures, system alerts, service degradation, and application outages. - Escalate critical issues appropriately through ITSM platforms such as ServiceNow. - Maintain accurate incident records, status updates, and post-incident documentation. Required Skills & Experience: - Solid experience with enterprise monitoring and alert management tools. - Hands-on experience with IT Service Management (ITSM) platforms, preferably ServiceNow. - Solid understanding of incident management processes and incident lifecycle management. - Experience working in production support environments with 24x7 operational coverage. - Exposure to mainframe systems, distributed applications, APIs, and integrated technology environments. - Knowledge of service restoration processes, escalation management, and operational support best practices. You should possess the ability to work effectively in high-pressure, mission-critical environments, strong analytical and problem-solving skills, excellent communication and stakeholder management skills, ability to coordinate across multiple teams, a strong sense of ownership and accountability, and the abilit

One address, no account. We’ll tell you when matching roles go live.

More at IntraEdge

Related open roles

View all roles