Padmi

L3 - Production Support (CEPD)

IndiaPosted 1 month ago
Software engineeringSeniorFull Time; Regular

Source description

About the role

View original

Key Responsibilities: ADVANCED PRODUCTION SUPPORT AND SERVICE RELIABILITY Act as the final technical escalation point for complex or high-impact issues across IVR, ACD, CTI, agent desktop, dialer, recording, reporting, workforce interfaces, and omnichannel services.Diagnose failures spanning application code, operating systems, databases, middleware, APIs, message queues, networks, SIP/VoIP, gateways, SBCs, carriers, and third-party integrations.Review platform health, capacity, performance, availability, failover readiness, and observability; define proactive monitoring and reliability improvements.Participate in the rotational shifts including night shifts and the on-call roster and provide expert support for critical incidents outside business hours. MAJOR INCIDENT, PROBLEM, AND ROOT CAUSE MANAGEMENT Lead technical recovery during P1/P2 incidents, establish the troubleshooting strategy, coordinate resolver teams, validate restoration, and advise incident leadership.Perform deep log, trace, packet, query, thread, heap, transaction, and call-flow analysis to isolate complex faults and identify systemic causes.Own or technically lead RCA for recurring and high-severity incidents; define corrective and preventive actions and drive permanent fixes through closure.Review L2 findings, identify diagnostic gaps, maintain known-error records, and improve escalation criteria and recovery procedures.Support vendor/OEM escalations with complete evidence, reproducible scenarios, impact details, and technical follow-up. CHANGE, RELEASE, CONFIGURATION, AND ENGINEERING SUPPORT Provide technical governance for deployments, upgrades, patches, hotfixes, certificate renewals, configuration changes, migrations, and environment refreshes.Review implementation, validation, rollback, dependency, risk, and post-change monitoring plans for complex or high-risk changes.Lead production-readiness reviews, smoke and regression validation, performance verification, and post-release stabilization.Diagnose product defects, collaborate with development/product teams on fixes, and validate solutions before production implementation.Maintain production baselines, configuration standards, version inventories, dependency maps, and technical debt registers. ARCHITECTURE, INTEGRATION, AND PERFORMANCE SUPPORT Troubleshoot and optimize CRM, ticketing, identity, API/web-service, database, reporting, messaging, CTI, and telecom integrations.Support complex SIP call flows, trunks, SBCs, gateways, codecs, DTMF, numbering plans, routing, media paths, and carrier interoperability.Develop and review advanced SQL queries, scripts, diagnostic utilities, and controlled automation without compromising production controls.Support IBM MQ or equivalent middleware, including queue managers, channels, listeners, clustering, persistence, transactions, security, and recovery.Analyze capacity and performance trends, identify bottlenecks, and recommend tuning, scaling, resilience, or architectural improvements. TECHNICAL LEADERSHIP AND CONTINUOUS IMPROVEMENT Mentor L1/L2 engineers, lead knowledge-transfer sessions, review technical documentation, and improve troubleshooting capability.Create and govern SOPs, runbooks, diagnostic guides, automation standards, known-error records, and disaster-recovery procedures.Drive observability, automation, self-healing, repeat-incident reduction, operational-risk remediation, and service-improvement initiatives.Contribute to architecture reviews, capacity planning, DR exercises, security and audit reviews, vendor governance, and service reviews. Key Responsibilities: ADVANCED PRODUCTION SUPPORT AND SERVICE RELIABILITY Act as the final technical escalation point for complex or high-impact issues across IVR, ACD, CTI, agent desktop, dialer, recording, reporting, workforce interfaces, and omnichannel services.Diagnose failures spanning application code, operating systems, databases, middleware, APIs, message queues, networks, SIP/VoIP, gateways, SBCs, carriers, and third-party integrations.Review platform health, capacity, performance, availability, failover readiness, and observability; define proactive monitoring and reliability improvements.Participate in the rotational shifts including night shifts and the on-call roster and provide expert support for critical incidents outside business hours. MAJOR INCIDENT, PROBLEM, AND ROOT CAUSE MANAGEMENT Lead technical recovery during P1/P2 incidents, establish the troubleshooting strategy, coordinate resolver teams, validate restoration, and advise incident leadership.Perform deep log, trace, packet, query, thread, heap, transaction, and call-flow analysis to isolate complex faults and identify systemic causes.Own or technically lead RCA for recurring and high-severity incidents; define corrective and preventive actions and drive permanent fixes through closure.Review L2 findings, identify diagnostic gaps, maintain known-error records, and improve escalation criteria and recovery procedu

One address, no account. We’ll tell you when matching roles go live.

More at IFTAS - Indian Financial Technology & Allied Services

Related open roles

View all roles