Source description
About the role
Job Description – NOC Engineer
We are looking for a proactive and technically skilled NOC Engineer to join our 24x7 Operations team. The ideal candidate should have hands-on experience in monitoring production environments, managing incidents, and ensuring service availability while working in a fast-paced operational environment.
Key Responsibilities
• Monitor production infrastructure and applications using Datadog to proactively identify, investigate, and respond to alerts.
• Handle production incidents across different severity levels and ensure timely resolution in accordance with defined SLAs.
• Act as an Incident Commander for high-severity (SEV) incidents by leading bridge calls, coordinating with cross-functional teams, driving incident resolution, and providing regular stakeholder updates.
• Perform incident triaging, impact assessment, escalation, and coordination until service restoration.
• Create, update, and manage JIRA tickets for incident tracking, assignment, and closure.
• Coordinate and facilitate Change Approval Board (CAB) calls by reviewing planned changes, validating implementation readiness, assessing operational risks, and ensuring adherence to change management processes.
• Monitor production systems during and after change implementations to identify and mitigate any operational impact.
• Provide timely and accurate communications to stakeholders during incidents and planned maintenance activities.
• Collaborate with Engineering, Development, Infrastructure, and Support teams to resolve production issues efficiently.
• Ensure proper shift handovers, maintain operational runbooks, and update knowledge base documentation.
• Participate in post-incident reviews (PIRs), root cause discussions, and continuous service improvement initiatives.
Required Skills
• Hands-on experience with Datadog monitoring, dashboards, and alert management.
• Strong experience in Incident Management within a Production Support or NOC environment.
• Proven experience handling Severity (SEV) incidents and serving as an Incident Commander.
• Experience coordinating and managing CAB (Change Approval Board) calls and change implementation activities.
• Hands-on experience with JIRA for incident, problem, and change tracking.
• Excellent verbal and written communication skills with the ability to coordinate across multiple technical and business teams.
• Strong analytical, troubleshooting, and decision-making skills with the ability to perform effectively under pressure.
• Familiarity with ITIL processes, including Incident, Change, and Problem Management, is preferred.
• Willingness to work in a 24x7 rotational shift environment, including weekends and holidays.
More at Ascendion