Padmi
Skylo logo
Skylo

satellite connectivity · non-terrestrial networks (NTN)

Staff Network Reliability Engineer, Incident Management

BangalorePosted 2 months ago
Infrastructure And DatabasesStaff+Full Time; Regular
Apply at Skylo

Opens the source posting on shine.com

Source description

About the role

View original

The world still has coverage blind spots. You could help eliminate them at Skylo. Skylo has pioneered a standards-based approach to satellite connectivity. We connect smartphones and IoT devices directly to satellites. No special hardware, no entirely new networks. Just billions of existing devices, suddenly reachable anywhere on Earth. We're not building toward this future. We're already in it. Our direct-to-device service is live on millions of activated devices across five continents, covering more than 72 million square kilometers, in partnership with leading satellite operators, mobile network operators, Tier-1 chipset makers, and OEMs worldwide. And we're just getting started. At the heart of it all is Skylo's commercial NTN vRAN: a 3GPP standards-based, cloud-native platform that seamlessly bridges terrestrial and satellite networks. It's the infrastructure that makes true anywhere, anytime connectivity possible. When you join Skylo, you'll work at the intersection of three markets reshaping how the world stays connected: mass-market consumer devices, automotive, and industrial IoT. Enabling people outdoors and critical workflows in the world's most remote places. This is a rare chance to work on technology that matters, at a company that's already proving it works Summary: How you will impact Skylo As a Staff Network Reliability Engineer, Core Operations, in the Global Product Support & Customer Success organization, you are the 5G Core domain authority within Skylos production NTN network. Where the Incident Manager coordinates the bridge, you own the technical outcome. You are the escalation target for every Core-domain Sev 12 event the engineer who diagnoses AMF registration failures, SMF session establishment drops, UPF forwarding anomalies, IMS SIP/Diameter failures, and IMSI provisioning breakdowns at a protocol level, and delivers a resolution or a definitive root cause. On a network where every subscriber is roaming over satellite, Core NF health is the difference between a connected device and a dark one. You own that health 247. You define the KPIs, write and own the runbooks, set the diagnostic standards the entire team operates against, and continuously surface toil and failure patterns as engineering requirements. You are not a first responder; you are the last technical stop before a Core problem becomes an engineering escalation. At Staff NRE level you are also a force multiplier: mentoring Senior NREs in Core domain depth, contributing to the automation backlog with operational requirements, and partnering with Product Engineering to ensure new Core releases meet operational readiness standards before they reach production. Key Responsibilities Core Network Operations & Health Ownership Own 247 5G Core health across Skylos production NTN stack: AMF/SMF/UPF/AUSF pod status, NAS/NG-AP signaling success rates, session establishment and tear-down metrics, subscriber registration KPIs, IMS registration state, and Core-layer SLA compliance. Monitor and triage Core NF alarms using OSS dashboards, Grafana/other inhouse telemetry, and Loki log correlation distinguish transient anomalies from systemic degradation before escalating or acting. Execute and own Core-domain runbooks for P2P4 fault categories: pod restarts, persistent storage recovery, certificate rotation, IMSI state reconciliation, and BSS-IIS cluster incident response without requiring engineering team involvement for covered fault classes. Maintain DMP certificate management procedures and own escalation to BOSS (BSS & OSS) for DMP outages, certificate rotation failures, and EMS alarm integration issues. Own IMSI lifecycle operations: activation, deactivation, KML file management, subscriber state reconciliation, and exception handling for provisioning failures through the OSS platform. 5G Core Incident Diagnosis & Escalation Authority Serve as the L3 escalation authority for all Core-domain incidents: take ownership from the Incident Manager, diagnose at the protocol level using NAS traces, NG-AP message flows, Diameter/SIP signaling captures, and NF-specific log analysis, and deliver a resolution or a decision-grade root cause. Lead Core-domain troubleshooting bridges: command the technical investigation, direct vendor and engineering participants, correlate signals across AMF, SMF, UPF, AUSF, PCF, and IMS NFs, and drive the bridge to a documented resolution or a clear engineering handoff. Diagnose and resolve Core failure modes: UE registration failures, PDU session establishment drops, handover interruptions, IMS registration and call setup failures, SEPP interconnect e The world still has coverage blind spots. You could help eliminate them at Skylo. Skylo has pioneered a standards-based approach to satellite connectivity. We connect smartphones and IoT devices directly to satellites. No special hardware, no entirely new networks. Just billions of existin

One address, no account. We’ll tell you when matching roles go live.

More at Skylo

Related open roles

View all roles