Staff Network Reliability Engineer, Incident Management
Department:
Job Summary
The world still has coverage blind spots. You could help eliminate them at Skylo.
Skylo has pioneered a standards-based approach to satellite connectivity. We connect smartphones and IoT devices directly to satellites. No special hardware no entirely new networks. Just billions of existing devices suddenly reachable anywhere on Earth. Were not building toward this future. Were already in it.
Our direct-to-device service is live on millions of activated devices across five continents covering more than 72 million square kilometers in partnership with leading satellite operators mobile network operators Tier-1 chipset makers and OEMs worldwide. And were just getting started.
At the heart of it all is Skylos commercial NTN vRAN: a 3GPP standards-based cloud-native platform that seamlessly bridges terrestrial and satellite networks. Its the infrastructure that makes true anywhere anytime connectivity possible.
When you join Skylo youll work at the intersection of three markets reshaping how the world stays connected: mass-market consumer devices automotive and industrial IoT. Enabling people outdoors and critical workflows in the worlds most remote places.
This is a rare chance to work on technology that matters at a company thats already proving it works
Summary: How you will impact Skylo
As a Staff Network Reliability Engineer Core Operations in the Global Product Support & Customer Success organization you are the 5G Core domain authority within Skylos production NTN network. Where the Incident Manager coordinates the bridge you own the technical outcome. You are the escalation target for every Core-domain Sev 12 event the engineer who diagnoses AMF registration failures SMF session establishment drops UPF forwarding anomalies IMS SIP/Diameter failures and IMSI provisioning breakdowns at a protocol level and delivers a resolution or a definitive root cause.
On a network where every subscriber is roaming over satellite Core NF health is the difference between a connected device and a dark one. You own that health 247. You define the KPIs write and own the runbooks set the diagnostic standards the entire team operates against and continuously surface toil and failure patterns as engineering requirements. You are not a first responder; you are the last technical stop before a Core problem becomes an engineering escalation.
At Staff NRE level you are also a force multiplier: mentoring Senior NREs in Core domain depth contributing to the automation backlog with operational requirements and partnering with Product Engineering to ensure new Core releases meet operational readiness standards before they reach production.
Key Responsibilities
Core Network Operations & Health Ownership
Own 247 5G Core health across Skylos production NTN stack: AMF/SMF/UPF/AUSF pod status NAS/NG-AP signaling success rates session establishment and tear-down metrics subscriber registration KPIs IMS registration state and Core-layer SLA compliance.
Monitor and triage Core NF alarms using OSS dashboards Grafana/other inhouse telemetry and Loki log correlation distinguish transient anomalies from systemic degradation before escalating or acting.
Execute and own Core-domain runbooks for P2P4 fault categories: pod restarts persistent storage recovery certificate rotation IMSI state reconciliation and BSS-IIS cluster incident response without requiring engineering team involvement for covered fault classes.
Maintain DMP certificate management procedures and own escalation to BOSS (BSS & OSS) for DMP outages certificate rotation failures and EMS alarm integration issues.
Own IMSI lifecycle operations: activation deactivation KML file management subscriber state reconciliation and exception handling for provisioning failures through the OSS platform.
5G Core Incident Diagnosis & Escalation Authority
Serve as the L3 escalation authority for all Core-domain incidents: take ownership from the Incident Manager diagnose at the protocol level using NAS traces NG-AP message flows Diameter/SIP signaling captures and NF-specific log analysis and deliver a resolution or a decision-grade root cause.
Lead Core-domain troubleshooting bridges: command the technical investigation direct vendor and engineering participants correlate signals across AMF SMF UPF AUSF PCF and IMS NFs and drive the bridge to a documented resolution or a clear engineering handoff.
Diagnose and resolve Core failure modes: UE registration failures PDU session establishment drops handover interruptions IMS registration and call setup failures SEPP interconnect errors subscriber provisioning mismatches and signaling loop conditions.
Engage vendors with technical specificity: reproduce failures with log evidence own the vendor ticket lifecycle enforce SLA response commitments and escalate vendor delays with full impact context.
Participate in the global 247 on-call rotation as the Core domain escalation tier reachable within defined SLA windows for Sev 1 events; function as the technical decision-maker not the first responder.
Root Cause Analysis & Post-Incident Ownership
Own Core-domain RCA end-to-end: lead the post-incident investigation document the complete causal chain from triggering condition through downstream NF impact and deliver systemic action items with owners timelines and measurable success criteria.
Deliver Initial RCA documentation within defined SLA windows post-incident closure; own the final RCA through engineering review and sign-off.
Identify systemic failure patterns configuration drift missing alarm coverage stale thresholds vendor software defects and translate them into engineering requirements with clear impact scope and acceptance criteria.
Contribute to the weekly and monthly Network Performance Report: Core NF availability session success rates MTTR by fault category top recurring issues and SLA deviation analysis.
KPI Ownership & Performance Assurance
Define baseline and continuously refine Core KPIs: NAS registration success rates PDU session establishment ratios NG-AP failure rates IMS registration and session success UPF throughput and latency and subscriber provisioning success calibrated to Skylos NTN subscriber population not vendor defaults.
Proactively track Core availability metrics against MNO SLA commitments flag degradation trends before they breach thresholds and initiate preventive action before a subscriber impact occurs.
Establish and maintain Core NF performance baselines; detect and investigate deviations that indicate emerging failures capacity constraints or configuration regressions.
Drive proactive network issue detection through OSS alarm tuning alert threshold calibration and correlation rule improvements reduce noise increase signal fidelity and eliminate alarm storms that mask real events.
Runbook Authorship & Operational Standards
Author own and maintain all Core-domain runbooks and SOPs every procedure must be tested before it is relied upon in production; runbooks are living documents not static artifacts.
Define the diagnostic decision tree for each known Core fault class: entry condition triage steps isolation method resolution action and escalation criteria written at the level where a Senior NRE can execute independently.
Review and accept runbook contributions from Senior NREs; maintain the Core runbook library as the authoritative operational reference for the team.
Identify runbook gaps from incident post-mortems and operational observations; prioritize gap closure based on incident frequency and MTTR impact.
Cross-Functional Collaboration & Team Development
Partner with Product Engineering on Core software release readiness: define observability and operational acceptance criteria for new NF versions before they reach production; flag missing alarm coverage changed default configurations and new failure modes.
Partner with BOSS (BSS & OSS) on OSS platform stability IMSI management interface SLAs and EMS alarm integration provide operational requirements as input to OSS roadmap planning.
Collaborate with Cloud Infrastructure NRE on Kubernetes-layer issues affecting Core NFs: PVC availability pod scheduling failures network policy changes and GKE upgrade impacts.
Surface toil and automation opportunities to the Service Assurance & Automation team document the procedure frequency and MTTR cost as structured input to the automation backlog; contribute to closed-loop automation requirements for P3/P4 Core events.
Mentor Senior NREs in Core domain depth: KPI interpretation NAS/SIP/Diameter trace reading log correlation patterns and escalation judgment.
Vendor Management & Engineering Interface
Work closely with Core NF vendors for incident resolution and RCA manage vendor ticket lifecycle from opening through closure ensure reproducible evidence is provided and escalate responsiveness failures.
Engage Skylos Core engineering team with full operational context when issues exceed operational resolution authority deliver a structured problem statement timeline NF log bundle and a clear question rather than a vague escalation.
Support MNO partner technical discussions on Core-layer SLA definitions IMSI provisioning workflows and subscriber management API behavior provide operational evidence and data to support partner conversations.
Qualifications
Required
810 years of experience in 5G/4G Core operations or Core engineering in a production 247 environment carrier or vendor side with direct ownership of live Core NF infrastructure.
Deep 3GPP Core expertise: working knowledge of TS 23.501/23.502 NAS protocol NG-AP/S1AP IMS architecture (SIP Diameter CSCF/TAS) IMSI lifecycle procedures and subscriber provisioning flows.
Production 5G Core NF troubleshooting: demonstrated ability to diagnose AMF registration failures SMF session drops UPF forwarding anomalies AUSF authentication failures IMS SIP/Diameter errors and SEPP interconnect issues using NF logs and signaling traces.
Packet capture and call flow analysis: proficiency with Wireshark or equivalent for NAS SIP Diameter and HTTP/2 (SBI) protocol-level analysis.
Kubernetes-native Core operations: pod health monitoring rolling restart procedures PVC and persistent storage management log aggregation (Loki ELK or equivalent) and kubectl proficiency.
Production observability: Prometheus/Grafana/other inhouse tools for Core NF KPI dashboarding; ability to build queries set alert thresholds and interpret metric-level degradation signals.
Runbook authorship: ability to write diagnostic procedures at the level where a less-experienced engineer can execute them independently under incident pressure.
ITSM proficiency (Jira ServiceNow or equivalent): incident lifecycle management RCA documentation and vendor ticket tracking.
Strong written and verbal communication: capable of delivering RCA documents engineering escalations with structured problem statements and MNO-facing technical summaries.
Preferred
Experience with NTN or satellite connectivity Core operations: IMSI provisioning for NTN devices DMP integration BSS/IIS platform familiarity.
IMS operational experience: SIP registration and session troubleshooting Diameter routing CSCF and TAS log analysis.
Experience with OSS/BSS platforms: FCAPS alarm management EMS integration SNMP trap handling and provisioning system diagnostics.
Familiarity with Core NF vendors: Druid Nokia (Core) Mavenir Samsung or Oracle Core including vendor CLI tools log formats and support escalation processes.
Scripting ability in Python or Bash sufficient to automate log parsing build diagnostic queries or create operational data exports.
Experience contributing to closed-loop automation requirements or service assurance platform development.
ITIL certification or demonstrated application of ITIL problem management and change management processes in a production environment.
What We Offer
With employees working across three continents Skylo is proud to be an equal opportunity employer dedicated to building an inclusive and diverse workforce. Our worldwide and inclusive culture encourages a flexible approach to work and we also offer an attractive range of benefits such as:
Competitive compensation packages including a stock option based equity program
Comprehensive benefits plans
Monthly allowances for wellness and education reimbursement
A generous time off policy holidays and the opportunity to temporarily work abroad
Once-in-a-lifetime opportunity to be part of developing and running the worlds first commercial live direct-to-device satellite network and service
Access to a world-class team and talent across tech domains: software hardware chipsets telecom satellite and network virtualization
Open transparent inclusive culture that blends Silicon Valley Nordic and South Asia characteristics
EEO Statement
Skylo is an equal-opportunity employer and we celebrate diversity. We do not discriminate on the basis of race religion color ancestry national origin caste sex sexual orientation parent or caregiver status political affiliation gender gender identity or expression age disability medical condition pregnancy genetic makeup marital status or military service consistent with applicable federal state and local laws.
We are also committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. Please let us know if you need assistance or accommodation due to a disability.
Required Experience:
Staff IC
About Company
We are an NTN service provider, offering a service that enables smartphones, wearables, sensors, and other devices to connect by satellite, without requiring any special hardware