Enter a job title or keyword

Sr. NOC Engineer (NOC, SRO, Scripting, Observability, AWS, Monitoring Tools)

Vertafore


Job Location:

Hyderabad - Pakistan

Monthly Salary: Not provided by the employer
Posted: 21 August 2026 (13 days ago)
Application Deadline: 18 November 2026
Vacancies: 1 Vacancy

Job Summary

Role Summary

We are seeking a Senior Reliability Operations Engineer to own day-to-day production operations service continuity operational readiness and restoration excellence for critical production services. This role is responsible for monitoring effectiveness incident coordination production change execution operational governance runbook maturity automation adoption and continuous operational improvement across AWS hybrid data centers and customer-hosted environments.

As part of the Global Command Center (GCC) this role partners with SRE engineering platform security product support and business teams to ensure services are observable supportable resilient and operationally mature. The role focuses on operational execution and service reliability outcomes while partnering with engineering on systemic reliability improvements.

Key Responsibilities
Service Reliability & Production Operations

Own operational execution for day-to-day service health including monitoring triage escalation restoration coordination operational risk tracking and service continuity across cloud and hybrid environments.

Lead operational readiness reviews for new services releases infrastructure changes and production onboarding; ensure support models runbooks dashboards alerts escalation paths rollback procedures access and validation checks are ready before handoff.

Operate and tune monitoring alerting dashboards and operational observability practices aligned with the Four Golden Signals; partner with SRE and engineering on telemetry and instrumentation standards.

Track SLO attainment SLA risk error budget burn operational KPIs recurring incidents and reliability trends; escalate service risk and recommend operational guardrails when thresholds are breached.

Monitor service performance infrastructure utilization capacity saturation and customer-impacting degradation to proactively identify and escalate operational risk.

Operational Excellence & Automation

Maintain an operational toil backlog and reduce repetitive manual work through automation scripting AI-assisted operations self-healing workflows AIOps capabilities and process simplification.

Plan execute coordinate and validate production changes including patching certificate renewals software releases infrastructure updates and maintenance using standardized change governance risk assessment rollback readiness stakeholder communication and post-change verification.

Investigate and troubleshoot complex production issues across applications infrastructure middleware databases and platform services; restore service while partnering with SRE/engineering on permanent remediation.

Develop and improve runbooks SOPs escalation frameworks recovery procedures production validation checks and knowledge documentation to ensure consistent support execution.

Identify recurring operational failure patterns and partner with SRE engineering platform and application teams to drive preventive controls and permanent corrective actions.

Incident Management Reporting & GCC Collaboration

Act as the operational incident lead during major incidents managing bridges stakeholder communications escalation paths restoration coordination event timelines and post-incident follow-up.

Facilitate blameless post-incident reviews and own corrective-action tracking for remediation items runbook updates automation opportunities preventive controls and operational improvements.

Own incident problem and service reliability trend reporting operational dashboards SLA/SLO tracking error budget burn visibility and leadership reviews.

Support the GCC operational model by driving globally standardized operational practices governance runbook maturity escalation frameworks and service management processes.

Collaborate with globally distributed SRE engineering cloud operations security support product and business teams to align operational priorities with reliability risks and service-continuity needs.

Mentor junior engineers and promote knowledge sharing while maintaining a customer-first mindset focused on service reliability responsiveness operational quality and business continuity.



Qualifications

3.5 - 5 years of hands-on experience in Production Operations Service Reliability Operations Infrastructure Operations NOC or related operational engineering roles.

Proven experience managing production environments with accountability for operational stability incident response and service continuity.

Strong understanding of operational readiness incident management production support operational governance change management problem management and reliability practices.

Practical experience with observability monitoring alert tuning dashboarding and platforms such as Datadog Splunk Grafana CloudWatch Dynatrace or similar tools.

Experience using SLO/SLA metrics error budget burn incident trends and operational KPIs to manage service risk and drive continuous improvement.

Hands-on experience supporting AWS Kubernetes CI/CD pipelines infrastructure platforms and hybrid environments.

Strong knowledge of Linux and Windows systems application hosting platforms middleware and relational databases.

Familiarity with automation and scripting using PowerShell Python Bash or similar technologies.

Exposure to ITIL operational governance globally distributed support or GCC/shared-services models is preferred.

Strong communication stakeholder management collaboration and problem-solving skills; bachelors/masters degree in computer science information systems or equivalent experience; participation in an on-call rotation for 24x7 GCC support is required.



Required Experience:

Senior IC


About Company

Company Logo

Looking to start your career in Technology? We have opportunities right here in mid-Michigan! Vertafore is looking for talented people to join our team in Michigan. Our dynamic environment provides professional development, fast upward mobility, and e

View Profile View Profile