Senior Site Reliability Engineer (Networking)
Job Summary
We are hiring a Senior Site Reliability Engineer (Networking) to join OCI Corporate Network and Security Operations. This role is responsible for improving the reliability scalability performance and operational excellence of Oracles global corporate network infrastructure. The engineer will collaborate with Network Engineering Security Software Development and Cloud Operations teams to automate operations improve service resilience resolve complex production incidents and ensure highly available network services supporting Oracles enterprise environment.
The successful candidate combines Site Reliability Engineering principles with strong expertise in enterprise networking automation cloud infrastructure and operational excellence. This role requires a proactive mindset focused on improving service reliability through automation observability capacity planning and continuous improvement while supporting mission-critical production environments.
Responsibilities
Key Responsibilities
Capacity Ingestion and Management
- Takes proactive steps to design and architect infrastructure and/or services according to reliability availability scalability and performance requirements.
- Designs and improves highly available enterprise network infrastructure supporting Oracle Corporate Network and Security Operations.
- Forecasts infrastructure and network capacity requirements and responds to capacity needs to ensure systems can support current and future workloads.
- Collaborates with Software Development Network Engineering and Security teams to develop reliable scalable and resilient infrastructure and services.
- Independently identifies opportunities for innovation and drives prototyping initiatives including onboarding new infrastructure services and automation capabilities.
Incident and Service Lifecycle Management
- Performs data collection triage technical analysis and issue resolution to maintain and optimize infrastructure and network reliability.
- Independently monitors production services and enterprise network infrastructure maintaining up-to-date knowledge of service health and performance.
- Supports Corporate Network and Security Incident Change Capacity and Problem Management processes.
- Leverages comprehensive knowledge to perform incident response root cause analysis (RCA) maintenance activities software upgrades security updates backup and recovery.
- Performs deep troubleshooting across enterprise and cloud network environments including Layer 2/Layer 3 technologies routing switching DNS VPN load balancing and firewall platforms.
- Provides health and performance reporting while proactively identifying trends and opportunities to improve operational reliability.
- Performs provisioning activities supporting infrastructure applications and network services.
- Performs standard and non-standard decommissioning activities when required.
- Drives corrective and preventive actions following production incidents to eliminate recurring operational issues.
Automation
- Identifies opportunities for automation and evaluates operational benefits.
- Develops automation solutions using Python Ansible Infrastructure as Code and related technologies to improve operational efficiency and network reliability.
- Develops automation tools and scripts to gather metrics monitor services analyze system behavior mitigate operational risks and remediate production issues.
- Supports Zero Touch Provisioning and automated infrastructure lifecycle management where applicable.
- Independently validates automation solutions through testing to ensure expected functionality and operational readiness.
Technical Communication and Guidance
- Communicates the scale capacity security performance characteristics and operational requirements of services and enterprise network infrastructure.
- Identifies and explains the operational impact of infrastructure tooling and service changes before production implementation.
- Partners with Network Engineering Security Cloud Operations and Software Engineering teams during production changes maintenance activities and operational improvements.
- Effectively communicates operational risks capacity constraints and reliability improvements to technical stakeholders.
Troubleshooting and Resolution
- Provides operational support for Oracle technology services escalating incidents and complex operational issues as appropriate.
- Resolves technical issues spanning multiple infrastructure and network services while maintaining established Service Level Objectives (SLOs).
- Troubleshoots complex production issues involving enterprise routing switching firewalls VPN technologies DNS load balancers Linux systems cloud networking and distributed infrastructure.
- Documents incidents performs comprehensive Root Cause Analysis (RCA) and independently executes post-incident reviews to prevent recurrence.
- Participates in a 24x7 on-call rotation supporting mission-critical enterprise network services.
Innovation and Continuous Improvement
- Evaluates emerging networking technologies cloud capabilities observability platforms and automation solutions to improve operational efficiency.
- Independently identifies and implements improvements addressing performance bottlenecks scalability challenges and infrastructure optimization opportunities.
- Continuously improves monitoring observability automation and operational tooling supporting enterprise network environments.
- Maintains knowledge of Site Reliability Engineering and enterprise networking trends sharing knowledge and best practices across the organization.
- Performs standard and non-standard production analysis to support operational excellence and informed business decisions.
Required Qualifications
- Bachelors degree in Computer Science Information Technology Engineering or a related field or equivalent practical experience.
- 610 years of experience supporting enterprise or cloud network environments and mission-critical production infrastructure.
- Strong knowledge of enterprise networking technologies including TCP/IP Layer 2/Layer 3 switching routing firewalls VPNs load balancers DNS/DHCP and enterprise network operations.
- Hands-on experience with Incident Change Capacity and Problem Management Linux administration Python Ansible infrastructure automation cloud networking (OCI preferred; AWS Azure or GCP) and monitoring and observability platforms.
- Strong analytical troubleshooting and problem-solving skills with the ability to collaborate effectively across cross-functional engineering teams.
- Excellent verbal and written communication skills and willingness to participate in a 24x7 on-call support rotation.
Preferred Qualifications
Experience with modern networking cloud and automation technologies such as OCI Networking Terraform Git Infrastructure as Code (IaC) Zero Touch Provisioning Grafana Prometheus Splunk ELK Cisco Nexus Juniper Palo Alto Networks F5 Load Balancers BGP and OSPF is highly desirable. Experience designing and implementing automation solutions to improve operational efficiency scalability and network reliability in large-scale enterprise or cloud environments is a plus.
Core Responsibilities
Planning & Execution
Independently manages work monitoring timelines and deliverables to ensure projects or initiatives stay on track and meet requirements. Proactively prioritizes work and adapts to resource or timeline shifts suggesting adjustments to maintain project efficiency.
Collaboration & Partnership
Collaborates across teams to align on expectations and achieve shared objectives. Builds and maintains a comprehensive understanding of business stakeholder and/or customer needs to build and support effective partnerships. Actively listens to diverse perspectives and asks questions to ensure understanding of others.
Problem Solving
Independently identifies and addresses standard and non-standard issues in accordance with standard practices escalating more complex issues as appropriate. Analyzes data and/or information from multiple sources to troubleshoot standard and non-standard errors. Contributes to knowledge sharing and best practices.
Continuous Learning
Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools and staying current with industry trends and best practices. Seeks out and leverages feedback and training to improve skills. Contributes to a culture of continuous learning and knowledge sharing with team members.
Continuous Improvement
Develops ideas and recommends updates to increase the efficiency and effectiveness of processes protocols and workflows within a team. Seeks input from team members on alternative approaches and methods for improving work.
Qualifications
Career Level - IC3
Required Experience:
Senior IC
About Company
As a world leader in cloud solutions, Oracle uses tomorrow’s technology to tackle today’s challenges. We’ve partnered with industry-leaders in almost every sector—and continue to thrive after 40+ years of change by operating with integrity. We know that true innovation starts when eve ... View more