Lead Principal Site Reliability Engineer

Oracle


Job Location:

Nashville, TN - USA

Yearly Salary: $ 96300 - 264100
Posted on: 5 hours ago
Vacancies: 1 Vacancy

Job Summary

Description
Serves as a consultant and leads the design and architecture of infrastructure and service ensuring alignment with reliability and functionality standards. Takes full ownership of forecasting of demands and responding to capacity needs. Owns collaborations with software development teams to develop reliable and scalable infrastructures. Recommends methods for performing data collection to maintain and optimize operations and reliability. Oversees incident response and/or maintenance tasks. Provides strategic future-oriented health and performance reporting. Contributes to strategies for automation and reviews the development and implementation of automation. Provides expert-level communication about services and anticipates analyzes and explains the impact of changes considering strategic goals. Serves as a role model in providing support for technology and reviews documentation for accuracy. Leads the implementation of innovative tools and provides expertise in site reliability trends.

Responsibilities

Key Responsibilities

Reliability Strategy and Technical Leadership

  • Define and drive the site reliability engineering strategy for large-scale distributed and business-critical platforms.
  • Establish reliability standards engineering practices and operational readiness requirements across multiple teams.
  • Serve as a senior technical authority for system reliability scalability resilience performance and production operations.
  • Influence architecture and design decisions to ensure systems are supportable observable fault tolerant and capable of meeting availability objectives.
  • Identify systemic reliability risks and lead cross-functional initiatives to address them.
  • Provide technical direction and mentorship to site reliability engineers software engineers platform engineers and operations teams.
  • Lead technical reviews and promote consistent engineering practices across the organization.

Service Reliability and Observability

  • Define and implement service-level indicators service-level objectives error budgets and operational health metrics.
  • Develop comprehensive monitoring logging tracing alerting and observability strategies.
  • Improve the quality and actionability of alerts while reducing unnecessary operational noise.
  • Establish dashboards and reporting mechanisms that clearly communicate service health performance capacity and risk.
  • Use production data and reliability trends to prioritize engineering investments and continuous-improvement initiatives.

Automation and Platform Engineering

  • Design and implement automation that reduces manual intervention operational toil and human error.
  • Build or enhance tools for deployment configuration management infrastructure provisioning incident response and service recovery.
  • Promote infrastructure-as-code policy-as-code automated testing and repeatable deployment practices.
  • Partner with development teams to improve continuous integration and continuous delivery pipelines.
  • Develop self-healing and automated remediation capabilities where appropriate.
  • Contribute production-quality software and reusable platform components using modern programming and scripting languages.

Incident Management and Problem Resolution

  • Provide technical leadership during complex high-severity production incidents.
  • Coordinate diagnosis containment recovery and stakeholder communication during service disruptions.
  • Lead blameless post-incident reviews and ensure that corrective actions address root causes rather than symptoms.
  • Identify recurring failure patterns and develop long-term engineering solutions.
  • Improve incident-management processes escalation procedures runbooks and recovery playbooks.
  • Participate in an on-call rotation or provide senior escalation support for critical services as required.

Capacity Performance and Resilience

  • Lead capacity planning performance analysis load testing and demand forecasting for critical platforms.
  • Identify performance bottlenecks and recommend architectural or operational improvements.
  • Design and validate high-availability disaster-recovery backup and business-continuity capabilities.
  • Lead resilience testing failure-mode analysis game days and controlled fault-injection exercises.
  • Ensure recovery-time and recovery-point objectives are defined tested and achievable.

Security and Operational Governance

  • Partner with security and compliance teams to embed security into infrastructure automation and operational practices.
  • Support vulnerability remediation access-control improvements audit readiness and secure configuration management.
  • Ensure production environments meet organizational standards for change management data protection and operational governance.
  • Balance reliability security delivery speed cost and business priorities when recommending technical solutions.

Required Qualifications

  • Extensive professional experience in site reliability engineering software engineering cloud infrastructure platform engineering systems engineering or a related technical discipline.
  • Demonstrated experience designing operating and improving highly available production systems at significant scale.
  • Deep knowledge of distributed systems cloud architecture networking operating systems storage databases and service dependencies.
  • Advanced experience with at least one major cloud platform such as Oracle Cloud Infrastructure Amazon Web Services Microsoft Azure or Google Cloud Platform.
  • Strong experience with containerization and orchestration technologies including Docker and Kubernetes.
  • Proven expertise with infrastructure-as-code and configuration-management technologies such as Terraform Ansible Chef Puppet or equivalent tools.
  • Experience implementing observability solutions using metrics logs traces dashboards and automated alerting.
  • Strong programming or scripting skills in one or more languages such as Python Go Java JavaScript Bash or similar.
  • Experience with continuous integration continuous delivery automated testing and modern release-management practices.
  • Demonstrated leadership during critical production incidents and complex technical investigations.
  • Ability to diagnose difficult system issues across applications infrastructure networks databases and cloud services.
  • Strong written and verbal communication skills including the ability to explain technical risk and recommendations to engineering leaders and business stakeholders.
  • Proven ability to lead cross-functional technical initiatives without relying solely on formal authority.

Preferred Qualifications

  • Experience supporting enterprise-scale cloud services software-as-a-service platforms or other high-availability customer-facing systems.
  • Experience defining and operating service-level objectives error budgets and reliability scorecards.
  • Knowledge of chaos engineering resilience testing and automated recovery techniques.
  • Experience with multi-region hybrid-cloud or multi-cloud architectures.
  • Familiarity with security frameworks compliance requirements and regulated operating environments.
  • Experience improving cloud cost efficiency capacity utilization or infrastructure performance.
  • Contributions to internal engineering standards technical communities open-source projects or industry publications.
  • Bachelors or advanced degree in computer science engineering information systems or a related field or equivalent practical experience.



Qualifications
Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements such as immunization/occupational health mandates and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $96300 to $264100 per annum. May be eligible for bonus equity and compensation deferral.


Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge skills experience market conditions and locations as well as reflect Oracles differing products industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:
1. Medical dental and vision insurance including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - IC5





Required Experience:

Staff IC

DescriptionServes as a consultant and leads the design and architecture of infrastructure and service ensuring alignment with reliability and functionality standards. Takes full ownership of forecasting of demands and responding to capacity needs. Owns collaborations with software development teams ...

About Company

Company Logo

As a world leader in cloud solutions, Oracle uses tomorrow’s technology to tackle today’s challenges. We’ve partnered with industry-leaders in almost every sector—and continue to thrive after 40+ years of change by operating with integrity. We know that true innovation starts when eve ... View more

View Profile View Profile