Enter a job title or keyword

Principal Site Reliability Engineer – Performance (A&D, Ultra HAExadata)

IFS


Job Location:

Ottawa - Canada

Monthly Salary: Not provided by the employer
Posted: 25 August 2026 (19 hours ago)
Application Deadline: 22 November 2026
Vacancies: 1 Vacancy

Job Summary

Job Description

The role of IFS Principal Site Reliability Engineer Performance Engineering (Principal-SRE) exists within the Unified Support organization/division and serves as a technical and strategic leader for Site Reliability Engineering practices. By being part of the shift operation and broader SRE leadership a Principal-SRE will drive 24x7x365 support excellence to IFS customers across the globe within the Aerospace & Defense (A&D) industry vertical while establishing and elevating SRE practices across the organization. This role sits within a dedicated performance engineering team supporting our Ultra High Availability (Ultra HA) on Exadata customers where sustained database and application performance is a contractual commitment rather than a best-effort goal.

As Principal SRE Performance Engineering you will be responsible for architecting and evolving the reliability frameworks automation platforms and operational strategies that underpin IFSs cloud-native products and infrastructure. This role involves setting technical direction leading cross-functional teams and partnering with R&D Product and Operations leadership to ensure that reliability is embedded at every level of service delivery. You will be an influential technical leader who drives organizational transformation through SRE principles automation and continuous improvement culture.

Strategic Impact Areas

  • Reliability Architecture & Strategy: Define and evolve the SRE operating model for the A&D vertical including incident management capacity planning and disaster recovery strategies. Architect resilience patterns and frameworks for cloud-based applications supporting critical Aerospace & Defense operations. Establish Service Level Objectives (SLOs) and Error Budgets aligned with customer contractual requirements and business objectives. Drive adoption of chaos engineering and resilience testing across production systems.
  • Technical Leadership & Platform Engineering: Lead the design and implementation of enterprise-scale monitoring observability and automated remediation platforms. Establish and maintain infrastructure-as-code standards and CI/CD pipelines for reliable deployments. Champion modern SRE tooling and methodologies across Unified Support and R&D organizations. Define technical architecture for Kubernetes multi-cloud orchestration and cloud-native application patterns. Architect solutions for performance optimization cost efficiency and security at scale (3200 concurrent users 6048 database connections).
  • Ultra HA & Exadata Performance Engineering: Own the performance and availability posture of Ultra HA environments running on Oracle Exadata including proactive database health monitoring workload and wait-event analysis and end-to-end application performance management. Establish deep observability across the Oracle stack and the application tiers it serves using Elastic Grafana and OpenTelemetry-based instrumentation to correlate database behaviour with user-facing latency. Lead performance troubleshooting for the most demanding customer workloads and drive tuning capacity and remediation decisions ahead of SLA impact.
  • Organizational Transformation & Culture: Drive automation-first culture eliminating toil and enabling teams to focus on high-value creative work. Lead post-incident reviews and establish blameless culture focused on systemic improvement. Mentor and develop SRE team members establishing career progression paths and technical mastery standards. Establish communities of practice for knowledge sharing across Unified Support and R&D organizations.

 

Key Responsibilities

Strategic Leadership & Architecture

  • Set technical direction and establish long-term reliability strategies for IFS cloud products supporting the A&D industry
  • Partner with R&D Product and Customer Success leadership to define reliability requirements and drive architectural decisions
  • Design and oversee implementation of enterprise-scale observability automation and incident response platforms
  • Lead architectural assessments and provide recommendations on infrastructure resilience capacity planning and disaster recovery
  • Drive adoption of SRE best practices across product teams establishing metrics and accountability for reliability

Performance Engineering & Ultra HA Operations

  • Own proactive performance monitoring for Ultra HA on Exadata customers database health workload profiles wait events and application response times and act on trends before they breach SLA
  • Lead deep performance investigations spanning the Oracle database application servers and infrastructure producing evidence-based tuning and capacity recommendations
  • Design and maintain the performance observability stack (Elastic Grafana OpenTelemetry/APM) including dashboards baselines and alert thresholds tied to customer SLOs
  • Run performance baselining load testing and pre-release validation for Ultra HA environments and quantify the impact of changes before they reach production
  • Produce customer-facing performance reporting and participate in service reviews for Ultra HA accounts

Incident Management & Operational Excellence

  • Lead resolution of long-running complex and critical incidents affecting A&D customers; establish escalation frameworks and decision-making protocols
  • Establish incident command systems on-call rotations and escalation procedures that balance response speed with team sustainability
  • Drive post-incident review processes that focus on identifying systemic improvements rather than individual blame
  • Establish incident trends analysis and drive prevention of recurrence through root cause elimination

Automation & Continuous Improvement

  • Champion the elimination of toil through intelligent automation leveraging AI to reduce redundant admin-heavy tasks
  • Design and implement self-healing systems automated remediation and predictive alerting to minimize manual intervention
  • Establish continuous improvement programs that measure track and drive down Mean Time To Detect (MTTD) and Mean Time To Recovery (MTTR)
  • Drive infrastructure-as-code adoption standardization and versioning across cloud platforms (Azure AWS GCP)
  • Establish runbook automation and GitOps practices for reliable auditable deployments

Documentation & Knowledge Management

  • Establish knowledge management frameworks and internal KBAs that guide support operations for cloud-based applications
  • Create and maintain architecture documentation disaster recovery playbooks and operational runbooks
  • Lead documentation standards that ensure knowledge is accessible accurate and actionable for all support tiers
  • Identify systemic knowledge gaps and drive closure through training documentation and process improvement

Cross-Functional Collaboration

  • Liaise with R&D and Unified Support Engineering to define automation requirements and platform capabilities for cloud applications
  • Partner with customer success and professional services teams to translate customer reliability requirements into technical strategies
  • Lead working groups and architecture reviews with multiple stakeholders to ensure reliability-first design decisions
  • Establish feedback loops between support operations and product development to drive product reliability improvements

Team Development & Mentorship

  • Mentor and develop SRE team members establishing technical mastery standards and career progression frameworks
  • Lead technical training programs on cloud operations Kubernetes performance engineering and SRE methodologies
  • Foster a culture of learning experimentation and continuous improvement within the SRE organization
  • Establish on-call culture that balances operational needs with team well-being and sustainable practices

Qualifications :

Required Qualifications

Education & Certifications (Core Requirements)

  • University degree or equivalent professional qualification in Software Engineering Computer Science Information Technology or similar discipline
  • Demonstrated expertise in modern ticket/service desk tooling (ServiceNow Jira Service Desk or equivalent)
  • ITIL ISO 20000 or equivalent IT service delivery framework certification; or demonstrated mastery through practical application

Experience (Mandatory)

  • Minimum 10 years of progressive experience in cloud computing services enterprise IT delivery or site reliability engineering
  • Minimum 7-8 years in hands-on SRE DevOps or cloud operations roles with demonstrated impact on reliability automation and operational efficiency
  • Proven experience leading SRE teams and establishing SRE practices across organizations
  • Deep understanding of low-level concepts in cloud computing (resource management networking storage security)
  • Demonstrated expertise in performance engineering including load testing capacity planning and system optimization
  • Experience supporting mission-critical high-availability systems (preferably in Aerospace Defense Finance or similar industries)
  • Hands-on experience operating or supporting Oracle Exadata platforms in production ideally under Ultra High Availability or comparable premium-SLA commitments
  • Demonstrated experience in a dedicated performance engineering or performance-analysis function owning latency throughput and capacity outcomes for named enterprise customers
  • Track record of designing and implementing large-scale automation platforms that reduce toil and improve MTTR
  • Demonstrated ability to establish and drive organizational change around reliability and operational excellence

Required Technical Skills

The successful candidate will demonstrate mastery in most of the following areas:

Cloud Infrastructure & Operations

  • Expert-level cloud service administration and operations (Azure AWS GCP)
  • Kubernetes/Docker architecture operations and troubleshooting at production scale
  • Network administration and troubleshooting (DNS Load Balancing VPN firewall rules)
  • Enterprise storage solutions block storage and object storage optimization

Oracle Database Monitoring Administration & Performance

  • Expert-level Oracle database administration and performance tuning
  • Expert-level Oracle database monitoring and diagnostics AWR ASH ADDM Statspack SQL trace/TKPROF and Oracle Enterprise Manager (OEM) with the ability to move from symptom to root cause under production pressure
  • Advanced Oracle performance troubleshooting: wait-event and contention analysis execution-plan regression SQL and PL/SQL tuning optimizer statistics partitioning and index strategy
  • Oracle high-availability architectures for Ultra HA workloads RAC Data Guard/Active Data Guard ASM and RMAN including failover testing and recovery validation
  • Understanding of connection pooling query optimization and capacity planning for high-concurrency systems (3000 connections)
  • Experience with database backup recovery and disaster recovery strategies
  • Hands-on Oracle Exadata administration and performance management (Smart Scan storage indexes IORM flash cache cell-level diagnostics) required for Ultra HA customer environments

Application & Web Server Administration

  • Web server administration and troubleshooting (Wildfly WebLogic Nginx)
  • Application performance monitoring (APM) and optimization across the full request path instrumentation transaction tracing and latency breakdown from web tier through middleware to the database
  • SSL/TLS certificate management and security best practices

Infrastructure-as-Code & Automation

  • Advanced proficiency in BASH PowerShell Python or Go scripting
  • Infrastructure-as-Code expertise (Terraform Ansible CloudFormation)
  • CI/CD pipeline design and implementation
  • GitOps and declarative infrastructure practices

Observability Monitoring & APM

  • Design and implementation of enterprise-scale monitoring solutions
  • Hands-on experience with Elasticsearch and the Elastic Stack (Elasticsearch Logstash/Beats Kibana) for log aggregation search and analysis at scale; equivalent experience with Splunk or DataDog also considered
  • Grafana dashboard design and metrics visualization including building performance dashboards that surface Oracle infrastructure and application metrics side by side (Prometheus or equivalent metrics backends)
  • OpenTelemetry-based instrumentation and distributed tracing or equivalent APM tooling (Dynatrace AppDynamics New Relic Elastic APM)
  • Ability to correlate database infrastructure and application telemetry into a single performance narrative and to build proactive alerting on leading indicators of degradation rather than outright failure

Debugging & Troubleshooting

  • Expert ability to debug complex multi-tier applications
  • Deep understanding of application server internals and JVM diagnostics
  • Network packet analysis and protocol debugging
  • Linux system troubleshooting and performance analysis

Additional Information :

Soft Skills

  • Strategic thinking and ability to translate business objectives into technical strategies
  • Executive communication skills with ability to present complex technical concepts to non-technical stakeholders
  • Team leadership and mentorship abilities with proven track record of developing technical talent
  • Change management and organizational influence skills
  • Problem-solving with ability to innovate new approaches to complex challenges
  • Decision-making under pressure with ability to weigh trade-offs between reliability cost and feature delivery
  • Emotional intelligence and ability to build trust across diverse international teams
  • Proactive ownership of work items and willingness to challenge status quo for improvement
  • Self-learning and ability to stay current with rapidly evolving cloud technologies and SRE practices
  • Excellent communication in English (verbal and written) for cross-functional collaboration

Desirable Qualifications & Nice-to-Have Skills

  • Advanced cloud security best practices (Zero Trust Identity & Access Management encryption)
  • Extensive experience supporting large Aerospace & Defense customers and understanding of industry-specific compliance requirements
  • Performance engineering mastery including advanced load testing capacity planning and cost optimization strategies
  • Experience with Disaster Recovery (DR) planning and execution including multi-region failover and RTO/RPO optimization
  • Expertise in cost optimization and FinOps practices for cloud infrastructure
  • Hands-on experience with event streaming platforms (Kafka RabbitMQ) and asynchronous architecture patterns
  • API design and management experience
  • Machine Learning/AI applications for observability anomaly detection or predictive remediation
  • Certifications such as AWS Solutions Architect Professional Azure Solutions Architect Expert or CKA (Certified Kubernetes Administrator)
  • Oracle certifications (OCP Database Administrator Exadata Database Machine Certified Implementation Specialist) or Elastic/Grafana certifications
  • Experience contributing to OpenTelemetry Elastic or Grafana ecosystems or building custom exporters and instrumentation libraries
  • Published articles conference talks or open-source contributions in SRE or DevOps domains
  • Contribution to SRE community through mentorship speaking or knowledge sharing

Role Scope & Impact

Reporting Structure

  • Reports to: Senior Leadership (Engineering Manager Technical Director or VP Engineering)
  • Team Leadership: Manages SRE team of 3-5 engineers; influences broader organization
  • Scope: Responsible for reliability and performance strategy across the A&D vertical touching 6000 customer deployments with direct ownership of performance outcomes for Ultra HA on Exadata accounts

Key Performance Indicators

  • System reliability metrics (uptime % SLA attainment error rates)
  • Performance metrics for Ultra HA accounts (application response time database wait time throughput at peak capacity headroom)
  • Observability coverage and alert quality (instrumented services false-positive rate share of issues detected before customer impact)
  • Operational efficiency (MTTR reduction incident frequency automation coverage)
  • Team development (mentee growth retention skill advancement)
  • Business impact (customer satisfaction churn reduction revenue protection)

Remote Work :

No


Employment Type :

Full-time


About Company

Company Logo

We are growing! At IFS we are constantly growing to deliver award-winning solutions to hundreds of partners and thousands of customers worldwide! We help companies who want to be their best when it matters most – at their #momentofservice. Visit https://ifs.link/IzM0px to find out mo ... View more

View Profile View Profile