SeniorPrincipal SRE Engineer
Job Summary
Our Mission
At Palo Alto Networks were united by a shared missionto protect our digital way of life. We thrive at the intersection of innovation and impact solving real-world problems with cutting-edge technology and bold thinking. Here everyone has a voice and every idea counts. If youre ready to do the most meaningful work of your career alongside people who are just as passionate as you are youre in the right place.
Who We Are
In order to be the cybersecurity partner of choice we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption Collaboration Execution Integrity and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest we invite you to join us!
We believe collaboration thrives in person. Thats why most of our teams work from the office full time with flexibility when its needed. This model supports real-time problem-solving stronger relationships and the kind of precision that drives great outcomes.Job Summary
Your Career:
Were hiring a Senior/Principal Site Reliability Engineer to own production reliability for Cortex Agentix Endpoint Security (following an acquisition of KOI Start Up) as it scales. Youll define and operate our SLOs and error budgets lead high-severity incident response and ensure our Kubernetes and AWS infrastructure stays stable under growth. Youll also build and supervise the AI agents that handle routine alert triage and monitor tuning focusing your own time on the reliability engineering that requires human judgment. This role is a strong fit for someone who treats reliability as an engineering discipline and enjoys ownership incident command and applying AI to operational work.
Your Impact:
- Own reliability as an engineering discipline - define SLIs set SLOs and run error-budget-based decision-making so how reliable are we becomes a number that governs how fast we ship.
- Own production incidents end-to-end - lead response mitigation and resolution for high-severity incidents and drive blameless postmortems that feed real fixes back into the system.
- Own the reliability and capacity of production infrastructure as we scale - forecasting headroom validating scaling behavior under load and keeping latency and error rates within SLO.
- Run and evolve Kubernetes environments so releases and infra changes are safe by default across hundreds of tenant apps.
- Own build and supervise our SRE AI agents that triage alerts review monitors resolves and summarize incidents. Set and expand the trust ladder that governs what the agents do autonomously what needs approval and what stays human. This is a core part of the role.
- Improve observability and incident response - raise signal quality cut alert noise and own the monitoring the triage agents depend on.
- Eliminate toil - relentlessly identify manual repetitive operational work and remove it through automation and agents protecting engineering time for reliability work that only humans can do.
- Analyze operational data across incidents alerts deployments infra health and cost to find reliability gaps capacity risks and automation opportunities.
- Evaluate and introduce new tools and AI-assisted approaches balancing innovation with reliability cost and operational simplicity.
Qualifications
- 5 years operating production cloud infrastructure with a strong reliability focus (SRE or DevOps/platform engineering with reliability ownership).
- Deep hands-on experience with Kubernetes Helm ArgoCD Terraform and CI/CD.
- Experience defining and operating SLIs SLOs and error budgets - or a clear grasp of the discipline and the drive to establish it from scratch.
- Strong observability and alerting experience in Datadog or comparable platforms including raising signal-to-noise in production.
- Proven incident-response instincts - comfortable owning high-severity incidents and a genuine believer in blameless postmortems.
- Proven ability to own platform and reliability projects end-to-end from design through production operation and ongoing improvement.
- Strong troubleshooting across distributed systems Kubernetes CI/CD and live incidents.
- Collaborative mindset - comfortable working across engineering security product and leadership.
- Comfort in a fast-paced high-ownership environment where priorities shift but production quality doesnt.
- Genuine interest in applying AI automation and intelligent workflows to operational work - and in building and supervising agents not just using them.
- Ownership-driven - You take responsibility for the reliability of the systems you build and operate from SLO definition through incident command and continuous improvement.
- Reliability as engineering - You treat reliability as a software problem to be solved with code measurement and automation - not an ops queue to be worked by hand.
- Collaboration - You work effectively across engineering security product and leadership to align on reliability priorities and drive shared outcomes.
- Innovation balanced with pragmatism - You actively explore new approaches particularly AI-assisted operations and agent supervision while weighing them against reliability maintainability and operational simplicity.
- Security mindset - You design and build with least privilege auditability and production safety as foundational principles rather than afterthoughts.
- Clear communication - You articulate reliability risk cost and security tradeoffs precisely to both technical and non-technical stakeholders.
Our Commitment
Were trailblazers that dream big take risks and challenge cybersecuritys status quo. Its simple: we cant accomplish our mission without diverse teams innovating together.
We are committed to providing reasonable accommodations for all qualified individuals with a disability. If you require assistance or accommodation due to a disability or special need please contact us at .
Palo Alto Networks is an equal opportunity employer. We celebrate diversity in our workplace and all qualified applicants will receive consideration for employment without regard to age ancestry color family or medical care leave gender identity or expression genetic information marital status medical condition national origin physical or mental disability political affiliation protected veteran status race religion sex (including pregnancy) sexual orientation or other legally protected characteristics.
All your information will be kept confidential according to EEO guidelines.
Is role eligible for Immigration Sponsorship No. Please note that we will not sponsor applicants for work visas for this position.Required Experience:
Staff IC
About Company
Palo Alto Networks is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information ... View more