Lead Site Reliability Engineering Network
Palo Alto, CA - USA
Job Summary
Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.
Job responsibilities
- Applies network reliability principles (Permit to Operate FMEA operational readiness) balancing feature delivery efficiency and stability.
- Partners with network engineering domains (Datacenter Firewall Proxies DMZ Load Balancing etc.) and Lines of Business to align goals and outcomes.
- Drives adoption of reliability best practices and observability demonstrating impact through stability/reliability metrics.; Bridges Engineering Operations DevOps and customers to build resilient scalable and secure network services.
- Provides Tier-3 network support leading major incident response rapid restoration RCA and follow-through on corrective actions.
- Leads reliability and stability initiatives using data-driven analysis to improve service levels and reduce recurring failure modes.
Defines SLI/SLOs and error budgets with stakeholders and customers ensuring measurable performance targets and trade-off clarity. Identifies and removes technical bottlenecks within core domains of expertise proactively preventing reliability and capacity risks.
Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage troubleshooting and post-incident analysis validating outputs and handling operational data according to sensitivity and security requirements.
- Runs blameless data-driven post-mortems and debriefs converting learnings (successes and failures) into actionable improvements.
- Fosters continuous improvement and strong knowledge sharing soliciting real-time feedback avoiding duplicated work and promoting innovation via internal communities.
- Produces and packages thought leadership with specialists/product/engineering teamsdocumenting best practices and lessons learned for internal assets and industry forums/conferences.
- Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g. CI/CD quality checks test/validation automation and operational readiness) ensuring traceability/auditability resiliency and security controls.
Required qualifications capabilities and skills
- Formal training or certification in network engineering concepts and 5 years of applied experience.
- 10 years of experience leading technologists to manage and solve complex technical items within your domain of expertise.
Advanced proficiency in network reliability engineering including Permit to Operate FMEA and operational readiness processes.
Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g. incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
Ability to evaluate AI-assisted operational recommendations for correctness and risk define appropriate guardrails for team usage and ensure outcomes align to resiliency and security expectations.
Experience leading technologists to manage and solve complex network issues at a firmwide level.
- Ability to influence team culture by championing innovation and change for success.
- Proficiency in SD-WAN cloud platforms (AWS Azure etc.) and major network technologies (Palo Alto Juniper F5 Broadcom Arista Cisco etc.).
Proficiency in observability and monitoring tools such as Grafana SevOne Prometheus Kibana ThousandEyes and Splunk.
- CCIE Load-balancing SD-WAN Observability tools eBPF Cloud certs
- Demonstrated proficiency in troubleshooting and supporting complex networking environments including Tier-3 operational support for major incidents.
Experience with continuous integration and delivery tools (e.g. Jenkins GitLab Terraform etc.).
- Experience in scalable networking design including high availability redundancy failover and load balancing.
- Experience troubleshooting networking protocols such as TCP/IP HTTPS and BGP.
- Experience in customer-facing migration including service discovery assessment planning execution and operations.
This position is subject to Section 19 of the Federal Deposit Insurance Act. As such an employment offer for this position is contingent on JPMorganChases review of criminal conviction history including pretrial diversions or program entries.
Required Experience:
IC
About Company
JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world’s most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans ov ... View more