Sr. Lead Infrastructure Engineer Storage SRA
Job Summary
Become a member of a team where you can contribute significantly to shaping the future of a world-renowned and influential company. Among top performers you can make a direct and meaningful impact.
As a Senior Lead Infrastructure Engineer at JPMorganChase within the Corporate Sector Infrastructure Platforms you exhibit both depth and breadth of knowledge regarding software applications and technical processes across multiple technical disciplines. You also have a specialization in a specific domain within infrastructure engineering to drive programs or initiatives consisting of multiple technologies and applications.
Job responsibilities
Applies deep technical expertise and problem-solving methodologies focused on analyzing complex data and systems anticipating issues and finding ways to mitigate risk
Works with other platforms to architect and implement changes required to resolve issues and modernize the organization and its technology processes
Be responsible for infrastructure engineering in accordance with business requirements
Executes work according to compliance standards risk and security and business objectives
Own end-to-end problem detection resolution and prevention; lead efforts to improve MTTD MTTM and MTTR through better observability triage and remediation practices.
Drive a workstream or project spanning one or more infrastructure engineering technologies includingtechnology lifecycle managementand dependency management acrossnetwork compute database and applicationplatforms.
Partner with adjacent platform teams toarchitect and implementchanges that resolve systemic issues reduce operational risk and modernize technology and operational processes.
Design and deliver creative scalable solutions for high-complexity engineering challenges including development of automation tooling and repeatable operational patterns.
Evaluate upstream/downstream impacts across systems data flows and integrations; proactively identify risks and provide clear mitigation and contingency recommendations.
Expertise to continuously improve telemetry and observability of storage products with tools like Grafana Dynatrace Prometheus Splunk Netcool etc.
Required qualifications capabilities and skills
Formal training or certification on site reliability engineering concepts and 5 years applied experience
Own reliability and operational excellence for IP-based storage services including availability performance capacity and operational risk reduction.
Act as the SME for enterprise storage platforms (e.g. Dell EMC PowerFlex NetApp SolidFire Pure Storage PMAX) including design input troubleshooting upgrades and lifecycle management. Provide 24x7 support coverage incident response and escalations; lead restoration activities during high-severity events.
Drive SRE best practices; Define measure and improve SLIs/SLOs error budgets and operational KPIs for storage services. Build and maintainobservability: metrics logs traces (where applicable) dashboards alerts and runbooks.
Reduce operational load by identifying and eliminating toil throughautomation; Automate repetitive tasks (provisioning health checks failover validation reporting hygiene tasks). Implement safe automation with guardrails change controls and rollback strategies. Expert knowledge of AAAS automation.
Perform root cause analysis (RCA)and problem management; implement corrective and preventive actions to prevent recurrence.
Maintain and continuously improverunbooks standard operating procedures on-call playbooks and knowledge articles. Collaborate with engineering network compute and platform teams onarchitecture reviews change planning and reliability improvements.
Support capacity managementand performance engineering: forecasting trending saturation analysis and proactive compliance with operational standards (change management risk controls documentation audit readiness).
Role participates in a 24x7 on-call rotation and is expected to respond to production incidents within defined SLAs. May require off-hours work for planned maintenance upgrades and risk-reduction activities.
Preferred qualifications capabilities and skills
Experience with storage networking and adjacent technologies (VLANs MTU/jumbo frames latency analysis QoS).
Understanding of resilience patterns: redundancy failure domains replication backup/restore DR testing.
Experience driving operational maturity initiatives (SLO rollout runbook standardization automation roadmaps).
Clear SLOs/SLIs and dashboards adopted by stakeholders; fewer false positives and more actionable alerts.
Improved platform stability and predictable performance/capacity posture.
Required Experience:
Senior IC
About Company
JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world’s most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans ov ... View more