Location: Mt. Laurel NJ / Philadelphia PA (Onsite)
Job Summary
STAFFXPERT LLC is seeking a Senior Site Reliability Engineer (SRE) Architect on behalf of our client in Mt. Laurel NJ / Philadelphia PA. This role is responsible for defining and governing operational architecture across cloud and on-premises environments ensuring high availability resilience observability automation and operational excellence. The ideal candidate will have extensive experience in telecom or large-scale enterprise environments and a strong background in Site Reliability Engineering infrastructure operations and service reliability.
Key Responsibilities
Define and govern operational architecture standards across cloud and on-premises infrastructure environments.
Lead infrastructure architecture reviews capacity planning initiatives network and firewall assessments and middleware operational reviews.
Design and manage SLO SLI and error budget frameworks to improve service reliability and operational performance.
Drive Site Reliability Engineering (SRE) best practices operational maturity programs and reliability improvement initiatives.
Lead disaster recovery (DR) business continuity planning (BCP) resiliency testing incident response processes and post-incident reviews.
Develop standards and governance for Infrastructure as Code (IaC) automation GitOps Terraform Ansible and related operational tooling.
Partner with cross-functional teams to enhance operational efficiency reduce alert fatigue and improve system stability.
Support operational governance lifecycle management security hardening initiatives and infrastructure modernization efforts.
Required Qualifications
10 years of experience in IT Operations Infrastructure Engineering Site Reliability Engineering (SRE) or related disciplines.
Experience working within telecommunications communications or other highly available enterprise environments.
Proven experience supporting carrier-grade or mission-critical platforms with stringent uptime and reliability requirements.
Strong expertise with observability and monitoring tools such as Splunk ELK Stack Dynatrace Grafana or similar platforms.
Hands-on experience designing and implementing SLO SLI and error budget frameworks.
Working knowledge of public cloud platforms including AWS Azure or GCP.
Experience with Infrastructure as Code and automation tools such as Terraform Ansible GitOps or equivalent technologies.
Strong understanding of ITIL-based operational processes and service management frameworks.
Excellent leadership stakeholder management and communication skills.
Preferred Qualifications
Knowledge of TM Forum frameworks including eTOM SID and TAM.
Experience with Kafka Spark Elasticsearch or other distributed data platforms.
Exposure to AIOps machine learning-based monitoring and anomaly detection solutions.
Experience leading multi-vendor programs and large-scale operational transformation initiatives.
Familiarity with regulatory and compliance requirements related to data governance privacy and telecommunications operations.
Why Join
This is an opportunity to play a strategic role in shaping the reliability scalability and operational excellence of complex enterprise environments while working alongside experienced engineering and operations teams on mission-critical initiatives.
Senior Site Reliability Engineer (SRE) Architect Location: Mt. Laurel NJ / Philadelphia PA (Onsite) Job Summary STAFFXPERT LLC is seeking a Senior Site Reliability Engineer (SRE) Architect on behalf of our client in Mt. Laurel NJ / Philadelphia PA. This role is responsible for defining and governing...
Senior Site Reliability Engineer (SRE) Architect
Location: Mt. Laurel NJ / Philadelphia PA (Onsite)
Job Summary
STAFFXPERT LLC is seeking a Senior Site Reliability Engineer (SRE) Architect on behalf of our client in Mt. Laurel NJ / Philadelphia PA. This role is responsible for defining and governing operational architecture across cloud and on-premises environments ensuring high availability resilience observability automation and operational excellence. The ideal candidate will have extensive experience in telecom or large-scale enterprise environments and a strong background in Site Reliability Engineering infrastructure operations and service reliability.
Key Responsibilities
Define and govern operational architecture standards across cloud and on-premises infrastructure environments.
Lead infrastructure architecture reviews capacity planning initiatives network and firewall assessments and middleware operational reviews.
Design and manage SLO SLI and error budget frameworks to improve service reliability and operational performance.
Drive Site Reliability Engineering (SRE) best practices operational maturity programs and reliability improvement initiatives.
Lead disaster recovery (DR) business continuity planning (BCP) resiliency testing incident response processes and post-incident reviews.
Develop standards and governance for Infrastructure as Code (IaC) automation GitOps Terraform Ansible and related operational tooling.
Partner with cross-functional teams to enhance operational efficiency reduce alert fatigue and improve system stability.
Support operational governance lifecycle management security hardening initiatives and infrastructure modernization efforts.
Required Qualifications
10 years of experience in IT Operations Infrastructure Engineering Site Reliability Engineering (SRE) or related disciplines.
Experience working within telecommunications communications or other highly available enterprise environments.
Proven experience supporting carrier-grade or mission-critical platforms with stringent uptime and reliability requirements.
Strong expertise with observability and monitoring tools such as Splunk ELK Stack Dynatrace Grafana or similar platforms.
Hands-on experience designing and implementing SLO SLI and error budget frameworks.
Working knowledge of public cloud platforms including AWS Azure or GCP.
Experience with Infrastructure as Code and automation tools such as Terraform Ansible GitOps or equivalent technologies.
Strong understanding of ITIL-based operational processes and service management frameworks.
Excellent leadership stakeholder management and communication skills.
Preferred Qualifications
Knowledge of TM Forum frameworks including eTOM SID and TAM.
Experience with Kafka Spark Elasticsearch or other distributed data platforms.
Exposure to AIOps machine learning-based monitoring and anomaly detection solutions.
Experience leading multi-vendor programs and large-scale operational transformation initiatives.
Familiarity with regulatory and compliance requirements related to data governance privacy and telecommunications operations.
Why Join
This is an opportunity to play a strategic role in shaping the reliability scalability and operational excellence of complex enterprise environments while working alongside experienced engineering and operations teams on mission-critical initiatives.