Senior Site Reliability Engineer Architect
Mount Laurel, NJ - USA
Job Summary
Location: Mt. Laurel NJ / Philadelphia PA (Onsite)
Job SummarySTAFFXPERT LLC is seeking a Senior Site Reliability Engineer (SRE) Architect on behalf of our client in Mt. Laurel NJ / Philadelphia PA. This role is responsible for defining and governing operational architecture across cloud and on-premises environments ensuring high availability resilience observability automation and operational excellence. The ideal candidate will have extensive experience in telecom or large-scale enterprise environments and a strong background in Site Reliability Engineering infrastructure operations and service reliability.
Key Responsibilities-
Define and govern operational architecture standards across cloud and on-premises infrastructure environments.
-
Lead infrastructure architecture reviews capacity planning initiatives network and firewall assessments and middleware operational reviews.
-
Establish enterprise observability frameworks encompassing metrics logs traces monitoring and alerting strategies.
-
Design and manage SLO SLI and error budget frameworks to improve service reliability and operational performance.
-
Drive Site Reliability Engineering (SRE) best practices operational maturity programs and reliability improvement initiatives.
-
Lead disaster recovery (DR) business continuity planning (BCP) resiliency testing incident response processes and post-incident reviews.
-
Develop standards and governance for Infrastructure as Code (IaC) automation GitOps Terraform Ansible and related operational tooling.
-
Partner with cross-functional teams to enhance operational efficiency reduce alert fatigue and improve system stability.
-
Support operational governance lifecycle management security hardening initiatives and infrastructure modernization efforts.
-
10 years of experience in IT Operations Infrastructure Engineering Site Reliability Engineering (SRE) or related disciplines.
-
Experience working within telecommunications communications or other highly available enterprise environments.
-
Proven experience supporting carrier-grade or mission-critical platforms with stringent uptime and reliability requirements.
-
Strong expertise with observability and monitoring tools such as Splunk ELK Stack Dynatrace Grafana or similar platforms.
-
Hands-on experience designing and implementing SLO SLI and error budget frameworks.
-
Working knowledge of public cloud platforms including AWS Azure or GCP.
-
Experience with Infrastructure as Code and automation tools such as Terraform Ansible GitOps or equivalent technologies.
-
Strong understanding of ITIL-based operational processes and service management frameworks.
-
Excellent leadership stakeholder management and communication skills.
-
Knowledge of TM Forum frameworks including eTOM SID and TAM.
-
Experience with Kafka Spark Elasticsearch or other distributed data platforms.
-
Exposure to AIOps machine learning-based monitoring and anomaly detection solutions.
-
Experience leading multi-vendor programs and large-scale operational transformation initiatives.
-
Familiarity with regulatory and compliance requirements related to data governance privacy and telecommunications operations.
This is an opportunity to play a strategic role in shaping the reliability scalability and operational excellence of complex enterprise environments while working alongside experienced engineering and operations teams on mission-critical initiatives.