SRE Platform Engineer
Washington, DC - USA
Job Summary
The Opportunity:
CACI is seeking a seasoned Site Reliability (SRE) Platform Engineer to support the Department of Homeland Security (DHS) Office of the Inspector General (OIG). This role offers a unique opportunity to ensure the reliability performance and availability of cutting-edge AI and data analytics platforms that strengthen national security oversight through investigative audit and inspection operations.
As an SRE Platform Engineer you will be the operational guardian of mission-critical systems including OIG Chata large language model assistantenterprise data platforms and AI-powered applications that federal oversight professionals depend on daily. You will monitor production environments implement comprehensive alerting and observability troubleshoot complex incidents optimize performance and costs and ensure high availability through proactive capacity planning and reliability engineering. From analyzing performance metrics to identify bottlenecks to supporting disaster recovery and operational readiness you will apply site reliability engineering principles to maintain service excellence. Working within secure Azure Government environments you will build operational resilience for a transformational program of national importance. Join us to make a meaningful impact by ensuring mission-critical AI and data capabilities are always available performant and reliable.
Responsibilities:
Monitor maintain and support production and non-production environments for OIGChat AI applications and enterprise data platforms to ensure availability performance service health and adherence to service level objectives (SLOs)
Implement comprehensive observability including alerting dashboards health checks synthetic monitoring and log analysis using Azure Monitor Application Insights Log Analytics or equivalent tools to enable proactive incident detection
Lead incident response and troubleshooting efforts including root cause analysis defect resolution dependency updates integration validation and coordination of emergency changes to restore service rapidly
Analyze performance and usage metrics across applications APIs AI model endpoints data pipelines and infrastructure to identify and remediate bottlenecks latency issues resource constraints and efficiency opportunities
Support capacity planning resource sizing autoscaling configuration and cost optimization for compute storage and AI model consumption to balance performance requirements with fiscal responsibility
Implement and maintain backup/restore processes disaster recovery procedures high availability architectures and business continuity capabilities to ensure data protection and operational resilience
Monitor data platform availability including Azure Databricks clusters data pipelines storage services and analytical workloads with alerting for pipeline failures job errors and performance degradation
Support deployment and operational readiness for pilot applications and new capabilities including pre-production validation performance testing runbook development and go-live coordination
Provide surge support for complex technical issues large-scale data collection analysis analytical environment optimization and specialized troubleshooting requiring deep platform knowledge
Develop and maintain operational documentation including runbooks troubleshooting guides architecture diagrams incident post-mortems and knowledge transfer materials to support sustainable operations
Qualifications:
Required
Bachelors degree 15 years of experience in site reliability engineering DevOps platform engineering systems administration or related field; equivalencies considered (Masters 12 years; 21 years with no degree; AA 17 years)
Must be able to obtain a Active DHS/ EOD Clearance as required.
Extensive experience with Site Reliability Engineering (SRE) principles including monitoring observability incident response capacity planning performance optimization and reliability engineering practices
Proven expertise with Azure cloud services including compute storage networking monitoring and platform-as-a-service (PaaS) offerings with deep understanding of operational best practices
Strong experience with monitoring and observability tools (Azure Monitor Application Insights Grafana Prometheus ELK stack) and implementing alerting dashboards and log aggregation
Demonstrated ability to troubleshoot complex technical issues across application platform and infrastructure layers with strong analytical and problem-solving skills
Desired
Experience operating AI/ML platforms large language model services (Azure OpenAI) data analytics platforms (Databricks Synapse) or high-scale cloud applications in production environments
Hands-on experience with Azure Government or other secure government cloud environments (AWS GovCloud) with understanding of compliance monitoring security operations and federal operational requirements
Background in federal government mission-critical systems or 24/7 operational environments with experience supporting incident response change management and operational excellence programs
What You Can Expect:
A culture of integrity.
At CACI we place character and innovation at the center of everything we do. As a valued team member youll be part of a high-performing group dedicated to our customers missions and driven by a higher purpose to ensure the safety of our nation.
An environment of trust.
CACI values the unique contributions that every employee brings to our company and our customers - every day. Youll have the autonomy to take the time you need through a unique flexible time off benefit and have access to robust learning resources to make your ambitions a reality.
A focus on continuous growth.
Together we will advance our nations most critical missions build on our lengthy track record of business success and find opportunities to break new ground in your career and in our legacy.
Pay Range:
There are a host of factors that can influence final salary including but not limited to geographic location Federal Government contract labor categories and contract wage rates relevant prior work experience specific skills and competencies education and certifications. Our employees value the flexibility at CACI that allows them to balance quality work and their personal lives. We offer competitive compensation benefits and learning and development opportunities. Our broad and competitive mix of benefits options is designed to support and protect employees and their families. At CACI you will receive comprehensive benefits such as; healthcare wellness financial retirement family support continuing education and time off benefits.
Since this position can be worked in more than one location the range shown is the national average for the position.
The proposed salary range for this position is:
$114600-$252100CACI is anEqualOpportunity Employer. All qualified applicants will receive consideration for employment without regard to race color religion sex pregnancy sexual orientation age national origin disability status as a protected veteran or any otherprotectedcharacteristic.
Required Experience:
IC
About Company
CACI employs a diverse range of talent to create an environment that fuels innovation and fosters continuous improvement and success. At CACI, you will have the opportunity to make an immediate impact by providing information solutions and services in support of national security miss ... View more