Senior Support Engineer (GovOps)
McLean, MD - USA
Job Summary
P-298
About Databricks
At Databricks we are passionate about enabling data teams to solve the worlds toughest problems from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the worlds best data and AI infrastructure platform so our customers can use deep data insights to improve their business. Founded by engineers and customer obsessed we leap at every opportunity to solve technical challenges from designing next-gen UI/UX for interfacing with data to scaling our services and infrastructure across millions of virtual machines. And were only getting started.
The Role
s enterprise data enclaves and Generative AI applications become mission-critical to national security and public sector modernization maintaining the reliability and operational integrity of the underlying platform is paramount. Databricks is seeking a Senior Support Engineer (GovOps) to join our Cleared Engineering team in McLean VA. Sitting at the vital intersection of core infrastructure operations and active platform reliability this team serves as the ultimate technical authority for our platform across GovCloud FedRAMP High and air-gapped environments.
In this role you will move beyond basic ticket resolution to handle the hardest technical escalations across our platform domainsincluding Compute Networking Storage IAM and OS internals. You will partner directly with Core Engineering conduct deep live troubleshooting safely execute automation tools and isolate complex system boundary blocks inside secure enclaves. If you are a systems specialist who thrives under pressure enjoys writing custom tooling to unblock platform bottlenecks and wants to safeguard the infrastructure powering next-generation federal AI workloads this opportunity provides unmatched scope and impact.
The Impact Youll Have
- The Impact Youll HaveDrive High-Side Platform Reliability: Serve as the technical authority for resolving the most complex platform escalations across Compute Networking Storage and OS domains within air-gapped and GovCloud enclaves.
- Partner with Core Engineering: Triage system defects perform root-cause analysis and collaborate directly with product engineering teams to drive permanent software bug fixes and fleet-wide supportability enhancements.
- Execute Advanced Automation & Tooling: Write and adapt custom Python Bash or Go scripts and diagnostic tools to automate platform triage and streamline operational recovery inside restricted boundaries.
- Manage Critical Incidents: Lead technical triage and live incident response during high-consequence platform events maintaining composure and clear stakeholder communications under intense pressure.
- Optimize Secure Environment Pipelines: Identify systemic environment blocks diagnose patch deployment failures and review Infrastructure-as-Code (IaC) configurations to ensure seamless platform operations.
- Uphold Federal Guardrails:Navigate evolving secure facility (SCIF) operational frameworks ensuring all platform interventions strictly align with FedRAMP High ITAR/IL5 and federal security compliance standards.
What Were Looking For
- Minimum Qualifications
- Active TS/SCI security clearance.
- Bachelors degree in Computer Science Systems Engineering Mathematics or equivalent practical experience.
- 5 years of experience in Systems Engineering Cloud Infrastructure Senior Support Engineering DevSecOps or a related technical operations role.
- Deep expertise in Linux/Unix operating-system internals networking process/resource debugging and distributed systems.
- Hands-on proficiency in Python Bash Go or Java with the ability to read source code debug software defects and write custom diagnostic scripts.
- Proven experience managing debugging and operating cloud infrastructure across public or federal cloud boundaries (AWS GovCloud Azure Government or GCP).
- Solid working knowledge of Infrastructure-as-Code (IaC) tools such as Terraform or CloudFormation along with container orchestration mechanics (Docker/Kubernetes).
- 3 years of experience writing clear succinct written and verbal technical communications including technical status reports incident summaries or cross-team escalations.
- Ability to work on-site out of our McLean VA secure facility (SCIF) including participation in an on-call rotation for critical incidents.
- Preferred Qualifications
- Active Counterintelligence (CI) or Full Scope Polygraph (FSP).
- Prior experience supporting AWS World Wide Public Sector (WWPS) isolated defense integration networks or FedRAMP High / ITAR environments.
- Hands-on experience operating inside low-egress or fully air-gapped networks where standard external diagnostics are restricted.
Required Experience:
Senior IC
About Company
The Databricks Platform is the world’s first data intelligence platform powered by generative AI. Infuse AI into every facet of your business.