Senior Systems and Platform Engineer
Bethesda, MD - USA
Job Summary
Our Partner company Vexterra Group is looking for a highly skilled platform engineer with deep expertise in operating systems hardware GPU and high-speed this role you will design develop and optimize Kubernetes clusters that power enterprise AI for the mission customers. The work location is in Bethesda at the Intelligence Community Campus. Primary Responsibilities Kubernetes Cluster Engineering: Design configure and maintain enterprise Kubernetes platforms. Collaborate with a multidisciplinary team to define and optimize Kubernetes architecture ensuring they meet performance efficiency and feature requirements. Infrastructure as Code: Develop and manage Infrastructure as Code (IaC) using tools such as Terraform Salt Ansible Bash Python or similar frameworks. Collaborate with development teams to design and implement secure automated and repeatable pipelines (e.g. GitLab CI/CD). Troubleshot complex systems issues across cloud network and platform layers. Compliance & Documentation: Maintain technical documentation architectural specifications and Linux best practices. Support ATO (Authority to Operate) and ensure compliance with federal security standards.
Qualifications :
Basic Qualifications
Requires a Bachelors degree and 10 years of relevant experience or Masters degree with 8 years of experience. Additional years of experience may be considered in lieu of a degree
5 years in Platform Engineering or System Engineering experience.
Strong expertise with Linux distributions. (RHEL Ubuntu Oracle Linux and Rocky).
Experience administering Kubernetes clusters including deploying scaling and maintaining containerized workloads.
Hands-on experience creating managing and troubleshooting Docker containers and container images throughout the software development lifecycle.
Experience with Kubernetes cluster management and AI/ML workflow orchestration (Argo Airflow and Kubeflow).
Strong track record with consuming and troubleshooting RESTful APIs for platform integration and automation. Excellent problem-solving skills and the ability to collaborate within a team.
Candidate must at a minimum meet DoD 8570.11- IAT Level II certification requirements (currently Security CE CCNA-Security GICSP GSEC or SSCP along with an appropriate computing environment (CE) certification). An IAT Level III certification would also be acceptable (CASP CCNP Security CISA CISSP GCED GCIH CCSP).
Security Clearance: TS/SCI with CI Poly is required for position or a TS/SCI and willingness to obtain a Poly. Preferred Qualifications
Experience in managing NVIDIA GPU data center platforms. (DGX HGX H200 H100 200 B300 L40S).
Experience with NVIDIA enterprise tools such as Base Command Manager Run:AI Nvidia AI Enterprise.
Knowledge of enterprise server components (storage/network controllers HBA SSDs).
Familiarity with GPU virtualization and cloud computing.
Experience developing and deploying infrastructure in AWS. Knowledge of distributed resource scheduling systems. (Slurm LSF Open MPIetc
Additional Information :
All your information will be kept confidential according to EEO guidelines.
Remote Work :
No
Employment Type :
Full-time
About Company
TechSilo is not just another large IT company where you’ll get lost in the crowd — we’re a small, agile team making big moves in Systems Engineering and IT support for government agencies. With over 30 years of expertise, we deliver creative, cost-effective solutions that make a real ... View more