Principal Software Engineer GATK
Job Summary
About the role
GATK HTSJDK and Picard are foundational infrastructure for the global genomics community with tens of millions of downloads tens of thousands of scientific citations and a documented footprint across hundreds of clinical trials spanning cancer ALS and other diseases.
The Broad Institute is resuming dedicated development and active maintenance of all three packages and were hiring an experienced engineer to serve as a primary maintainer. Youll take broad ownership of the work needed to clear the file-format standards backlog restore community confidence and build a sustainable decentralized contributor model amplified by modern AI coding tools and an empowered community of contributors. This is a high-autonomy role with significant latitude to shape priorities carrying ownership that was previously spread across a larger team.
The impact is real and immediate. Numerous computational groups across the Institute depend on this toolset to meet their commitments. Youll collaborate with both internal and external partners (such as those at the Wellcome Sanger Institute EBI/ENA Fulcrum and across GA4GH).
What youll do
The exact mix will flex with community and Institute needs but the role generally spans:
Standards & development
- Advance file-format standards support across the toolset (for example VCF 4.4) keeping the Java/JVM genomics ecosystem aligned with ratified GA4GH standards as they evolve.
- Author bug fixes and high-impact features across the repositories upholding the engineering quality and test coverage that have kept these tools relevant for over a decade.
Release engineering & maintenance
- Drive a regular release cadence across the repositories ensuring updates reach downstream channels.
- Triage issues and review pull requests working down a substantial backlog and prioritizing high-impact internal and community needs.
Community governance & communication
- Help establish and run a Trusted Developer Program recruiting and onboarding external contributors administering a tiered access model and maintaining governance documentation.
- Support the user community including monitoring and responding on the GATK forums.
- Represent the project at conferences and community meetings (e.g. GA4GH) and help align priorities across collaborators and Broad programs.
Tooling & acceleration
- Apply state-of-the-art AI/ML coding tools to accelerate review and development across large complex multi-language codebases.
- Over time explore LLM-based tooling to help answer recurring community questions from existing documentation.
Early focus areas
Sequencing and timelines will be set with the hiring manager but early work is likely to center on:
- Advancing priority standards work such as VCF 4.4.
- Standing up the Trusted Developer Program and re-establishing a regular release cadence.
- Actively updating the source repositories to modern AI-friendly standards.
What were looking for
- Excellentwritten and verbal communicationand a collaborative community-facing disposition.
- Substantial professional software engineering experience building and maintaining production-grade systems.
- Typically requires 10 years of related experience with Bachelors degree; or 8 years with Masters degree; or a PhD with 5 years experience; or equivalent experience.
- Proficiency inJava/JVM languages or a demonstrated track record of becoming rapidly productive in a large mature multi-language codebase.
- StrongPythonskills for tooling analysis and pipeline work.
- Experience writing testing and deploying bioinformatics workflows in WDL or NextFlow.
- Deep working knowledge ofgenomics/bioinformatics:sequencing data and core file formats (BAM/CRAM/VCF/BCF) variant representation and large-callset processing.
- Demonstrated ownership ofopen-source or shared infrastructurecode review release management and sustaining quality in a codebase many others depend on.
- Experience designing or operatinglarge-scale cloud data pipelines(GCP AWS or Azure).
- Comfort adoptingAI-assisted developmentworkflows.
- Interest in learningnew technologies and programming languages such asRust.
Strongly preferred
- Prior hands-on experience with theGATK ecosystem(GATK / HTSJDK / Picard) or other Broad genomics infrastructure (GenomicsDB Terra GCS-based workflows).
- Familiarity withgenomic file-format internals and compression().
- MLOpsbackground and ML/AI pipeline tooling experience.
- Experience inregulated/GxPclinical or production environments.
- Engagement withGA4GHor other genomics standards bodies.
Tech environment
Java/JVM Python GitHub-based open-source workflow Docker GCP GA4GH file-format standards AI-assisted development tooling
Required Experience:
Staff IC
About Company
Broad Institute is a multidisciplinary community of researchers on a mission to improve human health.