Senior Computational and Data Science Research Specialist 4
San Francisco, CA - USA
Job Summary
Involves a hybrid of multiple computational / data / Cyberinfrastrcture (CI) dominated fields such as but not limited to bioinformatics geological information services (GIS) data analytics and computational chemistry. Applies computational computer science data science and cyber infrastructure (CI) research and development principles with relevant domain science knowledge to perform research and technology integration and development. Responsibilities include research design development analysis operation and support of high performance computing (HPC) and data science research software tools and hardware resources. Develops data algorithms and performs computations statistical analyses interpretation and reporting of research. This specialty / function exists for those positions whose primary responsibility is to do research and use computational and data science technology as a tool to accomplish the research.
The Preclinical Design and Clinical Translation of Regimens for Tuberculosis (PReDiCTR-TB) Consortium is a 21st-century model-informed drug development (MIDD) platform integrating computational science translational pharmacology and global clinical insight to accelerate the design and delivery of transformative TB regimens.
Positioned at the intersection of data biology engineering econometrics and decision science PReDiCTR-TB functions as a strategic intelligence engine for regimen development linking preclinical evidence synthetic experiments mechanistic models and clinical data into a unified predictive framework. By embedding quantitative systems pharmacology (QSP) AI-driven analytics and probabilistic decision modeling throughout development the consortium enables real-time prioritization of regimens with the highest probability of clinical success.
Rather than advancing compounds in isolation PReDiCTR-TB takes a regimen-first translation-driven approach that optimizes combinations dosing strategies and treatment durations through iterative simulation-validation cycles grounded in human-relevant biology. This framework helps reduce reliance on costly empirical experimentation while increasing translational fidelity from bench to bedside.
The consortium delivers:
- Predictive insights that de-risk development earlier
- Optimized trial designs informed by mechanistic and statistical rigor
- Faster and more confident go/no-go decisions across the R&D lifecycle
PReDiCTR-TB is redefining how infectious disease regimens are designed prioritized and translated into clinical impact.
Custom Scope
Model-informed drug development (MIDD) and AI in Clinical Trials market is experiencing a massive shift projected to skyrocket from $1.35 billion to $2.74 billion by 2030. Data and model-driven efficiencies in patient selection drug repurposing and real-time trial monitoring are fundamentally compressing drug development timelines.
At the UCSF Savic Lab we arent just reacting to this shiftwe are driving it. We are moving past traditional isolated data storage models to establish a modern highly interconnected and secure data & model infrastructure.
We are seeking a solution-minded highly technical and mission-driven Distributed Data Pipeline this role you will design build lead implementation and operate a research-grade scale-distributed data architecture. This framework will serve as the foundation for advanced analytics multi-institution translational science and 21st-century accelerated drug development decision support.
Department Overview
The Savic Lab is a global leader in model-informed drug development for infectious diseases and serves as a quantitative innovation hub for translational pharmacology AI-enabled modeling and next-generation regimen design. The laboratory conducts groundbreaking research across TB HIV malaria pediatric infectious diseases translational PK/PD and systems pharmacology. Through its leadership role in PReDiCTR-TB the lab collaborates with international partners to integrate computational science mechanistic modeling and clinical translation into actionable strategies that improve global health outcomes.
UC San Francisco seeks candidates whose experience teaching research or community service has prepared them to contribute to our commitment to diversity and excellence. The University of California is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard to race color religion.
Responsibilities
Applies advanced HPC / data / CI research and development concepts to plan design develop modify debug deploy and evaluate highly complex HPC (software and / or hardware) or data science or computational science or CI software and technologies or combination thereof. Analyzes existing highly complex software scientific codes data science / analytics codes / algorithms and HPC related hardware or works to formulate logic for new and highly complex systems and devises new algorithms. Performs highly complex analysis and tests / debugs highly complex software and hardware. Applies highly complex programming principles. Initiates large and complex research projects with multi-institutional scope in HPC / data / CI areas. May involve collaboration with domain science experts.
Architect the Ecosystem: Design and implement distributed data platforms data fabrics and data mesh capabilities that safely unify complex datasets.
Specifies develops implements and executes highly complex software and hardware research and development plans. Performs or directs highly complex HPC computational and data modeling performance and integration testing. Works with research communities to develop implement and optimize computational and data analysis / analytics software / tools / algorithms / research codes with broad applicability.
Build Resilient Pipelines: Oversee the development of scalable data pipelines optimized for massive batch processing of pre-clinical and clinical data.
Initiates and contributes to HPC / data science / CI research proposals with partners internal and external to the institution in collaboration with other researchers and PIs. May lead a proposal of small to moderate size as a PI.
Enable Cross-Domain Sharing: Create secure cross-institution data frameworks that align with NIH Data Management and Sharing (DMS) policies and FAIR principles.
- Drive Bio-Informatics Decisions: Set the architectural direction for formatting and structuring data that directly feeds pharmacometrics biostatistics and AI/ML pipelines.
Understands and applies advanced research and development practices community standards and department policies and procedures. May serve as technical lead for multiple research and development projects of moderate to broad scope.
Lead and Translate: Convert high-level clinical research data goals into concrete technical roadmaps. Collaborate across software engineering machine learning cyber security clinical pharmacology teams and consortium participants.
- Brief Stakeholders: Author technical documentation lead architectural reviews and deliver executive briefings to internal leadership and external consortium partners.
Qualifications
Required Qualifications
- Experience: 5 years of hands-on experience in data engineering distributed systems or enterprise platform architecture.
- Systems Design: A proven track record of architecting deploying and maintaining production-grade distributed data architectures.
- Data Processing: Direct experience handling large-scale data ingestion multi-tenant databases and event-driven data flows.
- Communication: Exceptional communication and presentation skills with a demonstrated ability to explain complex technical concepts to non-technical executive stakeholders.
- Bachelors degree in Computer / Computational / Data Science or Domain Sciences with computer / computational / data specialization or equivalent experience.
Preferred Qualifications
- Architectures: Data mesh environments data fabric layers and semantic web modeling.
- Streaming & Compute: Apache Kafka Pulsar AWS Kinesis Apache Spark Flink or Apache Beam.
- Storage & Databases: Distributed object storage (S3-compatible API environments) and NoSQL systems (Cassandra DynamoDB MongoDB).
- Clinical Data Standards: Familiarity with NIH-preferred schemas and ontologies (e.g. BIDS for imaging OMOP common data models LOINC SNOMED CT or FHIR transfer protocols).
- Governance & Security: Data cataloging automated data lineage tools and strict data access control frameworks (HIPAA / NIST compliance).
- Masters degree in Computer / Computational / Data Science or Domain Sciences with computer / computational / data specialization preferred.
Required Experience:
Senior IC
About Company
About UCSF The University of California, San Francisco (UCSF) is a leading university dedicated to promoting health worldwide through advanced biomedical research, graduate-level education in the life sciences and health professions, and excellence in patient care. It is the only camp ... View more