Member of Technical Staff | ML Systems
Department:
Job Summary
At Avra every technical IC is a Member of Technical Staff (MTS). The title doesnt put anyone in a silo: you own systems and outcomes not steps in a function and you keep building depth in your area. Seniority shows up in your scope level and compensation not in titles.
In this role youll join our ML Systems team which owns Avras ML core and the governance of every model we ship. Research produces candidate models and evidence; you build the reliable path from data and training to a governed reproducible release that can run in our cloud or in any customer environment. ML Systems is an internal platform: its users are our researchers and platform engineers and its success is measured by the leverage it creates for them.
Build CUDA kernels and compute primitives for training and serving graph neural networks (GNNs).
Evolve Monad our sampler and distributed-training library including neighbor sampling and training performance.
Specify our binary data formats (Lance Arrow CSR/CSC) and own materializations and feature backfills for training and evaluation.
Define data contracts and consumption requirements with the teams that build our customer and proprietary datasets.
Build and operate experiment tracking checkpoints and evaluation infrastructure with reproducibility by default.
Own the model registry lineage versioning and compatibility across models embeddings and downstream models.
Define and run release gates so every model running in production batch or on-premise maps to a governed release.
Make it possible to audit exactly which data code configuration and evidence produced each release.
Time-to-experiment: how quickly a researcher goes from a hypothesis to materialized data compute and tracking.
Time-to-governed-release: how quickly a validated candidate becomes an authorized release.
Training throughput per GPU on our foundation model training runs.
100% of production models with complete release records and lineage no ad hoc models in any environment.
Every release reproducible from its registered data code and configuration.
Strong systems engineering skills and production-quality Python.
Experience with distributed training (e.g. Ray PyTorch distributed) and multi-node GPU workloads.
Experience with columnar data formats and large-scale data materialization.
Familiarity with ML lifecycle tooling: experiment tracking model registries evaluation and reproducibility.
A product mindset: you treat an internal platform as a product with real users. You dont need to be a data scientist.
CUDA kernel development or GPU performance optimization.
Graph neural networks or graph sampling at scale.
Lance Arrow or other columnar/indexed storage formats.
Multi-cloud GPU compute (e.g. SkyPilot).
Model governance or audit requirements in financial services or other regulated environments.
Required Experience:
Staff IC
About Company
Our foundation model helps our clients bring the right SME to the top of the funnel, hyper-personalize offers, and reduce default. Request a demo.