Principal Machine Learning Engineer
New York City, NY - USA
Job Summary
You will own the ML infrastructure that turns research into reliable real-time compliance enforcement systems driving model training evaluation and production serving. You will partner closely with research stakeholders and engineering peers to ship reproducible pipelines and low-latency serving; the role is Hybrid (3 days onsite) in the New York City Metro area.
- Build and own training pipelines: data preparation reproducible fine-tuning runs experiment tracking and release automation
- Build evaluation infrastructure: automated eval runs regression gates dashboards and dataset versioning
- Own model serving in production: low-latency inference batching optimization autoscaling and cost management
- Ship model updates safely with versioning canarying rollback and drift monitoring
- Create repeatable workflows to adapt models to new domains and customer needs
- Turn expert labels and reviewer feedback into clean training and evaluation data
- Set the engineering bar for ML infrastructure as the team grows
- 8 years of software engineering experience including 4 years building infrastructure for ML or LLM systems in production
- Hands-on experience with the modern LLM stack: PyTorch distributed training fine-tuning at scale (e.g. LoRA SFT) and inference engines such as vLLM or TensorRT-LLM
- Experience building eval harnesses regression gates or dataset pipelines; strong understanding of precision recall and calibration
- Proven ownership of production model serving with real latency reliability and cost constraints
- Strong fundamentals in Python containers CI/CD cloud infrastructure and observability
- Ability to scope work ship frequently and make pragmatic build-vs-buy decisions
- Experience collaborating tightly with research partners and defining clear interfaces
- Experience productionizing small or specialized language models
- Experience with structured-output serving or constrained decoding in production
- Prior work in regulated or high-stakes domains (fintech healthcare legal trust and safety)
- Experience deploying models into customer-controlled environments
$200K - $250K/year
This role may fill quickly. Submit your resume to be considered.
Required Experience:
Staff IC
About Company
Hire trusted candidates who BELONG STAY ADVANCE NextDeavor is a recruiting agency helping companies make more strategic hiring decisions. FIND YOUR NEXT GREAT HIRE Using AI technology to make the recruiting process more human AI speeds up, refines, and expands our initial search. This ... View more