Member of Technical Staff Research Software Engineer Safety Evaluations Infrastructure
New York City, NY - USA
Department:
Job Summary
Reflection is a research lab making intelligence open and accessible for everyone to use customize and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.
Reflection is a research lab making intelligence open and accessible for everyone to use customize and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.
As a Research Software Engineer on the Safety team you will design build and own the infrastructure used to run our most sensitive model evaluations including evaluations in CBRN (chemical biological radiological and nuclear) child safety and other dangerous-capability domains. These evaluations inform release decisions for our open models so the systems you build must be secure isolated reproducible and trustworthy under scrutiny.
This is a deeply technical high-ownership role at the intersection of platform engineering security and safety research. You will partner closely with domain experts legal and safety researchers to turn their evaluation needs into robust scalable infrastructure: sandboxed execution environments controlled data pipelines for sensitive material access controls audit logging and the tooling that lets researchers safely elicit and measure model capabilities in high-consequence areas.
Design and build secure sandboxed infrastructure for running sensitive model evaluations including CBRN and other dangerous-capability domains.
Build controlled data pipelines and storage for sensitive evaluation material applying least-privilege and need-to-know access role-based access control (RBAC) encryption at rest and in transit audit logging and data-minimization safeguards.
Partner with safety researchers and domain experts to translate evaluation designs into reliable reproducible and scalable systems.
Build eval-orchestration tooling and harnesses that let researchers run high-throughput evaluations against models and agents in isolated environments.
Develop infrastructure for measuring AI capability uplift in high-consequence domains and integrate results into the pipelines that inform release decisions.
Implement guardrails monitoring and compartmentalization so sensitive work stays appropriately siloed applying least-privilege need-to-know and defense-in-depth principles across compute data and tooling.
Write production-quality Python (and related tooling) for high-throughput data processing and evaluation systems.
Improve the reliability security posture and developer experience of the safety teams evaluation platform over time.
Strong software engineering skills particularly in Python with a track record of building reliable scalable infrastructure or platform systems.
Experience building sandboxed isolated or otherwise security-sensitive execution environments (e.g. containerization VM isolation secure compute) for Trust and Safety teams.
Solid grounding in security engineering fundamentals: principle of least privilege need-to-know access role-based access control (RBAC) secrets management encryption audit logging compartmentalization and defense-in-depth design.
Experience building data pipelines and handling sensitive or restricted data with appropriate safeguards.
Ability to own entire problems end-to-end including ambiguous cross-functional ones.
Comfort working on sensitive projects that require discretion integrity and sound judgment.
Thrive in a fast-paced high-agency startup environment with a bias toward action.
Experience building evaluation benchmarking or experimentation infrastructure for ML systems.
Experience working with LLMs agents or ML training/inference pipelines.
Familiarity with dangerous-capability or dual-use domains (CBRN cyber etc.) and the information-security considerations they involve.
Familiarity with compliance frameworks relevant to sensitive data handling.
We encourage you to apply even if you dont meet every qualification. Not all strong candidates will match every item listed.
We believe that to make intelligence open and accessible to all you need to start at the foundation. Joining Reflection means building from the ground up as part of a talent-dense team. You will help define our future as a company and help define the future of open foundational models.
We want you to do the most impactful work of your career with the confidence that you and the people you care about most are supported.
Top-tier compensation: Salary and equity structured to recognize and retain our talent globally.
Stock options: Everyone who joins and contributes to Reflections success gets to share in the upside through stock options.
Health & wellness: Comprehensive medical dental vision and life with an annual wellness allowance.
Meals: Lunch and dinner are provided in the office daily.
Life & family: 22 weeks paid parental leave for all new birthing and non-birthing parents including adoptive and surrogate journeys.
Vacation days: Unlimited paid time off in the U.S. and 30 days in the U.K.
Sponsorship support: We sponsor visas to help exceptional talent join our team and support long-term immigration pathways where applicable.
Team building: We have regular off-sites happy hours and team celebrations.
Export Control Notice: This position may require access to technology or source code subject to the U.S. Export Administration Regulations. Any offer of employment for this role may be conditioned on the Companys ability to provide the candidate with access to such technology or source code in compliance with applicable U.S. export control laws which may require the Company to seek government authorization.
Required Experience:
Staff IC