SeniorStaff Applied AI Engineer, Agent Harness
San Francisco, CA - USA
Job Summary
Confidential client a Pre-seed startup (1-10 employees) building AI coworkers for IT teams: a security-focused product where an agent registers as a governed identity in a customers directory requests scoped access per task escalates for human approval and drives systems it was never given an API for. Raised a $6M seed round backed by a strong bench of operator-angels from across IT and security. Several design partners today and a founding team of three. Full-Time In-person San Francisco CA. Experience: 4 years. Salary: $200000-$300000/yr. Visa sponsorship: H-1B O-1 OPT.
About the Role: This role builds the layer that turns model capability into systems that actually work for users. Youll develop the core agent harness (the execution loop tool-use strategies context construction and model-facing experimentation) and iterate on agent behaviors across real customer workflows and long-horizon tasks. A defining stance: the creative step happens once at authoring and what executes afterward is deterministic compiled type-checked code rather than stochastic tool-chaining. Youll own that boundary between what the model decides at runtime and what ships as code. Youll build and run evals against replicas of real customer environments rather than synthetic tasks reliable enough to gate a release; analyze production failures and attribute them to the right layer (model prompt tool contract environment state or retry logic); and extend the computer-use agent to systems with no API owning the guarantees that let it run in production at all: a fresh microVM per task credentials injected at session start and never seen by the model every action recorded any session killable mid-run.
What Youll Own:
- The core agent harness: execution loop tool-use strategies context construction and model-facing experimentation
- Agent behaviors across real customer workflows and long-horizon tasks
- The boundary between runtime model decisions and deterministic compiled type-checked execution
- Evals against replicas of real customer environments reliable enough to gate releases
- Production failure analysis and systematic robustness improvements attributed by layer
- Extending computer-use to API-less systems with production guarantees (per-task microVM model-blind credentials full action recording killable sessions)
- Feedback loops and data systems that get better real-task data into eval and training
Company stage: Pre-seed. Work type: In-person (San Francisco CA).