Member of Technical Staff
New York City, NY - USA
Job Summary
The role
As a Member of Technical Staff focused on Applied AI youll own our AI stack end to end. One framing we keep coming back to: agents are the primary users of our system of record. Everything we build (the data models the APIs the UI) has to work for a non-human user that operates at scale across every client all the time. Thats a different design constraint than most teams are used to.
What youll build (and own)
A general agent capable of complex multi-step tasks planning sandboxed code execution web search retrieval that powers a do anything experience for advisors.
Ambient agents that act on behalf of clients and advisors: triaging email processing meetings drafting communications surfacing what needs attention before anyone asks.
The agent harness that orchestrates LLMs context tools retrieval and business logic into something coherent and reliable.
Generative UI and human-in-the-loop interfaces where the agent and advisor genuinely collaborate not just take turns.
Evaluation infrastructure that holds two bars at once: high-correctness financial data and subjective judgment-heavy tasks.
Example problems youd work on
These arent hypothetical problems; were actively working on versions of all of these.
Human / AI collaboration that actually works in practice. An advisor is mid-call when the agent surfaces a multi-step recommendation: rebalance adjust the savings rate revisit the estate plan. The advisor takes two steps and overrides the third. Now what How does the agent update its model of what this advisor wants present reasoning the advisor can relay without sounding scripted and learn over time what to do autonomously versus flag This is the flywheel: the tighter the collaboration the more the agent can take on.
Memory systems that know a client the way a great advisor does. A good advisor remembers that a client gets anxious when markets drop cares more about the kids college fund than their own retirement and prefers plain-English summaries. Building memory that captures and evolves this understanding across yearsand surfaces the right context at the right momentis genuinely hard. The challenge is knowing what to retrieve whats still relevant and how to represent a persons relationship to money in a way an agent can use.
Generative UI as an agent architecture problem. When an agent views and updates the advisors screen in real timerendering scenarios adjusting visualizations mid-conversation surfacing recommendations inlinethe UI is constantly changing. The challenge is how the agent manages state across those changes and how you keep the experience from feeling unpredictable. When the visualization shows a clients actual retirement savings the advisor cant be surprised by their own screen.
Evaluations that work for financial services. Most evals are built for tasks with a single right answer. Financial advice isnt like that; the same recommendation can be right for one client and wrong for another and correct often depends on context the eval harness doesnt have. Youll build evaluation infrastructure that can hold two different bars simultaneously: high-accuracy financial data where errors have real consequences and judgment-heavy tasks where the right answer is subjective and the stakes are relational. Add to this the compliance requirements of financial services where auditability isnt optional and infrastructure has to handle large constantly changing datasets where a stale answer can be as harmful as a wrong one.
The kind of person who thrives here
Youre excited by ownership ambiguity and building things that matter.
Youre comfortable where correct isnt always obvious. Financial advice isnt deterministic and neither is evaluating it. Youre energized by probabilistic systems and rigorous about the evaluation infrastructure that tells you whether youre improvingyou trust experimentation more than your first instinct.
Youre ambitious in a way specific to this work. Youre building agents that touch real retirement savings and estate plans. A hallucination here isnt a product bug; its a wrong answer that affects someones financial future. That consequence makes you more careful not slower.
You move fast and you know speed and reliability arent in tension here. An agent that behaves differently in production than in eval is a liability. You treat evaluation and iteration as part of shipping not steps that come after.
You have taste and a high bar for what that means here. You spot AI slop instantly and wont let it through review. You know tech debt vs. shipping is a false tradeoff and act accordinglyleaving systems better instrumented and better understood than you found them.
Youre someone people actually want to be in the room with. This is a 12-person team one office five days a week working on problems that dont have clean answers. The people who thrive here to argue about agent architecture at lunch and then run an eval together in the afternoon. Kind direct and low-ego you can give candid feedback without being an asshole and youre genuinely energized by this environment (not just tolerant of it).
About Company
Nxt Level is a recruiting agency specializing in high-level technical recruiting in engineering, video games, and executive search.