Staff Software Engineer, Search Systems Code Data
San Francisco, CA - USA
Department:
Job Summary
Mercors mission is to organize human intelligence to power the AI economy. Were a leading AI data company building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercors APEX benchmark family measures AIs real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious fast-paced and deeply committed team. Youll work alongside researchers operators and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco NYC or London offices.
Some of the most valuable work on Mercors platform is codethe tasks problems and solutions that train and evaluate the worlds frontier coding models. As a Staff Software Engineer for Code Search & Retrieval youll own the architecture and algorithms behind how we search across code tasks: finding similar tasks routing them to the right models and turning natural-language questions into precise retrieval.
Searching over code is a genuinely hard problem. Identifying similar code tasksnot just similar textrequires retrieval that understands structure semantics intent and difficulty far beyond what generic embeddings capture. Youll own that end to end: the vision is to search across code tasks at scale identify and select the correct code-specific models for a given task and translate NLP questions into search queries that return the right code and tasks.
This is a hands-on technical leadership role not a management role. Youll design and build the retrieval systems that combine dense (code) embeddings and lexical (BM25) signals and make the tradeoffs that let search stay fast and affordable at scale. Because Mercor operates at the frontier of data and models youll also own the hardest part of the problem: continuously evolving embeddings models and search quality as newer state-of-the-art code models arrivewithout regressing what already works.
As a Staff engineer your impact extends well beyond your own commits. Youll set the technical direction the rest of the org builds on mentor and grow the engineers around you and raise the bar for how we build.
Own the architecture of Mercors code search and retrieval systems end to endhybrid retrieval combining dense code embeddings and BM25 candidate generation ranking and re-ranking over code tasks.
Solve the hard problem of identifying similar code tasksretrieval that captures code structure semantics intent and difficulty rather than surface text.
Build the system that identifies and selects the correct code-specific models for a given task and routes tasks to the right model.
Design natural-language-to-query translation that turns NLP questions into precise search over code and tasks.
Design and operate the indexing pipeline so the task index stays fresh and consistent as new tasks solutions and results arrive continuouslybalancing incremental updates full rebuilds and real-time ingestion.
Make the cost-and-speed tradeoffs that keep search fast and economical at scale: embedding dimensionality and quantization ANN index choice and parameters caching sharding and serving infrastructure.
Build the systems and evaluation harnesses that let us continuously evolve embeddings models and search qualitysafely swapping in new code models re-embedding corpora and A/B testing relevance as SOTA advances.
Define and drive the long-term technical strategy for code retrieval across the organization and lead the highest-stakes design reviews.
Establish evaluation metrics offline/online testing and quality guardrails so search improvements are measurable and regressions are caught before they ship.
Stay deeply hands-on: prototype critical systems ship production code and unblock teams on their hardest retrieval and infrastructure problems.
Mentor and grow engineersjunior and seniorthrough design reviews pairing and clear technical writing raising the technical bar across the org.
Partner with product researchers and engineering leadership on build-vs-buy decisions platform investments and technical hiring.
8 years of professional software engineering experience including 3 years operating at a Senior level or above with a Staff-level track record of org-wide technical impact.
Deep hands-on expertise building search and retrieval systems: dense-embedding retrieval lexical scoring (BM25/TF-IDF) hybrid ranking and re-ranking.
Good to have but not required: Experience with code search or code understandingretrieval over code code embeddings or working with code-specific modelsand an appreciation for why matching similar code tasks is harder than matching text.
Strong understanding of the search algorithms and index internalsvector/ANN indices (e.g. HNSW IVF product quantization) inverted indices and engines such as Elasticsearch/OpenSearch Lucene FAISS or vector databases.
A track record of making the right cost-vs-speed tradeoffs: latency budgets throughput memory footprint and infrastructure spend on high-QPS systems.
Familiarity translating natural-language questions into structured search queries (query understanding semantic parsing or LLM-assisted query generation).
Excellent systems fundamentals: distributed systems data modeling and API design at scale.
Demonstrated technical leadership and mentorshipyouve helped junior and senior engineers grow and level up an engineering team.
Genuine excitement for agentic development and new technology fluency with modern AI dev tools (e.g. Claude Code Cursor Copilot) and a deep passion for writing great code.
Excellent communicationable to make complex tradeoffs legible to both engineers and leadership. Strong opinions loosely held. High ownership pragmatism and a bias toward shipping.
Experience training or fine-tuning code embedding models or code-specific LLMs.
Experience with learning-to-rank semantic search or recommendation systems in production.
Familiarity with LLM-based retrieval RAG patterns and model routing/selection.
Background operating latency-critical services on modern cloud and orchestration infrastructure.
Impact: Own the code search systems that decide how well Mercor finds routes and evaluates the code tasks training the worlds frontier models at a company scaling faster than almost any in its category.
Ownership: Org-wide technical scope with a direct line to engineering leadership and real authority over technical direction.
Learning: Work alongside world-class engineers product leaders and AI researchers building at the frontier of AI.
Growth: Shape the engineering organization itselfits standards its architecture and its next generation of technical leaders.
Generous equity grant vested over 4 years
Up to $15K relocation bonus (if moving to the Bay Area)
A $10K housing bonus (if you live within 0.5 miles of our office)
A $1.5K monthly stipend for meals
Free Equinox membership
Health insurance
Mercor is an equal opportunity employer. We work in-person five days a week in our San Francisco office.
Required Experience:
Staff IC