Enter a job title or keyword

Lead NLP ML Researcher (LLM Evaluation & Agentic Systems)

Iris.ai


Job Location:

Sofia - Bulgaria

Monthly Salary: Not provided by the employer
Posted: 26 August 2026 (4 days ago)
Application Deadline: 23 November 2026
Vacancies: 1 Vacancy

Job Summary

Why

At were building an agentic AI platform that scales expert-level domain knowledge across entire organizations.

For more than a decade weve worked at the intersection of scientific research industrial data and applied AI helping researchers engineers and business teams reason over complex technical knowledge.

Our products - Neuralith Axion and RSpace - span the full GenAI lifecycle:

  • Data ingestion across text tables figures and technical formats
  • Advanced RAG and indexing pipelines
  • Agentic orchestration and reasoning
  • Rigorous LLM evaluation and governance

What makes us different: we care deeply about accuracy evaluation and responsibility. We dont optimize for demos and proof-of-concepts we optimize for systems that experts trust and use.

The Role

Were looking for a Senior/Lead LLM Researcher to lead research around LLM evaluation model interpretability and mechanistic interpretation of LLMs.

Youll define and drive research into how we evaluate understand and measure the reliability of LLMs - from answer quality and groundedness to uncertainty confidence and model behaviour.

This is a research and leadership role with a strong applied focus. Youll shape research directions lead experimentation and work closely with our engineering and product teams to turn research into products used within the platform.

Youll also participate in shaping projects for EU and national research funding leading and co-authoring grant proposals such as Horizon Europe and EIC.

What Youll Research:

Youll work on a focused set of highimpact research directions that sit at the core of modern applied NLPMLLLM and agentic systems. You will be making LLM-based systems measurable interpretable and trustworthy. The core aspects include:

  • LLM evaluation model grounding (in context following instructions following domain knowledge) faithfulness and task adherence.

  • Mechanistic interpretation - analyzing model behaviour during inference. All our metrics are focused on real-time analysis so they require model analysis during inference not another LLM call.

  • Model interpretability understanding and explaining model behaviour and outputs.

  • Uncertainty & confidence identifying when outputs can be trusted and when they cannot.

  • Evaluation frameworks developing scalable methods for comparing models prompts agents and whole agentic systems

  • Agentic reasoning & control understanding when models should reason stop reasoning or act including inferencetime steering

  • Translation & multilingual NLP evaluation and system design for modern LLMbased translation including lowresource languages

Your goal will be turning rigorous research into capabilities that real users can trust and use.

What Youll Do:
  • Lead research direction around LLM evaluation and interpretability
  • Design novel evaluation methods metrics and experimental frameworks
  • Run rigorous experiments benchmarking and ablation studies
  • Translate research into prototypes and production capabilities
  • Provide scientific guidance to researchers and engineers
  • Collaborate closely with product and engineering teams
  • Publish research and engage with the AI research community
  • Lead and co-author EU and national research grant proposals
  • Write and publish research articles
  • Supervise interns and master thesis students
Our Tech Stack
  • Languages: Python (strong OOP practices)
  • ML: PyTorch Transformers TensorFlow
  • LLMs: Hugging Face OpenAI custom and finetuned models
  • Systems: RAG pipelines Multi-agent frameworks Evaluation tools
  • Infra: AWS Docker HPC
  • Practices: Git CI/CD reproducible research workflows
What Were Looking For:
  • PhD in ML NLP Computer Science or a related field
  • Strong handson experience with R&D grants and proposal writing (e.g. Horizon Europe EIC national or international research funding)
  • 5 years of industry or applied research experience
  • Strong background in NLP (transformers semantic search RAG)
  • Handson experience with LLMs and their evaluation
  • Solid software engineering skills and experience with Python
  • Publications in ML/NLP conferences or journals
  • Able to work within European time zones

Why Join

If you want to do meaningful NLP work help secure funding for frontier AI research and grow in a culture built on trust rigor and fairness lets talk.

Were not your typical tech company. We believe in:

  • Real transparency information is shared context is open and questions are welcome.
  • Fairness designed in policies opportunities and growth are aligned across countries and teams.
  • Ownership and empoweredness to make decisions without micromanagement.
  • Metrics that guide us but they never replace human thinking or responsibility
Compensation & Ownership

Pay

  • Compensation that reflects your value. Our salaries are typically 25% above local market averages ensuring competitive fair pay across regions and roles. And we review it annually.

Equity

(Just imagine: Someone once bought a Tesla option for $1 its worth $400 today.)

Benefits

Weve built our benefits to reflect how we work: with trust fairness and room to grow.

  • 30 days paid vacation
  • 5 additional days paid vacation for Learning and Development
  • Private health insurance (premium coverage) and bi-annual health checks
  • Free MultiSport card for your physical well-being
  • Remote-first & flexible hours work where youre at your best
  • Personal annual learning budget for conferences courses or certifications
  • Personal equipment budget to choose the gear that suits your style
  • Charity and volunteer activities
  • Seasonal working camps (summer & winter) and team retreats
  • Ongoing growth through weekly tech deep dives mentorship pair coding and knowledge-sharing
Lets Build the Future of Responsible AI

If you care about building high-quality ethical AI guided by data and human judgment youll feel at home at .

Apply now or reach out with questions. Were transparent by default.