Enter a job title or keyword

Technical Lead, Senior AI Engineer VonHalsky (mfn)

InPost


Job Location:

Kraków - Poland

Monthly Salary: Not provided by the employer
Posted: 28 August 2026 (4 days ago)
Application Deadline: 25 November 2026
Vacancies: 1 Vacancy

Job Summary

Why this role exists

Von Halsky is InPosts conversational AI shopping assistant live in production and serving a growing share of our customers. What decides whether it wins is not the model but whether we can tell at release cadence that a change made conversations better. That is the Evaluations Platform and we are hiring the engineer who takes it to the next level.

What you will own

  • The LLM-judge pipeline. Our release-gating judge over real and golden conversations calibrated well enough that people act on its verdicts instead of arguing about it.
  • Eval datasets and the golden set. Real coverage across intents and categories including the Polish-language coverage generic benchmarks do not give us.
  • Root-cause analysis on real conversations. Making our conversation-mining stack diagnostic rather than descriptive on a stable issue taxonomy.
  • The sus detector. Abusive adversarial and anomalous sessions next to our guardrails and red-team work.
  • Evals-driven development. The eval comes before the feature and writing it is as cheap as writing the code.
  • The interface to product and business. Vague asks in measurable quality definitions out and results stakeholders can act on.

What leadership aspirations means here concretely

This is a technical lead role not a people-management role and you will not carry line-management duties on day one. You will set and defend the technical direction for the platform act as reviewer of record for the area mentor other engineers scope work with our PM and EM present results to stakeholders and hold the line against ad-hoc requests crowding out platform work.

Engineering management later or a Staff-level hands-on track are both paths we will build with you. Either way we need someone accountable for an area rather than for a ticket.

How we define success in this role

  • Judge pass rate is calibrated against human labels and is a number people trust and cite.
  • A stable versioned issue taxonomy is live and week-over-week trends are comparable.
  • Every release is gated by an eval run the team can reproduce.
  • The golden set has documented coverage and named blind spots.
  • At least two engineers besides you can operate and extend the platform.

Qualifications :

What we are looking for

Required

  • 5 years building and running production software with 2 years on LLM-based systems that real users hit.
  • Strong engineering fundamentals plus the habits that go with production ownership: testing CI/CD containers observability and working in cloud. We work primarily in Python.
  • Demonstrable experience evaluating generative systems not only building them: LLM-as-judge human-label calibration inter-annotator agreement regression suites offline-versus-online divergence. You should have opinions about what makes an eval worthless.
  • AI engineering fundamentals. Prompt and context engineering as a discipline (context-window budgeting structured outputs failure-mode taxonomies); agentic primitives in production (tool use multi-turn state MCP agent-to-agent integration patterns); and eval and LLM-observability tooling (LangFuse Braintrust Weave or equivalent including things you built yourself).
  • Comfort with data at scale: SQL working with a lake or warehouse and building a metric someone else can reproduce.
  • Fluency with AI-assisted development tooling (Claude Code Cursor Copilot). We use it daily and expect it.
  • Ability to make a technical argument to a non-technical audience and be understood.
  • English B2 and Polish. Our users converse in Polish and you will read their conversations; judging quality you cannot read is not possible.

Nice to have

  • Harness and loop engineering. Building the scaffolding around models rather than only calling them: agent loops retries and fallbacks tool-call orchestration deterministic replay and the plumbing that makes a non-deterministic system testable.
  • Auto-improving systems. Closing the loop from production signal back into the product: mining failures into cases using eval results to drive prompt retrieval and routing changes and automating the parts of that cycle that people do by hand today.
  • Adversarial robustness jailbreak testing red-teaming or abuse and fraud detection.
  • E-commerce search or recommendation domain experience.

Additional Information :

What we offer

  • A product that has already been released to millions of users with a real business case not a lab pilot.
  • Direct access to frontier models at committed capacity across multiple providers plus an open-source track we run ourselves.
  • A quality mandate with executive attention.
  • Ownership of a platform that is greenfield in practice inside a company with production traffic which is the rarest combination in this market.
  • Hybrid working from Warsaw or Kraków in a team growing fast enough that early hires shape how it works.
  • Fulfilling careers with a range of benefits for people and investing in providing training opportunities for their development. 
  • You will feel a part of the InPost community that makes an impact on sustainability convenient deliveries and the circular economy every day. 
  • Excellent working environment and flexible hours
  • We offer B2B type of contract

Remote Work :

No


Employment Type :

Contract


About Company

Company Logo

InPost S.A. – a company listed on Euronext Amsterdam - is the leading out-of-home e-commerce enablement platform in Europe. InPost Group operates across key geographies: Poland, the UK, Italy and the Iberian Peninsula, as well as France and Benelux, through its subsidiary, Mondial Rel ... View more

View Profile View Profile