Skip to main content
RecruoRecruo

UK engineering hire

Hire AI evals engineers in the UK — 5-day shortlist, AISI-aware.

UK enterprise and financial-services buyers now write AISI-aligned evaluation requirements directly into RFPs. We bring pre-validated CEE evals engineers — Inspect AI fluent, bias-testing aware, ICO and FCA-model-risk ready — into UK working hours at 41–50% lower total comp than London-local hires.

IR35 safe

Off-payroll working rules don't bite for genuine B2B CEE contractors.

Every engineer we shortlist operates their own registered business in Poland, Ukraine, Romania, or elsewhere in CEE — multiple clients, their own tools, their own hours. HMRC treats this as outside IR35 by default. For inside-IR35 cases we arrange an Employer-of-Record (Remote, Deel, Oyster) in the engineer's home country so your UK entity has zero payroll exposure.

UK GDPR compliant & EU AI Act aligned

Human-in-the-loop by design, not as an afterthought.

Every shortlist is reviewed and signed off by a human recruiter before reaching you. We sign a DPA with every engagement, store candidate data in EU regions by default, notify candidates up-front that AI is used, retain logs for 5 years, and run quarterly bias audits. The EU AI Act postponement to 2027 doesn't change our posture — we built to the standard, not to the deadline.

22% → typ. 15% agency fee

You save £21–36K on a single senior evals engineer, every year, for the life of the engagement.

London senior evals-engineer base compensation runs £120–150K in 2026, per Harnham UK AI Salary Guide 2026 cross-referenced against three Series C+ placements we monitored this quarter. A classical UK contingency agency bills 22–25% of that on placement; at the conservative 22% end that is £26.4–33K one-time. Recruo bills typically 15%, and our CEE evals candidates sit at £66–88K total-comp-equivalent (see band table). The fee delta alone saves £8–11K on a first-year placement; the total-comp delta saves £40–74K every year afterwards, and evals roles tend to be 18+ month engagements because the evaluation harness a senior engineer builds on month one is the same harness the compliance team audits on month fifteen.

UK salary bands by seniority: local salary, Recruo total comp, Recruo fee, classical agency fee and saving
SeniorityUK local salaryRecruo total compRecruo fee (typ. 15%)Classical 22% feeSaving
Mid (3–5y)£90–115K£52–68K£7.8–10.2K£19.8–25.3K£12–15K
Senior (5–8y)£120–150K£66–88K£9.9–13.2K£26.4–33K£16–20K
Staff / Eval Lead (8y+)£155–190K£88–115K£13.2–17.3K£34.1–41.8K£21–25K

Worked example

What a placement actually costs

Worked example, senior evals engineer into a London-HQ regulated fintech (placed 2026-Q1, FCA-supervised). Target role: UK-local compensation would have been £140K base + 15% overhead = £161K/yr fully loaded. Actual placement: CEE senior evals engineer at €86K/yr on B2B (≈£74K total-comp-equivalent). Recruo fee: 15% of £74K = £11.1K one-time. Classical UK agency alternative, at the conservative 22% end of the 22–25% UK standard: 22% of £140K = £30.8K. Year-1 delta: (£161K + £30.8K) − (£74K + £11.1K) = £106.7K in the client's favour. From year 2 onward the recurring delta holds at roughly £87K because the placement fee does not repeat, and the candidate's evaluation harness keeps compounding in value against the FCA model-risk review cycle.

UK-specific hiring context

Why UK evals-engineer hiring looks different from any other AI role in 2026

The UK is the only major Western market where evaluation is now a named, visible, budgeted function inside AI teams — not a subtask buried under ML engineering. The UK AI Safety Institute (AISI) released Inspect AI as an open-source evaluation framework in 2024 and has continued to expand it through 2025–2026, and the framework is now the default starting point for internal evaluation harnesses at UK AI labs, regulated fintechs, and the top tier of UK enterprise. This matters for hiring: when a UK buyer writes 'evals engineer' in a job spec, they increasingly mean 'someone who can extend Inspect AI, stand up task-specific evaluations, and hand those evaluations to a compliance reviewer intact'. The supply of engineers who have done this end-to-end in production is measured in low hundreds across the UK — we see the same twenty or thirty names rotating between DeepMind, Anthropic London, UK AISI itself, and the safety teams inside two or three large consultancies.

The hiring-geography story is tighter than for general LLM engineering. London dominates because the regulated buyers are there; Cambridge and Oxford concentrate the AI-safety-native candidates because of the university research ecosystem and the cluster of DeepMind-adjacent and ex-OpenAI spinouts that settled around Cambridge in 2024–2025. For UK clients outside that golden triangle — a Bristol scale-up, a Manchester insurer, an Edinburgh fintech — hiring a UK-local senior evals engineer means competing head-on with the same London salary bands and frequently losing. Our UK recruitment agency practice exists specifically to let those firms field a credible evals function without relocating to Zone 1.

Evals hiring also carries an unusual compliance-adjacent profile. Evals engineers are the practical interface between an AI engineering team and three UK oversight regimes at once: ICO AI guidance (bias and explainability), UK GDPR (data-protection-by-design on training and eval data), and, for FCA/PRA-regulated firms, the model-risk management framework published in SS1/23. A competent evals engineer turns each of these from a board-level anxiety into a repeatable test that runs on every model iteration. That is why UK enterprise buyers — especially those going through procurement due diligence with a regulated customer — treat the evals hire as schedule-critical, and why the shortlist quality matters more than for almost any other AI hire. Our EU AI Act and UK GDPR readiness checks are run on every candidate in the evals pipeline as a matter of course.

Time-to-hire

Recruo: 5 business days median to shortlist

UK market median: 78 days to hire (local UK senior AI-evals roles)

Source: Recruo internal (n=6 evals-engineer shortlists, 2025-Q4–2026-Q1); Harnham UK AI Salary Guide 2026; LinkedIn Talent Insights UK AI Safety sub-index (accessed 2026-04-10)

Visa & work-authorisation patterns

  • Polish, Romanian, Czech, Slovak, Bulgarian, Hungarian evals engineers: B2B contractor model, no UK visa required.
  • Ukrainian engineers: UK Homes for Ukraine scheme (extended through 2027) or B2B from Ukraine/Poland with verified Starlink-backed home-office setup.
  • Full-time UK employment via Employer-of-Record (Remote, Deel, Oyster) at 11–15% overhead for regulated-firm clients who want the evals function on UK payroll for SYSC / operational-resilience reasons.
  • Skilled Worker visa route is viable for Head-of-Evals or Eval-Lead hires only — adds 8–12 weeks plus ~£5K legal — we typically steer FCA-regulated clients toward EOR instead because the delivery timeline matters.

Seniority mix

Our UK evals-engineer placements in 2025–2026 skew heavily senior: 14% mid (3–5y), 58% senior (5–8y), 28% staff/lead. London financial services and AI-safety-adjacent enterprise consistently ask for seniors or above because the role sits in front of auditors and regulators, not behind a sprint board.

Remote setup

100% remote-first. UK working hours (9am–6pm BST/GMT) with 90-minute overlap reserved for joint eval-harness review sessions. CEE engineers at UTC+1/+2 deliver 7–8 hours of same-day overlap. Home-office, secondary internet and UK-aligned device posture are verified before shortlisting; for FCA-regulated placements we also verify ISO 27001-aligned personal security practices and, where required, managed-device enrolment.

Reviewed by

Oleh Datskiv

Oleh Datskiv

CEO & Co-founder

Oleh is CEO of Recruo and a 7-year AI engineer — Associate AI Lead at N-iX (2024–2026) leading GenAI/ML R&D prototypes, prior production computer vision and robotics at GlobalLogic and SoftServe. NeurIPS 2020 workshop co-author; MSc in Data Science from Ukrainian Catholic University. Every UK evals-engineer shortlist is screened by him for Inspect AI fluency, UK ICO AI-guidance literacy, and ability to sit opposite an FCA model-risk reviewer without a translator.

FAQ

Frequently asked questions

Inspect AI fluency is a named screening criterion on our evals pipeline, not an afterthought. We test it two ways. First, on the technical-interview task we ask candidates to extend or migrate an existing Inspect AI task — not to write one from a blank file — because production evals work is 90% extension. Second, we ask candidates to critique a deliberately weak Inspect scorer and propose three replacements; this distinguishes candidates who have shipped Inspect in production from candidates who have only read the docs. Approximately 35% of our current CEE evals shortlist has contributed to an Inspect AI repository (public or client) in 2024–2026; the remainder has stood up equivalent harnesses using Promptfoo, LangSmith, or in-house tooling and can migrate inside two weeks.

Yes, with the correct engagement pattern. SS1/23 is principles-based: it requires model-risk ownership, documented validation, and independent challenge — none of which are disturbed by the contractor being CEE-based rather than UK-local. In practice our FCA-regulated clients run one of two patterns. Pattern A: the evals engineer is on a B2B contract but the named senior-management-function (SMF) owner inside the firm retains the model-risk accountability; the evals engineer produces the artefacts, the SMF-holder signs them off. Pattern B: for firms that want the evals function on UK payroll for auditor comfort, we place through an Employer-of-Record (Remote, Deel, Oyster) so the engineer is a UK employee of record while remaining CEE-resident. We flag which of the two patterns a given candidate is compatible with before interview — not every CEE engineer wants the EOR route.

We run a 30-minute structured competency segment before any candidate reaches a client interview. It covers three areas the ICO's AI guidance specifically names: protected-characteristic proxies (candidates have to identify a proxy attack on a credit-risk LLM in under ten minutes), explainability of non-deterministic systems (candidates have to describe an honest explanation artefact for a RAG-over-policy system, not an invented one), and bias-testing methodology at the evaluation layer (candidates have to walk through how a disparate-impact test is implemented inside Inspect AI or an equivalent harness). Approximately 40% of the broader UK evals-candidate pool fails this segment — which is why we run it before the client meets the candidate, not after. Oleh reviews the competency notes personally on every UK evals shortlist.

Median UK evals engagement: senior evals engineer, £74K total-comp-equivalent on a B2B contract (home country: Poland, Romania, or Ukraine), 9am–6pm BST working hours with a reserved 90-minute afternoon window for joint eval-harness review, onsite in London quarterly, 90-day replacement guarantee. We have placed this six times for UK clients in 2025-Q4–2026-Q1 with a 100% 90-day retention rate and a 67% shortlist-to-offer conversion at client interview. Two of the six were FCA-regulated firms and went through the EOR pattern described in the SS1/23 FAQ above; the remaining four went through the straight B2B contractor pattern.

Yes, and we encourage it. Every client interview loop we run for UK evals roles includes an optional 'candidate runs your private eval harness' slot — typically a 90-minute working session where the candidate extends your internal harness in your environment, under observation. This shortens the diligence cycle materially: most UK clients go from offer to start date inside ten working days once the private-eval session has been passed. We sign an NDA covering the private-eval session before access is granted, and candidate machines are not required to retain any artefact after the session closes.

Get started

Get a shortlist of 3–5 vetted candidates in 5 days

We'll scope one open UK ai evals engineers role and deliver a shortlist of 3–5 vetted candidates in 5 business days. Success fee typically 15%, 90-day replacement guarantee, IR35-safe.

Pay on placement

No upfront fee on the Standard plan

90-day guarantee

Free re-search if hire leaves

EU AI Act aligned

Aligned by design

By submitting, you agree to our Privacy Policy. We'll never spam you.