Skip to main content
RecruoRecruo

Sample Scorecard

One scorecard per candidate. Every shortlist.

A real example of the AI hard-skills report we deliver as part of every shortlist. Structured evaluation, per-question scores, English assessment, and a clear recommendation — reviewed by a human recruiter before delivery.

Interview Details

Oleksandr K.Senior ML/LLM Engineer

ID: 6f842308-4290-4433-ade2-9be***

Export PDFExport PDF (No Scores)

Candidate Information

Name: Oleksandr K.

Email: o.k***@gmail.com

Status: completed

Created At: 3/3/2026, 9:53 PM

Position Details

Role: Senior ML/LLM Engineer

Field: AI & Machine Learning

Type: With Manager

Final Recommendation

87/100
PROCEED

Candidate demonstrates strong applied LLM engineering depth across RAG architecture, eval design, and retrieval tuning with an average score of 8.6/10. Shows excellent judgement on fine-tune vs prompt engineering decisions and a systematic approach to hallucination and regression testing. Minor gaps include limited hands-on distributed fine-tuning experience and could quote more concrete eval metrics from past projects. Current salary $70K, targeting $75-85K — within range for a senior ML engineer in CEE.

English Assessment

Communication & clarity signal

C1(8/10)

A1

A2

B1

B2

C1

C2

Key takeaway

Communicates technical concepts clearly and fluently. Can lead architecture discussions and explain complex tradeoffs to both technical and non-technical stakeholders. Minor grammatical imprecisions that do not impede understanding.

Question-by-Question Evaluation

Overall Performance Summary

The candidate demonstrated strong applied knowledge across all LLM engineering topics with consistent performance scores of 8-9. Technical understanding was evident throughout, with particularly strong answers on RAG architecture and eval design.

General Feedback

The candidate shows excellent practical knowledge of retrieval pipelines, eval methodology, fine-tuning decisions, and production LLM monitoring. They consistently identified key tradeoffs and appropriate solutions across all areas. Their responses demonstrated real-world experience shipping LLM features to production.

Strengths

Demonstrates comprehensive practical knowledge of RAG architecture, eval design, and retrieval tuning with strong awareness of production-readiness concerns like regression suites, canary rollouts, and cost-latency budgets.

Areas of Improvement

Should provide more specific examples from past fine-tuning work, particularly around distributed training and LoRA hyperparameter selection, and could quote more concrete metrics from previous eval projects.

Overall Assessment

The candidate shows senior-level working knowledge suitable for an ML/LLM engineering role with demonstrated ability to reason about retrieval, eval, and deployment tradeoffs at scale.

Individual Question Evaluations

Candidate Answer:

The first trade-off is chunking: smaller chunks improve retrieval precision but lose context, so I'd start around 512 tokens with overlap and tune against a labelled query set. For retrieval I'd go hybrid — BM25 plus dense embeddings — because pure vector search misses exact terms like product codes. At 2M documents you need a reranker: retrieve top-50 cheaply, rerank to top-5 with a cross-encoder. Then there's the context budget — stuffing more chunks raises cost and can hurt answer quality, so I'd measure groundedness against context size. Finally, incremental indexing for freshness rather than full re-embeds.

Evaluation:

Excellent RAG design answer covering all critical components. The candidate correctly framed chunking, hybrid retrieval, and reranking as measurable trade-offs rather than fixed recipes, and anchored each decision to evaluation against a labelled query set. The mention of incremental indexing and context-budget cost shows practical experience running RAG at scale.

Key Takeaway:

Strong RAG architecture skills with clear reasoning about retrieval precision, reranking, and cost tradeoffs. Production-ready thinking.

This is the AI hard-skills report we deliver per candidate, alongside the recruiter soft-skills notes and CV analysis. Generated as part of our 6-step shortlist process — every shortlist is signed off by a human recruiter.

Premium validation tools

For senior roles where cheating costs more than missing a good hire.

Two layers of defense: AI Skills Validation tests how candidates actually use AI tools. Recruo Secure Browser ensures the test results you see are real.

78% of engineering teams now use AI daily

GitHub Copilot, Cursor, ChatGPT — AI-assisted development is the default. But most hiring pipelines still test for pre-AI skills only.

AI proficiency ≠ copy-pasting prompts

The best engineers know when to use AI, when to override it, and how to validate its output. That judgment is what separates a 2x engineer from a 10x one in 2026.

Bad AI habits cost more than no AI at all

Blindly trusting AI-generated code leads to security vulnerabilities, hallucinated logic, and tech debt that compounds silently for months.

AI Skills Validation

A dedicated module added to the technical interview: real coding tasks with AI tools available, followed by probing questions about the candidate's AI-assisted workflow. The result is a separate AI Proficiency Score alongside the standard technical evaluation.

How it works

The candidate gets access to Copilot/Cursor during part of the interview. We observe how they prompt, validate, and iterate — then score their AI fluency on a structured rubric.

AI-assisted coding

Can they use Copilot/Cursor effectively while catching errors?

Prompt engineering

Do they write precise prompts, or do they brute-force trial and error?

AI output validation

Can they spot when an LLM hallucinates, introduces a vulnerability, or produces subtly wrong logic?

AI-native architecture

Do they know when to reach for an LLM vs. a deterministic solution?

Recruo Secure Browser

Our proprietary anti-cheat interview environment. The candidate joins the AI technical interview through our locked browser session — preventing tab switching, screen sharing to a second device, and silent ChatGPT usage.

Why it matters in 2026

73% of candidates admit to using AI assistants during remote technical interviews. Without proctoring, a great-looking technical score may just be a great-looking ChatGPT prompt.

Locked browser session

Candidate cannot open new tabs, navigate away, or use external apps during the interview.

Copy-paste detection

Every paste from outside the session is flagged with timestamp and content fingerprint.

Eye-tracking & focus checks

Webcam-based attention monitoring detects looking off-screen for sustained periods.

Screen-share & process monitoring

Detects screen sharing to a second device, suspicious processes, and AI assistants running in the background.

Both included by default in Retained mandates. Available as add-ons for Standard and Fixed-Fee plans.

Most used when hiring LLM engineers, RAG engineers, Evals engineers, ML platform engineers, Agentic AI developers, Generative AI engineers or AI/ML engineers.

Get started

Get a scorecard like this for your candidates

Scope one open role on a 30-min call — we deliver a shortlist of 3–5 vetted candidates with full scorecards in 5 business days. No upfront fee on Standard, 90-day replacement guarantee.

Pay on placement

No upfront fee on the Standard plan

90-day guarantee

Free re-search if hire leaves

EU AI Act aligned

Aligned by design

By submitting, you agree to our Privacy Policy. We'll never spam you.