| What the role owns | The measurement apparatus around a model-backed system: the dataset and how it was sampled, the rubric annotators apply, the deterministic scorers, the model judge and its validation against human labels, the adversarial and regression suites, the statistical claim attached to every number, the CI gate that blocks a bad release, and the written report a decision maker reads. |
|---|---|
| Licence or credential required | None. There is no licence, board exam, registration or protected title for AI evaluations work. The practical gate is evidence: an eval suite you built, a rubric you wrote, a judge you validated, and a failure taxonomy that came from reading real outputs rather than from a blog post. |
| Degree expectations | A degree in something quantitative or language-heavy is common and none is formally required at most employers. Frontier labs and third-party assurance organisations skew toward graduate degrees for the research-flavoured evals roles, often from outside computer science: psychometrics, survey methodology, experimental psychology, linguistics and education measurement. Product-side evals teams hire on portfolio. |
| Time to become hireable | From an adjacent technical role (QA automation, data analyst, support engineering, annotation lead), roughly three to nine months of deliberate work: one public eval suite with a labelled set, a validated judge, a written failure taxonomy and a short report. From no technical background, longer, and the realistic first step is annotation or domain-grader work that pays while you learn the vocabulary. |
| Certificates and their real weight | No certificate carries weight on its own. Cloud and vendor eval credentials, LLM evaluation courses and prompt engineering certificates rarely move a screen for this title. The exception is adjacent governance credentials such as ISO/IEC 42001 lead auditor training or the IAPP AI governance certification, which matter at assurance firms, consultancies and regulated enterprises, not at product teams. |
| Who actually hires | Frontier model providers and AI labs; large product companies shipping assistants and agents; eval and observability tooling vendors; data, annotation and RL environment vendors; third-party evaluation and assurance organisations; government AI institutes; banks, insurers, health systems and defence contractors building model risk and assurance functions. |
| Interview loop shape | At product companies, usually four to five stages over two to five weeks: hiring manager screen, take-home or live eval design exercise, sampling and statistics round, a harness-building coding round at a moderate bar, and a written memo or judgment discussion. Large technology employers may still run their standard engineering coding bar on top. Lab loops add research-style rounds and run longer. |
| Where to get real pay numbers | No BLS occupation code exists for this title. Triangulate from the OES codes the work gets filed under, namely 15-2051 Data Scientists, 15-1252 Software Developers, 15-1221 Computer and Information Research Scientists and 15-1253 Software Quality Assurance Analysts and Testers; then read ranges in postings from jurisdictions with pay transparency rules, and the US Department of Labor's public labour condition application disclosure files, which list filed salaries by employer and job title. |
What an AI evaluations engineer actually does, and the three versions of the job
An AI evaluations engineer builds the measurement system for software whose core behaviour comes from a model. The deliverable is not a model and not a feature. It is a number somebody is willing to bet a release on, plus the argument for why that number means what you say it means. Everything in the job follows from that: if the number is not trustworthy, nothing you built matters, and if nobody uses it to make a decision, you built a dashboard rather than an eval.
The work is unusual because the system under test is non-deterministic and the correct answer is often contested. Ordinary software testing assumes a specification: given this input, expect that output. A model-backed system has no such specification across most of its surface. Two reasonable people disagree about whether an answer was good, the same prompt returns different text on two runs, and the thing you most want to measure (did this actually help the user) is not visible in the response at all. So the job is partly engineering and partly measurement science: you have to define the construct before you can measure it, and defining the construct is where most teams fail.
Three versions of this job are hiring in 2026-27. They share a vocabulary and almost nothing else about the day. The first is model evaluation at a lab or model provider: measuring capability and safety properties of the model itself, running pre-deployment tests, building structured red-teaming exercises, and contributing the evidence that goes into a model card or system card. The second, and where most of the open postings sit, is product evaluation at a company building on somebody else's model: does this support assistant resolve the ticket, is the answer grounded in the retrieved document, did the agent call the right tool with the right arguments, and did last Tuesday's prompt change quietly break the Spanish-language cases. The third is platform and vendor work: building eval tooling, running evaluation as a service, third-party assurance of other people's systems, and producing the labelled data and task environments that everyone else's evals depend on.
Pick which one you are applying to before you write a line of your resume. A lab evals interview will probe threat modelling, elicitation (are you sure the model cannot do the thing, or did you just fail to ask properly) and research judgment. A product evals interview will probe sampling from real traffic, failure taxonomies, CI gating and the politics of telling a product manager the feature is not ready. A vendor interview will probe annotation quality control and throughput. The same candidate can do all three, but not with the same first paragraph.
The thing most newcomers get wrong is treating this as manual QA with a model attached. A good week looks like this: pull a stratified sample of production traces from the last fortnight, read two hundred of them by hand, notice that a cluster of the bad outcomes shares a cause nobody had named, write a rubric precise enough that two annotators apply it the same way, measure whether they actually do, build a deterministic scorer for the cases that can be checked exactly and a model judge for the cases that cannot, check the judge against the human labels and discover it is systematically lenient on one category, fix it, wire the suite into CI with a per-category breakdown, then write the two-page memo that makes a director change their release plan. The reading of raw outputs by hand is not a junior task you graduate from. It is where the findings come from.
- The artefacts you own, roughly in the order they get asked about: the labelled dataset and its sampling method; the rubric and annotation guideline; the scorers, deterministic and model-based; the judge validation against human labels; the adversarial suite; the regression gate in CI; the failure taxonomy; the report.
- Offline evals: fixed datasets scored before release. Fast, repeatable, and the only thing you can gate a deploy on. Their weakness is that they go stale and drift away from what users actually send.
- Online evals: measurement on live traffic, including explicit feedback, implicit signals (regeneration rate, edit distance between the suggestion and what the user kept, abandonment, escalation to a human), and A/B tests with guardrail metrics. Slower, truer, and the only thing that settles an argument about whether the offline number mattered.
- Red teaming and adversarial suites: structured attempts to make the system do the thing it must not do, including prompt injection arriving through retrieved content and tool output rather than only through the user's message.
- Human labelling operations: recruiting and calibrating annotators, writing guidelines, running adjudication on disagreements, and measuring agreement rather than assuming it. At domain-specific products this means paying clinicians, lawyers or accountants to grade, and designing a task they can do in minutes rather than hours.
- Cost and latency as quality dimensions, not afterthoughts. A change that improves answer quality and triples token spend or adds four seconds of latency is a trade, and the eval report is where that trade gets quantified instead of argued.
- The report: a short document with the claim, the denominator, the uncertainty, the per-category breakdown and an explicit recommendation. Eval work that never becomes a document does not change anything.
Who hires evals engineers, and the titles the job hides behind
This title barely existed in 2022. That has two consequences for your search, both in your favour if you know about them. First, job boards are unreliable: a keyword search for "AI evaluations engineer" shows you a fraction of the open roles, because the same job is posted under half a dozen names. Second, recruiters frequently do not have a working keyword list for it, so the hiring manager screens the pipeline personally and a specific, legible artefact travels further than it would in a mature title with a tidy funnel.
Search these strings on job boards and company career pages, not just the canonical title: evals engineer, LLM evaluation engineer, AI quality engineer, model evaluation engineer, research engineer evaluations, member of technical staff evaluations, AI test engineer, applied evaluation scientist, model behaviour analyst, AI red team engineer, trust and safety engineer (model), alignment evaluations, benchmark engineer, AI assurance engineer, RL environment engineer, and AI quality analyst. In regulated enterprises the same work appears inside model risk management and validation teams under titles that never mention AI at all.
The employer types differ in what they will pay you for and what they will interview you on. Frontier labs and model providers want elicitation skill, threat modelling and research taste, and they hire a small number of people slowly. Product companies shipping assistants and agents are the volume market: they want somebody who can turn a messy production log into a gate that stops regressions, and they increasingly want that person embedded in the product team rather than in a central quality function. Eval tooling vendors want engineers who have felt the pain and can build the abstraction. Data, annotation and task-environment vendors want operations-strong people who can design a grading task and control quality at scale, and this is the most accessible entry point for people without a software background.
Two employer categories are growing and are consistently under-searched by candidates. One is third-party evaluation and assurance: independent organisations that test other people's models and systems, and the audit and consultancy arms building AI assurance practices because their enterprise clients now have to show evidence to somebody. The other is the public sector and defence: national AI institutes and government technology units run evaluation programmes, and some of that work is gated on citizenship and security clearance, which narrows the field dramatically in your favour if you already hold one.
Regulated industry is the quiet third growth area. Banks have had formal model validation functions for years under supervisory guidance on model risk management (in the US, SR 11-7 and OCC 2011-12), and those functions are now being asked to cover generative systems, which they are not staffed for. Health systems and medical device companies have their own device-software and change-control expectations. Insurers, credit providers and employers face sector rules on automated decisions in several jurisdictions. If you have a background in one of these industries and you learn evals, you are a rarer candidate than a generalist engineer who learned evals, and the job is more durable because the obligation is not optional.
One channel candidates miss: a lot of lab and frontier-model eval work is contracted out rather than hired in. Data and environment vendors recruit domain graders, task authors and environment builders on contract, staff them against a named customer, and the work is real eval work under a different employment shape. It pays less than a staff role, it is often part-time and remote, and it is the fastest way to get production-grade eval experience on a resume when your current job offers none.
- Product companies post the most of these roles and interview on real traffic, taxonomies and CI gating.
- Labs and model providers hire fewer, interview on elicitation and threat modelling, and often expect a public artefact or a referral before the first conversation.
- Tooling vendors (eval platforms and LLM observability) hire engineers who have used the tools in anger and can explain what their competitors get wrong.
- Annotation, data and RL environment vendors are the most accessible entry point and the least glamorous. Domain graders with real credentials (nursing, law, accounting, trades) are in demand and the work is a legitimate on-ramp.
- Assurance, audit and consultancy practices hire people who can write, because their deliverable is a report with a defensible method section.
- Government and defence evaluation work exists, pays less than industry, and is gated on nationality and clearance. Hold a clearance and you compete against a very small field.
- Regulated enterprises hire into model risk, validation and AI governance functions. The AI part is learnable; the industry part is not, which is why insiders convert well.
What actually gates this job: no licence, and the three bars that screen you
Nothing legally gates the title. There is no licence, no board exam, no registration, no protected designation and no continuing education requirement. That is unlikely to change for the individual engineer: obligations in AI regulation land on the organisation deploying a system, not on the person measuring it. Anyone can call themselves an AI evaluations engineer, which means every claim you make will be probed in the room.
Bar one is measurement literacy: can you look at a number and say what would have to be true for it to be wrong. This is what separates evals engineers from everyone else who has run an eval. It shows up as questions about sampling, denominators, uncertainty and bias, and it eliminates most software engineers who applied because the title sounded adjacent to their job. You do not need a statistics degree. You need to be comfortable saying "that difference is inside the noise on a set this size" and then showing the arithmetic.
Bar two is engineering competence at a moderate level. You will write Python, handle a few hundred thousand model calls without melting the budget, deal with async batching, retries, rate limits and caching, version datasets and prompts, pin model snapshots, and wire a suite into continuous integration so it runs on every change. At most employers for this title the coding round looks like "build a small harness that scores these outputs and reports per-category results" rather than a graph puzzle. Two caveats worth knowing before you prepare: large technology companies often still run their standard engineering coding bar regardless of the role, and lab research-engineer titles can run harder. Ask what the coding round actually is.
Bar three is not advertised but behaves like one: writing. More than in almost any other engineering role, your output is consumed as prose by people who will not read your code. A candidate who can write a clear, short, honest report with an explicit recommendation beats a stronger engineer who cannot, because the second one's findings never reach anybody. Several employers now include a written exercise for exactly this reason, and the ones that do not are reading your take-home write-up for it anyway.
There is one genuine credential effect and it is narrow. In assurance, audit, consultancy and heavily regulated enterprises, management-system and governance credentials are taken seriously because the client or the regulator takes them seriously: the AI management system standard ISO/IEC 42001, its risk companion ISO/IEC 23894, the NIST AI Risk Management Framework and its generative AI profile, and AI governance certifications from professional bodies. Note what this is. It is a credential for the governance wrapper, not for the measurement. It gets you taken seriously in a procurement conversation. It will not get you through an eval design round.
One more thing functions as a gate at larger companies and nobody warns candidates about it: data handling. Evals run on production traffic, which means real user content, and in most organisations that puts you inside privacy review. You will deal with redaction and de-identification, retention limits on the eval set, regional restrictions on where data can be processed, and whether a third-party judge model is an approved sub-processor. Candidates who have never thought about this get stopped at the point where the interesting data lives.
- No licence. No board. No mandatory certification.
- Bar one, measurement: sampling, denominators, uncertainty, bias, and the honesty to say when a result is not significant.
- Bar two, engineering: Python, data handling, batched and cached model calls at controlled cost, dataset and prompt versioning, CI integration. Confirm the coding format before preparing.
- Bar three in practice, writing: a one or two page report with a claim, a method, a limitation and a recommendation.
- Underrated, data governance: redaction, retention, regional processing limits, and whether your judge model is an approved sub-processor for customer content.
- Governance frameworks worth knowing by name for enterprise and assurance roles: the NIST AI Risk Management Framework and its generative AI profile, ISO/IEC 42001, ISO/IEC 23894, and in US banking the long-standing model risk management guidance SR 11-7 and OCC 2011-12.
- On regulation and dates: obligations exist and timelines have been amended more than once, including in the EU. Describe the obligation (for example, requirements governing training and validation data quality, or that an evaluation be documented before deployment) and say the current timetable should be checked against the official text. Never quote a compliance date from memory in an interview.
- Security clearance is a real gate for a slice of government and defence evaluation work, and an advantage nowhere else.
Which backgrounds convert, and the specific gap each one has to close
Almost nobody in this job trained for it. The useful question is not "am I qualified" but "which half do I already have, and what is the one gap that will get me rejected". Five backgrounds convert reliably, and each has a predictable failure mode in interviews.
QA and test engineering is the most natural fit and the most common conversion. You already think in coverage, test design, flakiness, regression suites and release gates, and those concepts transfer directly. The gap is non-determinism and statistics: your instincts say a test passes or fails, and here a suite returns a proportion with an interval around it, two runs of the same suite differ, and variation is the expected state rather than a defect. Close it by learning proportions and uncertainty properly, by learning to set a threshold with a stated tolerance rather than demanding green, and by building one suite where you can explain how many items you needed to detect the size of change you cared about. Do not lead your resume with manual test case counts or a framework migration. Lead with a defect taxonomy you built and a gate you owned.
Data analysis and data science converts on the measurement half. You already have sampling, stratification, significance, intervals and the discipline of looking at distributions rather than means. The gap is engineering and shipping: evals are software that runs on every pull request, not a notebook you ran once. Close it by building and deploying a harness, putting it in CI, and dealing with the unglamorous parts (retries, cost, caching, dataset versioning). A second gap is tolerance for qualitative work: this job requires reading several hundred text outputs by hand and inventing categories, which is closer to qualitative coding than to a regression.
Linguistics, annotation and content work is an underrated and strong path, and the people who come from it are often the best in the room at the hardest part. Writing an annotation guideline that two strangers apply identically is a specialised skill, and so is building an error taxonomy that is exhaustive and mutually exclusive. If you have run annotation projects, measured inter-annotator agreement, adjudicated disagreements or written quality rubrics, you have the core of the job. The gap is code: you need enough Python to build and run a harness yourself rather than requesting one. That is a few months of focused work, not a degree, and it pays off immediately because it turns you from the person who defines quality into the person who ships the measurement of it.
Research backgrounds convert on construct validity, and the relevant research is often not machine learning. Psychometrics, education measurement, survey methodology, experimental psychology and clinical trial design are all disciplines about measuring things that cannot be observed directly, which is exactly the problem here. If that is your training, say so in those words: construct validity, reliability, item analysis, inter-rater reliability, blinding, order effects. Hiring managers who understand what you are saying will move you forward, because most of their candidates have never heard of the failure modes you were trained to avoid. The gap is engineering and speed: academic measurement is careful and slow, and a product team needs an answer this week with a stated limitation rather than a perfect answer next quarter.
Domain experts converting into evals are the fastest-growing group and the least well advised. A nurse, a lawyer, an accountant, a tax preparer, an electrician, a radiographer or a teacher who learns the eval vocabulary becomes the person who can say whether the model's answer is actually correct in a field where the generalists are guessing. Many people enter this way through grading and domain-expert data work, which is now a real contract market, and the move from grader to evals engineer is a matter of taking ownership of the rubric and the harness rather than only applying them. If that is you, the thing to build is a domain-specific eval set with a defensible correctness standard and a write-up of where generalist models fail in your field. That artefact is rarer and more interesting than another coding benchmark.
Two backgrounds convert less well than people expect. Pure prompt engineering experience reads thin by 2026, because everyone has done it and it says nothing about measurement. And machine learning research experience, which sounds like the closest fit, often is not: training models and evaluating products are different crafts, and a research candidate who has only reported benchmark scores can be weaker at building a labelled set from messy production traffic than a QA engineer who has never trained anything.
- From QA or SDET: keep coverage thinking and CI discipline, add proportions, uncertainty and tolerance-based gates. Rewrite the resume around failure taxonomies and release gates, not test counts.
- From data analysis: keep sampling and statistics, add shipped software, CI and cost control. Show that you can read two hundred outputs by hand and come back with categories.
- From linguistics or annotation: keep rubric writing and agreement measurement, add enough Python to own the harness. Your guideline-writing skill is the scarcest thing in the room.
- From research measurement disciplines: name construct validity, reliability and inter-rater agreement explicitly. Add speed and a willingness to ship a result with a stated limitation.
- From a professional domain: build one domain eval set with a defensible correctness standard, and write up where general models fail in your field. This is the most defensible portfolio piece available to anyone.
- From support or operations: you have the failure data nobody else has seen. Turn a year of tickets into a taxonomy and an eval set, with the customer content properly redacted, and you have a portfolio piece built from work you already did.
- Weak on its own: prompt engineering experience, certificate stacks, a tour of benchmark names, and "I have used an eval platform".
How hiring for this role actually works in 2026-27
Start with an honest read of the market, because the two common narratives are both wrong. The pessimistic one says evals is a fad that tooling will absorb. The optimistic one says every company is desperate to hire evals engineers. What is true is narrower: the number of companies that have shipped a model-backed feature is now very large, the number that can tell you whether it works is much smaller, and the gap between those two is the job. Demand is real and specific rather than universal, and it concentrates in organisations where somebody has already been burned by a silent regression, or has an obligation to show evidence to a customer, a regulator or a board.
The practical consequence is that postings are fewer than the work. A large share of these roles are created rather than advertised: a team hits the wall, somebody argues for a dedicated hire, and the first name considered is the person whose eval writeup the hiring manager read. That is why public work converts unusually well in this niche compared with more mature titles. A single well-written public artefact, such as a careful teardown of a popular benchmark's flaws, an open-source eval suite for a specific domain, or a merged contribution to one of the open eval frameworks, generates inbound interest that resume polishing does not.
Who screens you depends on the employer type. At product companies the hiring manager usually reads the pipeline personally, because the recruiter has no reliable keyword filter for a title this new. At labs there is often a research-style process with a take-home and multiple technical conversations, and a referral or public work is close to necessary to start it. At vendors and consultancies the process is more conventional and faster. At regulated enterprises the posting may sit inside a validation or risk function and the process will include a stakeholder round with compliance or legal.
A typical loop at a product company runs four to five stages over two to five weeks. A hiring manager screen, often not a recruiter screen, about what you have measured and what you found. An eval design exercise, either a take-home over real data or a live session, which is the decisive stage. A statistics and sampling conversation. A coding round at a moderate bar, building or extending a harness. And a final round that is partly judgment and partly writing, often including a stakeholder scenario where you have to say no to a launch. Lab loops add elicitation and threat-modelling rounds and a research discussion. Enterprise loops add governance and documentation.
Expect the take-home to be realistic and messy on purpose. The common shapes: here are several hundred model outputs from our product, tell us what is wrong with them and how you would measure it; here is a product spec, design the eval suite you would build before launch; here is an existing eval suite with a flaw in it, find the flaw; here is a benchmark result claiming an improvement, tell us whether you believe it. Ask what the expected time box is, because a well-scoped evals take-home is a few hours of work and a badly scoped one tells you something about the team.
- Who screens: hiring manager at product companies, research team at labs, recruiter-first at vendors and consultancies, risk or validation leadership at regulated enterprises.
- Stage that decides the offer: the eval design exercise. Most of the rest is confirmation.
- Stage candidates underestimate: the writing. Your take-home write-up is read as a work sample of the memos you will produce in the job.
- Stage that is usually easier than feared: the coding round, which is harness building at a moderate bar at most employers. Confirm the format first, because large technology employers may still run their standard engineering bar.
- Ask early which version of the job it is: model evals, product evals, or tooling. The answer changes what you should prepare, and sometimes reveals the team has not decided, which is useful to know.
- Ask who consumes the eval results and what decision they make with them. If nobody can answer, the role is a dashboard role and the measurement will not have teeth.
- Ask what access you will have to production traffic and how long privacy review takes. An evals role without data access is a role that cannot do the work.
- Timeline: two to five weeks at product companies and vendors. Labs and government roles run longer, sometimes considerably, and clearance processes add months.
The eval design round, where the offer is won or lost
One round decides this hire. You are handed a product, or a pile of outputs, and asked how you would know whether it works. Weak candidates start naming metrics and tools. Strong candidates start by refusing to answer the question until it has been made answerable, which is the actual skill being tested.
The structure that lands runs in this order. First, what decision will this number be used to make, and by whom. Second, what does success mean to the user, stated as an observable behaviour rather than a quality adjective. Third, where will the items come from, and what population do they represent. Fourth, who or what assigns the label, and how do we know the labeller is right. Fifth, what uncertainty is attached to the result and how large a change could we actually detect. Sixth, how is it broken down, because an aggregate hides exactly the thing you are looking for. Seventh, how does it run on every change and what does it block. Walk that ladder out loud and you will outperform people with more impressive backgrounds.
Expect hard follow-ups on the dataset, because that is where most eval suites are quietly broken. Where did the items come from: scraped from a benchmark, written by you, or sampled from production. How was it stratified, and does it over-represent the easy cases because those are the ones that logged cleanly. How do you know the correct answers are correct. How do you keep it from going stale as the product changes and users change what they ask. How do you keep it out of training data and prompt context, and what is your holdout policy. What did you deliberately exclude, and why. If the data is real user content, what is redacted, how long you keep it, and who approved the processing. The answer "we generated the test cases with a model" is not disqualifying, but it invites the question of whether you validated them against anything real, and most people have not.
Then the judge. Model judges are the default in 2026 and calling one is not a skill. The skill is validating one: measuring your judge's agreement with human labels on a held-out subset, knowing which biases to look for (preference for longer answers, preference for the first option presented in a pairwise comparison, leniency toward outputs in its own style), checking whether it is systematically wrong on a specific category rather than uniformly noisy, and deciding what agreement level is good enough for the decision at hand. The strongest single sentence you can say in this round is that your judge disagreed with humans on one named category, you found out because you checked, and here is what you did about it.
A statistics follow-up almost always arrives, and it is not adversarial. Typical: your pass rate moved from one figure to another, is that real. You should be able to say that it depends on the set size and on whether the items are paired, that on a small set a few points is routinely noise, that comparing two versions on the same items is far more sensitive than comparing two independent samples, and that you would attach an interval rather than a point estimate. Knowing the names signals literacy: bootstrap intervals, McNemar's test for paired pass or fail outcomes on the same items, Cohen's kappa for two annotators, Krippendorff's alpha or Fleiss' kappa for more. Being able to say when you would simply collect more items, and roughly how many more, signals judgment.
Expect at least one question about agents, because that is where the hard evaluation problems are now. A single response can be graded on its text. A trajectory cannot: the agent took eleven steps, called four tools, recovered from one error, and ended somewhere. You need an opinion on outcome-level versus step-level scoring, on partial credit, on how to check that a tool call was correct in its arguments rather than merely well formed, on whether you grade the environment state at the end, and on how to make runs reproducible when the environment itself changes. Say plainly that outcome-only grading tells you little about why a failure happened, and that step-level grading without an outcome check rewards agents that do everything right and still fail.
Finally, the "so what" question: you ran the suite and it says the new version is worse on one category and better overall. What do you recommend. This tests whether you can carry a result into a decision. Good answers name who is affected by the regressed category, whether that category is a small slice of traffic or the most valuable customers, whether the regression is recoverable with a targeted fix, and what you would ship behind a flag while you check. Bad answers either defer entirely to the product manager or treat the aggregate number as the decision.
- Open with the decision the number serves, not with a metric. Everything downstream follows from it.
- Separate the retrieval question from the generation question in any system with retrieval. A wrong answer because the right document was never retrieved is a different bug from a wrong answer despite it, and one aggregate score hides which you have.
- Name the sampling method and the population it represents. Production traffic with stratification beats a synthetic set, and you should say why.
- Break results down by category, segment, language, input length and conversation turn. The aggregate is what hides the problem you were hired to find.
- Validate the judge against human labels and say the agreement. A judge nobody checked is an opinion with a decimal point on it.
- For agents, have a position on outcome versus step-level scoring, partial credit, tool-call argument correctness, cost caps and environment reproducibility.
- Have an answer ready for how you handle real user content: redaction, retention, approved processing locations, and which judge models are allowed to see it.
- Close with a recommendation and a limitation. The limitation is what makes the recommendation credible.
The resume and the evals portfolio: what lands and what is ignored
The unit of credibility on an evals resume is a measured decision, not a tool. Write each bullet as a sentence that answers five things: what system, what the dataset was and where it came from, how it was labelled and checked, what the result was with its denominator, and what changed because of it. If the last clause is missing, the bullet reads as work nobody used.
Numbers must have denominators and provenance. "Improved accuracy by 30 percent" is the single most damaging line you can write for this title, because the entire job is being the person who does not accept a number like that. A bullet shaped like "pass rate on a 900-item labelled set sampled from six months of production tickets, stratified by customer tier, moved from one figure to another, with remaining failures concentrated in multi-turn clarification" is credible even when the figures are unexciting, because it is shaped like something real. If you cannot state the set size and where the items came from, do not state the number at all. Describe the mechanism instead.
The portfolio matters more here than almost anywhere else, because the title is young and the artefact is small enough to actually finish. One good eval suite in public beats a long list of courses, and it is a weekend-scale project if you pick a narrow domain you know. What a reviewer looks for in the first three minutes: a README that states what is being measured and what decision it serves; a dataset with a documented sampling method rather than a hundred cases a model invented; a rubric specific enough that a stranger could apply it; a judge validation section with agreement against human labels; a per-category breakdown rather than one number; and a write-up that admits what the suite does not cover.
Build it on a domain you actually know. An eval suite for general chat quality is a commodity and reads as a tutorial. An eval suite for discharge summary accuracy, or lease clause extraction, or whether a model gets UK versus US payroll rules right, or whether it can read a wiring diagram, is a portfolio piece that starts a conversation, because the labels are hard to produce and you are the reason they exist. The failure analysis is the part almost nobody writes and the part hiring managers read first. If you build it from work data, use public or synthetic material or get written permission, because publishing customer content is a faster way to end a candidacy than any technical mistake.
Learn the public benchmark landscape well enough to be critical of it, and never lead with it. You should be able to name the common ones and say what each actually measures and where it breaks. Broad knowledge and reasoning sets such as MMLU and GPQA, grade-school and competition maths sets such as GSM8K, coding sets from HumanEval through SWE-bench and its human-verified subset, agent and tool-calling suites such as the Berkeley Function Calling Leaderboard, tau-bench, WebArena, OSWorld and GAIA, and the public preference arenas. The useful thing to say is why a number from any of them is weak evidence about your product: contamination is the assumed state of anything public for long, saturation compresses the top of the range, and none of them contain your users. Saying that with a specific example is a stronger signal than reciting scores.
Know the tooling by name and be opinionated rather than loyal. The open harnesses and frameworks (EleutherAI's lm-evaluation-harness, Stanford CRFM's HELM, the UK AI Security Institute's Inspect, promptfoo, Ragas, DeepEval) and the commercial eval and observability platforms (Braintrust, LangSmith, Weights and Biases Weave, Arize Phoenix, Galileo, plus the evaluation services inside the major clouds) are all reasonable answers to "what have you used". The good answer explains what you would use for what: a harness for reproducible benchmark-style runs, a tracing platform for sampling real production traces into a review queue, and your own thin code for the scorers specific to your product. Mentioning that traces follow the OpenTelemetry generative AI semantic conventions, which are still evolving, signals that you have thought about not being locked in.
- Lands: an eval suite with a stated set size, a sampling method, and a named source of items. Provenance is the thing that cannot be faked.
- Lands: a judge validation sentence. Agreement with human labels, measured, on a held-out subset, with a named category where the judge was wrong.
- Lands: a failure taxonomy with category names that sound like they came from reading outputs rather than from a framework. "Answers a different question than asked when the user's message contains two questions" is specific. "Hallucination" is not.
- Lands: a gate. The suite runs in CI, blocks on a stated threshold, and reports per category. Say what it blocked at least once.
- Lands: scope honesty. "I built the dataset and the scorers, a colleague owned the tracing pipeline." This survives the deep dive. Inflated ownership does not.
- Lands: a red team or adversarial result, especially indirect prompt injection through retrieved content or tool output rather than through the user's message.
- Lands: the cost and runtime of your suite. Knowing your cost per full run and per month marks you as someone who has operated a suite rather than built one.
- Ignored: a list of eval platforms you have logged into. Everyone has.
- Ignored: benchmark scores you reproduced from a paper, and leaderboard positions. They show you can run a script.
- Ignored or harmful: percentage improvements with no denominator, "reduced hallucinations significantly", and "improved model performance" with no statement of what was measured on what set.
- Ignored: certificate stacks and prompt engineering courses at the top of the page. Move them to the bottom or leave them off.
What this role pays, and how to get numbers you can trust
Anybody quoting you a single band for AI evaluations engineer is averaging across a frontier lab, a Series B startup, an annotation vendor and a bank's model validation group, which are four different jobs with different pay structures. Do not take a band from a content site. Build your own picture from sources that can be checked.
There is no occupation code for this title, which is why the aggregate sites are unreliable here. The nearest official anchors in the US are the Bureau of Labor Statistics Occupational Employment and Wage Statistics series, where this work is filed under whichever of these the employer's HR function picked: 15-2051 Data Scientists, 15-1252 Software Developers, 15-1221 Computer and Information Research Scientists, or 15-1253 Software Quality Assurance Analysts and Testers. Those four codes have meaningfully different medians, and the spread between them tells you something real: the same work is paid like data science at one employer and like QA at another, and which one it is depends largely on where the role sits in the organisation chart. That is a negotiation lever, not a trivia point.
The most useful source for a specific employer is the one most candidates never open. Several US jurisdictions require a pay range in the posting, so a national search filtered to those locations gives you ranges for roles that are otherwise opaque. And in the US, employers sponsoring work visas must file a labour condition application, and the Department of Labor publishes those disclosure files: they list employer, job title, worksite and the wage offered. For a named company you are interviewing with, that is public, checkable and better than an estimate. Levels-reporting sites are useful for large technology employers specifically, where the level structure is public and self-reported data is dense.
What moves the number for this title, in rough order of effect: the employer category, with labs and large technology firms at the top, vendors and annotation companies at the bottom, and regulated enterprise in the middle with better stability; whether the role is scoped as research or as quality assurance, which is often decided by a job architecture accident rather than by the work; whether you can own the harness yourself or depend on an engineer to build it, which is the clearest step change on the individual contributor ladder; and domain depth in a field where correctness is expensive to judge. Equity is a large share of total compensation at labs and startups and close to none at consultancies and in government.
One structural note worth having in mind when you weigh an offer. Evals work is cheap to start and expensive to sustain, which means some teams fund a burst of it before a launch and then let it decay. Ask what happened to the last eval suite the team built, who maintains it now, and whether anybody has ever delayed a release because of it. A team that has said no to a launch at least once has a real function. A team that has never blocked anything is paying you to produce reassurance, and that job is both less interesting and less durable.
- Start from BLS OES codes 15-2051, 15-1252, 15-1221 and 15-1253 to understand the range the work is coded into, then find out which one your target employer uses.
- Filter job searches to jurisdictions with pay transparency requirements to see real posted ranges for comparable roles.
- Use the US Department of Labor's public labour condition application disclosure data for filed salaries at a named employer and title.
- Expect large employer-category effects: labs and big technology at the top, annotation and data vendors at the bottom, regulated enterprise in between with more stability and less equity.
- The clearest individual contributor step change is owning the harness and the gate rather than only producing labels and analysis.
- Before accepting: ask whether an eval result has ever blocked a release. The answer tells you whether the function is real.
What an AI evaluations engineer has to know about AI in 2026-27
For this role the usual advice ("learn to use AI tools") is useless, because AI is the subject under test. The useful question is narrower: what changed between the 2023 version of evals work and the 2026 version, what will an interviewer assume you have internalised, and where has the hype outrun the reality.
The biggest change is that the centre of gravity moved from static public benchmarks to task-grounded, private evaluation. In 2023 much of the eval conversation was about leaderboard positions on broad knowledge and reasoning sets. By 2026 the long-standing ones are saturated, contaminated or both: scores cluster near the ceiling on sets like MMLU, GSM8K and HumanEval, the questions have been in training data for years, and a high number tells you very little about whether a product works. The response has been a shift toward environment-based and task-based evaluation, where a model is given a real task in a sandbox and graded on the end state, toward human-verified subsets of older benchmarks such as SWE-bench Verified and GPQA Diamond, and above all toward private held-out sets built from an organisation's own traffic. The practical implication for you is that the valuable skill is building a set nobody else has, not running a set everybody has.
The second change is that model judges became the default, which moved the skill up one level. Three years ago using a model to grade another model was a slightly controversial shortcut. Now it is standard practice across product teams, and the differentiator is no longer whether you use a judge but whether you validated one. Interviewers assume you can prompt a judge. They are checking whether you measured its agreement with human labels, whether you know its biases (length preference, position effects in pairwise comparison, leniency toward outputs that look like its own writing), whether you tested it per failure category rather than in aggregate, and whether you know when the honest answer is that a judge cannot carry this decision and a human has to.
The third change is agents, and it broke the most existing practice. When the product was a single response, evaluation was text scoring. When the product is a multi-step agent with tools, the unit of evaluation is a trajectory with a state at the end, and most of what teams built for the single-response era does not transfer. This brought a set of problems that are now standard interview material: outcome versus step-level scoring, partial credit, whether a tool call was correct in its arguments and not merely syntactically valid, how to make a run reproducible when the environment mutates, how to attribute a failure to the step that caused it rather than the step where it surfaced, how to cap cost so a suite run does not become a budget incident, and how to evaluate recovery, since an agent that errs and corrects itself may be better than one that never errs on easy cases. Public agent and tool-calling suites exist (the Berkeley Function Calling Leaderboard, tau-bench, WebArena, OSWorld, GAIA among them) and the serious work is still mostly bespoke, which is why the demand is here.
The fourth change created a second labour market next door. Because post-training now leans heavily on graded tasks and simulated environments, the same skills that build an eval set also build a training environment: a task, a verifier, and a defensible notion of correct. Vendors staffing environment-building and expert-grading work recruit continuously, usually on contract, often part-time and remote, and they want the same thing an evals hiring manager wants: somebody who can specify what correct means in a narrow domain and defend it. Treat it as an on-ramp and as an alternative buyer for the same skill, not as a separate career.
The fifth change is less technical and more consequential for employment: evals became an external artefact. Model and system cards, enterprise procurement questionnaires, customer security and AI reviews, third-party assurance, internal audit and sector regulators now ask organisations to show that they evaluated a system before deploying it, and to show the method. That turned eval results from an engineering convenience into a document that leaves the building, which is why writing matters so much in this role and why assurance and regulated-industry hiring is growing. Be careful with specifics. Obligations and timetables in AI regulation have been amended more than once, including deferrals, so in an interview describe the obligation and say the current date should be checked against the official text rather than quoting a deadline from memory. Getting a date wrong in the one room that matters is worse than not naming one.
Now the honest counterweight, because overstating disruption is its own failure. Several things did not change and are not about to. You still cannot automate away defining what good means for your product. Synthetic test generation helps with volume, not with the construct. Human labels are still the ground truth everything else is calibrated against, and the economics shifted rather than vanished: easy labelling got cheap and expert labelling got more expensive, which is why qualified domain graders are now a real labour market. A leaderboard score still does not predict whether your product works. And the highest-return hour in the job is still a person reading real outputs and noticing something, which no tool has replaced. If a candidate tells you evals are becoming automated, ask them who writes the rubric.
Dataset construction from real traffic, with provenance and a holdout policy
The dataset is where most eval suites are silently broken, and it is the first thing a competent interviewer attacks. Sets assembled from public benchmarks are contaminated, sets written by a model are untethered from what users actually send, and sets built from whatever logged cleanly over-represent the easy cases. A suite with a bad set produces confident numbers about the wrong population, which is worse than no suite because people act on it.
Show it: State the size, the source, the sampling method and the stratification: how many items, drawn from what window of production traffic, stratified by what dimension and why. Say how correctness was established and by whom. Say what your holdout policy is and how you keep items out of prompt context and training data. Name one thing you deliberately excluded and the reason, because the exclusion is the detail that proves you built it rather than downloaded it.
Rubric and annotation guideline writing, with measured agreement
Every score traces back to a judgment about what good means, and a rubric two people apply differently makes every downstream number noise. This is the scarcest skill on most teams because it is tedious and unglamorous, and it is exactly where people from linguistics, annotation, measurement research and the professions have an advantage. Teams that skip it discover six months later that their labels were never consistent and their trend line was an artefact.
Show it: Show a real rubric with the categories, the boundary cases, and the tie-break rules for the disagreements you actually hit. Report an agreement figure between annotators and name the statistic you used (percent agreement is not enough when classes are imbalanced: Cohen's kappa for two raters, Krippendorff's alpha or Fleiss' kappa for more). Describe your adjudication process for disagreements and one guideline revision the disagreements forced.
Judge validation: agreement, bias and per-category checks
Model judges are now the default scoring mechanism for anything that cannot be checked deterministically, so the integrity of the whole suite rests on whether the judge is right. Known failure modes are specific and testable: preference for longer answers, position bias in pairwise comparison, leniency toward outputs that resemble its own style, and systematic blindness in one category while looking fine in aggregate.
Show it: Give the method, not the prompt: a held-out subset with human labels, agreement measured overall and broken out per category, a named category where the judge was systematically wrong, and what you did about it (a rubric change, a different judge model, pairwise instead of pointwise, or routing that category to humans). Say what agreement level you decided was good enough for the decision at hand, and why.
Deterministic scorers wherever they are possible
Model judges are expensive, slow and fallible, and a surprising share of what teams judge with a model can be checked exactly. Schema validation, executing generated code against unit tests, checking a tool call's arguments against the API contract, verifying a cited span actually exists in the retrieved document, numeric tolerance checks and constraint checks are all deterministic, cheap, instant and reproducible. Reaching for a judge first is a reliable sign that somebody has not thought about the problem.
Show it: Describe which parts of your suite are deterministic and which required a judge, and say why the split fell where it did. A concrete example works best: schema and tool-argument validation done exactly, factual groundedness checked by verifying the cited passage exists and supports the claim, and only tone and helpfulness left to a judge.
Sampling and statistical claims: denominators, intervals and paired comparisons
The central function of this role is being the person who will not accept a number at face value. Teams routinely celebrate movements that are inside the noise on a set of eighty items, and routinely miss real regressions because the aggregate hid them. Knowing how large a set you need to detect the size of change you care about is what converts eval work from theatre into a gate.
Show it: Attach intervals to results rather than reporting point estimates. Compare two versions on the same items rather than on independent samples, say that you did, and name the test you used for paired pass or fail outcomes, which is usually McNemar's. Be able to answer "is that difference real" with arithmetic rather than opinion, and be willing to say a result is not significant and you need more items, which is the answer that builds the most trust.
Agent and trajectory evaluation
This is where the hard open problems are in 2026-27, and where teams most often need to hire rather than retrain. A multi-step agent cannot be graded on its final text: the interesting information is in which step went wrong, whether a tool call had correct arguments, whether the environment ended in the intended state, and whether a failure was unrecoverable or self-corrected. Suites built for single responses do not transfer, and most candidates have only ever evaluated single responses.
Show it: Have an explicit position on outcome-level versus step-level scoring and on partial credit, and say what each misses on its own. Describe how you made runs reproducible when the environment has state, how you attributed a failure to the causing step rather than the surfacing step, and how you capped cost (step limits, token budgets, wall-clock timeouts) so a suite run does not become a budget incident.
Retrieval evaluation kept separate from generation evaluation
In any system with retrieval, a large share of what gets reported as hallucination is a retrieval miss: the model was never shown the right passage. A single end-to-end quality score cannot tell you whether the answer was wrong because the document was missing or wrong despite it, which means it cannot tell you which team should fix it. Collapsing these two is one of the most common structural flaws in real eval suites.
Show it: Report retrieval and generation separately: recall at k against a set of questions with known-relevant documents on one side, and groundedness or attribution checking on the other. Say what share of your failures were retrieval failures, because that number is usually surprising and nobody can produce it without having done the work.
Adversarial testing and prompt injection suites
Once a system retrieves documents or calls tools, untrusted text reaches the model from places the user did not type: a web page, an email, a support attachment, a tool response. This is the attack surface security reviewers ask about most in agent deployments, and teams that can show a maintained injection suite clear enterprise security reviews that others stall in. It is also a distinguishing skill, because most candidates test only what the user types.
Show it: Describe a maintained suite rather than a one-off exercise: a taxonomy of attack categories, cases that arrive through retrieved content and tool output rather than the user message, tests for data exfiltration and unauthorised tool use, and a regression run so a fixed attack cannot silently return. Name one attack you found that the team had not anticipated.
Online measurement and the link to user outcomes
Offline suites are the only thing you can gate a deploy on, and they drift from reality. The teams that stay honest connect offline scores to online behaviour: whether the offline improvement showed up as fewer escalations, fewer regenerations, more accepted suggestions, shorter handling time. Being able to close that loop is what gets an evals engineer taken seriously by product leadership rather than treated as a checkpoint.
Show it: Name the implicit signals you used and what each actually indicates (regeneration rate, edit distance between a suggestion and what the user kept, abandonment, escalation to a human, repeat contact within a window). Describe one case where the offline number improved and the online number did not, and what you learned about your eval set from that gap. That story is the strongest single thing you can tell in a senior interview.
Keeping a suite alive: model deprecation, drift and maintenance
An eval suite is not a deliverable, it is a running system, and it decays in specific ways. Pinned model snapshots get retired and your baseline disappears. Prompts change underneath you. The traffic distribution moves, so a set sampled last spring measures a population that no longer exists. Judge models get upgraded and your historical trend line silently rebases. Teams that do not plan for this discover their suite is measuring the past at exactly the moment they need it.
Show it: Say how you version datasets, prompts and judge configurations together, so a result is reproducible as a tuple rather than a number. Describe what you do when a pinned snapshot is deprecated, including re-baselining and keeping the old and new judge in parallel for a window. Say how often you resample from traffic and what triggers it.
Writing the report, and the governance vocabulary around it
Eval results are consumed as documents by people who will not read your code, and increasingly those documents leave the building: into system cards, procurement questionnaires, customer AI reviews, internal audit and assurance engagements. An engineer who cannot write a short, honest, decision-shaped report has findings that do not reach anybody. In regulated settings this is not a soft skill, it is the deliverable.
Show it: Produce a real two-page example: the claim, the method, the denominator and uncertainty, the per-category breakdown, the limitations, and an explicit recommendation. Know the governance frameworks by name (the NIST AI Risk Management Framework and its generative AI profile, ISO/IEC 42001, ISO/IEC 23894, and model risk management expectations in banking) and talk about obligations rather than dates, noting that timetables in AI regulation have been amended and should be checked against the current official text.
What a screen is looking for
These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.
- AI evaluations engineer
- Evals engineer
- LLM evaluation
- Model evaluation
- AI quality engineer
- Evaluation harness
- Eval suite
- Benchmarking
- LLM-as-a-judge
- Judge validation
- Human evaluation
- Annotation guidelines
- Rubric design
- Inter-annotator agreement
- Inter-rater reliability
- Cohen's kappa
- Krippendorff's alpha
- Ground truth labelling
- Labelled dataset
- Golden dataset
- Data sampling
- Stratified sampling
- Holdout set
- Benchmark contamination
- Regression testing
- Continuous integration
- CI gating
- Failure taxonomy
- Error analysis
- Root cause analysis
- Non-determinism
- Reproducibility
- Confidence intervals
- Bootstrap resampling
- Paired comparison
- McNemar's test
- Statistical significance
- Power analysis
- Construct validity
- Reliability
- Psychometrics
- Measurement design
- Python
- pandas
- pytest
- Async batching
- Rate limiting
- Caching
- Cost optimisation
- Prompt versioning
- Dataset versioning
- Model snapshot pinning
- Retrieval augmented generation (RAG)
- Recall@k
- Groundedness
- Attribution checking
- Hallucination detection
- Agent evaluation
- Trajectory evaluation
- Tool-call correctness
- Function calling
- Sandboxed environments
- RL environments
- Red teaming
- Adversarial testing
- Prompt injection
- Indirect prompt injection
- Jailbreak testing
- Guardrails
- Trust and safety
- Responsible AI
- AI governance
- Model risk management
- Model validation
- SR 11-7
- NIST AI Risk Management Framework
- ISO/IEC 42001
- ISO/IEC 23894
- Model cards
- System cards
- PII redaction
- Data retention
- Observability
- Tracing
- OpenTelemetry
- A/B testing
- Online metrics
- Guardrail metrics
- User feedback signals
- lm-evaluation-harness
- HELM
- Inspect
- promptfoo
- Ragas
- DeepEval
- LangSmith
- Braintrust
- Weights & Biases Weave
- Arize Phoenix
- MLflow
- SWE-bench
- SQL
- Data analysis
- Technical writing
- Stakeholder communication
Mistakes that cost people this job
Reporting one aggregate score. "Our eval pass rate is 87 percent" is the number least likely to contain the information anybody needs, because the whole purpose of an eval is to find the slice that is broken and an average is designed to hide it.
Break every result down by failure category, customer segment, language, input length and conversation turn, and report the breakdown as the headline with the aggregate as a footnote. The sentence that gets you hired sounds like "overall it looks flat, but the multi-turn Spanish cases dropped, and that slice is small in volume and includes our two largest accounts".
Using a model judge without ever validating it. The judge is treated as an oracle, nobody compares it to human labels, and six months of trend data turns out to measure the judge's preferences rather than product quality.
Keep a human-labelled held-out subset and measure agreement against it, broken out per category rather than in aggregate. Know the standard biases to probe for (length preference, position effects in pairwise comparison, leniency toward its own style). Be able to name one category where your judge was wrong and what you changed, because that detail proves the validation happened.
Building the eval set from a public benchmark or from cases a model invented, then claiming it represents your users. The set is clean, convenient, and measures a population that does not exist.
Sample from real traffic, stratified deliberately, and state the window and the method. If you genuinely have no traffic yet, say so, use synthetic items as scaffolding only, and state the limitation in the report. Replace them with real items as soon as traffic exists, and say in the interview that you planned to.
Claiming an improvement with no denominator, as in "improved accuracy 30 percent" or "reduced hallucinations significantly". In any other role this is vague. In this one it is disqualifying, because interrogating exactly that sentence is the job.
Always give the set size, the provenance of the items, the before and after figures, and the uncertainty. If the honest answer is that the set was small and the change was inside the noise, say that, and say how many more items you would have needed. Candour about a weak result reads as competence. A confident unsupported number reads as the problem you were hired to prevent.
Building an eval suite nobody uses. It runs nightly, it posts to a channel, nothing is ever blocked by it, and after two quarters it has drifted and everybody ignores the red.
Wire it to a decision from day one: a threshold that blocks a merge or a deploy, with an owner and a documented process for overriding it deliberately. Then protect it by keeping the suite fast and the variance low, because a slow or noisy gate gets switched off. In interviews, lead with what your suite blocked, not with what it measured.
Treating this as pass or fail testing inherited from QA. Demanding a green suite on a non-deterministic system produces either a trivially easy suite or a permanently red one, and both get ignored.
Set thresholds with tolerances and track trends with uncertainty. Decide in advance what size of drop in which category constitutes a block, and accept that run-to-run variation exists by pinning what you can (model snapshot, temperature, seed where available) and measuring what you cannot.
Chasing public benchmark scores as evidence of quality, either for a model choice or on your own resume. Saturation and contamination make a high score close to uninformative about your product.
Use public benchmarks for coarse screening of candidate models only, then decide on your own private set drawn from your traffic. On a resume, lead with a set you built and labelled, not with a benchmark you reproduced. Being able to explain why a popular benchmark is unreliable is a far stronger signal than citing a score from it.
Skipping the hand review. The engineer builds the harness, never reads the outputs, and invents categories from a taxonomy they found online. The suite then measures the failures somebody else had.
Read several hundred real outputs before you write a single scorer, and let the categories come from what you see. Keep doing it periodically, because the failure distribution moves as the product and the users change. This is the highest-return hour in the job and the one candidates most often admit they have never spent.
Ignoring the cost and runtime of the eval suite itself. A full run across several models and thousands of items with a judge call each becomes expensive and slow enough that the team quietly stops running it on every change.
Tier the suite: a small fast smoke set on every change, the full set nightly or on release candidates. Cache aggressively, keep deterministic scorers where possible, and know your cost per full run and per month. Being able to state those two numbers marks you as someone who has operated a suite.
Treating production traffic as freely usable. Pulling real user content into an eval set without redaction, retention limits or approval, or publishing a portfolio suite built from an employer's customer data, is the fastest non-technical way to lose a job or an offer.
Learn the constraints before you need them: what must be redacted, how long an eval set may be retained, which regions the data may be processed in, and whether your judge model is an approved sub-processor. For public portfolio work, use open data, your own data, or synthetic material, and say in the README which it is.
Over-claiming regulatory knowledge, especially naming compliance dates. Timetables in AI regulation have been amended, including deferrals, and a candidate who confidently states a wrong date in a room with a compliance lead has damaged their credibility on everything else they said.
Describe obligations rather than deadlines: what a framework requires regarding evaluation evidence, data quality or documentation. Then say the current timetable should be checked against the official text. In this role, being precise about the limits of your knowledge is part of what is being assessed.
Writing a resume around tools rather than decisions. A list of eval platforms, model names and frameworks occupies the exact space where a measured decision should be, and every candidate has the same list.
Write each bullet as system, dataset and provenance, labelling method, result with denominator, and what changed as a result. Four bullets of that shape beat a page of tooling. If a bullet has no last clause, it describes work nobody used, and it should be cut.
Questions people ask
What does an AI evaluations engineer actually do?
An AI evaluations engineer builds the measurement system that tells an organisation whether a model-backed product works. The work has six recurring parts: building labelled datasets sampled from real traffic; writing rubrics precise enough that two annotators apply them the same way; building scorers, both deterministic checks and validated model judges; running adversarial and red-team suites; wiring a regression gate into continuous integration that blocks a release when a change makes something worse; and writing the report a decision maker reads. A large share of the real findings come from one unglamorous activity: reading several hundred real model outputs by hand and noticing a pattern nobody had named. A useful test of whether a team is doing evals or theatre is simple: has an eval result ever stopped a release.
Do you need a degree, a certificate or a PhD to become an AI evaluations engineer?
No. There is no licence, board exam, registration or protected designation behind the AI evaluations engineer title, and no certificate carries weight on its own at product companies. A degree in something quantitative or language-heavy is common and is not a formal requirement at most employers. Frontier labs and third-party assurance organisations skew toward graduate degrees for the research-flavoured roles, and those degrees are often from outside computer science, in fields such as psychometrics, survey methodology, experimental psychology or linguistics, because those disciplines train people to measure things that cannot be observed directly. The one place credentials genuinely help is governance-adjacent work in regulated industry and assurance, where AI management system and governance certifications are taken seriously by clients and auditors. They will not carry you through an eval design interview.
Which backgrounds convert into AI evaluations work most easily?
Five backgrounds convert reliably into AI evaluations engineer roles. QA and test engineering brings coverage thinking, regression discipline and CI gating, and has to add statistics and tolerance for non-determinism. Data analysis and data science brings sampling and uncertainty, and has to add shipped software and willingness to read raw text by hand. Linguistics, annotation and content work brings rubric writing and agreement measurement, which is the scarcest skill in the field, and has to add enough Python to own the harness. Measurement-heavy research backgrounds such as psychometrics and survey methodology bring construct validity and reliability, and have to add speed. Domain experts in a profession bring the ability to say whether an answer is actually correct, and convert through grading work into ownership of the rubric and the suite. Pure prompt engineering experience converts poorly, and machine learning research converts less automatically than people expect, because training models and measuring products are different crafts.
What does an AI evaluations portfolio look like?
An AI evaluations engineer is hired off one narrow, finished eval suite in public, built on a domain you actually know. A reviewer looks for six things in the first three minutes: a README stating what is measured and what decision it serves; a dataset with a documented sampling method and real provenance rather than cases a model invented; a rubric specific enough that a stranger could apply it; a judge validation section reporting agreement against human labels per category; results broken down by category rather than a single number; and a written failure analysis that admits what the suite does not cover. The failure analysis is the part almost nobody writes and the part hiring managers read first. A suite for a specific domain such as clinical discharge summaries, lease clause extraction or payroll rules in a named jurisdiction beats a general chat quality suite, because the labels are expensive and you are the reason they exist. Use open, synthetic or personal data, never an employer's customer content.
What is the difference between an AI evaluations engineer and an AI engineer?
An AI engineer builds the model-backed system: retrieval, tool-calling, prompt assembly, latency and cost. An AI evaluations engineer builds the measurement around it and is institutionally positioned to say the system is not ready. The two overlap heavily, because any good AI engineer builds evals and many evals engineers can build the product, but the deliverable differs: the AI engineer's artefact is a running service, the evals engineer's artefact is a trustworthy number plus the argument for why it means what it says. At smaller companies one person does both. At scale they separate, and the separation is deliberate, because the person who built the feature is the worst-placed person to judge whether it works.
Is LLM-as-a-judge actually trusted, or is it a shortcut?
It is the default scoring mechanism an AI evaluations engineer reaches for in 2026 whenever something cannot be checked deterministically, and it is trusted exactly as far as it has been validated. Using a judge is not a skill. Validating one is. That means measuring agreement with human labels on a held-out subset, breaking that agreement out per failure category rather than reporting it in aggregate, and probing for known biases such as preference for longer answers, position effects in pairwise comparison, and leniency toward outputs written in its own style. It also means knowing when a judge cannot carry the decision and a human has to, which is usually when the stakes are high, the domain is specialised or the correct answer is genuinely contested. A judge nobody checked is an opinion with a decimal point on it.
How do I get experience in evals if my current job has none?
Build the artefact that gets an AI evaluations engineer hired, using data you can legitimately use. Three routes work. If you work in support, operations or a profession, turn a year of real cases into a labelled eval set and a failure taxonomy, redacted or recreated so no customer content leaves your employer. If you are already technical, pick an open model and a narrow task, build a suite with a documented sampling method and a validated judge, publish it with a failure write-up, and contribute a task or a fix to one of the open eval frameworks such as Inspect or lm-evaluation-harness. If you have no technical background yet, domain grading, annotation and task-environment work is a real paid contract market and a legitimate on-ramp; the step up is taking ownership of the rubric and then of the harness. In all three cases the thing to avoid is a tutorial reimplementation of a public benchmark, because it demonstrates only that you can run a script.
What does the interview loop look like, and which stage decides the offer?
An AI evaluations engineer loop at a product company runs roughly four to five stages over two to five weeks: a hiring manager screen (frequently not a recruiter, because the title is too new for a reliable keyword filter), an eval design exercise as a take-home or live session, a sampling and statistics conversation, a coding round at a moderate bar building or extending a harness, and a final round mixing judgment under ambiguity with writing. The eval design exercise decides the offer and the rest is largely confirmation. Lab loops add elicitation, threat modelling and research-style discussion, and run longer. Enterprise and assurance loops add a governance and documentation round. Two preparation mistakes are common: grinding competitive algorithm problems for a coding round that is actually about building a scoring harness, and assuming that is true everywhere, since large technology employers may still run their standard engineering bar. Ask what the coding round is before you prepare for it.
What does an AI evaluations engineer get paid?
There is no dedicated occupation code for AI evaluations engineer, which is why aggregate salary sites are unreliable here. Build your own picture: use the US Bureau of Labor Statistics Occupational Employment and Wage Statistics for the codes this work is actually filed under, namely 15-2051 Data Scientists, 15-1252 Software Developers, 15-1221 Computer and Information Research Scientists and 15-1253 Software Quality Assurance Analysts and Testers; read posted ranges from jurisdictions that require pay transparency in postings; and for a specific employer, read the US Department of Labor's public labour condition application disclosure files, which list filed salaries by employer and job title. The largest driver is employer category, with labs and large technology firms at the top, annotation and data vendors at the bottom, and regulated enterprise in between with more stability and less equity. The clearest individual step change is moving from producing labels and analysis to owning the harness and the gate.
Is AI evaluations a durable career, or will tooling automate it away?
The tooling has absorbed the plumbing of an AI evaluations engineer's work and not the judgment, and that split looks stable. What platforms now do well: running suites, storing traces, orchestrating judge calls, displaying breakdowns, managing annotation queues. What they do not do: decide what good means for your product, build a labelled set that represents your users, write a rubric two strangers apply identically, notice the failure category nobody had named, validate that the judge is not systematically wrong, or carry a result into a release decision with a recommendation attached. Those are the parts hiring managers interview on, and they get harder as systems become agentic rather than easier. The real risk to this career is not automation. It is joining a team that funds a burst of eval work before a launch and lets it decay afterward. Ask in the interview whether an eval result has ever blocked a release, because the answer tells you whether the function is real.
Put this on a resume in about a minute
Paste your history once and point it at the AI Evaluations Engineer posting you are looking at. No account, no card.
Build my resume free More roles