AI & Machine Learning

How to get hired as an AI engineer in 2026-27

The short answer

An AI engineer builds production software whose core behaviour comes from a foundation model they call rather than train, and getting hired in 2026-27 turns on one thing: a system you shipped that you can describe with numbers, covering what it does, how you measured whether it actually works, what it costs per request, and its p95 latency. No licence exists for this title and no PhD is required outside research organisations. The real bar at nearly every employer is a competent production software engineer, usually in Python and often TypeScript as well, who has built retrieval, bounded tool-calling agents and an evaluation harness, and who can read a trace to find which step produced a wrong answer. The loop is typically a recruiter screen, a hiring manager conversation, a practical build or pairing round whose AI tooling policy you must ask about, an LLM systems design session, an evals and debugging round where most offers are decided, and a deep dive on one project you claim. Resume screens reward named systems, measured outcomes and numbers with units, not lists of model names, frameworks or certificates.

What the role ownsProduction software whose core behaviour comes from a foundation model you call rather than train. You own the context that goes in (retrieval, ranking, permissions, prompt assembly), the tools the model can call, the checks on what comes out, the evals that say whether a change helped, and the cost and latency of every request.
Licence or credential requiredNone. No licence, board exam, registration or protected title gates AI engineering anywhere in the world. A degree is not a formal gate either, outside research organisations and some enterprise or government requisitions. The practical gate is a production software engineering bar plus evidence of one shipped model-backed system.
Is a PhD neededNot for AI engineer. A PhD is the norm for research scientist roles at frontier labs, which is a different job with a different interview loop. Confusing those two titles is the most common reason capable candidates never apply to roles they would pass.
Certificates and their real weightThe main ones are Microsoft Certified: Azure AI Engineer Associate (exam AI-102), the AWS machine learning and AI exams, Google Cloud Professional Machine Learning Engineer, and the Databricks and NVIDIA generative AI credentials. Exam names and codes change, so check the vendor page before paying. At product companies they rarely move a screen. At cloud-partner consultancies and systems integrators they can be a staffing or billing requirement.
Typical time to transitionFrom a backend, platform or data engineering job with real production experience: roughly three to nine months of deliberate work to ship one reviewable system with an eval suite, then interview. From a non-engineering start, the software engineering bar comes first, and that path is measured in years rather than in one bootcamp.
Typical loopRecruiter screen, hiring manager conversation, practical coding or build round (ask the AI tooling policy), LLM systems design session, evals and failure debugging round, past-project deep dive. Commonly three to six weeks at startups and longer at labs and large enterprises. Small teams sometimes compress it to two conversations plus a paid trial.
Pay: where to get a real numberThere is no US Bureau of Labor Statistics occupation code for AI engineer. The nearest published series are software developers (OES 15-1252), data scientists (15-2051) and computer and information research scientists (15-1221). For a figure that is current and about you, read posted ranges under state and EU pay-transparency rules for the specific companies you are targeting, and ask which engineering level the role maps to. Level placement moves the number far more than the title does.
What changed by 2026Model access commoditised for most product workloads, so differentiation moved to context, evaluation and integration. Agents work in bounded, verifiable loops and still fail in open-ended ones. Per-token prices fell while the tokens one user action consumes rose steeply, so cost work shifted from choosing a cheaper model to making fewer and smaller calls. Evaluation went from a nice-to-have to the central skill the interview tests.

What an AI engineer actually builds in 2026, and the titles it gets confused with

An AI engineer builds production software whose core behaviour comes from a foundation model they call rather than train. The model is an input, not the deliverable. What you own is everything around it: the context that goes into each request, the tools the model is allowed to call, the verification applied to what comes back, the evidence that a change made things better rather than merely different, and the seconds and dollars each request spends. Most of the code is ordinary backend code, including queues, retries, caching, streaming, idempotency, authorisation and schema migrations, arranged around one non-deterministic function call.

Five kinds of work recur week to week, and a resume that shows four of them reads as real. Context assembly: chunking and indexing a corpus, hybrid retrieval, reranking, carrying the source system's permissions through to query time, and templating what the model sees. Tool and agent orchestration: defining tool schemas, bounding loops with step and token budgets, handling partial failure and non-idempotent side effects, and deciding where a human has to approve. Evaluation: building a labelled set, writing scorers, validating an LLM judge against human labels, and gating deploys on a regression suite. Cost and latency engineering: token accounting, prompt and prefix caching, model routing, streaming, batching, and cutting the number of model calls rather than only their price. And production debugging: reading traces to find which step in a chain produced the wrong answer, which is a different skill from reading a stack trace.

The title is used loosely, which is why reading the posting's body matters more than reading its title. Some postings labelled AI engineer are a senior backend role with a chat feature attached. Some labelled software engineer are the real thing. Search the body text for retrieval, evaluation, agent, latency, inference and fine-tuning, and for the names of specific tools. The verbs tell you what you would actually do.

What actually gates this job: no licence, no PhD, and the two bars you must clear

Nothing legally gates the title. There is no licence, no board exam, no registration, no protected designation and no continuing-education requirement. Anyone can call themselves an AI engineer, which is exactly why employers lean so hard on demonstrated systems. In the absence of a credential, the evidence is the credential.

A degree is not a formal gate at most employers either. It is a filter in two situations worth naming precisely: research organisations hiring research scientists, where a PhD and a publication record is the norm, and large enterprises or government contractors whose requisitions carry a hard degree field. Everywhere else, including the applied and product teams inside frontier labs, the requirement reads "or equivalent practical experience" and is meant. Do not self-reject from an applied AI engineering role because you read somewhere that AI needs a doctorate. Read the posting's own words.

There are two real bars, and candidates usually fail the one they were not preparing for. The first is the software engineering bar: can you design, build, test, deploy and operate a service that other people depend on. This is the bar most people underestimate when coming from notebooks and courses, and it is non-negotiable, because a model-backed feature is still a service with a database, a deploy, an on-call rota and a security review. The second is the LLM systems bar: have you built something where you had to decide how context is retrieved, what the model is allowed to do, how you knew it worked, and what it cost. This is the bar strong backend engineers underestimate, because it is tempting to treat the model as just another API dependency. It is not. It is a dependency that is confidently wrong sometimes, costs money per call, changes behaviour between versions, has rate limits and outages you have to design around, and can be steered by text it reads out of your own database.

Certificates occupy a narrow and honest place. The Azure AI Engineer Associate exam, the AWS machine learning and AI exams, Google Cloud's Professional Machine Learning Engineer, and the Databricks and NVIDIA generative AI credentials all exist and all teach something real. At product companies they rarely change a screening outcome, because the hiring manager wants to see a system rather than a transcript. At consultancies and cloud-partner firms they can be a staffing requirement or a partner-tier obligation, which makes them worth having if that is your target. Decide on the employer type, not on the course marketing, and check the current exam name on the vendor's page because these get renamed and retired regularly.

Two context-specific gates do exist. US defence, intelligence and some federal AI work requires a security clearance, which is sponsorship-dependent and slow, taking months to years, and is a real reason a posting that looks open is not. And in the EU, the AI Act entered into force in August 2024 with obligations phasing in on a staged timetable: the prohibited-practice rules came first, general-purpose model obligations followed, and the heavier provider duties for higher-risk systems around documentation, data governance, logging and human oversight sit later in the schedule. That schedule has been subject to active amendment, so check the current position with the employer or the official text rather than trusting any article's dates, including this one. What hiring has already absorbed is the expectation that an engineer building a consequential system can describe its intended use, its known failure modes, what is logged, and where a human decides. If you have written that document once, say so.

How hiring for this role actually works in 2026-27

Two things about this market should be kept apart, because merging them produces a tidy story that is wrong. Demand for engineers who can ship model-backed features has been strong and visible since 2023 and remains so. These are among the few requisitions that survived general engineering headcount discipline, because they attach to budget somebody has already committed. At the same time the applicant pool grew enormously, both because the title became desirable and because AI-assisted applying means a single posting now draws a very large volume of competently tailored applications. Those are different forces. The first means the jobs exist. The second means the bottleneck is being distinguishable, not being qualified.

The practical consequence is that tailoring is now table stakes rather than an advantage, because almost every application looks tailored. Two things still cut through. A referral, which routes you past the volume problem entirely and is worth more effort than another twenty applications. And specificity that a general-purpose rewrite cannot fake: naming the corpus, the request volume, the eval set size, the per-request cost before and after. A bullet that says "built a RAG pipeline using LangChain and a vector database" is indistinguishable from thousands of others. A bullet in this shape, with your own real figures in place of these, is a person: retrieval over a named corpus with its size, hybrid BM25 plus embeddings with a reranker, recall at 10 measured on a labelled question set of a stated size, p95 time to first token in seconds, and dollars per resolved conversation.

Expect the first pass on your resume to be a model summarising it for a recruiter, with a human then reading the summary. That rewards plain structure, explicit nouns, and numbers with units sitting in short sentences. It punishes design-led two-column layouts that scramble into nonsense when parsed, skills buried inside paragraphs, and anything expressed only as an icon or a progress bar. Use a single-column layout and put the system names and the numbers where neither a parser nor a tired human can miss them.

The coding round has changed more than any other stage, and the change is uneven enough that you must ask about it rather than assume. Because current models solve typical algorithm puzzles instantly, many teams moved to a practical build: extend a small existing repo, add a feature to a working agent, fix a failing retrieval path. Sometimes AI coding tools are explicitly permitted and your usage of them is observed. Sometimes they are explicitly banned and the environment is proctored. Both exist at serious companies right now. Ask the recruiter directly what the tooling policy is for the technical round. Nobody will hold the question against you, and guessing wrong in either direction is fatal.

Here is the loop, stage by stage. Smaller companies drop stages rather than inventing new ones, and labs add a round and a bar raiser.

What AI engineers are paid, and what moves the number

Anyone quoting you a single band for AI engineer is averaging across a frontier lab, a seed-stage startup, a bank's platform team and a consultancy, which are four different jobs. Get your own figures instead. The US Bureau of Labor Statistics does not publish a separate occupation for AI engineer. The nearest series are software developers under OES code 15-1252, data scientists under 15-2051, and computer and information research scientists under 15-1221, each with a national median and a tenth to ninetieth percentile spread by state and metropolitan area. Use those for the shape of the market, not for your offer.

For current, company-specific numbers, read the posted range. Pay-transparency rules in California, Colorado, New York, Washington, Illinois, Minnesota, Maryland, New Jersey, Vermont, Massachusetts and others require ranges on the posting, and EU employers are moving the same way under the Pay Transparency Directive. Self-reported aggregators such as levels.fyi carry company-attributed bands and are useful for calibrating the step between levels, with the caveat that the data is submitted rather than audited. The posted ranges for the actual companies you are targeting are the only figures that are both current and about you.

Five things move the number more than the title does.

The resume: the system sentence, and what gets ignored

The unit of credibility on an AI engineer resume is not a tool and not a responsibility. It is a system described with numbers. Write each one as a single compact sentence that answers five things: what it does, who uses it and at what volume, how you measured whether it works, what it costs and how fast it is, and what you changed to get there. Two or three of those sentences with the specifics intact outperform a page of responsibilities and a forty-item skills grid.

Numbers need units and denominators or they read as decoration. "Improved accuracy by 23%" with no eval set behind it is a liability, because it invites the one question you cannot answer. A claim in this shape invites a conversation you will enjoy instead: pass rate on a labelled set of a stated size rose from one figure to another, and the remaining failures are mostly one named category that you chose not to support. Latency belongs in percentiles, p50 and p95, and for streaming interfaces time to first token, which is what the user actually feels. Cost belongs per unit of work: per thousand requests, per conversation, per document processed, per resolved ticket.

Keep the document structurally plain. One column. Standard section headings. Dates in a consistent format. Two pages is fine with eight or more years of experience and one is better below that. Put a short stack block at the bottom for the keyword screen and do not let it become the resume. Link one deployed thing and one repository, and check that both still work the week you apply, because a dead link is worse than no link.

The evals round, where the offer is won or lost

If one interview separates candidates with the same stack on paper, it is the one about evaluation. The reason is structural: a model-backed system cannot be tested the way ordinary software is tested, so the engineer's ability to establish whether a change improved things is the whole quality mechanism. Teams have learned this the expensive way, shipping prompt changes that felt better and quietly regressed a category of users, and they now interview for it directly.

What is being graded is whether you have actually built a measurement loop, and that is easy to detect. Candidates who have will reach for a labelled set, a scorer, a baseline and a gate without being prompted. Candidates who have not will say "we tested it with some examples" and then, pressed, "we kept an eye on user feedback".

A strong answer has a recognisable shape. Start from the failure modes rather than the metric, listing the ways the system is wrong: wrong retrieval, correct retrieval but hallucinated synthesis, refusing something it should answer, right answer in an unusable format, tool called with bad arguments, loop that never terminates. Build a set of real cases drawn from production traffic rather than invented, deliberately over-sampling the failure categories. A few dozen well-chosen cases is enough to start and a few hundred is enough to gate on, and both beat tens of thousands of easy ones. Keep it as a flat file in version control, one case per line, with the input, the expected behaviour and a category label, so a diff shows what changed. Score each case the cheapest way that is valid: exact match or a schema check where the answer is structured, a deterministic check where you can write one, and a model as judge only where you must. When you use a judge, validate it against human labels on a subset and report the agreement, because an unvalidated judge is a metric that measures itself. Then wire the suite into CI as a gate on prompt, model and retrieval changes, with the per-category breakdown visible, because a flat aggregate hides the regression that matters. And keep a small human review queue regardless, because your eval set only contains failures you already know about.

Expect follow-ups about the awkward parts, and have real answers. How do you test a stochastic system in CI without a flaky build. What do you do when the vendor ships a new model version and your pass rate moves two points, and is that signal. How do you evaluate a multi-step agent where the final answer is right but the path was wasteful, or wrong only on step three of seven. How do you measure retrieval separately from generation, so you know which half to fix. How do you handle the fact that your labelled set ages while your production distribution drifts away from it.

Two specifics are worth having ready, because they come up constantly. First, evaluate retrieval on its own with recall at k against known-relevant documents, before any generation is involved. Most reported hallucination in internal assistants is actually a retrieval miss, and candidates who conflate the two cannot debug either. Second, be able to describe tracing: spans per step, inputs and outputs captured, token counts and latency per call, and a way to pull up the exact trace behind one user complaint. Several tools do this, including Langfuse, LangSmith, Braintrust, Arize Phoenix and OpenTelemetry's GenAI conventions, and naming the one you used is fine. Naming none while claiming you debugged a production agent is not believable.

The systems design session, prompt by prompt

The design round is usually one broad prompt, such as design an assistant over our documentation and support history, or design an agent that handles a class of tickets end to end, followed by an hour of follow-ups. You are graded on the questions you ask before you draw anything, on whether you name trade-offs and the conditions under which you would reverse a decision, and on whether evaluation, permissions and cost arrive without being dragged out of you. The usual tell of inexperience is drawing the happy path confidently and quickly.

Ask first. Who are the users, and are they internal or external, which is really a question about blast radius. What is the corpus, how big, in what formats, how often does it change, and who is allowed to see which parts. What does a wrong answer cost: mild embarrassment, a bad refund, or a clinical or financial decision. What latency is acceptable, and is the interface streaming. Is there a budget per request or per month. Does the answer need a citation, and does the system need to be able to say it does not know. Those seven answers change the design completely, and asking them is half the grade.

On retrieval, the depth expected in 2026 is specific. Chunking is a decision with reasons: structure-aware splitting on headings or clauses, chunk size tuned against your own recall measurements, overlap, and a parent-document or late-chunking strategy when the useful unit for the model is bigger than the useful unit for search. Retrieval should be hybrid, because lexical BM25 catches acronyms, part numbers, error codes and names that embeddings reliably smear, while vectors catch paraphrase. A reranker over a wider candidate set buys more accuracy per dollar than almost any other single change. Query rewriting and decomposition matter for real user phrasing. Metadata filters on product, version, region and date prevent the most common support failure, which is a confident answer from documentation three releases out of date. And access control has to reach query time: permissions live in the source system, your index is a copy, and nothing propagates by itself. Store group identifiers or ACLs per chunk, filter at search, re-sync when source permissions change, and be able to say which corpus you deliberately excluded because its permission model could not be honoured. The exclusion is the part that sounds lived. On long context, say the true thing: large context windows reduced how much retrieval plumbing a prototype needs, and did not remove retrieval from production, because stuffing a corpus into every request is expensive, slower, and still degrades when the relevant passage is buried among many distractors.

On agents, the mark of someone who has shipped one is that they talk about bounds and failure rather than autonomy. Tools get narrow, typed schemas with validation on the way in and the way out, because the model will eventually pass a plausible wrong argument. Every loop gets a step limit, a token budget and a wall-clock timeout. Side effects get idempotency keys, because the retry is coming. Anything irreversible, such as money moving, an email sending or a record deleting, gets either a confirmation step or a reversible staging area. State is explicit and inspectable, so a failed run can be resumed or replayed rather than restarted. And the architecture stays as flat as the problem allows, because a single well-instrumented loop with good tools beats a committee of specialised agents in most real deployments: every handoff adds latency, cost and a new place for context to be lost. If you propose a multi-agent topology, be ready to justify it against a flat one. "It was in the framework" is not a justification.

On cost and latency, the shape of the problem changed. Per-token prices fell substantially while agentic patterns multiplied the tokens a single user action consumes, so the lever moved from choosing a cheaper model to making fewer and smaller calls. Know the moves: prompt and prefix caching, which rewards putting the stable material at the front of the prompt and is the highest-leverage change in most chat systems; routing, where a small model handles or classifies the easy majority and escalates the rest, with the escalation rate measured; cutting chain length, since each hop adds a full round trip; streaming, so perceived latency drops even when total time does not; parallel tool calls where the calls are independent; batch APIs for anything that does not need to be synchronous; and trimming retrieved context, since more passages stops being more accuracy past a point you can find by measurement. Also know your failure path: rate limits, quota exhaustion and provider outages are operational events, so say what degrades, what queues and what falls back.

On security, the one framing worth internalising is that an agent becomes dangerous when three things meet: access to private data, exposure to text it did not write, and a way to send information outward. Retrieved documents, web pages, tool results, ticket contents and uploaded files are all untrusted input, and the model will treat it as instruction if it reads like instruction. Defences are layered and none is complete: scope tool permissions to the minimum and per-user rather than per-service, require human confirmation on consequential actions, treat retrieved text as data rather than instructions in your prompt structure, constrain outbound channels and strip or allowlist URLs and markdown images that can carry data out, validate tool arguments against a schema, log everything, and assume any filter you add will eventually be bypassed. Say explicitly that you cannot fully solve prompt injection with a prompt. Interviewers in regulated environments are listening for exactly that sentence. If data residency or confidentiality rules out a hosted provider, be ready to discuss open-weight models served in your own environment and what that costs you in capability and operations.

The portfolio, the transition paths, and where these jobs are posted

One deployed system with an evaluation suite beats five notebook tutorials, and it is not close. A reviewer gives your repository minutes, not hours. What earns the next few minutes: a README that states what the system does, what it is measured on and where it fails; an evals directory with a real labelled set and the scorer; a trace or screenshot of a failure you diagnosed, with the diagnosis written down; a note on cost and latency per request with the numbers; and a deployed URL that works for a stranger who has no API key of their own. The failure analysis is the single most differentiating file in the repository, because almost nobody includes one.

Pick a problem with a messy, permissioned, real corpus rather than a clean one. Public filings, municipal meeting minutes, a hobby community's twenty years of forum archives, scanned manuals for discontinued equipment, or your own employer's documentation if you can get permission. The difficulty is the point, because clean corpora make every design decision look equally good, which is why a thousand identical chat-with-a-PDF projects teach the builder nothing and tell the reviewer nothing. Then go get ten real users: post it in the community whose corpus you used, or hand it to the team inside your company that owns the documents. A week of other people's traffic gives you the one thing most candidates lack entirely, which is production failures you caused and fixed.

Transition paths differ in what they already clear. From backend or platform engineering, you have the software bar, and what you need is one shipped model-backed system with evals plus the willingness to learn that retrieval quality is an empirical discipline rather than a configuration choice. This is the shortest and most common path in, and it runs through backend engineering far more often than through research. From data engineering, you have the corpus, pipeline and permissions instincts, and what you need is service ownership, API design and the online serving mindset. From data science or research, you have the measurement instincts, which is a real advantage in the evals round, and what you need is production software practice: tests, deploys, on-call, code someone else maintains. Expect that to be the bar you are actually assessed against. From machine learning engineering you are close, so emphasise serving, latency and evaluation over training pipelines, and resist leading with fine-tuning when the posting is about retrieval. From frontend there is a real and underrated niche in agent and assistant interfaces: streaming, interruption, showing tool use legibly, letting a user correct a wrong step. From support, operations or a domain role, the path runs through becoming the person who automates your own team's work with real systems and then transferring internally, which is slower but starts from a problem you understand better than any outside hire.

On where to look, the title is an unreliable index and the body text is a good one. Search job boards for retrieval, RAG, evaluation, agent, LLM, inference and vLLM inside the description rather than filtering on the title field, because a large number of genuine AI engineering jobs are posted as software engineer, platform engineer or machine learning engineer, since that is what the requisition system allows.

Two channels outperform applications and almost nobody works them properly. The first is contributing to the open-source projects in this space: retrieval libraries, eval frameworks, inference servers, agent runtimes, MCP servers. A merged pull request that fixes something real is a work sample, a referral and a conversation starter in one, and maintainers get hired by the companies that depend on their work. The second is writing up one thing you learned in enough technical detail that it is useful to someone doing the same work: a measured comparison, a failure post-mortem, a benchmark you ran honestly including the result that was inconvenient for your thesis. You are not building an audience. You are creating one artefact a hiring manager can read that proves you think in evidence.

Finally, apply narrower than feels comfortable. Twenty applications with a specific first line referencing something real about that company's product, aimed at teams whose posting body describes work you have actually done, beat two hundred generic ones, because the generic two hundred compete against an effectively unlimited supply of generated applications while the twenty compete against the few people who also bothered.

Working with AI in this role

What an AI engineer has to know about AI in 2026-27

This is the one role where "what do you need to know about AI" cannot be answered with "learn to use the tools", because the tools are the job. So the useful question is narrower: what actually changed between the 2023 version of this role and the 2026 version, and what will an interviewer assume you have internalised.

The headline shift is that the model layer commoditised at the API boundary for most product workloads. Several providers ship capable models at similar prices, swapping between them is a configuration change plus a re-run of your eval suite, and raw capability stopped being a defensible advantage for anyone who is not training the models. Frontier capability still separates providers on the hardest tasks, which is why the eval re-run matters rather than being a formality, but for the majority of product features the choice is no longer the interesting decision. Differentiation moved to three places: the context you can assemble, which depends on your data and permissions; the evaluation loop that lets you improve deliberately rather than by vibes; and the integration surface, meaning the unglamorous work of making the thing fit an existing product, an existing identity system and an existing support process. Interviews followed that shift. You will be asked far less about model internals than you expect and far more about retrieval quality, measurement and failure handling.

The second shift is that agents moved from demo to narrow production, and the boundary is sharper than the discourse admits. Agentic loops work well where the task has a verifier: code that either compiles and passes tests or does not, an extraction that can be checked against a schema or a source document, a triage decision a human confirms, a search that can be re-run and compared. They still fail in open-ended multi-step work with no ground truth available along the way, where errors compound silently across steps. The honest answer to "are agents working now" is that in bounded, verifiable loops with budgets and human gates they are, and companies are getting real value from them, while in open-ended autonomy they are not reliable, and the teams claiming otherwise are usually not measuring. Saying that clearly in an interview reads as experience. Enthusiasm without the boundary reads as a blog reader.

The third shift is economic and counterintuitive. Per-token prices fell sharply while the tokens consumed per user action rose steeply, because agentic patterns make many calls and carry large contexts. So cost engineering did not go away when models got cheap. It moved from picking the cheap model to making fewer calls, caching the stable prefix, routing the easy majority to a small model, and shortening the chain. Alongside that, the binding constraint on user experience is now latency and reliability more often than quality. A correct answer in nine seconds loses to a good answer in two.

What has not changed is the engineering judgment, and a candidate who says so credibly stands out. Models now write plausible code quickly, including the plumbing of a retrieval pipeline or an agent loop. They will not tell you which of two sources is authoritative when they disagree, what the acceptable failure rate is for your users, which irreversible action needs a human gate, or whether the measured improvement you are about to ship is real. What shifted is the ratio: generated code arrives faster than review capacity grew, so the ability to verify, constrain and test what lands matters more than the ability to author it. Employers now ask how you use coding agents, not as a trap, but because an engineer who cannot use them is slow and an engineer who ships their output unreviewed is dangerous. Have a specific, unembarrassed answer about where you let them drive, where you do not, and how you review what they produce.

Retrieval as an empirical discipline, measured rather than configured

Most reported hallucination in internal assistants is a retrieval miss: the model was never shown the right passage. Teams that treat chunking, hybrid search and reranking as settings to pick rather than parameters to measure plateau at mediocre and cannot explain why. This is the first thing a competent interviewer probes, because it separates people who built something from people who assembled a tutorial.

Show it: Describe the measurement, not the stack. A labelled set of real questions with known-relevant documents, recall at k before and after, and the specific change that moved it: structure-aware chunking, adding BM25 alongside embeddings so acronyms and part numbers stop smearing, a reranker over a wider candidate set, a metadata filter that stopped answers coming from a deprecated release. Name the one change that did nothing despite being popular advice.

Evaluation infrastructure: labelled sets, validated judges, CI gates

A model-backed system has no unit test for correctness, so the eval suite is the entire quality mechanism. Teams have shipped prompt changes that felt better and silently regressed a whole category of users. This is the most reliable discriminator between candidates with identical stacks, and the round where most offers are decided.

Show it: Give the set size and its provenance from production traffic, the mix of deterministic scorers and model judges, and the agreement you measured between your judge and human labels. Say that the suite gates deploys and that results are broken out per failure category rather than reported as one aggregate. Then name a case you removed from the set and why, because that detail cannot be faked.

Agent loops with real bounds: budgets, idempotency, and reversibility

The difference between a demo agent and a production one is almost entirely failure handling. Retries duplicate side effects, loops fail to terminate, a plausible but wrong tool argument gets accepted, and an irreversible action fires on a hallucinated premise. Interviewers who have operated an agent ask about exactly these. Candidates who have not talk about autonomy.

Show it: Name the bounds you set, including maximum steps, token budget, wall-clock timeout and a per-session spend alert, and the idempotency mechanism on anything with a side effect. Describe one irreversible action you put behind a human confirmation or a reversible staging area, and one run that failed halfway and was resumed from explicit state rather than restarted.

Knowing when a flat loop beats a multi-agent topology

Multi-agent architectures spread faster than the evidence for them. Every handoff adds latency, token cost and a place for context to be lost, and plenty of production systems that started as five specialised agents ended as one well-instrumented loop with better tools. An engineer who can say that, with a reason, signals judgment. One who proposes an org chart of agents by default signals framework-following.

Show it: Describe a case where you collapsed a multi-step topology into a single loop, or deliberately kept one boundary and can say why, which is usually a genuinely different tool permission scope, a different model, or a verification step that must not see the generator's reasoning. Quote the latency or cost difference if you measured it.

Cost and latency as engineered properties, in units

Model-backed features are the rare software where each request has a visible marginal cost, and agentic patterns multiplied the tokens per user action even as prices fell. Teams get surprised by the bill or by a nine-second response, and the engineer who can name the levers and the measurements is the one trusted with the production system.

Show it: Report per-unit figures, such as dollars per thousand requests or per resolved case, p50 and p95, and time to first token, with the intervention behind each improvement: prefix caching with the stable material moved to the front of the prompt, a small model handling the easy majority with a measured escalation rate, a four-call chain cut to two, streaming, parallel independent tool calls, and batch processing for anything asynchronous.

Prompt injection and the private-data, untrusted-text, outbound-channel trio

Retrieved documents, web pages, tool results and uploaded files are text your system did not write, and the model will follow instructions found in them. Once private data, untrusted content and a way to send information outward meet in one agent, the system is exfiltratable. Security reviewers in regulated industries hold a veto and this is what they ask about.

Show it: Describe layered defences and name their limits: per-user least-privilege tool scopes rather than a service account with everything, human confirmation on consequential actions, retrieved text structurally separated as data, outbound URLs and markdown images allowlisted or stripped, schema validation on tool arguments, full logging. Then say plainly that you cannot solve injection with a prompt, because that sentence is the thing being listened for.

Observability for non-deterministic systems: traces, not logs

When a user complains about one answer, you need the exact chain that produced it: which documents were retrieved, what the prompt actually contained after assembly, which tools were called with what arguments, and how many tokens and milliseconds each step took. Without that, debugging is guesswork, and "we looked at the logs" is not an answer to a failure in a seven-step chain.

Show it: Say how you traced it and what you could pull up from a single complaint. Naming a specific tool is fine and concrete, whether Langfuse, LangSmith, Braintrust, Arize Phoenix or OpenTelemetry's GenAI semantic conventions. Then walk one real diagnosis: the symptom, the span where it went wrong, and the fix.

Tool and context interoperability, including MCP and its authorisation problem

Connecting models to systems converged on a common protocol layer, and the Model Context Protocol is the one most employers now name. The interesting part is not the wire format, which is easy, but identity and authorisation: whose permissions does a tool call run as, how is a token scoped and rotated, and what stops a server you do not control from reading more than it should.

Show it: Describe a tool or server you built or integrated, the scope of credentials it ran with, and whether calls executed as the end user or as a service identity, plus what you did about the gap. If you evaluated a third-party server and rejected it on permission or supply-chain grounds, that decision is worth a sentence.

An honest, specific position on fine-tuning versus everything cheaper

Fine-tuning is the most over-indexed skill on AI engineer resumes relative to how often teams do it. Most problems presented as "the model does not know our domain" are retrieval or prompt-structure problems, and most presented as "the output format is wrong" are solved with constrained decoding or a schema. Interviewers use the question to test whether you reach for the expensive answer first.

Show it: State the ladder you actually work through: prompt and context structure, retrieval quality, few-shot examples, constrained or structured output, then distillation of a large model's behaviour into a small one for cost or latency, and only then full fine-tuning. Name the condition under which you would climb it. If you have fine-tuned, give the reason it beat the cheaper options and the eval that proved it.

Using coding agents well, and being able to say how

By 2026 this is asked directly in most engineering interviews, including for this role. The failure modes are symmetrical: an engineer who refuses the tools is slow, and one who ships generated code without review is a liability. Employers are trying to find out which you are, and an evasive answer reads as the second.

Show it: Be specific about the division. Where you let an agent drive, such as scaffolding, test generation, mechanical refactors, unfamiliar API exploration and migration drudgery. Where you do not, such as the schema, the security boundary, and anything whose failure mode is silent. And how you review: tests first, reading the diff rather than the summary, and at least one example of generated code you rejected and why.

What a screen is looking for

These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.

Mistakes that cost people this job

Submitting a portfolio of notebooks and demo scripts with no running service, no tests and no deploy. It reads as someone who has studied the model layer and never carried a system.

Ship one small system end to end and let the reviewer use it: a URL that works for a stranger, a repository with an evals directory, a cost-and-latency note with real numbers, and a written failure analysis. One deployed thing with measurement beats five tutorials, and the failure analysis is the file almost nobody else includes.

Claiming an improvement with no eval set behind it, such as "increased accuracy 30%" or "reduced hallucinations significantly". It invites the one follow-up you cannot survive, and interviewers ask it every time.

Give the denominator and the residual: pass rate on a labelled set of a stated size, built from real questions, moved from one figure to another, and the remaining failures are mostly one named category you chose not to support. If you genuinely had no eval set, say that and say what you would build now. Honesty recovers, a bluff does not.

Leading with fine-tuning, model training or Kaggle results when the posting is about retrieval, agents and shipping. These are credible signals for machine learning engineer or research roles and they displace the evidence this screen wants.

Read the posting's verbs. Build, ship, integrate, latency, evaluate, agent and retrieval means applied AI engineering, so lead with systems and measurement. Train, model, dataset, experiment and offline metric means ML engineering, so lead with pipelines and modelling. The same career supports both resumes, but not the same document.

Proposing a multi-agent architecture by default because the framework has one, and being unable to defend it against a single loop with good tools.

Start flat and justify every boundary you add: a genuinely different permission scope, a different model, or a verifier that must not see the generator's reasoning. If you have collapsed a multi-agent design into one loop and measured the latency and cost difference, lead with that story, because it is a strong judgment signal.

Walking into the technical round without asking the AI tooling policy, then either being caught using a coding agent where it was banned or freezing in a proctored environment because you have not written code unassisted in a year.

Ask the recruiter directly whether tools are permitted, observed or prohibited, then practise in that mode. Nobody penalises the question. Also prepare a concrete, unembarrassed answer to how you use coding agents, covering where you let them drive, where you do not, and how you review the diff.

Treating the model as just another HTTP dependency. Strong backend engineers do this and it surfaces immediately in design: no eval story, no cost ceiling, no handling of a confidently wrong response, no thought about text the model reads out of your own database.

Name the ways this dependency is unlike others and what you did about each. It is non-deterministic, so you need an eval suite rather than assertions. It costs money per call, so you need a per-request budget. It changes under you when a version ships, so you re-run evals against the candidate. It rate-limits and has outages, so you need a degradation path. And it follows instructions found in untrusted input, so tool scopes are narrow and consequential actions are gated.

Skipping permissions in the design round. Drawing an index box over "all company documents" and moving on is the most common unforced error in the systems session, and in regulated industries it ends the interview.

Say early that permissions live in the source system and the index is a copy, then describe the mechanism: ACLs or group identifiers stored per chunk, filtered at query time, re-synced when source permissions change. Name one corpus you excluded because its permission model could not be honoured.

Applying at volume with a generically tailored application. Because almost every application is now competently tailored, the generic hundreds compete against an effectively unlimited supply and convert at close to nothing.

Apply narrower and deeper: twenty applications aimed at teams whose posting body describes work you have actually done, each opening with one specific true sentence about their product or problem. Spend the recovered time on referrals, an open-source pull request in a project they depend on, and one honest technical write-up.

Self-rejecting from AI engineer postings because you do not have a PhD or a research background, and applying only to junior or AI-adjacent roles instead.

Separate the two job families. Research scientist at a lab is usually PhD-gated. AI engineer is gated on production software skill plus one shipped model-backed system. If the posting says "or equivalent practical experience", it means it. Apply to the level your software experience supports, not the level your AI experience feels like.

Being unable to name what your system does badly. Candidates present a flawless summary, and the interviewer concludes either that it was never under real load or that you were not close enough to it to know.

Bring one failure you understand deeply and one you chose not to fix, with the reasoning: frequency, severity, cost of the fix, and what you shipped instead. This single prepared answer does more for a senior-level read than any additional item on the stack list.

Quoting a salary expectation from an aggregated AI engineer average without knowing which engineering level the role maps to.

Ask in the first or second conversation which level the role is and what the band for that level is, then read the posted range under pay-transparency rules for the specific companies you are targeting. Level placement moves your number far more than negotiation does, and a startup equity grant needs the strike price, the total preferred raised and the percentage before it means anything.

Questions people ask

What does an AI engineer actually do?

An AI engineer builds production software whose core behaviour comes from a foundation model they call rather than train. The day-to-day work is five things: assembling the context the model sees, covering retrieval, ranking, permissions and prompt structure; defining and bounding the tools it can call; building the evaluation suite that says whether a change helped; engineering cost and latency down to a per-request budget; and debugging production failures from traces. Most of the code is ordinary backend code, including queues, retries, caching, streaming and authorisation, arranged around one non-deterministic function. A useful working test: if your main artefact is a trained model you are doing machine learning engineering, and if it is a running service whose quality you can measure you are doing AI engineering.

Do you need a PhD to become an AI engineer?

No. A PhD is the norm for research scientist roles at frontier labs, which is a different job with a different interview loop, but applied AI engineering is not degree-gated at most employers, where postings say "or equivalent practical experience" and mean it. The actual gate is two bars. The software engineering bar: you have designed, deployed and operated a service other people depend on. And the LLM systems bar: you have shipped something where you decided how context was retrieved, what the model could do, how you measured it, and what it cost. Confusing the research role with the engineering role is the most common reason capable candidates never apply.

What is the difference between an AI engineer and a machine learning engineer?

A machine learning engineer trains, serves and monitors models, owning feature pipelines, training runs, offline to online metric parity and drift, and the model is the deliverable. An AI engineer treats an existing foundation model as an input and owns the system around it: retrieval, tools, evaluation, guardrails, cost and latency. The two overlap at the serving boundary and diverge in everything else, because ML engineering leans on statistics and data engineering while AI engineering leans on distributed systems and product judgment. The common route into AI engineering runs through backend or platform engineering rather than through research.

What skills get an AI engineer past resume screening in 2026?

Production software skill in Python, often with TypeScript, plus demonstrated depth in four areas. Retrieval measured empirically: chunking, hybrid lexical plus vector search, reranking, recall at k, document-level permissions. Evaluation infrastructure: a labelled set from real traffic, scorers, a validated model judge, a CI gate. Bounded agent engineering: tool schemas, step and token budgets, idempotency, human gates on irreversible actions. And cost and latency work expressed in units: tokens per request, p95, time to first token, dollars per resolved case. What does not get an AI engineer past screening is a list of model names, "prompt engineering" as a standalone bullet, and certificate stacks at the top of the page.

What does the AI engineer interview process look like?

An AI engineer loop is typically five to six stages over three to six weeks: a recruiter screen matching vocabulary and logistics; a hiring manager conversation about judgment and whether your problems resemble theirs; a practical coding or build round, such as extending a repo, fixing a retrieval path or making a flaky agent testable, where the AI tooling policy varies by company and must be asked about; an LLM systems design session on a shared canvas; an evals and debugging round where you are shown a failure or asked how you would know a change helped; and a past-project deep dive where a staff engineer pressure-tests one resume claim. Small teams compress this to two conversations plus a paid trial. Frontier labs add a round and raise the software bar.

What do AI engineers get paid?

There is no US Bureau of Labor Statistics occupation code for AI engineer, so the nearest published series are software developers (OES 15-1252), data scientists (15-2051) and computer and information research scientists (15-1221), each with medians and percentile spreads by state and metro area. For a usable number, read the posted range: pay-transparency laws in California, Colorado, New York, Washington, Illinois, Minnesota, Massachusetts and several other states require it, and EU employers are moving the same way under the Pay Transparency Directive. The most important question is which engineering level the role maps to, because AI engineer is almost never a separate pay scale. It is a software engineering level with a specialism attached, and level placement moves the number more than negotiation does.

Are AI certifications worth it for an AI engineer job?

It depends entirely on who is hiring. At product companies and startups, certificates such as Microsoft's Azure AI Engineer Associate (AI-102), the AWS machine learning exams or Google Cloud's Professional Machine Learning Engineer rarely change a screening outcome, because the hiring manager is looking for a system you shipped. At consultancies, systems integrators and cloud-partner firms they can be a staffing or billing requirement, which makes them genuinely useful if that is your target. Nothing legally requires any certification, because there is no licence for this role. If you have limited time, build and deploy one measured system instead, since that outperforms any certificate at the employers that pay the most.

How do I show LLM work on my resume if I have only built side projects?

Describe the side project the way you would describe production work, with real numbers, and make it inspectable. Name the corpus and its size and messiness, the retrieval approach and the recall you measured, the eval set size and where the cases came from, the per-request cost and p95 latency, and one failure you diagnosed with the diagnosis written down. Deploy it so a stranger can use it without their own API key, and keep the link alive while you are applying. A side project with a labelled eval set and a failure analysis reads as more credible than a professional bullet that says "built a RAG pipeline" with no measurement attached.

Are AI agents actually working in production in 2026, or is it hype?

Both, and the boundary is specific enough to state. Agentic loops are delivering real value where the task has a verifier along the way: code that compiles and passes tests, extraction checkable against a schema or source document, triage a human confirms, retrieval that can be re-run and compared. They remain unreliable in open-ended multi-step work with no ground truth available mid-run, where small errors compound silently. The engineering that makes the difference is unglamorous: step and token budgets, idempotency on side effects, reversible staging for consequential actions, explicit resumable state, and tracing. Saying this clearly in an interview reads as experience, while unbounded enthusiasm reads as having read about it.

What is the fastest realistic path into an AI engineer job from a backend role?

Roughly three to nine months of deliberate work, because a backend engineer already clears the harder bar. Pick a problem with a messy, permissioned real corpus rather than a clean one, build and deploy a service around it, then spend most of your effort on the part nobody else does: a labelled eval set built from real questions, scorers, a CI gate, tracing, measured recall and measured per-request cost, and a written analysis of what still fails. Get ten real users if you can, so you have production failures you caused and fixed. Then apply narrowly, searching job descriptions for retrieval, evaluation, agent and inference rather than filtering on the title, and work referrals and one open-source contribution in parallel.

Put this on a resume in about a minute

Paste your history once and point it at the AI Engineer posting you are looking at. No account, no card.

Build my resume free More roles