AI & Machine Learning

How to get hired as a machine learning engineer in 2026-27

The short answer

To get hired as a machine learning engineer in 2026-27, show one model or model-powered system you put in production, the evaluation harness you built to know whether it was working, and the latency and cost numbers it ran at. The title has shifted: most ML engineering jobs now build systems around models somebody else trained, covering retrieval, evaluation, serving, agents and tool use and selective fine-tuning, rather than training architectures from scratch. The loop still opens with a coding round, which at large technology companies is a standard software-engineering round at medium to medium-hard difficulty and eliminates more ML candidates than anything else, but ML system design and an evaluation-focused round are where the offer is decided. Candidates lose on the same two things: no evidence anything reached real users, and no way to say how they knew the system was good.

What the role is in 2026Building, shipping and operating systems that make predictions or generate text, where the model is one component among retrieval, features, evaluation, serving, guardrails and cost control. At most companies the model is bought or downloaded; the engineering is everything around it. Training a novel architecture from scratch is a small minority of these jobs.
Closest confusionsAn AI engineer builds LLM product features and rarely trains anything. An applied scientist is measured on a modelling result. A data scientist is measured on an answer or a decision, not a running service. An MLOps or platform engineer owns the infrastructure other people's models run on. ML engineer sits in the middle: owns a model's behaviour and the system it runs inside.
Typical loopRecruiter screen, hiring-manager call, a coding round, an ML system design round of 45 to 60 minutes, an evaluation or modelling-depth round, and a deep dive on a project you shipped. Four to eight weeks. A take-home of four to eight hours appears mostly at startups and mid-size companies; large employers favour live rounds.
Entry requirementsNo licence and no mandatory certification. A bachelor's in computer science, maths, statistics or a quantitative engineering field is the common floor; a master's is common and helps most at research-adjacent employers. A PhD is the norm for research scientist and is usually preferred for applied scientist, though several large employers hire applied scientists at master's level with a strong applied record. It is not required for machine learning engineer.
PayNo single band, and the US Bureau of Labor Statistics has no SOC code for 'machine learning engineer'. Bracket it with OES 15-2051 (data scientists) and 15-1252 (software developers), read the metropolitan-area tables rather than the national median, then use levels.fyi for the large-technology ladder and the ranges employers publish under state pay-transparency laws in Colorado, California, Washington, New York, Illinois and a growing list of other states. Which of the five job types you are in moves pay further than the title does.
Resume lengthOne page for under five years of experience, two pages beyond that. A model with no outcome attached takes as much space as a model with one and is worth far less: cut the project list before you cut the numbers.
Evidence that landsA model in production with a measured business or product outcome; an eval harness with the dataset size and what it gated; p50 and p95 latency; cost per thousand requests or per resolved task; training-data volume and how labels were obtained; what you decided not to do and why.
What changed by 2026Pretraining consolidated into a handful of labs, so the leverage moved to evaluation, retrieval quality, inference economics and agent reliability. Fine-tuning became cheap enough that the skill is now knowing when not to do it. Post-training and reinforcement learning against a company's own eval suite emerged as a hiring category of its own. Classical tabular and ranking ML did not go away; in most industries it is still the majority of models in production.

What a machine learning engineer actually does in 2026, and the adjacent titles people apply to by mistake

A machine learning engineer owns a model's behaviour in production and the system that makes that behaviour useful. That means the data that goes in, the features or the retrieval, the choice of model, the evaluation that says whether it works, the serving path it runs on, the guardrails around its failures, and the latency and cost it runs at. The unit of work is a system that users or another service depend on, not a notebook result.

The title moved hard between 2022 and 2026, and the market has not finished relabelling. The old centre of gravity, design an architecture, train it on your own data, tune it, publish or ship it, is now a minority of postings, concentrated in a small number of labs and in domains where no general model exists. The new centre of gravity is building around models somebody else trained: retrieval over your own corpus, an evaluation harness you can trust, serving that holds a latency budget, fine-tuning only where the case for it is made, agents that call tools without looping forever, and an inference bill that does not grow faster than usage.

That shift matters for applications because the same posting title now covers two quite different jobs, and because several adjacent titles sit close enough that candidates routinely apply to the wrong one, get rejected, and learn nothing from it. The distinction between them is not seniority. It is what you are measured on.

The five kinds of ML engineer job: read the posting's nouns

Before you tailor a single bullet, decide which of five jobs the posting describes. They want different evidence, interview differently, and pay differently. The nouns in the first three responsibilities tell you which one it is.

Most candidates have genuine experience in one and a reading knowledge of the others. The failure mode is a single generic resume that leads with whichever project is most recent, so a recommendations team reads a RAG chatbot and a platform team reads a Kaggle notebook. The same career supports more than one of these resumes. It does not support the same document.

How the hiring loop works in 2026-27, round by round

The structure stabilised after the 2023-24 churn, and it is now fairly predictable across company sizes. What changed is the weighting: two rounds, ML system design and evaluation depth, became where the offer is won or lost. The coding round did not get easier. It stopped being where the decision is made while remaining the place most candidates are eliminated.

Expect four to eight weeks end to end. Large employers run more rounds and schedule them further apart; startups compress to one or two days and lean harder on a take-home or a working session. Where a company uses a hiring committee, the written feedback from the system design and project deep-dive rounds is what the committee actually reads.

Round by round, and what each one is really testing:

The ML system design round: what is actually being graded

This round is not a trivia test and it is not a software system design interview with a model bolted on. It is a test of whether you reason from the problem to the system, in the order a practitioner does: what decision is being made, what the cost of each kind of mistake is, what data exists, what the simplest thing that could work is, how you would evaluate it, then how you would serve it and what it would cost.

The single most common failure is starting at the architecture. A candidate hears 'design fraud detection' and begins with a model family and a feature store diagram, never having asked what happens when the system is wrong, whether a false positive blocks a customer's rent payment or merely queues a review. Written feedback on this round is overwhelmingly about framing rather than diagram quality: a candidate who spends the first five minutes on the decision usually outscores one who draws a better architecture.

A workable shape for 45 minutes, which also works as the structure of an answer in a take-home write-up:

The evaluation round, the one most candidates are unprepared for

Evaluation moved from a chapter at the end of the project to the centre of the job, and hiring followed. Many loops now dedicate a full round to it, and even where they do not, the evaluation questions inside the system design round carry disproportionate weight. The reason is practical: when the model is a commodity someone else trained, your only durable advantage is knowing, faster and more precisely than a competitor, whether a change made things better.

Questions in this round are concrete and they punish vagueness. 'Your summariser sometimes invents a policy number. How do you measure that, and how do you know your fix worked?' 'Accuracy went from 0.91 to 0.93 and complaints went up. Explain.' 'You have 200 labelled examples and no budget for more. What do you do?' 'Your LLM judge agrees with your preferred answer almost every time. Why is that not reassuring?'

What a strong answer has in it, regardless of the model type:

What machine learning engineers are paid, and what moves it

Be sceptical of any single number for this title, including the ones on salary aggregators, because the title spans a fraud-model engineer at an insurer and a frontier-lab infrastructure engineer, and because a large share of total compensation at the top end is equity whose value is not a salary at all.

Where to get numbers you can defend. The US Bureau of Labor Statistics has no SOC code for machine learning engineer; the honest approach is to bracket it with the two OES codes that absorb most of these jobs, 15-2051 (data scientists) and 15-1252 (software developers), and to read the metropolitan-area tables rather than the national median, because geography moves this more than almost anything else. For the large-technology ladder, levels.fyi has the most useful level-by-level breakdown, and the published level names let you ask a recruiter precisely which band a role sits in. For a specific employer, the ranges posted under state pay-transparency laws, which now cover Colorado, California, Washington, New York, Illinois and a growing list of other states, are the most reliable public figures available, because they are what that company is legally representing about that role.

What actually moves the number, in rough order of effect:

The resume: what lands, line by line

ML resumes fail in a specific and consistent way: they describe technique and omit consequence. 'Built a customer churn model using XGBoost and SHAP' tells a reader nothing about whether anything happened. Compare an illustrative rewrite of the same work: 'Replaced a rules-based churn list with a gradient-boosted model; retention worked the top 2% of scores; 1.8 percentage point retention lift against a holdout over two quarters.' Same project, and now the reader can tell you have been in the room where a model was judged. Use your own numbers, not these.

If the real number is confidential, keep the unit and give the shape: 'cut manual review volume by roughly a third at fixed precision', or 'moved top-of-funnel conversion by low single-digit percentage points'. Vagueness about magnitude is survivable. Absence of any outcome is not.

Four kinds of evidence do most of the work. If your resume has all four, it will clear nearly every screen for the job type you are targeting; if it has none, no amount of tool listing will substitute. Write bullets in this shape: what the system does, what it replaced, the measured outcome with units, and the constraint it ran under. One line per system, not three.

Portfolios that work, and the paths that convert

The portfolio question is asked in almost every loop, in one form or another: show me something you built. The answer that works is one deployed system with a public write-up, not five notebooks. Depth is the signal, because depth is what a notebook cannot fake. You only hit retrieval failures on acronyms, a cache that returns stale features, or a token bill that triples at 10% more traffic by running something past real inputs.

A portfolio project that gets interviews looks like this: a real dataset or a real corpus, a baseline you beat, an eval harness with a held-out set and a metric you argue for, a deployed endpoint with measured p95 latency and a cost-per-thousand figure, a README with an error taxonomy counted from actual failures, and one paragraph on what you would do with another month. Write the failures down. Interviewers trust a project that admits where it is weak far more than one that claims to work.

On the paths in. Four converge on this role, and each has a specific thing to fix:

Running the search in 2026-27, including the no-production-experience trap

Everything above assumes you get read. In 2026 that is the binding constraint for most readers of this page. A posting at a recognisable employer collects applicants in the hundreds within days, most screens reject on title and keyword match before a human reads a line, and the single most common blocker is circular: the job wants a model you shipped to production, and your current employer will not let you ship one. Treat those as two separate problems, because they have two different fixes.

On channels, the honest shape rather than a made-up conversion rate: cold applications convert badly enough that the only way to make them work is volume and precise targeting, while a referral or a direct note to the hiring manager converts well enough to be worth most of your effort. Do not quote yourself a percentage from anywhere. Just allocate your week accordingly: a few hours on applications, most of it on the five companies you actually want and the people inside them.

A referral request that works is short, names the requisition, and does the referrer's work for them. Two sentences on the one system you shipped that matches that posting, a line on why that company specifically, and a paste-ready paragraph they can forward. 'Let me know if you hear of anything' asks a stranger to do your thinking and reliably gets nothing.

On the production-experience trap, in order of cost:

Working with AI in this role

What a machine learning engineer must know about AI in 2026-27

For this role the AI question is not whether you know about AI, because the job has always been AI. The question is whether you have the 2026 version of the craft, which is a different discipline from the 2021 version. Four things moved: evaluation became the actual work, inference economics became an engineering constraint with a line in the budget, fine-tuning became cheap enough that the skill is restraint, and agents turned reliability into a research-shaped problem that someone has to engineer anyway.

Be precise about what did not change, because overclaiming here is as damaging as underclaiming. Training large models from scratch consolidated into a small number of organisations, but classical supervised learning on tabular data is still the majority of models running in the economy, and nothing about transformers made label leakage, calibration or a careful holdout less important. A candidate who dismisses gradient boosting as legacy will fail an interview at most companies that actually have ML in production.

Also be precise about coding assistants. They changed how everyone writes code, including ML code, and some employers now test that directly: you may be handed an assistant and graded on whether you direct it well and catch what it got wrong. What they did not change is the thing you are hired for. An assistant will happily produce a training loop with a leaking split, an eval that scores the model against its own output, or a prompt change that improves the example you tested and degrades the other thousand. The judgement is the job; the typing was never the job.

What interviewers are probing for, underneath the specific questions, is whether you optimise for the system's real objective rather than for the model's metric: whether you will reach for retrieval before a fine-tune, a smaller model before a bigger one, a rule before a classifier, and a measurement before any of them. Tool familiarity is assumed and cheap. Judgement is what gets graded, and judgement only reads as real when it comes attached to a decision you made, a number you measured, and an alternative you rejected.

Evaluation as the primary deliverable, not the final step

When the model is bought rather than built, the thing that compounds is your ability to tell quickly and reliably whether a change helped. Teams without a trustworthy eval loop ship on vibes, regress silently, and cannot say why last month's prompt was replaced. Hiring managers know this, which is why evaluation has become its own interview round and, in some companies, its own job title.

Show it: Describe an eval harness you built as an artefact with dimensions: the number of labelled examples, how they were sourced and split, the metrics and why those metrics matched the cost of being wrong, where it runs (a CI gate on every prompt or model change), and a specific regression it caught before users did. Then describe the error taxonomy you built by reading failures, with counts, and which group you fixed first.

LLM-as-judge, used with its limits stated

Grading generated text at volume requires a model judge, and an uncalibrated judge manufactures confidence. Judges favour longer answers, favour the first option shown, and favour text from their own family, so a judge that nearly always agrees with you may simply be agreeing with the style you prompted for. Using one without a human-labelled calibration set is the most common evaluation mistake in LLM teams right now.

Show it: Say that you kept a human-labelled subset and measured agreement between judge and humans, that you used pairwise comparison rather than absolute 1-to-5 scores where you could, that you randomised position, and that you did not use the same model as generator and judge. Naming one bias you measured and corrected is worth more than naming five tools.

Inference economics: latency and cost as design constraints

Inference, not training, is the recurring bill for most ML systems in 2026, and it scales with usage rather than with headcount. Reasoning models made this sharper: thinking effort is now a tunable with a direct line to both latency and cost, and a single user request can fan out into many model calls. Engineers who can hold a latency budget and a unit-cost target are hired over engineers who can only raise quality, because finance now asks for cost per request and someone has to be able to answer.

Show it: Quote the numbers you operated against: p50 and p95, time to first token where streaming mattered, tokens per second, and cost per thousand requests or per resolved task, which is the unit that still means something once a task spans many calls. Name the lever you pulled and what it cost you in quality: a quantised or distilled smaller model, continuous batching, prompt and prefix caching, a cheap retrieval or classifier stage in front of an expensive generator, a shorter context, speculative decoding, routing by the reasoning depth a request actually needs rather than by model size alone, or a cap on thinking tokens with a measured quality delta. A before-and-after pair with a quality number next to it is the single most persuasive line in this area.

Knowing when not to fine-tune

LoRA and hosted tuning made fine-tuning cheap, which made it the default reach for candidates and a running cost for teams: a tuned checkpoint has to be re-tuned when the base model is deprecated, evaluated separately, served separately, and explained to whoever inherits it. Most of the time the problem was retrieval, the prompt, or the task definition. Interviewers use this question specifically to separate tool familiarity from judgement.

Show it: State the order you work in and why: define the task and build the eval first, then prompt and few-shot, then retrieval, then fine-tune only when a measured gap survives all three. Name the cases that do justify it, such as a fixed output format or schema the base model keeps violating, a domain idiom or language it handles badly, distilling a large model's behaviour into a small one to hit latency or unit cost, or a tone or policy requirement. Name the case that does not: adding facts, which belongs in retrieval. Best answer of all is a time you fine-tuned, measured no real gain, and removed it.

Retrieval quality, treated as a measurable engineering problem

Most reports that the model is hallucinating are retrieval failures: the right passage was never in the context. Pure dense vector search misses exactly what people search internal corpora for, meaning part numbers, clause references, error strings, acronyms and ticket IDs, and a retrieval system without its own metrics cannot be improved, only fiddled with.

Show it: Describe hybrid retrieval, meaning lexical BM25 alongside dense embeddings, with a re-ranking stage, chunking tied to document structure rather than a fixed token count, and metadata filters. Then give retrieval its own numbers: recall@k measured on a labelled query set, and the gap between 'the answer was retrievable' and 'the answer was given'. That decomposition is how you prove you know which half of the system is broken.

Agents and tool use, framed as reliability engineering

Multi-step agents are where 2026 budgets are going and where they are most often wasted. The failure modes are engineering failures: a tool called twice because a retry was not idempotent, a loop that burns tokens without progress, a step that silently succeeded with the wrong argument, no budget cap, no human handoff. Teams hire ML engineers to make these systems boring.

Show it: Talk about task success rate on a fixed set of real tasks, cost and step count per completed task, trajectory-level evaluation rather than only final-answer scoring, idempotent tool design, explicit step and spend budgets, and a defined handoff when the agent is stuck. Describe the tool schema you wrote and the ambiguity you removed from it after watching failures.

Post-training literacy, even if you never do it

Supervised fine-tuning, preference optimisation and reinforcement learning on verifiable tasks became a hiring category of their own, and the vocabulary leaks into ordinary ML engineer interviews because the people building the systems you consume use it. You are not expected to have trained a reward model unless you are applying for that job. You are expected to know what a preference dataset is, why a reward model is gameable, and why an eval suite used as a training signal stops being an honest measurement.

Show it: Be able to say what each stage buys and costs: SFT for behaviour and format from demonstrations, preference methods for the choices demonstrations cannot express, and reinforcement learning on tasks with a verifiable answer where a checker exists. Then name the trap: once your eval becomes the optimisation target it is no longer a measurement, so you hold back a set the training never sees. If you are targeting these roles, bring the data pipeline, meaning how pairs were collected, how annotator disagreement was handled, and how you detected reward hacking.

Observability and tracing for model-powered systems

You cannot run what you cannot reproduce. A bad output a user reports is only debuggable if you logged the inputs, the retrieved context, the prompt version, the model version and the parameters. Teams that skipped this spend their days guessing, and interviewers can hear the difference in how a candidate describes debugging.

Show it: Describe what you logged per request and how you sampled it, how you versioned prompts and model configurations so an output maps to an exact configuration, and how a production trace became a test case in the eval suite. Name the drift signals you watched for the kind of model you ran, such as input distribution shift, feature nulls after an upstream change, score distribution movement or a fall in cache hit rate, and the one that actually fired.

Classical ML discipline, unchanged and still decisive

Most companies with ML in production are running tabular and ranking models, and the errors that cost them money are old errors: a feature computed with information from after the prediction time, a split that leaks a customer across train and test, an uncalibrated score fed to a cost-based decision, an offline gain that evaporates online. These questions are in the loop for almost every ML engineer job, including LLM-flavoured ones.

Show it: Be fluent and specific: point-in-time correct feature computation, grouped and time-ordered splits, training-serving skew and how sharing feature code prevents it, PR-AUC for imbalanced problems, probability calibration before a threshold is set against a business cost ratio, and label delay. One story about a leak you caught, and how you caught it, is worth a page of technique names.

The specific rules that now reach the engineering, named

By 2026 model documentation stopped being a compliance afterthought in regulated and enterprise settings. Employers in finance, insurance, healthcare and hiring now ask engineers questions they used to ask lawyers: what data trained or tuned this, what is it permitted to be used for, how would you prove it, and how does an affected person get an explanation or a review. Candidates who can answer calmly stand out, because most cannot, and naming the actual instrument instead of gesturing at 'AI regulation' is the whole signal.

Show it: Know which one applies to the employer in front of you. SR 11-7 is the US Federal Reserve and OCC guidance on model risk management: it governs fraud, credit and pricing models at banks, and it requires a documented model inventory and validation independent of the people who built the model. The EU AI Act applies if your system touches the EU, with general-purpose model obligations already in force and the high-risk duties phasing in across 2026 and 2027, and it bites hardest on hiring, credit, education and safety uses. NYC Local Law 144 requires a published bias audit for an automated hiring tool. In clinical work, HIPAA governs the data and the FDA's software-as-a-medical-device pathway governs the model itself. Then describe the records you kept, covering dataset provenance and licensing, what went into a fine-tune, model and prompt versions tied to deployments, and evaluation results retained per release, plus one design decision you made for provenance reasons, such as excluding a corpus whose licence you could not honour or keeping a reviewable audit trail for an automated decision.

What a screen is looking for

These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.

Mistakes that cost people this job

Applying to machine learning engineer, AI engineer, applied scientist and data scientist postings with one resume, because the titles look adjacent.

Read the posting's nouns. Training data, labels, offline metrics and retraining means ML engineer. Product feature, prompts, latency and no training means AI engineer. Publications, a novel method and a research-depth round means applied scientist. Stakeholders, experiments and recommendations means data scientist. Keep one document per target and reorder the first two bullets; the underlying experience is the same.

Treating the coding round as a formality because ML judgement is what the job is about.

At large technology companies and quantitative firms it is a standard software-engineering round at medium to medium-hard difficulty, graded on reaching the optimal complexity in roughly 35 to 40 minutes of coding, and it eliminates more ML candidates than any other round. Drill it to that bar if those employers are on your list, and ask the recruiter whether an AI assistant is permitted so you practise the right mode.

A resume of techniques with no consequences: 'built a churn model using XGBoost and SHAP', repeated six times.

One line per system: what it decides, what it replaced, the measured outcome with a unit, and the constraint it ran under. If the number is confidential, keep the unit and give the shape rather than dropping the outcome altogether.

Leading with Kaggle rank, coursework or a certification stack as the main evidence.

Lead with one deployed system: an endpoint, an eval suite, a p95 figure, a cost per thousand requests, a README with a counted error taxonomy. A leaderboard position proves you can optimise a fixed metric on a clean dataset, which is the part of the job that was automated away.

Opening the ML system design round with a model architecture.

Spend the first five minutes on the decision being made, the cost of a false positive versus a false negative in the business's units, the latency budget and the volume. Then name the baseline you would have to beat. Feedback on this round is overwhelmingly about framing rather than diagram quality.

Treating evaluation as the last slide: 'we used accuracy and it was 94%'.

Bring a dataset with provenance, a metric chosen against the cost asymmetry, an error taxonomy you counted by reading failures, and the eval suite's place in CI. Have a story ready about a time the offline metric improved and the online metric did not, and what you found.

Reaching for a fine-tune as the first answer to any quality problem.

State the order, which is task definition and eval, then prompting, then retrieval, then fine-tune only against a measured residual gap. Name what fine-tuning is actually for: format compliance, domain idiom, distilling into a smaller model for latency or cost. Facts belong in retrieval. A story about removing a fine-tune that bought nothing is stronger than one about adding it.

Quoting quality numbers with no latency or cost numbers beside them.

Carry p50 and p95, and a unit cost, for every system you claim. Name one optimisation you made, the lever, and what it cost in quality. Teams are now budgeted on inference and screen for people who think in it, including reasoning-token budgets and cost per resolved task.

Dismissing classical ML as legacy because the interesting work is LLMs.

Stay fluent in leakage, grouped and time-ordered splits, calibration, threshold choice against a cost ratio, and label delay. Most companies with models in production are running tabular and ranking models, and these questions appear in LLM-team loops too.

Under-reporting data work because it feels unglamorous: annotation design, pipeline building, cleaning, label sourcing.

Write it up as engineering with volumes and process: how many labels, how they were obtained, what the annotation guidelines fixed, the leak you found, the pipeline cadence. It is often the majority of the job and the most transferable evidence you have.

Preparing six projects to a shallow depth for the deep-dive round.

Prepare two to the point where you can defend the data, the baseline, the alternative you rejected, the production failure and the numbers. The round is designed to push until something breaks; breadth breaks faster than depth.

Over-engineering a take-home to chase the best score, and ignoring the stated time cap.

Produce a clear baseline, one honest improvement, an eval with a defended metric, a counted error analysis, and a README saying what you would do next and what you deliberately left out. Respect the cap visibly; reviewers read restraint as seniority.

Waiting for permission to get production experience while applying to jobs that require it.

Find the decision at your current employer that a rule or a spreadsheet makes today, scope the smallest useful model for it, build the eval first and ship behind a flag. Failing that, take the unowned ML-adjacent work (label pipeline, drift monitoring, inference cost), ask for an internal loan to an ML team, or put a portfolio system in front of a handful of real users and log its failures.

Spraying cold applications and never asking for a referral, then concluding the market is closed.

Spend most of the week on the few companies you want: a short note naming the requisition, two sentences on the system you shipped that matches it, and a paste-ready paragraph the referrer can forward. Track which stage you lose at, because a screen loss and a system design loss need completely different fixes.

Skipping the question of what the role actually consumes versus trains, then preparing for the wrong rounds.

Ask the recruiter directly: does this team train, fine-tune or post-train models, or build around hosted ones? Which rounds are in the loop, what difficulty is the coding round, and is an AI assistant permitted? Three questions, and your preparation changes completely.

Questions people ask

What does a machine learning engineer do in 2026?

A machine learning engineer owns a model's behaviour in production and the system that makes it useful: the data and labels going in, the features or retrieval, the model choice, the evaluation that proves it works, the serving path, the guardrails on its failures, and the latency and cost it runs at. At most employers the model itself is bought or downloaded rather than trained from scratch, so the engineering is everything around it. The unit of work is a running system other people depend on, not a notebook result.

What is the difference between a machine learning engineer and an AI engineer?

An AI engineer builds product features on top of models nobody on the team trained, covering prompting, retrieval, tool use, streaming, guardrails, latency and cost, and typically does no training or label work. A machine learning engineer usually owns those things too but also owns model quality: training or fine-tuning where it is justified, offline evaluation infrastructure, label pipelines and retraining. The reliable tell in a posting is whether training data, labels, offline metrics or retraining are mentioned at all. If they are absent, it is an AI engineer job whatever the title says.

Do I need a PhD to be a machine learning engineer?

No. A PhD is not required for machine learning engineer, where you are measured on a system that works in production. It is the norm for research scientist and usually preferred for applied scientist, though several large employers run applied scientist ladders that hire at master's level with a strong applied modelling record. For ML engineer, a bachelor's in computer science, maths, statistics or a quantitative engineering field is the common floor and a master's is common. Strong software engineers with one well-evidenced production model are hired into these roles routinely, and at some employers the easiest door is the 'software engineer, machine learning' posting, which runs a standard engineering loop with one ML round.

How is the machine learning engineer interview loop structured?

Typically a recruiter screen, a hiring-manager call, a coding round of 45 to 60 minutes, an ML system design round, an evaluation or modelling-depth round, and a deep dive on a project you shipped, running four to eight weeks end to end. The system design and evaluation rounds carry the weight and the project deep dive usually decides level, but the coding round is still where most candidates are eliminated: at large technology companies and quantitative firms it is a standard software-engineering round at medium to medium-hard difficulty, while at startups and non-technology employers it is more often Python and SQL data manipulation or an ML-flavoured implementation. Take-homes of four to eight hours are common at startups and mid-size companies and rare at large ones.

What do machine learning engineers get paid?

There is no single credible band, and the US Bureau of Labor Statistics has no SOC code for the title. Bracket it with OES 15-2051 (data scientists) and 15-1252 (software developers) and read the metropolitan-area tables rather than the national median, then use levels.fyi for the large-technology ladder level by level, and the ranges employers publish under pay-transparency laws in Colorado, California, Washington, New York, Illinois and a growing list of other states for a specific company. Employer type moves pay more than the title does: a frontier lab, a trading firm, a bank, a retailer and a hospital system can differ by more than a factor of two for the same nominal job, mostly through equity rather than base.

Is machine learning engineering still a good career now that models are commoditised?

Yes, but the valuable half moved. Pretraining consolidated into a handful of organisations, so training novel architectures is a narrow speciality. What grew is everything required to make a bought model work on a specific company's data and budget: evaluation, retrieval quality, serving and inference cost, agent reliability, post-training on proprietary data, and the governance records regulated employers now demand. Classical tabular and ranking ML also did not shrink; in most industries it remains the majority of models in production. The jobs are there, and they reward production judgement over technique collection.

What should an ML engineering portfolio project look like?

One deployed system, not five notebooks: a real dataset or corpus, a baseline you beat, an eval harness with a held-out set and a metric you argue for, a live endpoint with a measured p95 latency and a cost per thousand requests, and a README containing an error taxonomy counted from actual failures plus what you would do with another month. Write the weaknesses down, because interviewers trust a project that names where it fails far more than one claiming to work. Depth is the signal, since only real traffic produces the problems that prove you have operated something.

How do I prepare for the ML system design interview?

Practise a fixed order out loud and on a timer: frame the decision and the cost of each kind of error, state the latency budget and volume, name the baseline you would have to beat, go through data and labels including label delay and leakage before any model, choose the model with a stated trade-off, define offline and online evaluation, then serving, fallbacks, monitoring and rough cost, and finish by naming one alternative you rejected and the condition that would make it right. Rehearse three canonical prompts, namely fraud detection, support-ticket deflection over a help centre, and marketplace search ranking, because most questions are variants of those. For material, Chip Huyen's Designing Machine Learning Systems and AI Engineering cover the system-level frame, Aminian and Xu's Machine Learning System Design Interview matches the round's format, and the engineering blog of a company with your target job type supplies worked trade-offs.

What is the evaluation round and how do I prepare for it?

It is a dedicated round, increasingly common, on how you would know a model-powered system works and whether a change improved it. Expect concrete questions: how to measure a summariser that invents policy numbers, why accuracy rose while complaints rose, what to do with 200 labels and no budget, why your LLM judge agreeing with you almost every time is not reassuring. Prepare one real eval harness you can describe by its dimensions, covering dataset size and provenance, splitting, metrics and why, human calibration of any model judge, CI gating and a regression it caught, plus one story about an offline gain that did not survive online and the reason you found.

How do I get a machine learning job when every posting wants production experience I have not been allowed to get?

Create the production experience where you already have access rather than waiting for a title. The cheapest version is the decision at your current employer that a rule, a spreadsheet or a manual queue makes today: scope the smallest useful model for it, agree the metric with whoever owns the decision, build the eval before the model, ship behind a flag at low traffic, and write down what changed. Next cheapest is the unowned ML-adjacent work, meaning the label pipeline, drift monitoring, an eval harness or an inference cost reduction, then an internal loan or transfer to an ML team, which is a smaller field to compete in than the external market. If none of that exists, deploy a portfolio system for real users and log a counted failure taxonomy, and contribute to an open-source eval or serving project so a stranger can verify your work.

How has AI coding assistance changed what ML engineers are hired for?

It changed the typing, not the judgement. Assistants write training loops, dataloaders and eval scripts quickly, and some employers now permit one in the coding round and grade how well you direct it and whether you catch its errors, so ask which mode you will face. What they cannot do is notice that the split leaks a customer across train and test, that an eval is scoring a model against its own output, or that a prompt change improved your three examples and degraded the other thousand. The hiring bar moved toward exactly those judgements, which is why evaluation and system design now carry the loop.

Put this on a resume in about a minute

Paste your history once and point it at the Machine Learning Engineer posting you are looking at. No account, no card.

Build my resume free More roles