| What the role is in 2026 | Building, shipping and operating systems that make predictions or generate text, where the model is one component among retrieval, features, evaluation, serving, guardrails and cost control. At most companies the model is bought or downloaded; the engineering is everything around it. Training a novel architecture from scratch is a small minority of these jobs. |
|---|---|
| Closest confusions | An AI engineer builds LLM product features and rarely trains anything. An applied scientist is measured on a modelling result. A data scientist is measured on an answer or a decision, not a running service. An MLOps or platform engineer owns the infrastructure other people's models run on. ML engineer sits in the middle: owns a model's behaviour and the system it runs inside. |
| Typical loop | Recruiter screen, hiring-manager call, a coding round, an ML system design round of 45 to 60 minutes, an evaluation or modelling-depth round, and a deep dive on a project you shipped. Four to eight weeks. A take-home of four to eight hours appears mostly at startups and mid-size companies; large employers favour live rounds. |
| Entry requirements | No licence and no mandatory certification. A bachelor's in computer science, maths, statistics or a quantitative engineering field is the common floor; a master's is common and helps most at research-adjacent employers. A PhD is the norm for research scientist and is usually preferred for applied scientist, though several large employers hire applied scientists at master's level with a strong applied record. It is not required for machine learning engineer. |
| Pay | No single band, and the US Bureau of Labor Statistics has no SOC code for 'machine learning engineer'. Bracket it with OES 15-2051 (data scientists) and 15-1252 (software developers), read the metropolitan-area tables rather than the national median, then use levels.fyi for the large-technology ladder and the ranges employers publish under state pay-transparency laws in Colorado, California, Washington, New York, Illinois and a growing list of other states. Which of the five job types you are in moves pay further than the title does. |
| Resume length | One page for under five years of experience, two pages beyond that. A model with no outcome attached takes as much space as a model with one and is worth far less: cut the project list before you cut the numbers. |
| Evidence that lands | A model in production with a measured business or product outcome; an eval harness with the dataset size and what it gated; p50 and p95 latency; cost per thousand requests or per resolved task; training-data volume and how labels were obtained; what you decided not to do and why. |
| What changed by 2026 | Pretraining consolidated into a handful of labs, so the leverage moved to evaluation, retrieval quality, inference economics and agent reliability. Fine-tuning became cheap enough that the skill is now knowing when not to do it. Post-training and reinforcement learning against a company's own eval suite emerged as a hiring category of its own. Classical tabular and ranking ML did not go away; in most industries it is still the majority of models in production. |
What a machine learning engineer actually does in 2026, and the adjacent titles people apply to by mistake
A machine learning engineer owns a model's behaviour in production and the system that makes that behaviour useful. That means the data that goes in, the features or the retrieval, the choice of model, the evaluation that says whether it works, the serving path it runs on, the guardrails around its failures, and the latency and cost it runs at. The unit of work is a system that users or another service depend on, not a notebook result.
The title moved hard between 2022 and 2026, and the market has not finished relabelling. The old centre of gravity, design an architecture, train it on your own data, tune it, publish or ship it, is now a minority of postings, concentrated in a small number of labs and in domains where no general model exists. The new centre of gravity is building around models somebody else trained: retrieval over your own corpus, an evaluation harness you can trust, serving that holds a latency budget, fine-tuning only where the case for it is made, agents that call tools without looping forever, and an inference bill that does not grow faster than usage.
That shift matters for applications because the same posting title now covers two quite different jobs, and because several adjacent titles sit close enough that candidates routinely apply to the wrong one, get rejected, and learn nothing from it. The distinction between them is not seniority. It is what you are measured on.
- AI engineer: measured on a shipped product feature powered by a model nobody on the team trained. Prompting, retrieval, tool use, streaming, guardrails, latency, cost. Little or no training. If a posting never mentions training data, labels or offline metrics, it is this job regardless of what the title says.
- Applied scientist and research engineer: measured on a modelling result: a better loss, a better ranking metric, a method that works where the previous one did not. A PhD is the norm for research scientist and usually preferred for applied scientist, though several large employers run an applied scientist ladder that hires at master's level with a strong applied record. The loop includes a research-depth round on your own work, and it is more competitive than ML engineer at the same level.
- Data scientist: measured on an answer, a decision or a measurement. Experiment design, causal inference, forecasting, metric definition. May build models, but is not usually on call for one. The tell is the posting's emphasis on stakeholders, experiments and recommendations rather than services and service levels.
- MLOps and ML platform engineer: measured on whether other people's models can be trained, deployed and observed. Serving infrastructure, GPU scheduling, feature store, CI for models, drift monitoring, cost controls. Owns the paved road, not what drives on it.
- Machine learning engineer: measured on whether a specific model-powered system works in production and keeps working. Owns the model's quality and the engineering that delivers it. Writes production code most days, reads papers selectively, and is on call for something.
- One tactic worth knowing before you read any further: 'software engineer, machine learning' is a real and growing posting at large employers, where a strong software engineer is embedded in an ML team and hired on a standard engineering loop with one ML-flavoured round. If you are strong at engineering and medium at modelling, it is often the easier door into the same work.
The five kinds of ML engineer job: read the posting's nouns
Before you tailor a single bullet, decide which of five jobs the posting describes. They want different evidence, interview differently, and pay differently. The nouns in the first three responsibilities tell you which one it is.
Most candidates have genuine experience in one and a reading knowledge of the others. The failure mode is a single generic resume that leads with whichever project is most recent, so a recommendations team reads a RAG chatbot and a platform team reads a Kaggle notebook. The same career supports more than one of these resumes. It does not support the same document.
- Classical and tabular ML in a product. Fraud, credit risk, churn, pricing, demand forecasting, propensity, lead scoring, insurance, ad targeting, supply chain. Gradient-boosted trees still win most of these. Postings say feature store, label leakage, point-in-time correctness, calibration, drift, AUC or PR-AUC, model monitoring. Probably the largest category of ML jobs outside technology companies, and the one candidates chasing LLM work overlook. Interviews probe data leakage, class imbalance, threshold choice against a cost ratio, and the gap between offline metrics and online results.
- Ranking, recommendation and search. Marketplaces, media, commerce, ads, feeds. Two-stage retrieval-then-rank architectures, embeddings, candidate generation, features computed at request time, counterfactual evaluation, position bias, interleaving and online tests. Postings say nDCG, recall@k, click-through rate, engagement, two-tower, re-ranker, real-time features. The hard parts are evaluation under feedback loops and serving a ranker inside tens of milliseconds.
- LLM systems. Retrieval over internal documents, assistants, extraction and classification pipelines, agents that call tools, summarisation at scale. Postings say RAG, vector database, embeddings, evals, prompt engineering, function or tool calling, guardrails, token cost, time to first token. Overlaps heavily with AI engineer; the ML engineer version usually also owns fine-tuning or distillation, offline eval infrastructure, and the inference stack.
- Post-training and reinforcement learning. Supervised fine-tuning at scale, preference optimisation (DPO and the GRPO-style policy methods), reward modelling, reinforcement learning on tasks with a verifiable answer, and the company's own eval suite used as the reward surface rather than as a report. Postings say SFT, preference data, reward model, rollouts, RLHF or RLVR, annotation pipeline, training infrastructure. It is at the labs, at a handful of model-serving companies, and increasingly at product companies with enough proprietary data to tune on. The most competitive of the five: it usually wants prior experience training at scale, and the interview goes deep on the pipeline that collected the preference data and on why your eval is not gameable.
- Domain and edge ML. Computer vision for manufacturing inspection, robotics and autonomy, medical imaging, speech, defence, geospatial, embedded and on-device inference. Postings say ONNX, TensorRT, quantisation, latency on a named device, data collection rigs, annotation pipelines, sensor calibration, regulatory validation. Here training from scratch or heavy fine-tuning is still routine, because general models do not cover the data distribution. These teams care about your data pipeline as much as your model.
How the hiring loop works in 2026-27, round by round
The structure stabilised after the 2023-24 churn, and it is now fairly predictable across company sizes. What changed is the weighting: two rounds, ML system design and evaluation depth, became where the offer is won or lost. The coding round did not get easier. It stopped being where the decision is made while remaining the place most candidates are eliminated.
Expect four to eight weeks end to end. Large employers run more rounds and schedule them further apart; startups compress to one or two days and lean harder on a take-home or a working session. Where a company uses a hiring committee, the written feedback from the system design and project deep-dive rounds is what the committee actually reads.
Round by round, and what each one is really testing:
- Recruiter screen (20 to 30 minutes). A vocabulary and level match. Have a 90-second answer to 'what have you put in production' with one number in it, and a direct answer on which of the five job types you are strongest in. Ask here whether the role trains models or consumes them, because the answer rewrites your preparation.
- Hiring-manager call (45 minutes). Scope and judgement. What you owned versus what your team owned, why a project was worth doing, what you would do differently. This is also your best chance to find out what the team is actually struggling with, which tells you which stories to lead with for the rest of the loop.
- Coding round (45 to 60 minutes), and the difficulty depends on the employer. At large technology companies and quantitative trading firms this is a standard software-engineering coding round at medium to medium-hard difficulty, graded on whether you reach the optimal complexity inside roughly 35 to 40 minutes of actual coding. It is a filter rather than a decider, but it is the filter that fails more ML candidates than any other round, and preparing for it as though it were easy is the most expensive mistake in this article. At startups and non-technology employers it is more often data manipulation in Python and SQL, or an ML-flavoured implementation: write a training loop, implement k-means or a k-NN search, write an evaluation loop over a dataset, implement attention, debug a dataloader that silently shuffles labels. Some employers now let you use an AI assistant and grade how you direct it and whether you catch what it got wrong; others still ban it. Ask which, and practise the mode you will face.
- ML system design (45 to 60 minutes). 'Design a system that flags fraudulent transactions.' 'Design support-ticket deflection over our help centre.' 'Design the ranking for a two-sided marketplace's search.' The highest-weighted round for mid-level and above. Details below.
- Evaluation and modelling depth (45 to 60 minutes). How you would know the system works, offline and online. This is the round most candidates have not prepared for, and the one that most reliably separates people who have operated a model from people who have only built one. Details below.
- Project deep dive (45 to 60 minutes). One project, pushed until it breaks. Expect: why that model, what the baseline was, what the data looked like, what went wrong in production, what you measured, what you rejected. Prepare two projects to this depth, covering architecture, numbers, failures and alternatives, rather than six to a shallow depth.
- Take-home (four to eight hours, when used). Most common at startups and mid-size companies. Typically a messy dataset with a loose goal, or a small retrieval or extraction task with a held-out set. The scoring rarely rewards the best score; it rewards a clear baseline, an honest error analysis, a stated metric with a reason, and a README that says what you would do next and what you did not have time for. Ask whether it is paid and what the time cap is, and respect the cap visibly.
- Behavioural and cross-functional round. Usually a product manager or a partner engineer. The real question is whether you can explain a model's limits to someone who will have to sell or support it, and whether you will say 'this will not work' early rather than late.
The ML system design round: what is actually being graded
This round is not a trivia test and it is not a software system design interview with a model bolted on. It is a test of whether you reason from the problem to the system, in the order a practitioner does: what decision is being made, what the cost of each kind of mistake is, what data exists, what the simplest thing that could work is, how you would evaluate it, then how you would serve it and what it would cost.
The single most common failure is starting at the architecture. A candidate hears 'design fraud detection' and begins with a model family and a feature store diagram, never having asked what happens when the system is wrong, whether a false positive blocks a customer's rent payment or merely queues a review. Written feedback on this round is overwhelmingly about framing rather than diagram quality: a candidate who spends the first five minutes on the decision usually outscores one who draws a better architecture.
A workable shape for 45 minutes, which also works as the structure of an answer in a take-home write-up:
- Frame the decision (5 minutes). Who or what consumes the output, and what action follows. Is this a ranked list, a binary block, a score a human reviews, or generated text a user reads? What are the costs of a false positive versus a false negative, in the business's own units? What is the latency budget, and is it an online request path or a batch job? What volume?
- Name the baseline you would have to beat. Rules a domain expert could write, last week's value, a popularity ranking, a keyword search, a single prompt with no retrieval. Saying out loud that the baseline might be good enough is a strong signal, not a weak one. A candidate who proposes a transformer before mentioning a baseline reads as someone who has not had to justify a quarter of work.
- Data before model. Where the labels come from and how delayed they are: fraud labels arrive weeks later via chargebacks, and that single fact reshapes the whole design. Whether the label is actually the thing you care about or a proxy. How you avoid leakage and training-serving skew, through point-in-time correctness, the same feature code in training and serving, and no feature that is only knowable after the event you are predicting.
- Then the model, with a reason. Gradient-boosted trees for tabular, a two-tower retriever plus a re-ranker for recommendations, hybrid lexical-plus-dense retrieval with a cross-encoder re-ranker for document search, a hosted general model with retrieval for open-ended text, a small fine-tuned or distilled model where latency or unit cost rules the hosted one out. Say what you are giving up with each.
- Evaluation, offline and online, before you draw the serving diagram. The offline set and where it comes from; the metric and why it matches the cost asymmetry you named; the online test and its guardrail metrics; what you would ship behind a flag and at what traffic share.
- Serving and the numbers. Online or batch; where features come from at request time and how stale they may be; caching; what you do when the model or a dependency times out; the fallback that keeps the product usable. Give a rough latency breakdown and a rough cost per thousand requests, and say which number you are least sure of.
- Failure and monitoring. What drift looks like for this system and how you would detect it before a user does; what you log so a bad output can be reproduced; the retraining cadence and what triggers it; the kill switch and who is allowed to pull it.
- Say what you rejected. Name one alternative you considered and the condition under which it would have been the right call. This is the single cheapest way to read as senior, and most candidates never do it.
- What to practise from, since 'practise system design' is useless advice on its own. Chip Huyen's Designing Machine Learning Systems and AI Engineering (both O'Reilly) give you the system-level frame and the LLM-era vocabulary. Ali Aminian and Alex Xu's Machine Learning System Design Interview is organised the way the round is run, which makes it good for pacing. Then read the engineering blog of two companies whose job type matches your target, because a worked write-up of a real ranking or retrieval system teaches trade-offs no book can abbreviate. Rehearse out loud and on a timer; the round rewards verbal order, not knowledge.
The evaluation round, the one most candidates are unprepared for
Evaluation moved from a chapter at the end of the project to the centre of the job, and hiring followed. Many loops now dedicate a full round to it, and even where they do not, the evaluation questions inside the system design round carry disproportionate weight. The reason is practical: when the model is a commodity someone else trained, your only durable advantage is knowing, faster and more precisely than a competitor, whether a change made things better.
Questions in this round are concrete and they punish vagueness. 'Your summariser sometimes invents a policy number. How do you measure that, and how do you know your fix worked?' 'Accuracy went from 0.91 to 0.93 and complaints went up. Explain.' 'You have 200 labelled examples and no budget for more. What do you do?' 'Your LLM judge agrees with your preferred answer almost every time. Why is that not reassuring?'
What a strong answer has in it, regardless of the model type:
- A dataset with provenance. Where the examples came from, how many there are, how they were split, and why the split is honest: grouped by user or document so near-duplicates cannot straddle it, and time-ordered when the real task is predicting the future. 'I held out 20%' with no further detail is a weak answer; 'I split by customer and by week because the same ticket appears in several forms' is a strong one.
- A metric tied to the cost of being wrong. PR-AUC and a chosen operating threshold rather than accuracy on an imbalanced problem; recall at a fixed precision when a human reviews every flag; nDCG or recall@k for ranking; for generated text, a decomposed set, asking whether it is grounded in the retrieved source, whether it answers the question asked, and whether it follows the required format, rather than one fuzzy quality score.
- Error analysis by hand. Say that you read failures, grouped them into a taxonomy, counted the groups, and fixed the largest one first. Interviewers notice this immediately because it is the habit that separates people who improve systems from people who tune them. Bring a real taxonomy from real work if you have one: 'of 150 failures, 60 were retrieval misses on acronyms, 40 were the model answering a question the user did not ask, 30 were formatting, 20 were genuinely ambiguous.'
- Honest handling of LLM-as-judge. It is the standard tool for grading generated text at volume, and it is only as good as its calibration: a human-labelled subset, measured agreement between judge and humans, position and verbosity bias controlled, pairwise comparison rather than absolute scoring where possible, and a different model as judge than as generator. A candidate who uses a judge without a calibration set, or who cannot name a bias it introduces, is a candidate who has not been burned yet.
- Evals wired into CI. A regression suite that runs on every prompt, retrieval or model change and blocks a merge on a named threshold. This is the artefact most likely to make an interviewer lean forward, because most teams want it and fewer have it.
- Online and offline connected. A shadow deployment or canary, an A/B test with a pre-registered primary metric and guardrails, and a story about a time the offline number improved and the online number did not, with the reason you found. The offline-to-online gap is the real craft, and naming yours is credibility.
- Cheap-label strategies when data is scarce. Weak supervision and heuristic labelling, active learning on the model's uncertain cases, synthetic examples for format and edge cases with a human check, a small gold set you defend fiercely rather than a large noisy one you half-trust.
What machine learning engineers are paid, and what moves it
Be sceptical of any single number for this title, including the ones on salary aggregators, because the title spans a fraud-model engineer at an insurer and a frontier-lab infrastructure engineer, and because a large share of total compensation at the top end is equity whose value is not a salary at all.
Where to get numbers you can defend. The US Bureau of Labor Statistics has no SOC code for machine learning engineer; the honest approach is to bracket it with the two OES codes that absorb most of these jobs, 15-2051 (data scientists) and 15-1252 (software developers), and to read the metropolitan-area tables rather than the national median, because geography moves this more than almost anything else. For the large-technology ladder, levels.fyi has the most useful level-by-level breakdown, and the published level names let you ask a recruiter precisely which band a role sits in. For a specific employer, the ranges posted under state pay-transparency laws, which now cover Colorado, California, Washington, New York, Illinois and a growing list of other states, are the most reliable public figures available, because they are what that company is legally representing about that role.
What actually moves the number, in rough order of effect:
- Employer type, not title. A frontier lab, a large technology company, a quantitative trading firm, a well-funded AI startup, a bank, a retailer, a hospital system and a defence contractor can differ by more than a factor of two for the same nominal job, mostly through equity rather than base. Compare offers on the equity terms, not the headline.
- Which of the five job types you are in. Post-training, LLM infrastructure and ranking at scale pay above internal tabular ML at a non-technology employer, for the same years of experience.
- Demonstrated production ownership. Candidates who can show a system they ran, with latency and cost numbers and an on-call history, negotiate from a different position than candidates whose evidence is projects and papers.
- Level, which is negotiable in a way salary often is not. The level decision is made from the system design and deep-dive rounds. Pushing on level before the offer stage, with scope evidence, is usually worth more than haggling over base afterwards.
- Geography and remote policy. Many employers now band remote pay by location. Ask which band applies before the offer conversation, not during it.
- Scarce domain depth. Hardware-constrained inference, speech, medical imaging under regulatory validation, autonomy, or a specific regulated data environment command premiums because the hiring pool is genuinely small.
The resume: what lands, line by line
ML resumes fail in a specific and consistent way: they describe technique and omit consequence. 'Built a customer churn model using XGBoost and SHAP' tells a reader nothing about whether anything happened. Compare an illustrative rewrite of the same work: 'Replaced a rules-based churn list with a gradient-boosted model; retention worked the top 2% of scores; 1.8 percentage point retention lift against a holdout over two quarters.' Same project, and now the reader can tell you have been in the room where a model was judged. Use your own numbers, not these.
If the real number is confidential, keep the unit and give the shape: 'cut manual review volume by roughly a third at fixed precision', or 'moved top-of-funnel conversion by low single-digit percentage points'. Vagueness about magnitude is survivable. Absence of any outcome is not.
Four kinds of evidence do most of the work. If your resume has all four, it will clear nearly every screen for the job type you are targeting; if it has none, no amount of tool listing will substitute. Write bullets in this shape: what the system does, what it replaced, the measured outcome with units, and the constraint it ran under. One line per system, not three.
- A model in production with a measured outcome. The non-negotiable item. Name the decision it affected, the baseline it beat, and the number: revenue, loss avoided, tickets deflected, conversion, review hours saved, defects caught.
- An evaluation harness. 'Built the offline eval suite for the extraction service: 1,200 human-labelled documents, per-field precision and recall, run on every prompt and model change in CI, blocking merges below a set threshold.' This line is rarer than it should be and it reads as senior immediately.
- Latency and cost numbers. p50 and p95 for the serving path, and unit economics: cost per thousand requests, cost per resolved ticket, GPU-hours per training run, or the before-and-after of an optimisation with the lever named, whether that was quantisation, batching, caching, a smaller distilled model, a reasoning-effort cap, or a cheaper retrieval stage in front of a dearer generator.
- Data work, stated as engineering. Volume and shape, where labels came from, the annotation process you ran or designed, the leakage you found, the pipeline you built and its cadence. Data work is often the majority of the job, and candidates under-report it because it feels unglamorous. It is the most transferable evidence you have.
- Keep a short, honest tools block at the bottom, for the keyword screen only. Python, PyTorch, scikit-learn, SQL, XGBoost or LightGBM, a serving stack you have actually used, an orchestrator, a cloud, a tracking or eval tool. A block that lists everything you have ever touched reads as a list of things read about; keep it to what you could be questioned on for ten minutes.
- Cut the Kaggle rank and the coursework unless you are a new graduate, in which case lead with one deployed thing instead. A leaderboard position answers a question no interviewer asked: it says you can optimise a fixed metric on a clean dataset, which is the part of the job that no longer exists.
- Name the model family, not just 'machine learning'. Screens, human and automated, match on specifics: gradient boosting, two-tower retrieval, cross-encoder re-ranking, LoRA fine-tune, distillation, sequence labelling, time-series forecasting, preference optimisation. Generic phrasing matches nothing.
- Mirror the posting's job type in your first two bullets. The same underlying experience can lead with a ranking system for a marketplace and with an eval harness for an LLM team. Reorder; do not reinvent.
Portfolios that work, and the paths that convert
The portfolio question is asked in almost every loop, in one form or another: show me something you built. The answer that works is one deployed system with a public write-up, not five notebooks. Depth is the signal, because depth is what a notebook cannot fake. You only hit retrieval failures on acronyms, a cache that returns stale features, or a token bill that triples at 10% more traffic by running something past real inputs.
A portfolio project that gets interviews looks like this: a real dataset or a real corpus, a baseline you beat, an eval harness with a held-out set and a metric you argue for, a deployed endpoint with measured p95 latency and a cost-per-thousand figure, a README with an error taxonomy counted from actual failures, and one paragraph on what you would do with another month. Write the failures down. Interviewers trust a project that admits where it is weak far more than one that claims to work.
On the paths in. Four converge on this role, and each has a specific thing to fix:
- From software engineering. The strongest path and the most underrated. You already have the half most candidates lack: production code, testing, on call, systems thinking. Fix: one end-to-end model with offline and online evaluation, and enough statistical literacy to defend a metric choice. Target 'software engineer, machine learning' and the platform-adjacent postings first; they screen on your existing strengths.
- From data science or analytics. You have the modelling and the metric judgement. Fix: ship something that serves live traffic. Take the model you built that sat in a notebook, put it behind an endpoint, write the eval suite, add monitoring, and describe it as a service with latency and cost. Learn enough engineering discipline, meaning version control, tests, containers and CI, that a reviewer does not have to wonder.
- From research or a PhD. You can read and reproduce anything. Fix: show that you can hold a constraint. One project where the limit was 50 milliseconds or a cent a request, and where you chose a worse model on purpose for a stated reason. Also decide honestly whether you want applied scientist instead; that loop rewards your depth more directly.
- From MLOps or platform. You own the road. Fix: own one model's quality, not just its deployment. Build the eval suite and the error taxonomy for a system your platform hosts, and be able to argue about a threshold in business units.
- Where the jobs are. Go to company boards directly. Greenhouse, Lever and Ashby pages are where postings appear first and where descriptions are complete, and they are indexable. Then the hiring pages of labs and AI startups, the monthly 'Who is hiring' threads, specialist ML job boards, and the engineering blogs of companies whose problems you find interesting, because a referenced blog post in a cover note gets read. Avoid the aggregator reposts; the descriptions are truncated exactly where the job type is specified.
- On certifications. There is no licence and no credential that gates this work. Cloud ML certifications occasionally help a resume clear a procurement-minded screen at a consultancy or an enterprise, and they help nobody in the system design round. Spend the same weeks on a deployed system with an eval report instead.
Running the search in 2026-27, including the no-production-experience trap
Everything above assumes you get read. In 2026 that is the binding constraint for most readers of this page. A posting at a recognisable employer collects applicants in the hundreds within days, most screens reject on title and keyword match before a human reads a line, and the single most common blocker is circular: the job wants a model you shipped to production, and your current employer will not let you ship one. Treat those as two separate problems, because they have two different fixes.
On channels, the honest shape rather than a made-up conversion rate: cold applications convert badly enough that the only way to make them work is volume and precise targeting, while a referral or a direct note to the hiring manager converts well enough to be worth most of your effort. Do not quote yourself a percentage from anywhere. Just allocate your week accordingly: a few hours on applications, most of it on the five companies you actually want and the people inside them.
A referral request that works is short, names the requisition, and does the referrer's work for them. Two sentences on the one system you shipped that matches that posting, a line on why that company specifically, and a paste-ready paragraph they can forward. 'Let me know if you hear of anything' asks a stranger to do your thinking and reliably gets nothing.
On the production-experience trap, in order of cost:
- Find the model inside your current job. The cheapest production experience available is the decision at your employer that is currently made by a rule, a spreadsheet or a queue nobody has looked at in two years. Scope the smallest useful version, define the metric with whoever owns the decision, build the eval before the model, ship it behind a flag at a small traffic share, and write down what it changed. This is the same artefact the interview asks for, and it does not require a title change.
- Take the ML-adjacent work your team is avoiding. The label pipeline, the drift monitoring, the eval harness for a feature someone else prompt-engineered, the cost reduction on an inference bill. It is all legible as ML engineering evidence, and it is usually unowned.
- Internal transfer beats an external move when you lack production evidence. You are competing against a smaller field, the hiring manager can verify your engineering directly, and six months on an ML team internally makes the external search a different search. Ask for a rotation or a loan before you ask for a transfer.
- If none of that is available, build the portfolio system described above and put it in front of real users, even two. A side project with ten real users and a logged failure taxonomy beats a benchmark result with none, because the failures are what the interview is about.
- Open source is the other verifiable route: a merged contribution to an eval, serving or data-tooling project is public, dated and reviewable by a stranger, which a private side project is not.
- Fix the title mismatch on your own documents. If your title is 'Data Scientist' and you are targeting ML engineer, the headline under your name should say what you do (for example, 'Data Scientist, production ML systems'), your first two bullets should be the service you shipped, and the skills block should use the posting's nouns. Do not invent a title you never held; reframe the one you did.
- Track where you lose, not just whether you lost. Losing at the recruiter screen is a documents and targeting problem. Losing at the coding round is a drilling problem. Losing at system design is a framing and rehearsal problem. Losing at the deep dive means your two projects are not prepared to depth. Candidates who do not record the stage repeat the same fix for months. Expect the search to run in months rather than weeks, and keep several processes alive at once so no single loop carries your timeline.
What a machine learning engineer must know about AI in 2026-27
For this role the AI question is not whether you know about AI, because the job has always been AI. The question is whether you have the 2026 version of the craft, which is a different discipline from the 2021 version. Four things moved: evaluation became the actual work, inference economics became an engineering constraint with a line in the budget, fine-tuning became cheap enough that the skill is restraint, and agents turned reliability into a research-shaped problem that someone has to engineer anyway.
Be precise about what did not change, because overclaiming here is as damaging as underclaiming. Training large models from scratch consolidated into a small number of organisations, but classical supervised learning on tabular data is still the majority of models running in the economy, and nothing about transformers made label leakage, calibration or a careful holdout less important. A candidate who dismisses gradient boosting as legacy will fail an interview at most companies that actually have ML in production.
Also be precise about coding assistants. They changed how everyone writes code, including ML code, and some employers now test that directly: you may be handed an assistant and graded on whether you direct it well and catch what it got wrong. What they did not change is the thing you are hired for. An assistant will happily produce a training loop with a leaking split, an eval that scores the model against its own output, or a prompt change that improves the example you tested and degrades the other thousand. The judgement is the job; the typing was never the job.
What interviewers are probing for, underneath the specific questions, is whether you optimise for the system's real objective rather than for the model's metric: whether you will reach for retrieval before a fine-tune, a smaller model before a bigger one, a rule before a classifier, and a measurement before any of them. Tool familiarity is assumed and cheap. Judgement is what gets graded, and judgement only reads as real when it comes attached to a decision you made, a number you measured, and an alternative you rejected.
Evaluation as the primary deliverable, not the final step
When the model is bought rather than built, the thing that compounds is your ability to tell quickly and reliably whether a change helped. Teams without a trustworthy eval loop ship on vibes, regress silently, and cannot say why last month's prompt was replaced. Hiring managers know this, which is why evaluation has become its own interview round and, in some companies, its own job title.
Show it: Describe an eval harness you built as an artefact with dimensions: the number of labelled examples, how they were sourced and split, the metrics and why those metrics matched the cost of being wrong, where it runs (a CI gate on every prompt or model change), and a specific regression it caught before users did. Then describe the error taxonomy you built by reading failures, with counts, and which group you fixed first.
LLM-as-judge, used with its limits stated
Grading generated text at volume requires a model judge, and an uncalibrated judge manufactures confidence. Judges favour longer answers, favour the first option shown, and favour text from their own family, so a judge that nearly always agrees with you may simply be agreeing with the style you prompted for. Using one without a human-labelled calibration set is the most common evaluation mistake in LLM teams right now.
Show it: Say that you kept a human-labelled subset and measured agreement between judge and humans, that you used pairwise comparison rather than absolute 1-to-5 scores where you could, that you randomised position, and that you did not use the same model as generator and judge. Naming one bias you measured and corrected is worth more than naming five tools.
Inference economics: latency and cost as design constraints
Inference, not training, is the recurring bill for most ML systems in 2026, and it scales with usage rather than with headcount. Reasoning models made this sharper: thinking effort is now a tunable with a direct line to both latency and cost, and a single user request can fan out into many model calls. Engineers who can hold a latency budget and a unit-cost target are hired over engineers who can only raise quality, because finance now asks for cost per request and someone has to be able to answer.
Show it: Quote the numbers you operated against: p50 and p95, time to first token where streaming mattered, tokens per second, and cost per thousand requests or per resolved task, which is the unit that still means something once a task spans many calls. Name the lever you pulled and what it cost you in quality: a quantised or distilled smaller model, continuous batching, prompt and prefix caching, a cheap retrieval or classifier stage in front of an expensive generator, a shorter context, speculative decoding, routing by the reasoning depth a request actually needs rather than by model size alone, or a cap on thinking tokens with a measured quality delta. A before-and-after pair with a quality number next to it is the single most persuasive line in this area.
Knowing when not to fine-tune
LoRA and hosted tuning made fine-tuning cheap, which made it the default reach for candidates and a running cost for teams: a tuned checkpoint has to be re-tuned when the base model is deprecated, evaluated separately, served separately, and explained to whoever inherits it. Most of the time the problem was retrieval, the prompt, or the task definition. Interviewers use this question specifically to separate tool familiarity from judgement.
Show it: State the order you work in and why: define the task and build the eval first, then prompt and few-shot, then retrieval, then fine-tune only when a measured gap survives all three. Name the cases that do justify it, such as a fixed output format or schema the base model keeps violating, a domain idiom or language it handles badly, distilling a large model's behaviour into a small one to hit latency or unit cost, or a tone or policy requirement. Name the case that does not: adding facts, which belongs in retrieval. Best answer of all is a time you fine-tuned, measured no real gain, and removed it.
Retrieval quality, treated as a measurable engineering problem
Most reports that the model is hallucinating are retrieval failures: the right passage was never in the context. Pure dense vector search misses exactly what people search internal corpora for, meaning part numbers, clause references, error strings, acronyms and ticket IDs, and a retrieval system without its own metrics cannot be improved, only fiddled with.
Show it: Describe hybrid retrieval, meaning lexical BM25 alongside dense embeddings, with a re-ranking stage, chunking tied to document structure rather than a fixed token count, and metadata filters. Then give retrieval its own numbers: recall@k measured on a labelled query set, and the gap between 'the answer was retrievable' and 'the answer was given'. That decomposition is how you prove you know which half of the system is broken.
Agents and tool use, framed as reliability engineering
Multi-step agents are where 2026 budgets are going and where they are most often wasted. The failure modes are engineering failures: a tool called twice because a retry was not idempotent, a loop that burns tokens without progress, a step that silently succeeded with the wrong argument, no budget cap, no human handoff. Teams hire ML engineers to make these systems boring.
Show it: Talk about task success rate on a fixed set of real tasks, cost and step count per completed task, trajectory-level evaluation rather than only final-answer scoring, idempotent tool design, explicit step and spend budgets, and a defined handoff when the agent is stuck. Describe the tool schema you wrote and the ambiguity you removed from it after watching failures.
Post-training literacy, even if you never do it
Supervised fine-tuning, preference optimisation and reinforcement learning on verifiable tasks became a hiring category of their own, and the vocabulary leaks into ordinary ML engineer interviews because the people building the systems you consume use it. You are not expected to have trained a reward model unless you are applying for that job. You are expected to know what a preference dataset is, why a reward model is gameable, and why an eval suite used as a training signal stops being an honest measurement.
Show it: Be able to say what each stage buys and costs: SFT for behaviour and format from demonstrations, preference methods for the choices demonstrations cannot express, and reinforcement learning on tasks with a verifiable answer where a checker exists. Then name the trap: once your eval becomes the optimisation target it is no longer a measurement, so you hold back a set the training never sees. If you are targeting these roles, bring the data pipeline, meaning how pairs were collected, how annotator disagreement was handled, and how you detected reward hacking.
Observability and tracing for model-powered systems
You cannot run what you cannot reproduce. A bad output a user reports is only debuggable if you logged the inputs, the retrieved context, the prompt version, the model version and the parameters. Teams that skipped this spend their days guessing, and interviewers can hear the difference in how a candidate describes debugging.
Show it: Describe what you logged per request and how you sampled it, how you versioned prompts and model configurations so an output maps to an exact configuration, and how a production trace became a test case in the eval suite. Name the drift signals you watched for the kind of model you ran, such as input distribution shift, feature nulls after an upstream change, score distribution movement or a fall in cache hit rate, and the one that actually fired.
Classical ML discipline, unchanged and still decisive
Most companies with ML in production are running tabular and ranking models, and the errors that cost them money are old errors: a feature computed with information from after the prediction time, a split that leaks a customer across train and test, an uncalibrated score fed to a cost-based decision, an offline gain that evaporates online. These questions are in the loop for almost every ML engineer job, including LLM-flavoured ones.
Show it: Be fluent and specific: point-in-time correct feature computation, grouped and time-ordered splits, training-serving skew and how sharing feature code prevents it, PR-AUC for imbalanced problems, probability calibration before a threshold is set against a business cost ratio, and label delay. One story about a leak you caught, and how you caught it, is worth a page of technique names.
The specific rules that now reach the engineering, named
By 2026 model documentation stopped being a compliance afterthought in regulated and enterprise settings. Employers in finance, insurance, healthcare and hiring now ask engineers questions they used to ask lawyers: what data trained or tuned this, what is it permitted to be used for, how would you prove it, and how does an affected person get an explanation or a review. Candidates who can answer calmly stand out, because most cannot, and naming the actual instrument instead of gesturing at 'AI regulation' is the whole signal.
Show it: Know which one applies to the employer in front of you. SR 11-7 is the US Federal Reserve and OCC guidance on model risk management: it governs fraud, credit and pricing models at banks, and it requires a documented model inventory and validation independent of the people who built the model. The EU AI Act applies if your system touches the EU, with general-purpose model obligations already in force and the high-risk duties phasing in across 2026 and 2027, and it bites hardest on hiring, credit, education and safety uses. NYC Local Law 144 requires a published bias audit for an automated hiring tool. In clinical work, HIPAA governs the data and the FDA's software-as-a-medical-device pathway governs the model itself. Then describe the records you kept, covering dataset provenance and licensing, what went into a fine-tune, model and prompt versions tied to deployments, and evaluation results retained per release, plus one design decision you made for provenance reasons, such as excluding a corpus whose licence you could not honour or keeping a reviewable audit trail for an automated decision.
What a screen is looking for
These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.
- Python
- PyTorch
- scikit-learn
- SQL
- XGBoost
- LightGBM
- Gradient boosting
- Feature engineering
- Feature store
- Point-in-time correctness
- Training-serving skew
- Label leakage
- Class imbalance
- PR-AUC
- Probability calibration
- Recommendation systems
- Learning to rank
- Two-tower retrieval
- Cross-encoder reranking
- nDCG
- Recall@k
- A/B testing
- Shadow deployment
- Canary release
- Model monitoring
- Data drift
- Model retraining
- MLflow
- Weights & Biases
- Airflow
- Docker
- Kubernetes
- CI/CD
- AWS SageMaker
- Google Vertex AI
- Azure Machine Learning
- Model serving
- vLLM
- Triton Inference Server
- TensorRT
- ONNX
- p95 latency
- Time to first token
- Quantisation
- Model distillation
- Inference cost optimisation
- Large language models
- Retrieval-augmented generation (RAG)
- Embeddings
- Vector database
- Hybrid search
- BM25
- Chunking strategy
- Prompt engineering
- Tool calling
- Agents
- Guardrails
- Fine-tuning
- LoRA
- Supervised fine-tuning
- Direct preference optimisation (DPO)
- Reward modelling
- Eval harness
- LLM-as-judge
- Golden dataset
- Error analysis
- Groundedness
- Observability
- Tracing
- EU AI Act
- SR 11-7 model risk management
- Time-series forecasting
Mistakes that cost people this job
Applying to machine learning engineer, AI engineer, applied scientist and data scientist postings with one resume, because the titles look adjacent.
Read the posting's nouns. Training data, labels, offline metrics and retraining means ML engineer. Product feature, prompts, latency and no training means AI engineer. Publications, a novel method and a research-depth round means applied scientist. Stakeholders, experiments and recommendations means data scientist. Keep one document per target and reorder the first two bullets; the underlying experience is the same.
Treating the coding round as a formality because ML judgement is what the job is about.
At large technology companies and quantitative firms it is a standard software-engineering round at medium to medium-hard difficulty, graded on reaching the optimal complexity in roughly 35 to 40 minutes of coding, and it eliminates more ML candidates than any other round. Drill it to that bar if those employers are on your list, and ask the recruiter whether an AI assistant is permitted so you practise the right mode.
A resume of techniques with no consequences: 'built a churn model using XGBoost and SHAP', repeated six times.
One line per system: what it decides, what it replaced, the measured outcome with a unit, and the constraint it ran under. If the number is confidential, keep the unit and give the shape rather than dropping the outcome altogether.
Leading with Kaggle rank, coursework or a certification stack as the main evidence.
Lead with one deployed system: an endpoint, an eval suite, a p95 figure, a cost per thousand requests, a README with a counted error taxonomy. A leaderboard position proves you can optimise a fixed metric on a clean dataset, which is the part of the job that was automated away.
Opening the ML system design round with a model architecture.
Spend the first five minutes on the decision being made, the cost of a false positive versus a false negative in the business's units, the latency budget and the volume. Then name the baseline you would have to beat. Feedback on this round is overwhelmingly about framing rather than diagram quality.
Treating evaluation as the last slide: 'we used accuracy and it was 94%'.
Bring a dataset with provenance, a metric chosen against the cost asymmetry, an error taxonomy you counted by reading failures, and the eval suite's place in CI. Have a story ready about a time the offline metric improved and the online metric did not, and what you found.
Reaching for a fine-tune as the first answer to any quality problem.
State the order, which is task definition and eval, then prompting, then retrieval, then fine-tune only against a measured residual gap. Name what fine-tuning is actually for: format compliance, domain idiom, distilling into a smaller model for latency or cost. Facts belong in retrieval. A story about removing a fine-tune that bought nothing is stronger than one about adding it.
Quoting quality numbers with no latency or cost numbers beside them.
Carry p50 and p95, and a unit cost, for every system you claim. Name one optimisation you made, the lever, and what it cost in quality. Teams are now budgeted on inference and screen for people who think in it, including reasoning-token budgets and cost per resolved task.
Dismissing classical ML as legacy because the interesting work is LLMs.
Stay fluent in leakage, grouped and time-ordered splits, calibration, threshold choice against a cost ratio, and label delay. Most companies with models in production are running tabular and ranking models, and these questions appear in LLM-team loops too.
Under-reporting data work because it feels unglamorous: annotation design, pipeline building, cleaning, label sourcing.
Write it up as engineering with volumes and process: how many labels, how they were obtained, what the annotation guidelines fixed, the leak you found, the pipeline cadence. It is often the majority of the job and the most transferable evidence you have.
Preparing six projects to a shallow depth for the deep-dive round.
Prepare two to the point where you can defend the data, the baseline, the alternative you rejected, the production failure and the numbers. The round is designed to push until something breaks; breadth breaks faster than depth.
Over-engineering a take-home to chase the best score, and ignoring the stated time cap.
Produce a clear baseline, one honest improvement, an eval with a defended metric, a counted error analysis, and a README saying what you would do next and what you deliberately left out. Respect the cap visibly; reviewers read restraint as seniority.
Waiting for permission to get production experience while applying to jobs that require it.
Find the decision at your current employer that a rule or a spreadsheet makes today, scope the smallest useful model for it, build the eval first and ship behind a flag. Failing that, take the unowned ML-adjacent work (label pipeline, drift monitoring, inference cost), ask for an internal loan to an ML team, or put a portfolio system in front of a handful of real users and log its failures.
Spraying cold applications and never asking for a referral, then concluding the market is closed.
Spend most of the week on the few companies you want: a short note naming the requisition, two sentences on the system you shipped that matches it, and a paste-ready paragraph the referrer can forward. Track which stage you lose at, because a screen loss and a system design loss need completely different fixes.
Skipping the question of what the role actually consumes versus trains, then preparing for the wrong rounds.
Ask the recruiter directly: does this team train, fine-tune or post-train models, or build around hosted ones? Which rounds are in the loop, what difficulty is the coding round, and is an AI assistant permitted? Three questions, and your preparation changes completely.
Questions people ask
What does a machine learning engineer do in 2026?
A machine learning engineer owns a model's behaviour in production and the system that makes it useful: the data and labels going in, the features or retrieval, the model choice, the evaluation that proves it works, the serving path, the guardrails on its failures, and the latency and cost it runs at. At most employers the model itself is bought or downloaded rather than trained from scratch, so the engineering is everything around it. The unit of work is a running system other people depend on, not a notebook result.
What is the difference between a machine learning engineer and an AI engineer?
An AI engineer builds product features on top of models nobody on the team trained, covering prompting, retrieval, tool use, streaming, guardrails, latency and cost, and typically does no training or label work. A machine learning engineer usually owns those things too but also owns model quality: training or fine-tuning where it is justified, offline evaluation infrastructure, label pipelines and retraining. The reliable tell in a posting is whether training data, labels, offline metrics or retraining are mentioned at all. If they are absent, it is an AI engineer job whatever the title says.
Do I need a PhD to be a machine learning engineer?
No. A PhD is not required for machine learning engineer, where you are measured on a system that works in production. It is the norm for research scientist and usually preferred for applied scientist, though several large employers run applied scientist ladders that hire at master's level with a strong applied modelling record. For ML engineer, a bachelor's in computer science, maths, statistics or a quantitative engineering field is the common floor and a master's is common. Strong software engineers with one well-evidenced production model are hired into these roles routinely, and at some employers the easiest door is the 'software engineer, machine learning' posting, which runs a standard engineering loop with one ML round.
How is the machine learning engineer interview loop structured?
Typically a recruiter screen, a hiring-manager call, a coding round of 45 to 60 minutes, an ML system design round, an evaluation or modelling-depth round, and a deep dive on a project you shipped, running four to eight weeks end to end. The system design and evaluation rounds carry the weight and the project deep dive usually decides level, but the coding round is still where most candidates are eliminated: at large technology companies and quantitative firms it is a standard software-engineering round at medium to medium-hard difficulty, while at startups and non-technology employers it is more often Python and SQL data manipulation or an ML-flavoured implementation. Take-homes of four to eight hours are common at startups and mid-size companies and rare at large ones.
What do machine learning engineers get paid?
There is no single credible band, and the US Bureau of Labor Statistics has no SOC code for the title. Bracket it with OES 15-2051 (data scientists) and 15-1252 (software developers) and read the metropolitan-area tables rather than the national median, then use levels.fyi for the large-technology ladder level by level, and the ranges employers publish under pay-transparency laws in Colorado, California, Washington, New York, Illinois and a growing list of other states for a specific company. Employer type moves pay more than the title does: a frontier lab, a trading firm, a bank, a retailer and a hospital system can differ by more than a factor of two for the same nominal job, mostly through equity rather than base.
Is machine learning engineering still a good career now that models are commoditised?
Yes, but the valuable half moved. Pretraining consolidated into a handful of organisations, so training novel architectures is a narrow speciality. What grew is everything required to make a bought model work on a specific company's data and budget: evaluation, retrieval quality, serving and inference cost, agent reliability, post-training on proprietary data, and the governance records regulated employers now demand. Classical tabular and ranking ML also did not shrink; in most industries it remains the majority of models in production. The jobs are there, and they reward production judgement over technique collection.
What should an ML engineering portfolio project look like?
One deployed system, not five notebooks: a real dataset or corpus, a baseline you beat, an eval harness with a held-out set and a metric you argue for, a live endpoint with a measured p95 latency and a cost per thousand requests, and a README containing an error taxonomy counted from actual failures plus what you would do with another month. Write the weaknesses down, because interviewers trust a project that names where it fails far more than one claiming to work. Depth is the signal, since only real traffic produces the problems that prove you have operated something.
How do I prepare for the ML system design interview?
Practise a fixed order out loud and on a timer: frame the decision and the cost of each kind of error, state the latency budget and volume, name the baseline you would have to beat, go through data and labels including label delay and leakage before any model, choose the model with a stated trade-off, define offline and online evaluation, then serving, fallbacks, monitoring and rough cost, and finish by naming one alternative you rejected and the condition that would make it right. Rehearse three canonical prompts, namely fraud detection, support-ticket deflection over a help centre, and marketplace search ranking, because most questions are variants of those. For material, Chip Huyen's Designing Machine Learning Systems and AI Engineering cover the system-level frame, Aminian and Xu's Machine Learning System Design Interview matches the round's format, and the engineering blog of a company with your target job type supplies worked trade-offs.
What is the evaluation round and how do I prepare for it?
It is a dedicated round, increasingly common, on how you would know a model-powered system works and whether a change improved it. Expect concrete questions: how to measure a summariser that invents policy numbers, why accuracy rose while complaints rose, what to do with 200 labels and no budget, why your LLM judge agreeing with you almost every time is not reassuring. Prepare one real eval harness you can describe by its dimensions, covering dataset size and provenance, splitting, metrics and why, human calibration of any model judge, CI gating and a regression it caught, plus one story about an offline gain that did not survive online and the reason you found.
How do I get a machine learning job when every posting wants production experience I have not been allowed to get?
Create the production experience where you already have access rather than waiting for a title. The cheapest version is the decision at your current employer that a rule, a spreadsheet or a manual queue makes today: scope the smallest useful model for it, agree the metric with whoever owns the decision, build the eval before the model, ship behind a flag at low traffic, and write down what changed. Next cheapest is the unowned ML-adjacent work, meaning the label pipeline, drift monitoring, an eval harness or an inference cost reduction, then an internal loan or transfer to an ML team, which is a smaller field to compete in than the external market. If none of that exists, deploy a portfolio system for real users and log a counted failure taxonomy, and contribute to an open-source eval or serving project so a stranger can verify your work.
How has AI coding assistance changed what ML engineers are hired for?
It changed the typing, not the judgement. Assistants write training loops, dataloaders and eval scripts quickly, and some employers now permit one in the coding round and grade how well you direct it and whether you catch its errors, so ask which mode you will face. What they cannot do is notice that the split leaks a customer across train and test, that an eval is scoring a model against its own output, or that a prompt change improved your three examples and degraded the other thousand. The hiring bar moved toward exactly those judgements, which is why evaluation and system design now carry the loop.
Put this on a resume in about a minute
Paste your history once and point it at the Machine Learning Engineer posting you are looking at. No account, no card.
Build my resume free More roles