| What the role is in 2026 | Owning the path a model takes from a training run to a versioned artefact to a served, monitored, billable, rollbackable production dependency, and making that path repeatable for teams who are not you. You are measured on what other people can ship on your infrastructure and on whether it stays up. You are not measured on model accuracy. |
|---|---|
| Titles the same job is posted under | ML Platform Engineer, AI Platform Engineer, ML Infrastructure Engineer, AI Infrastructure Engineer, ML Systems Engineer, Inference Engineer, GPU Infrastructure Engineer, Production Engineer (ML), and plain DevOps Engineer or Platform Engineer with model serving in the requirements. Searching only the string MLOps misses most of the market. |
| Licence or credential gate | None. No licence, no registration, no mandatory certification, no required degree. Because there is no gate, the resume screen is the real filter, which is why measured outcomes matter more here than in gated professions. |
| Certifications that actually move a screen | CKA (Certified Kubernetes Administrator), the CNCF exam administered by the Linux Foundation, is the one platform leads respect, because it is hands-on in a live cluster rather than multiple choice. Check the current fee, exam length and renewal period on the Linux Foundation's own page rather than any figure you read second-hand, including this one. Then one cloud credential matching your target employer's cloud: AWS Certified Machine Learning Engineer Associate (MLA-C01), Google Professional Machine Learning Engineer, or Microsoft Azure AI Engineer Associate (AI-102). HashiCorp's Terraform Associate is cheap and signals the right habits. None of them substitute for one system you have operated, and a certification list stacked above thin experience reads as a substitute. |
| Realistic prep time from an adjacent job | From DevOps or SRE: roughly three to six months of evenings to add the model lifecycle vocabulary, one serving stack and one deployed artefact you can demo. From data engineering: six to nine months, because the missing piece is Kubernetes depth and service ownership including on-call. From data science or ML: often longer than people expect, because platform teams screen hard for operational maturity. The internal transfer, taking over the deployment path for the data team inside your current employer, converts faster than any external application. |
| Typical loop | Recruiter or platform-lead screen, hiring-manager call, a Python and Linux scripting round, a systems design round on model serving or ML platform architecture of 45 to 60 minutes, a live debugging round on a broken pipeline or a regressed deployment, and often an infrastructure-as-code review or a four to eight hour take-home. Three to six weeks end to end, commonly faster than modelling roles. At large technology companies the coding round may still be a standard software-engineering round at medium difficulty; ask the recruiter which it is. |
| Where to find real pay numbers | There is no dedicated BLS occupation code for MLOps. Employers classify these roles under BLS OES 15-1252 (Software Developers) or 15-2051 (Data Scientists), so use the metro-level OES figures for those codes as a lagging floor rather than a market read. For a current read, use the posted range in pay-transparency jurisdictions (including California, Colorado, Washington, New York, Illinois, Minnesota, Maryland, Massachusetts, New Jersey and Vermont), Levels.fyi for public-company bands, and the US Department of Labor OFLC disclosure data, which publishes actual offered base salaries by employer and job title and is free. |
| Work shape | On-call is normal and should be assumed unless a posting says otherwise, because you own running services. Mostly hybrid or remote. A meaningful share of roles are the only platform person at a company with models in production, which is the fastest learning and the heaviest pager. A growing share of AI platform roles involve no accelerators at all, because the models are bought through provider APIs and the job is gateway, quota, cost and evaluation work. |
What an MLOps engineer owns in 2026, and the titles it hides under
The unit of work is a platform, not a model. An MLOps engineer owns the path a model travels: training or fine-tuning compute, the data and model artefacts and their versions, the registry and promotion rules, the serving layer, the autoscaling and capacity, the observability that says whether it is working, the cost it incurs, and the rollback when it is not. The test of whether you did the job well is whether a data scientist who has never met you can get a model into production on Tuesday and you can prove on Friday which version answered a given request.
The title confuses people because it was coined for a 2019 to 2021 problem: notebooks that could not be deployed, models that trained once and rotted, no reproducibility. The tooling for that problem is now something you buy rather than build, although plenty of companies still have the problem because they never adopted the tooling. The 2026 problem is different and harder: serving large models on scarce and expensive accelerators at a defensible cost per unit of work, keeping non-deterministic multi-step systems observable, governing models you did not train and cannot inspect, and proving to an auditor or a customer what model version, prompt version and retrieval index produced an output. If you prepare for the 2021 version of this role you will interview badly.
Be clear about what you do not own, because candidates lose credibility by claiming it. You do not own whether the model is accurate. You own whether the number the model produced in production matches what the offline evaluation said it would, whether it arrived inside the latency budget, and whether anyone would notice if it quietly stopped being true. That distinction is the single clearest signal of whether someone has actually done this work.
The posting title is unreliable, and this costs people weeks of search time. Set a saved search on each of these, not just on MLOps:
- ML Platform Engineer and AI Platform Engineer. Usually the same job at a company with an internal platform team and multiple model-owning teams.
- ML Infrastructure Engineer and AI Infrastructure Engineer. Tilted towards compute, networking, storage and the cluster rather than the developer experience.
- Inference Engineer and GPU Infrastructure Engineer. The sharp end: throughput per accelerator, batching, quantisation, cost per million tokens. The variant whose postings most often carry the top of the range, and the smallest candidate pool.
- ML Systems Engineer and Production Engineer (ML). Often an SRE-flavoured framing, heavier on reliability and on-call.
- Platform Engineer or DevOps Engineer where the requirements mention model serving, GPUs, SageMaker, Vertex AI, Databricks, Kubeflow or vLLM. A large number of real MLOps jobs are posted this way because the hiring manager is a platform lead who does not use the word MLOps.
- LLMOps. Rare as a title, common as a section inside a posting. Treat it as a signal that the job is inference and evaluation rather than batch training.
- Data Platform Engineer with ML in the scope. Check carefully: this is sometimes a data engineering job with a fashionable paragraph attached, which is a legitimate job but a different one.
The five kinds of MLOps job, and the solo hire that is often the best first one
Read the nouns in the posting before you tailor anything. Five clusters account for almost all of the market, they want different things, and a single resume cannot win all five. The sixth entry below is not a cluster but a company situation, and for a career changer it is usually the best door.
The honest distribution matters here. Despite where the attention is, a large share of MLOps postings are still about classical models: batch scoring, feature pipelines, tabular models served on CPU, retraining schedules. Fraud, credit, pricing, forecasting, churn, recommendation and logistics systems did not get replaced by language models, and they are where most production ML revenue still sits. Candidates who can only talk about LLM serving screen out of these jobs, and candidates who dismiss gradient boosting as legacy screen out of them loudly.
- Classical and batch ML platform. Tells: feature store, point-in-time correctness, Airflow or Dagster, retraining, batch scoring, data drift, Spark or Databricks, model registry, SLA on a nightly job. What they test: pipeline reliability, data correctness, orchestration, idempotency and backfills, cost on data infrastructure. Most common cluster, especially in finance, insurance, retail, logistics and healthcare.
- Model-gateway AI platform, with no accelerators in it. Tells: Amazon Bedrock, Google Vertex AI or provider APIs from OpenAI and Anthropic, a gateway or proxy such as LiteLLM or Portkey, API keys and quotas per team, token budgets and chargeback, response caching, fallback routing when a provider degrades, prompt and config versioning, evaluation suites, guardrails, PII redaction, tracing. What they test: multi-tenancy, cost attribution, graceful failure when an upstream provider is slow or down, and evaluation discipline. This is a large and fast-growing share of enterprise AI platform work, it hires from a wider pool than self-hosted inference, and candidates who assume every AI platform role involves GPUs never see these postings for what they are.
- Self-hosted LLM inference platform. Tells: vLLM, SGLang, TensorRT-LLM, Triton, KServe, Ray Serve, time to first token, tokens per second, prefix or KV cache, quantisation, autoscaling GPUs, gateway, rate limits, cost per million tokens, evaluation gates, tracing. What they test: throughput and latency engineering under a cost ceiling, and whether you can measure before you tune. Highest pay, smallest candidate pool, and the cluster where a published benchmark of your own is worth more than any certification.
- Training and GPU infrastructure. Tells: multi-node training, NCCL, InfiniBand or EFA, Slurm or Kueue or Volcano, MIG, checkpointing, fault tolerance over long runs, driver and CUDA image management, capacity reservations, spot interruption handling. What they test: whether you have debugged a distributed job that hung at 70 percent and whether you know where the money goes. Concentrated in AI labs, autonomy and robotics companies, large platforms and HPC-adjacent employers.
- Lifecycle, governance and compliance platform. Tells: model inventory, lineage, approval workflow, audit trail, model risk management, SR 11-7, NIST AI Risk Management Framework, ISO/IEC 42001, EU AI Act readiness, retention, access control, documentation. What they test: whether you can make a reproducible, evidenced trail without stopping delivery. Common in banks, insurers, pharma, health systems and anyone selling into the EU. Often pays well and is less contested because it reads as unglamorous.
- The solo hire, which is a situation rather than a cluster. A company of 40 to 300 people with several models in production, data scientists deploying by hand, a provider or GPU bill nobody can explain, and no platform team. The posting may say MLOps Engineer or may say DevOps. You will do all five clusters badly at first and learn faster than in any specialised seat. The interview is usually short, the scope is enormous, and the evidence you leave with is worth more than the title. If you are converting from an adjacent job, this is the highest-probability door.
The doors in, and exactly what each background is missing
Nobody studies MLOps and then gets an MLOps job. Everyone arrives from somewhere, and hiring managers have a clear mental model of what each origin gets wrong. Name your own gap in the interview before they find it; it reads as self-awareness rather than weakness, and it is the fastest way to get the conversation onto what you can do.
One general rule before the specifics. Platform teams hire for operational maturity above tool knowledge. Having been woken at 03:00 by something you owned, having written the postmortem, having said no to a request that would have broken the platform: these are the stories that convert. Tool lists do not convert, because tools can be learned in a fortnight and judgement cannot.
- Door one, DevOps or SRE. The shortest bridge and the most common hire. You already have Kubernetes, Linux, Terraform, CI/CD, observability, incident practice and cost instincts, which is most of the job. What you are missing is that a model is an artefact that goes stale silently: no error, no alert, just a slowly worsening answer. Learn the lifecycle vocabulary and mean it: training-serving skew, point-in-time correctness, data drift versus concept drift, offline versus online evaluation, shadow deployment, champion and challenger, retraining triggers. Then deploy one real model yourself end to end so you can speak from a deployment rather than a diagram. Add enough accelerator literacy to discuss batching and utilisation. Three to six months of evenings is realistic.
- Door two, data engineering. You bring pipelines, orchestration, data correctness, schema evolution, backfills and SQL at scale, all directly relevant. What you are missing is service ownership: Kubernetes beyond using it, request latency as a first-class concern, autoscaling, SLOs, and carrying a pager for something that serves traffic. CKA is worth the money here precisely because it forces hands-on cluster work. Build and operate one always-on inference service, load test it, and break it on purpose. Six to nine months.
- Door three, data science or ML. You understand models, evaluation and why a metric moved, which is genuinely valuable and the reason this door exists. What you are missing is the production half: infrastructure as code, networking and DNS, container internals, secrets, least privilege, cost, and the discipline of changing production in small reversible increments. This conversion is harder than people expect, because platform leads worry that a modelling hire will treat the platform as a means to their own experiments. Counter it with evidence of operational work rather than modelling results: a Terraform module, a Helm chart, an on-call rotation you joined, an incident you ran.
- The side door, backend software engineering. Strong Python or Go, API design, systems thinking, and usually a cleaner coding round than the other doors. Missing: both the lifecycle and the cluster. Convert by taking the serving path at your current job, and note that the gateway cluster above is close to ordinary backend work with evaluation and cost attribution bolted on, which makes it the easiest first move from here.
- The inside door, and the one with the best odds. Stay where you are and take the work. Volunteer to own the deployment path for the data science team, put their worst manual process into CI, instrument the model or GPU spend, write the runbook. In nine to eighteen months you are not a career changer applying cold, you are the person who already does the job, with internal references and measured outcomes. Re-titling in place is a common origin story among people now doing this job, and it skips the screen entirely.
The stack employers list in 2026, and where depth actually pays
Postings list twenty tools. You cannot be deep in twenty and you do not need to be. Breadth gets you past the screen because keyword matching is real; depth in two or three areas is what wins the design and debugging rounds. The grouping below is what shows up in postings, with an honest note where the hype and the job diverge.
Choose your depth deliberately: one orchestrator, one serving stack, and the economics of whatever runs the model, meaning accelerators if you host and tokens if you buy. That combination maps to the fastest-growing variants of the role, and it is defensible in an interview because it is where measurement is possible.
- Non-negotiable base, assumed rather than listed as a differentiator: Linux, networking and DNS, containers and image building, Kubernetes, Helm, Terraform, Git, CI/CD and GitOps delivery (Argo CD or Flux), one major cloud, and Python good enough to write and test a service rather than a script. Weakness in any of these is disqualifying regardless of how much ML vocabulary you have.
- Orchestration and workflows: Airflow, Dagster, Prefect, Flyte, Metaflow, Argo Workflows, Kubeflow Pipelines. Airflow remains the most common thing you will inherit; Dagster shows up more often in newer stacks. Learn one properly, including retries, idempotency, backfills and what happens when an upstream dependency is four hours late.
- Experiment tracking, registry and lineage: MLflow, Weights and Biases, DVC. MLflow is the default you will be asked about. The concept being tested is whether a prediction from three months ago can be reproduced: which code, which data snapshot, which hyperparameters, which artefact hash.
- Serving and inference: vLLM, SGLang, TensorRT-LLM, NVIDIA Triton Inference Server, KServe, Ray Serve, BentoML, ONNX Runtime, and TorchServe on legacy stacks, which is no longer actively maintained and is something you migrate off rather than towards. Depth here pays more than anywhere else in 2026. Know continuous batching, paged and prefix KV caching, tensor parallelism, speculative decoding, quantisation formats and their quality cost, and the difference between latency, throughput and goodput under concurrency. The current frontier, worth being able to discuss even if you have not run it, is disaggregating prefill from decode and routing requests by KV cache locality, which is the organising idea behind newer stacks such as NVIDIA Dynamo and llm-d.
- Model gateways and provider operations, for the large share of jobs where the model is bought rather than hosted: LiteLLM or an equivalent proxy, Amazon Bedrock or Vertex AI as the provider layer, per-team keys and quotas, token budgets and chargeback, caching, retries and fallback routing across providers, and a rate limit strategy that fails gracefully instead of cascading. Nobody lists this in a tidy section, but it is what the job is in most enterprises.
- GPU and accelerator operations: scheduling and queueing with Kueue, Volcano or Slurm, MIG partitioning and time-slicing on accelerators that support it, node autoscaling with Karpenter or cluster-autoscaler, spot and reserved capacity strategy, DCGM metrics for real utilisation, NCCL for multi-GPU and multi-node, and the driver, CUDA and container-toolkit version matrix that breaks things. This is the scarcest skill set in the market because it is hard to practise without a budget, which is why even a small documented experiment stands out.
- Feature stores: Feast, Tecton, Databricks Feature Store, Hopsworks. Honest note: far fewer companies run a dedicated feature store than the discourse implies, and some that bought one regret it. The concept you must own is point-in-time correctness and offline-online consistency, with or without a product. Interviewers ask the concept and use the product name as shorthand.
- Vector and retrieval stores: pgvector first, because most teams start there and many should stop there, then Qdrant, Milvus, Weaviate, Pinecone, and OpenSearch or Elasticsearch for hybrid keyword plus vector search. The operational questions are index build time, memory footprint, recall after metadata filtering, and the reindex you must run when the embedding model changes. Vendor preference is not the question.
- Observability: Prometheus and Grafana, OpenTelemetry, Datadog, plus ML-specific layers for drift and data quality such as Evidently, Arize or WhyLabs, plus trace-level tooling for language-model systems such as Langfuse, Phoenix or LangSmith. OpenTelemetry's GenAI semantic conventions exist and are still marked experimental, which is worth knowing because it is why teams hand-roll span attributes. What matters is that one request's trace carries latency, token counts, cost, tool calls, retrieval hits and an evaluation score, and that you can slice it by model version.
- Managed platforms: Amazon SageMaker, now branded SageMaker AI, Google Vertex AI, Azure Machine Learning, Databricks. A large share of jobs are mostly making the managed platform work inside an enterprise, with its IAM, networking and quota realities. Candidates who treat this as beneath them screen out of a big chunk of the market, including some of the better-paid regulated-industry roles.
- Security and supply chain, increasingly in postings rather than implied: safetensors over pickle for untrusted weights, image signing and provenance, SBOMs, scanning model artefacts, secrets management, egress restriction and least-privilege credentials for anything a model-driven agent can call.
- Governance frameworks, for the regulated cluster: NIST AI Risk Management Framework, ISO/IEC 42001, SR 11-7 for US banks, and the EU AI Act, whose general-purpose model obligations began applying in 2025 while the high-risk timetable has been amended more than once. Check the official Commission timeline against today's date rather than trusting any summary, including this one, on exact dates.
How the hiring loop runs in 2026-27, round by round
Who screens you is the thing to understand first. Because there is no credential gate and no standard curriculum, resume screening is often done by the platform lead rather than by a recruiter working from a checklist. That is good news: a resume with measured infrastructure outcomes gets read properly by someone who recognises them. It also means vague bullets die instantly, because the reader has done the job.
The loop is shorter than an ML engineering loop, typically three to six weeks, and more weighted towards practical rounds. Startups compress to two or three conversations in a week. Large technology companies add rounds and may keep a standard software-engineering coding round, which is worth asking about explicitly on the recruiter call because it changes your preparation completely.
Round by round, and what each is actually testing:
- Recruiter or platform-lead screen, 20 to 30 minutes. A scope and vocabulary match. Have a 90-second answer to "what have you run in production" with one number in it. Ask three questions that reveal which cluster this is: how many models are in production, whether they are hosted or called through a provider API, and who carries the pager for them. The answers tell you what to prepare and sometimes tell you to withdraw.
- Hiring-manager call, 45 minutes. Ownership and judgement. What you owned versus what your team owned, the worst incident you handled, something you refused to build and why. This is where operational maturity is assessed, and it is weighted more heavily than most candidates assume.
- Scripting and coding round, 45 to 60 minutes. At most employers this is practical: parse and transform data in Python, write a small client with retries and backoff, fix a failing container build, debug a Linux problem, write a Kubernetes manifest or a Terraform resource correctly, use kubectl to find why a pod will not schedule. Some teams use Go. At large technology companies it may instead be a standard algorithms round at medium difficulty, which eliminates more platform candidates than anything else in the loop, so confirm which it is.
- Systems design round, 45 to 60 minutes. The prompt is concrete: design the serving stack for an open-weight model in the 70-billion-parameter class for a chat product with a stated concurrency, a time-to-first-token target and a monthly cost ceiling; design an ML platform for 30 data scientists and 60 models; design a gateway that lets eight teams call three providers with per-team budgets and a sane failure mode when one provider slows down; design the retraining and promotion path for a fraud model that must never be worse than the one it replaces. On the inference prompt the move that earns immediate credit is refusing to design until you have the token length distribution, the input to output ratio and whether streaming counts, because requests per second alone does not determine the hardware. The grading rubric is in the next section, because the same rubric decides the debugging round.
- Infrastructure-as-code review or take-home, four to eight hours when it appears. Common forms: write a Helm chart or Terraform module for a model service, containerise and deploy a provided model with autoscaling and a health check, or review a pull request seeded with real mistakes such as a hardcoded secret, a missing resource limit, a readiness probe that lies, or an autoscaler that will thrash. Take-homes appear mostly at startups and mid-size companies.
- Debugging round, and the one that decides the offer. Treated separately below.
- Incident and on-call round, 30 to 45 minutes. Walk through a real outage: how you noticed, what you did first, what you communicated, the root cause, and the change you made so it could not recur. Strong answers roll back before diagnosing when users are affected and say so. Weak answers are heroic stories with no follow-up action.
- Occasionally a cost round, almost always in the inference, GPU and gateway clusters. Expect to estimate out loud: cost per million tokens given a model size, an hourly accelerator rate, a batch size and a measured throughput, or the monthly bill for a given provider price and request volume, then argue which lever to pull first. Nobody expects precision; they expect the arithmetic to be attempted and the units to be right.
The debugging round, where the offer is decided, and the rubric behind it
Platform hiring has converged on this round because it is the only one that cannot be studied around. You are handed something broken that you did not build, with incomplete information, and asked to work. Interviewers are not grading whether you reach the planted root cause. They are grading method.
The rubric, which is also the rubric for the design round, is consistent across employers: do you form a hypothesis and state what evidence would kill it; do you check the cheap and likely thing before the clever one; do you know where state lives in the system; do you protect users first and diagnose second; do you say what you would have instrumented beforehand so this was visible; and do you notice what you do not know and ask instead of inventing.
The scenarios recur, and preparing actual chains of reasoning for these four covers most of what you will be asked:
- "The p99 latency of our inference service tripled after a deploy. The model did not change." Good answers ask what else shipped, because something always did: the serving image, the batching or max-concurrency configuration, a tokenizer version, the request mix including longer prompts, the cache hit rate, a node pool or driver change, a sidecar, the autoscaler's target, CPU-side pre-processing, or a partial rollout where only some replicas are bad so the mean looks fine and the tail does not. Then they say which metric or log would separate those within two minutes, and they mention rolling back while investigating.
- "A training job that succeeded yesterday now gets OOM-killed." Hypotheses worth naming: more data or a longer sequence-length tail, a changed base image with a different CUDA or framework version, gradient accumulation or batch-size config drift, checkpoint sharding, a memory leak in a dataloader, a co-scheduled pod on the same node, or a resource limit that was always too tight and only now exceeded. The strong move is to ask what changed in the artefact chain rather than to start tuning memory settings.
- "Our GPU fleet is at 15 percent utilisation and the monthly bill is enormous." This round separates people who have operated accelerators from people who have read about them. First, measure utilisation correctly, because the headline figure from nvidia-smi is not the same as SM activity or memory bandwidth utilisation from DCGM, and a service can look busy while doing almost no arithmetic. Then: batch size and continuous batching, concurrency at the client, replicas per GPU, MIG partitioning for small models, consolidating many under-loaded services onto fewer accelerators, scale-to-zero for development endpoints, queueing batch work onto idle capacity, right-sizing the accelerator to the model, and only then changing the purchasing strategy towards spot or reservations.
- "Predictions got quietly worse. No code changed." The best answers go to the data path before the model: a feature pipeline running late so the service reads stale values, an upstream schema change silently coerced to null, a unit change, a join that started dropping rows, a training-serving skew introduced when someone fixed a bug in the training transform and not the serving one. In a bought-model system add the ones specific to it: a provider changed the model behind a version alias, a prompt or system message was edited without review, the retrieval index was rebuilt with a different embedding model. Then ask the uncomfortable question: how long had this been true before anyone noticed, and what alert should have fired.
- One counter-question earns credit in every version of this round: "who is affected right now, and is there a safe rollback?" Candidates who triage before they theorise are read as people who have actually been on call.
What MLOps engineers are paid, and where to look it up instead of trusting a band
Any single salary number published for this title is close to meaningless, because the title spans a solo hire at a 60-person insurer and an inference engineer at an AI lab, and because aggregator figures are scraped from postings and self-reports with no normalisation. Naming the source beats naming a number, so here are the sources that are actually checkable.
Start with the official floor. There is no dedicated Standard Occupational Classification code for MLOps, so employers classify these roles under BLS Occupational Employment and Wage Statistics code 15-1252 (Software Developers) or 15-2051 (Data Scientists). Pull the metro-level figures for those codes in your area and treat them as a lagging, blended floor rather than a market read for this specialty.
Then get current numbers. In pay-transparency jurisdictions, including California, Colorado, Washington, New York, Illinois, Minnesota, Maryland, Massachusetts, New Jersey and Vermont, the posting itself must carry a range, so searching the same title filtered to those locations gives you real employer-stated bands. Levels.fyi is reliable for public-company levelling and total compensation structure. The US Department of Labor's Office of Foreign Labor Certification publishes disclosure data listing offered base salaries by employer and job title for sponsored roles, which is free, dull and unusually honest. For contract rates, ask two recruiters for their last three placements at your level.
What actually moves the number for this role, in rough order of effect:
- Cluster. Self-hosted inference and GPU work sits at the top, because the skill is scarce and the cost savings are directly attributable. Gateway and classical batch platform work pays in line with senior DevOps or data engineering at the same employer. Governance-heavy roles pay better than their reputation suggests in regulated firms and are less contested.
- Employer type. AI-native companies and large technology platforms pay well above enterprise IT for the same nominal title, with a bigger share in equity and a wider spread in outcome. Regulated enterprises pay less cash at the top end but often offer a more predictable ladder and better hours.
- Scale you can evidence. Owning a fleet, a budget and a set of internal customers moves level more than years of experience does. "Platform used by 9 teams, 60 models, 12 million predictions a day" changes the conversation, where that sentence is yours and true.
- Accountability. Carrying the pager and owning the accelerator or provider budget are both compensated, and both should be negotiated explicitly rather than assumed.
- Whether you are the only one. Solo platform roles at small companies usually pay less than the responsibility deserves, which is a reason to treat them as a two-year evidence-building move rather than a destination.
The resume when the work is invisible, the portfolio that proves it, and running the search
The central resume problem for this role is that good platform work is invisible by design. Nothing broke, deploys were boring, the bill went down. The fix is mechanical: every bullet is a before-and-after with a unit attached. Not "responsible for the ML deployment pipeline" but "cut commit-to-serving time from 9 days to 4 hours by replacing manual handoffs with a promotion pipeline gated on an eval suite". The reader is a platform lead who has done this work; a number they recognise buys more credibility than a paragraph of adjectives.
Use units that exist in this job: p50, p95 and p99 latency, time to first token, tokens per second per accelerator, requests per second, cost per million tokens or per thousand predictions, accelerator utilisation measured properly, deploy frequency, lead time from commit to served, change failure rate, mean time to recovery, pipeline SLA hit rate, training job success rate, number of teams and models served, requests per day, and absolute or percentage spend reduced. The bullets below are sentence shapes, not benchmarks: the bracketed figures are placeholders to be replaced with your own measurements, because a platform lead will ask how you measured any number you keep.
What gets ignored or actively hurts: a wall of tool logos with no outcome attached, certifications listed above experience, model accuracy figures that were not yours, "collaborated with cross-functional stakeholders", Kaggle placings, course completion certificates, and any sentence containing the word passionate. Add a short Platform section stating scale, because scale is the real credential here: cluster size, accelerator count, models in production, requests per day, teams served, annual infrastructure spend owned. Mirror the posting's exact tool nouns in the body text, since keyword matching is still how most applications are filtered, but only tools you would survive a question about.
On the portfolio, the trap is the toy project. A notebook, a Streamlit demo, a tutorial repo and an architecture diagram all prove nothing, because the thing being assessed is operation rather than construction. What counts is one repository and one short write-up covering a single small open model taken the whole distance: containerised, deployed to a real cluster (k3s on a cheap VM, or a single spot GPU node), served behind an API with vLLM or KServe, autoscaled, load tested with a published table of concurrency against latency and throughput, a dashboard screenshot from Prometheus and Grafana, an evaluation gate in CI that blocks promotion of a worse model, a canary and a rollback you actually executed, and a cost figure per unit of work with the arithmetic shown. Publish the bill. The whole thing can be built across four weekends for the price of a few rented GPU hours, and it answers in evidence the question every interviewer is trying to answer about you.
Two cheap additions multiply its value. Write a short teardown of an incident you caused on purpose, such as deliberately starving the autoscaler or shipping a tokenizer mismatch, describing how you detected it and what you changed. Platform leads read that and believe you. And contribute something small to one of the serving or orchestration projects you claim, because the code and the review are public, permanent and verifiable in a way that no bullet point is.
Running the search itself: keep saved searches on all the alternate titles above, apply hardest at companies that visibly have models in production and no platform team, and treat the solo hire and the regulated mid-size employer as the two highest-probability doors for a career changer. Cold outreach to a platform lead with one measured artefact and three sentences outperforms fifty applications. And interview them back: how many models in production, hosted or bought, who is on call for them, how long from commit to serving, who owns the model budget, and what happened the last time a model had to be rolled back. The answers tell you whether this is MLOps or a rebranded data engineering job, and asking them is itself a hiring signal.
- Shape, serving rebuild: "Rebuilt model serving on [stack]: p95 latency from [before] to [after] and cost per million tokens down [measured percentage] at equal quality, verified with a held-out eval suite run on every promotion."
- Shape, pipeline reliability: "Replaced hand-run nightly scoring with orchestrated DAGs with retries and backfills: SLA hit rate from [before] to [after] over two quarters."
- Shape, release safety: "Introduced shadow deployment and single-command rollback; change failure rate for model releases fell from [before] to [after] across [n] releases."
- Shape, cost: "Cut idle accelerator spend by consolidating [n] under-loaded endpoints onto MIG partitions and scaling dev endpoints to zero, saving [amount] a month against the verified bill."
- Shape, governance: "Built lineage so any prediction can be reproduced from artefact hash, data snapshot and prompt version; cleared a model-risk audit with no findings."
- Shape, gateway work where no models are hosted: "Put [n] teams behind one model gateway with per-team budgets, caching and provider fallback: spend attributable per team for the first time, and a provider outage degraded instead of failing."
What an MLOps engineer must know about AI in 2026-27
The AI question for this role is not whether you know about AI, since the job has always been AI infrastructure. It is whether you have the 2026 version of the craft. Start with the honest part, because overclaiming here is detected quickly: the core of the job changed less than the discourse suggests. It is still Linux, Kubernetes, networking, storage, infrastructure as code, CI/CD, observability, cost and on-call. A good platform engineer from 2021 is still a good platform engineer. Nobody is hiring an MLOps engineer who can discuss attention mechanisms but cannot explain why a pod is stuck in Pending.
Also honest: most models running in production are still small, and many organisations host nothing at all. Tabular classifiers, forecasters, rankers and recommenders served on CPU remain the majority of the fleet at most employers, and where large models are in use they are usually called through a provider API rather than served on accelerators the company owns. That means a lot of AI platform work in 2026 is gateway, quota, cost attribution, evaluation and tracing work with no GPU in sight. The strongest candidates can operate the hosted and the bought world and say plainly which they are deeper in.
What genuinely changed is the workload, and with it the economics. The expensive, scarce, politically contested resource became either the accelerator or the provider invoice, so a meaningful part of the job became capacity planning, budget enforcement and cost attribution, which was barely in the 2021 job description. Serving changed from a stateless CPU request-response problem into a throughput problem with a cache in it, where batching strategy, KV cache behaviour, sequence-length distribution and quantisation choices move cost by multiples rather than percentages. Observability changed from a handful of drift metrics into tracing non-deterministic multi-step systems where a single user request may involve retrieval, several model calls, tool invocations and a retry. Release engineering gained artefact types with no clean diff and no unit test: prompts, retrieval indexes, tool definitions, model versions. And security gained two new surfaces that land on the platform team by default, namely the model supply chain and the blast radius of a model that can call things.
What is being automated around you, and what is not, is worth being precise about. Coding assistants now write most first-draft Terraform, Helm charts, Dockerfiles, dashboard queries and runbook scaffolding, and that is fine; some employers will hand you an assistant in the interview and grade how you direct it and whether you catch what it got wrong, which is worth practising deliberately. What is not automated is every decision that carries consequence: what capacity to commit to, what the trust boundary is, which service may hold which credential, what to retire, what the SLO should be, and who is in command during an incident. An assistant will confidently generate a readiness probe that always returns healthy, an autoscaler configuration that oscillates, and a Kubernetes RBAC role far broader than needed. Reviewing that is now a named skill, not a bonus.
One piece of hype to handle carefully in interviews. Many postings now describe an agent platform, and in most companies that means a gateway, a tool registry, an evaluation suite, tracing and tight credentials, not an autonomous system doing novel work. Describing it in those plain terms is a credibility signal, because the hiring manager is usually the person who had to build it. The practical framing for the whole loop: tool familiarity is assumed and cheap; judgement under cost and reliability constraints is what gets graded; and judgement only reads as real when it comes attached to a decision you made, a number you measured, and an alternative you rejected.
Inference and token economics as an engineering discipline
In the fastest-growing variants of this role, the model bill is the single largest controllable line in the budget. Where you host, the levers are configuration rather than code: batching strategy, concurrency, KV and prefix cache reuse, sequence-length distribution, quantisation format, replicas per accelerator, model right-sizing. Where you buy, they are caching, routing a request to the cheapest model that passes evaluation, prompt and context size, and stopping one team from spending everyone's budget. Engineers who can move cost per unit of work by a multiple without degrading measured quality are scarce and know it.
Show it: One benchmark you ran yourself, published as a table: concurrency against throughput, p95 latency and cost per million tokens, across at least two configurations, with the quality check that proved the cheaper one was not worse. State the hardware or the provider, the model and the arithmetic. In the interview, be able to estimate aloud from an hourly accelerator rate or a published token price and a measured throughput figure.
Accelerator capacity, scheduling and real utilisation
Where a company hosts its own models, GPU supply is the constraint the roadmap actually hits, and the person who understands queueing, partitioning and measurement becomes central quickly. It also catches a common illusion: a fleet can look fully utilised by the naive metric while doing very little arithmetic, which means somebody is paying for idle silicon that appears busy.
Show it: Be able to distinguish the headline utilisation figure from SM activity and memory-bandwidth utilisation, and name the tooling you used to see it. Describe one decision you made between MIG partitioning, time-slicing, consolidation, scale-to-zero, queueing batch work onto idle capacity, and changing the purchase mix of spot against reserved capacity, and say what it saved.
Evaluation gates in the release path
When the model is bought rather than trained, the platform's contribution to quality is the gate: no artefact, prompt or index reaches production without passing a suite that someone trusts. This is one of the clearest dividing lines between an MLOps engineer who has worked on language-model systems and one who has only read about them, and it is the piece that makes promotion safe rather than hopeful.
Show it: A CI pipeline in your portfolio where promotion of a model, prompt or index version is blocked by a held-out evaluation suite with a stated threshold, including what happens on a regression, who can override, and how the baseline is versioned. Say out loud that you do not define quality, you enforce the definition the model owners agreed, and you make the result reproducible.
Tracing and observability for non-deterministic multi-step systems
A single request may fan out into retrieval, several model calls, tool use and retries, and the failure modes are rarely exceptions. They are a plausible but wrong answer, a slow tail, a retry storm, a cost spike, a cache that stopped being hit. Request-count and error-rate dashboards do not see any of that, so the platform has to carry richer telemetry than a conventional service does.
Show it: Show a trace design: one span tree per request carrying token counts, cost, latency per step, tool calls, retrieval hits, cache hit or miss, and an evaluation or guardrail score, all sliceable by model and prompt version. Name the stack, for example OpenTelemetry into your existing backend plus a trace-level tool such as Langfuse or Phoenix, and say which alert you would set on it first and why.
Model and prompt supply chain security, and agent blast radius
Downloaded weights are untrusted artefacts, pickle-format checkpoints can execute code on load, and a system that lets a model call tools has effectively granted a non-deterministic component your service account. Security teams have started asking about exactly this, and the questions arrive at the platform team because nobody else owns the deployment path.
Show it: State your defaults: safetensors for untrusted weights, scanning and provenance for model artefacts, signed images, SBOMs, secrets never in prompts or images, outbound egress restricted by default, and least-privilege short-lived credentials for every tool a model-driven system can invoke. Add the governance view: who is allowed to change a production prompt, through what review, with what audit record.
Reproducibility and audit as a product of the platform
Regulated employers, enterprise customers and incident reviews all eventually ask the same question: what exactly produced this output, and can you prove it. Meeting that with evidence rather than reconstruction is a platform capability, and it is what makes the governance cluster of jobs hire readily from candidates who can design it. The NIST AI Risk Management Framework, ISO/IEC 42001, SR 11-7 in US banking and the EU AI Act all ask for artefacts that infrastructure produces.
Show it: Describe the chain end to end for one system: model artefact hash or provider model version, training data snapshot, code commit, config, prompt version, retrieval index version, request and response logging with a retention policy, and an inventory entry with an owner. Mention that compliance deadlines move and that you check the official timeline rather than a summary, which signals that you have worked near this rather than memorised an acronym.
Directing coding assistants on infrastructure, and reviewing what they produce
Assistants now generate most of the first-draft YAML, Terraform and glue in this job, which raises the value of review and lowers the value of typing. The failure modes are specific and dangerous in platform code: health checks that lie, missing resource limits, over-broad permissions, autoscaler settings that oscillate, and plausible configuration for a provider version you are not running. Some employers test this directly in the coding round.
Show it: Practise the mode you will be tested in: generate, then review out loud against a checklist you can name, then verify by running it. In interviews, say which classes of assistant output you distrust by default in infrastructure code, and give one example of a generated configuration you caught before it shipped.
What a screen is looking for
These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.
- MLOps
- ML platform engineering
- ML infrastructure
- Kubernetes
- CKA
- Helm
- Docker
- Terraform
- Infrastructure as code
- CI/CD
- GitOps
- Argo CD
- Argo Workflows
- Linux
- Python
- Go
- Bash
- AWS
- Google Cloud Platform
- Microsoft Azure
- Amazon SageMaker
- Google Vertex AI
- Azure Machine Learning
- Databricks
- Amazon Bedrock
- Kubeflow
- KServe
- Ray
- Ray Serve
- vLLM
- SGLang
- TensorRT-LLM
- NVIDIA Triton Inference Server
- NVIDIA Dynamo
- llm-d
- BentoML
- ONNX Runtime
- Model serving
- Model gateway
- LiteLLM
- Continuous batching
- KV cache
- Prefix caching
- Prefill decode disaggregation
- Speculative decoding
- Quantization
- Tensor parallelism
- Time to first token
- p95 latency
- Throughput per GPU
- Cost per million tokens
- Token budgets
- Cost attribution
- GPU scheduling
- MIG
- Time-slicing
- Kueue
- Volcano
- Slurm
- Karpenter
- Spot instances
- Capacity reservations
- DCGM
- NCCL
- CUDA
- Distributed training
- Checkpointing
- MLflow
- Weights & Biases
- DVC
- Model registry
- Model versioning
- Artifact promotion
- Lineage
- Reproducibility
- Airflow
- Dagster
- Prefect
- Flyte
- Metaflow
- Feature store
- Feast
- Tecton
- Point-in-time correctness
- Training-serving skew
- Data drift
- Concept drift
- Model monitoring
- Batch scoring
- Retraining pipeline
- Shadow deployment
- Canary release
- Blue-green deployment
- Rollback
- A/B testing
- Eval harness
- LLM-as-judge
- Evaluation gate
- Guardrails
- Retrieval-augmented generation (RAG)
- Vector database
- pgvector
- Qdrant
- Milvus
- Weaviate
- Pinecone
- OpenSearch
- Hybrid search
- Embeddings
- Reindexing
- Prometheus
- Grafana
- OpenTelemetry
- Datadog
- Langfuse
- Arize
- Evidently
- Distributed tracing
- SLO
- SLI
- Error budget
- On-call
- Incident response
- Postmortem
- MTTR
- Change failure rate
- Autoscaling
- Horizontal Pod Autoscaler
- Scale to zero
- Resource limits
- RBAC
- Secrets management
- Least privilege
- safetensors
- Image signing
- SBOM
- Supply chain security
- Prompt injection
- Cost optimization
- FinOps
- NIST AI Risk Management Framework
- ISO/IEC 42001
- SR 11-7
- EU AI Act
- Model inventory
- Audit trail
Mistakes that cost people this job
Searching only for the title MLOps Engineer, and concluding the market is small.
Run saved searches on ML Platform Engineer, AI Platform Engineer, ML Infrastructure Engineer, AI Infrastructure Engineer, ML Systems Engineer, Inference Engineer, GPU Infrastructure Engineer, Production Engineer (ML), and Platform or DevOps Engineer postings whose requirements mention model serving, GPUs, SageMaker, Vertex AI, Bedrock, Databricks, Kubeflow or vLLM. Many of the best-paid versions of this job are never posted with MLOps in the title, because the hiring manager is a platform lead who does not use the word.
Preparing for the 2021 version of the role: notebooks that cannot be deployed, reproducibility, and a diagram of a generic ML pipeline.
Prepare for what is scarce in 2026: serving throughput and latency under a cost ceiling, accelerator capacity and scheduling, evaluation gates in the release path, cost attribution across teams, and tracing multi-step non-deterministic systems. Keep the lifecycle fundamentals, because they are still assumed, but do not make them your headline.
Assuming every AI platform job involves self-hosted models on accelerators you own.
Read the posting for whether models are hosted or bought. A large share of enterprise AI platform work is a gateway in front of provider APIs: per-team keys and quotas, token budgets and chargeback, caching, fallback routing, evaluation and tracing, with no GPU anywhere. It hires from a wider pool and is the easiest first move from backend or DevOps, and candidates who only prepared for vLLM and MIG interview badly for it.
Leading with a wall of tool names because the posting listed twenty of them.
Mirror the posting's nouns in the body for keyword matching, but build the resume around before-and-after outcomes with units: latency, cost per unit of work, deploy lead time, SLA hit rate, utilisation, teams and models served. The screener is usually a platform lead who has done the job and will ask about any tool you list, so only claim what survives a question.
Claiming model accuracy improvements as your own achievement.
Claim the platform contribution, which is what you were actually hired for: that the number offline matched the number in production, that it arrived inside the latency budget, that a regression was blocked before promotion, and that a bad release was reversible in one command. Taking credit for someone else's model quality is one of the fastest ways to lose a platform lead's trust.
Treating the managed platform jobs, meaning SageMaker, Vertex AI, Azure ML and Databricks work inside an enterprise, as beneath a real engineer.
Take them seriously, because a large share of the market and some of the better-paid regulated-industry roles are exactly this, with genuinely hard IAM, networking, quota and governance problems inside the constraints. Disdain for managed platforms is audible in interviews and screens you out of that whole segment.
Building a portfolio project that constructs something but never operates it: a notebook, a Streamlit demo, a tutorial repo, an architecture diagram.
Take one small open model the whole distance and operate it: containerised, deployed to a real cluster, served with autoscaling, load tested with a published concurrency against latency and throughput table, dashboarded, gated in CI by an evaluation suite, canaried, and rolled back for real. Publish the cost per unit of work with the arithmetic, and publish the bill.
Going into the debugging round intending to identify the planted root cause quickly.
Perform the method, because that is the rubric. State a hypothesis and what evidence would disprove it, check the cheap and likely cause first, ask who is affected and whether a safe rollback exists before theorising, and finish by naming the instrumentation that would have made this visible in two minutes instead of two hours.
Assuming the coding round is a formality because this is an infrastructure job.
Ask the recruiter directly which kind it is. At most employers it is practical Python, Linux, Kubernetes and Terraform work, sometimes Go. At large technology companies it may still be a standard algorithms round at medium difficulty, which eliminates more platform candidates than any other round in the loop. Also ask whether an AI assistant is permitted, so you practise in the mode you will be graded in.
Quoting a salary band from an aggregator during negotiation.
Cite checkable sources: employer-stated ranges in pay-transparency jurisdictions for the same title and level, Levels.fyi for public-company levelling, US Department of Labor OFLC disclosure data for actual offered base salaries by employer, and BLS OES 15-1252 or 15-2051 for the regional floor. Then negotiate on the two things specific to this role that are routinely compensated and routinely forgotten: carrying the pager, and owning the model or accelerator budget.
Questions people ask
What does an MLOps engineer actually do?
An MLOps engineer owns the path a machine learning model takes from a training run to a production dependency: the compute it trains on, the versioned data and model artefacts, the registry and promotion rules, the serving layer and its autoscaling, the observability that shows whether it is still working, the cost it incurs, and the rollback when it is not. The unit of work is a platform other teams ship on, not a model. The role does not own model accuracy; it owns whether production behaviour matches what the offline evaluation promised, inside the latency and cost budget, reversibly. Where a company buys models through provider APIs rather than hosting them, the same job becomes gateway, quota, budget, evaluation and tracing work.
Do I need a degree or certification to become an MLOps engineer?
No. There is no licence, no registration and no mandatory certification, and no specific degree is required, though most people hired into these roles have a technical degree or equivalent production experience. Because there is no credential gate, the resume screen is the real filter and measured outcomes do the work a credential would do elsewhere. Two certifications help a career changer get read: CKA, the CNCF Kubernetes exam administered by the Linux Foundation, because it is hands-on in a live cluster, and one cloud ML certification matching your target employer's cloud, such as AWS Certified Machine Learning Engineer Associate, Google Professional Machine Learning Engineer, or Azure AI Engineer Associate. Neither replaces one system you have operated end to end, and certifications listed above thin experience read as a substitute for it.
How do I move from DevOps into MLOps?
It is the shortest available bridge, because Kubernetes, Linux, Terraform, CI/CD, observability, cost control and on-call are most of the job already. The missing piece is that a model is an artefact that goes stale silently, with no error and no alert, so learn the lifecycle properly: training-serving skew, point-in-time correctness, data drift versus concept drift, offline versus online evaluation, shadow deployment, champion and challenger, retraining triggers. Then deploy one real model yourself end to end so you speak from a deployment rather than a diagram, and add enough accelerator literacy to discuss batching and utilisation. Three to six months of evenings is realistic, and taking over the deployment path for the data team inside your current job converts faster than applying cold.
Which stack do employers actually list for MLOps roles in 2026?
Assumed as a base: Linux, networking, containers, Kubernetes, Helm, Terraform, Git, CI/CD with GitOps delivery, one major cloud and solid Python. Then one orchestrator, usually Airflow or Dagster; tracking and registry, usually MLflow; a serving stack, where vLLM, KServe, Ray Serve and NVIDIA Triton dominate and TensorRT-LLM or SGLang appear in performance-sensitive shops; or, where models are bought rather than hosted, a gateway layer such as LiteLLM in front of Amazon Bedrock, Vertex AI or provider APIs with per-team quotas and budgets. Add accelerator operations covering scheduling, MIG, autoscaling, spot and reserved capacity and DCGM metrics; a vector store, most often pgvector before a specialised engine; and observability built on Prometheus, Grafana and OpenTelemetry with a drift or trace-level layer on top. Managed platforms, meaning SageMaker, Vertex AI, Azure ML and Databricks, appear in a large share of enterprise postings. Go deep in one orchestrator, one serving stack and the economics of whatever runs the model; breadth gets the screen, depth gets the offer.
What does the MLOps interview loop look like, and which round decides it?
Recruiter or platform-lead screen, hiring-manager call, a practical scripting round in Python and Linux, a systems design round on model serving or ML platform architecture, a live debugging round on a broken pipeline or a regressed deployment, and often an infrastructure-as-code review or a four to eight hour take-home, plus an incident and on-call conversation. Three to six weeks end to end, commonly faster than machine learning engineering loops. The debugging round decides the offer, and it is graded on method rather than on finding the planted root cause: state a hypothesis and the evidence that would kill it, check the cheap and likely cause before the clever one, ask who is affected and whether a safe rollback exists before theorising, and name the instrumentation that would have made the problem visible in two minutes. Recurring prompts are a p99 latency jump after a deploy where the model did not change, a training job that suddenly gets OOM-killed, a GPU fleet at 15 percent utilisation with a large bill, and predictions that quietly got worse with no code change. At large technology companies the coding round may be a standard algorithms round rather than a practical one, so ask which it is before you prepare.
How is MLOps different from machine learning engineering?
A machine learning engineer owns a model's behaviour and the system it runs inside, including model choice, features or retrieval, and offline evaluation quality. An MLOps or ML platform engineer owns the infrastructure that other people's models run on: the deployment path, the serving layer, capacity, observability, cost and reliability, measured in what other teams can ship and whether it stays up. The reliable tell in a posting is whose work is being described. If training data, labels, offline metrics and model quality are the deliverables, it is ML engineering. If clusters, pipelines, serving, SLOs, utilisation and spend are the deliverables, it is MLOps whatever the title says.
What should an MLOps engineer put on a resume when the work is invisible infrastructure?
Convert every responsibility into a before-and-after with a unit. Use the units that exist in this job: p95 and p99 latency, time to first token, tokens per second per accelerator, cost per million tokens or per thousand predictions, accelerator utilisation measured properly, deploy frequency, lead time from commit to served, change failure rate, mean time to recovery, pipeline SLA hit rate, teams and models served, and spend reduced. One worked example: "cut commit-to-serving time from 9 days to 4 hours by replacing manual handoffs with a promotion pipeline gated on an eval suite" says more than any paragraph of adjectives. Add a short Platform section stating scale, because scale is the credential here: cluster size, accelerator count, models in production, requests per day, infrastructure spend owned. Drop tool walls without outcomes, certifications stacked above experience, model accuracy figures that were not yours, and the word passionate. Only keep a number you can explain how you measured.
What portfolio project proves I can do MLOps work?
One repository that takes a single small open model the whole distance and operates it: a container, a deployment to a real cluster such as k3s on a cheap VM or one spot GPU node, serving behind an API with vLLM or KServe, autoscaling, a load test published as a table of concurrency against latency and throughput, a Prometheus and Grafana dashboard, an evaluation gate in CI that blocks promotion of a worse model, and a canary plus a rollback you actually executed. Include cost per unit of work with the arithmetic shown, and publish the bill. Then add a short teardown of an incident you caused on purpose, describing how you detected it and what you changed, because that is the document platform leads believe. A notebook, a Streamlit demo or an architecture diagram proves construction, and the thing being assessed is operation.
How much do MLOps engineers earn, and where can I check?
Any single published band is close to meaningless, because the title spans a solo hire at a small insurer and an inference engineer at an AI lab. Check it yourself from four sources: employer-stated ranges in postings in pay-transparency jurisdictions such as California, Colorado, Washington, New York, Illinois, Minnesota, Maryland, Massachusetts, New Jersey and Vermont; Levels.fyi for public-company levelling and total compensation structure; the US Department of Labor OFLC disclosure data, which lists actual offered base salaries by employer and job title; and BLS Occupational Employment and Wage Statistics codes 15-1252 (Software Developers) and 15-2051 (Data Scientists) for a regional floor, since MLOps has no dedicated occupation code. What moves the number most is which variant of the role you are in, with self-hosted inference and GPU work at the top, followed by employer type, the scale you can evidence, and whether you carry the pager and own the model budget.
Is AI making MLOps obsolete, or is the role growing?
The hype in both directions is wrong. Coding assistants now write most first-draft Terraform, Helm charts, Dockerfiles and dashboard queries, which reduced the typing but not the job, because the decisions that carry consequence are unchanged: what capacity to commit to, where the trust boundary sits, which service holds which credential, what the SLO should be, and who is in command during an incident. Meanwhile the workload the role operates got more expensive and harder to observe, which pushed demand up rather than down. The honest caution is that the core skill set did not transform: Kubernetes, Linux, infrastructure as code, cost and on-call are still the bulk of it, most models in production are still small classical models served on CPU, and plenty of AI platform jobs involve no accelerators at all because the models are bought through an API.
Put this on a resume in about a minute
Paste your history once and point it at the MLOps Engineer posting you are looking at. No account, no card.
Build my resume free More roles