| Licence required: none | There is no licence, no legally required degree and no mandatory certification to work as a DevOps engineer. Degree filters still exist in graduate schemes at large employers and in some skilled-worker visa routes, and they fade once you have two or three years of production experience. What replaces a credential is a system you personally operated: you can name its failure modes, its monthly cost, who got paged when it broke and what you changed so it stopped breaking. |
|---|---|
| The certifications that actually move a screen | Three tiers. Hands-on and respected: Certified Kubernetes Administrator (CKA), and its siblings CKAD and the security-focused CKS, all performance-based exams sat in a live terminal against running clusters. Useful mainly for passing recruiter and automated filters: AWS Solutions Architect Associate, AWS Certified DevOps Engineer Professional, Microsoft AZ-104 and AZ-400, Google Professional Cloud DevOps Engineer, HashiCorp Terraform Associate. Largely inert for this role: generic Agile, Scrum Master and ITIL Foundation certificates, which no DevOps hiring manager reads as technical evidence. A certification gets you read; it does not get you hired. |
| Time to each certification, and the prerequisites | Terraform Associate: a multiple-choice exam of about an hour, realistically a few weeks of evenings if you already write HCL, credential valid two years. CKA: two hours in a live terminal solving tasks against running clusters, open book against the official Kubernetes documentation, realistically one to three months of deliberate practice because speed under time pressure is what is being tested, valid two years, and the exam environment tracks a specific Kubernetes version that moves through the year. CKS requires an active CKA before you can sit it. Cloud associate certifications: typically four to eight weeks of study, valid three years, and the Azure expert-level AZ-400 assumes an associate certification first. Confirm the current exam version, duration, prerequisites and curriculum on training.linuxfoundation.org, developer.hashicorp.com or the cloud vendor's certification page before you book, because exam contents are revised on their own schedule. |
| The one credential gate that is real: cleared work | US defence and federal contracting is the exception where paperwork genuinely gates the job. A security clearance (Secret, Top Secret, TS/SCI, sometimes with a polygraph) must be sponsored by an employer, so you cannot obtain one in advance, and an existing active clearance commands a premium and a much shorter hiring process. Privileged roles on Department of Defense systems also carry baseline certification requirements under the department's cyber workforce directive, with CompTIA Security+ the usual entry answer; check the current approved list for the exact work role rather than trusting a recruiter's summary, because that list is revised. Expect STIG hardening, FIPS-validated cryptography, air-gapped or classified environments and hardened internal container registries to appear in the interview. |
| What screening actually looks for | Five evidence types, in roughly this order of weight: infrastructure as code you wrote and maintained (Terraform or OpenTofu modules, Pulumi, CloudFormation, Ansible) with remote state, locking and module versioning handled properly; a CI/CD pipeline you built or rebuilt, with a build-time or lead-time figure attached; container and orchestration work where you can say clearly who ran the control plane and who ran the workloads; observability you instrumented, including at least one SLO you defined and defended; and an incident you were paged for, diagnosed and permanently fixed. A tool list with no scale and no outcome is the most common reason a technically capable candidate gets filtered. |
| The loop, and where people get cut | Recruiter screen (cloud, orchestrator, IaC tool, on-call expectation, salary), a technical screen on Linux, networking and troubleshooting, then a practical stage that is either a take-home repository (a module, a pipeline, a Dockerfile and a README) or a live exercise in a deliberately broken environment, then an infrastructure design round, an incident and postmortem round, and a hiring manager conversation. Two to five weeks end to end. The two stages that reject most candidates are the live troubleshooting round, where people guess instead of gathering evidence, and the design round, where people recite components instead of sizing the system and naming tradeoffs. |
| The numbers to carry out of your current job | Before you leave any job, write down: deploys per day or week and how that changed, lead time from merge to production, change failure rate, time to restore service after a failed deployment, availability against the SLO you actually published, p99 latency for the service you owned, monthly cloud spend and what you cut it to, cluster and node counts, number of services and number of engineers your platform served, and pipeline duration before and after your work. These are the figures the resume needs, and you cannot reconstruct them once you lose access to the dashboards and the billing console. |
| Pay: cite the source, not an average | The US Bureau of Labor Statistics has no DevOps engineer occupation code, which is why aggregator figures for this title vary so widely. Employers most often code the work to SOC 15-1252 Software Developers, 15-1244 Network and Computer Systems Administrators, or 15-1299 Computer Occupations All Other, so read the OES tables for all three plus your metro area. For live ranges, read postings from pay-transparency jurisdictions (Colorado, California, Washington, New York, Illinois, Minnesota, Maryland, New Jersey, Massachusetts, Hawaii, Vermont and Washington DC among them, and the list keeps growing, so check what applies where you are looking) and filter on title and seniority. In the UK, ITJobsWatch publishes permanent and contract medians by individual skill, which is unusually useful for pricing a Kubernetes or Terraform specialism. For large technology employers, levels.fyi carries self-reported total compensation by level; treat it as directional rather than audited. |
DevOps engineer is four different jobs: read the posting before you write a word
The title covers work that has very little in common day to day. One posting is a platform job: you build and run the internal tooling that two hundred application engineers deploy through, and your customers are developers. Another is a cloud infrastructure job at a company with twelve engineers, where you are the only person who understands the account structure, the network and the bill. A third is a build and release job in a regulated bank, where most of the difficulty is change control, audit evidence and separation of duties rather than technology. A fourth is embedded with one product team, writing pipelines and Terraform alongside the people whose code you deploy. The interviews, the resume and the pay differ, so read the posting for four things before you decide it is for you.
First, the stack, named exactly. Which cloud and which services, which orchestrator (managed Kubernetes, Nomad, ECS, plain virtual machines, or a serverless platform), which IaC tool and whether modules are versioned or everything lives in one root module, which CI system, and which observability vendor. Note whether the posting says Terraform or OpenTofu: the licence change pushed some employers to the fork, and postings now name one, the other, or both, which tells you something about how recently the team revisited its tooling. A posting that lists all three clouds and six CI tools is usually describing a wish rather than a system, and it tends to mean the team does not know what it has.
Second, the on-call line, which candidates skip and then regret. Ask whether there is a rota, how many people are on it, the frequency (one week in four is comfortable, one week in two is a staffing problem), whether it is compensated and how, the typical number of out-of-hours pages per rotation, and whether DevOps carries the pager for application services or only for the platform. A team that pages the platform engineer for every application error has an ownership problem that no amount of tooling will fix.
Third, who the customer is. If the posting says "internal developer platform", "golden path", "paved road", "Backstage" or "self-service", the job is product work with engineers as users, and the interview will ask how you learn what developers need and how you measure adoption. If it says "support the infrastructure", "ticket queue" or "maintain the pipelines", you are closer to operations, and the interview will be about response and stability rather than platform design.
Fourth, the maturity of what exists. There are three very different situations: greenfield (you choose, and you own the consequences for years), stable and mature (you optimise, upgrade and reduce cost, and the hard part is changing things without breaking them), and rescue (there is one Jenkins server nobody can rebuild, a hand-built cluster, no state locking, and the previous engineer has left). Rescue jobs teach you more in a year than anything else and are brutal if you walk in expecting mature. Ask directly: what is the oldest thing in production that nobody wants to touch, and how did the last person in this role leave.
One practical note on your search. Restricting yourself to the exact title "DevOps Engineer" hides a large share of the jobs you want. The same work is advertised as Platform Engineer, Site Reliability Engineer, Cloud Engineer, Infrastructure Engineer, Build and Release Engineer, Production Engineer, Systems Engineer and, increasingly, platform or infrastructure roles with an AI or ML qualifier. Search all of them, and read the responsibilities rather than the title.
- Decode the posting: "paved road", "golden path", "internal developer platform" and "self-service" mean platform product work. "Ticket queue", "support requests" and "maintain" mean operations. "Greenfield" and "first infrastructure hire" mean you will own the architecture and the bill.
- Ask the on-call questions in the first call, not at offer stage: rota size, frequency, compensation, typical out-of-hours pages per rotation, and what the worst week in the last six months looked like.
- A posting that lists AWS, Azure and GCP together with Jenkins, GitLab CI, GitHub Actions, Terraform, Puppet, Chef and Ansible is usually a list of everything anyone there has ever used. Ask which of them were touched in the last quarter.
- Regulated environments (finance, health, defence, payments) are worth taking seriously. Shipping daily inside a SOX-controlled or FedRAMP-scoped change process is a scarce and well-paid skill, and the experience transfers upward.
- Managed service providers and consultancies hire more volume than product companies and expose you to many environments fast. The pace is punishing, the learning is real, and one or two years there is a legitimate way into a better internal job.
- Contract and day-rate work is a larger share of this market than of most engineering roles, particularly in the UK and Europe. It pays more per day, carries no notice protection, and recruiters price it by individual skill, which is why skill-level pay data is useful here.
What actually gates this job: no licence, and a short list of credentials that help
Nothing stops you calling yourself a DevOps engineer. There is no board, no registration, no continuing education requirement and no legal standard of competence. That makes the hiring process the entire filter, and it makes it heavily evidence-driven: hiring managers assume the resume is optimistic and design the loop to find out what you have actually run.
Certifications are worth being precise about, because the advice sold to beginners is inflated. The Certified Kubernetes Administrator is the one with the most credibility, for a structural reason: it is a performance-based exam sat in a live terminal where you fix real clusters against the clock, so it cannot be passed by memorising a question bank. CKAD covers the application-developer side of Kubernetes and CKS the security side, which is the one to hold if you want to be the person who owns admission control and cluster hardening; CKS requires an active CKA first. A cloud certification (AWS Solutions Architect Associate, AZ-104, or the Google Professional Cloud DevOps Engineer) is mostly useful for getting read: it answers the recruiter's stack-match question and satisfies partner-status requirements at consultancies, which is why consultancies will often pay for it. The HashiCorp Terraform Associate is cheap and quick, and proves familiarity rather than depth.
What a certification does not do is survive contact with the technical round. If you hold a CKA and cannot explain why a pod is stuck Pending, what a readiness probe does differently from a liveness probe, or what happens to running workloads when you drain a node with a PodDisruptionBudget in the way, the certificate works against you. Hold them for the doors they open, and build the artefacts that justify them.
Degrees matter less here than in almost any other engineering role, and many strong DevOps engineers came through help desk, networking, QA, broadcast engineering, the military or a trade. The exceptions where a degree still binds: graduate programmes at large employers, some government and defence contracts, and skilled-worker visa routes where a degree or an assessed equivalent is part of the eligibility calculation. If you are routing through a visa, check that specific scheme's requirement rather than assuming.
The thing that functions as the real credential is production accountability. Interviewers are listening for signals that you have been responsible when something was down: you can describe a page that woke you up, what you looked at first, what the cause turned out to be (usually not what it looked like), what you did to restore service, and the structural change you made afterwards. In most rooms, two years of genuine on-call with no certifications beats five certifications with no production history, and hiring managers will tell you so directly.
- Worth holding, roughly in order of impact for this role: CKA, a cloud certification matching the employer's cloud, Terraform Associate, CKS if you are targeting security-adjacent platform work, Security+ if you are targeting US defence contracting.
- Worth little here: Scrum Master, SAFe, generic Agile certificates, ITIL Foundation on its own, and vendor certificates in tools the employer does not use.
- Red Hat certifications (RHCSA, RHCE) still carry weight in enterprise Linux, telecoms and government shops running RHEL and OpenShift. They are performance-based too, which is why they travel.
- If you are choosing one thing to spend three months on with no DevOps job yet, build and document a real system rather than take a second certification. The second certification adds little; the first real artefact changes your screening rate.
- Open source contributions carry unusual weight in this ecosystem because so much of the stack is open. A merged fix to a Terraform provider, a Helm chart, an operator, or the Kubernetes documentation is checkable evidence a recruiter can click.
- A conference talk or a published incident writeup functions similarly. Hiring managers for platform roles read writing as a core skill, because most of the job is persuading other engineers to adopt something.
How DevOps hiring actually works in 2026-27
The loop is practical and getting more so. Whiteboard algorithm puzzles have largely left this discipline, and what replaced them is work sampling: you will be asked to fix something, build something small, or design something sized. Expect two to five weeks start to finish, and expect at least one stage where you share a screen and touch a terminal.
The recruiter screen is a stack-match screen. Which cloud, which orchestrator, which IaC tool, how big was the estate, are you on call now, where are you located and what is your number. Answer with specifics and scale rather than adjectives: "AWS, EKS, four clusters, about ninety services, Terraform with versioned modules and remote state, GitHub Actions, Datadog, one week in five on call" tells a recruiter everything in one breath and gets you forwarded. Vague answers get filtered because the recruiter cannot map them to the requisition.
The first technical screen is usually Linux, networking and troubleshooting, conducted as a conversation. Common ground: what happens between a request leaving a browser and reaching a container, how DNS resolution works and how it fails, what you look at when a host shows high load average but low CPU, how TLS termination and certificate renewal are arranged, what connection pool exhaustion looks like from the outside, how you would find what is filling a disk, and what a process in uninterruptible sleep means. The interviewer is checking whether your knowledge is causal or memorised. "I would check the logs" is the answer that ends interviews; naming which log, what you expect to see, and what each possibility would rule out is the answer that continues them.
The practical stage takes one of three shapes. A take-home repository is the most common: provision something with Terraform or OpenTofu, containerise a supplied application, build a pipeline, write a README. A live exercise is increasingly used instead because it is harder to outsource: you are dropped into a broken environment with a cluster that will not schedule, a pipeline that fails at a specific step, or a service returning errors, and you have forty-five to ninety minutes to diagnose it while narrating. A third variant, which a growing number of teams now use, asks you to review a pull request containing infrastructure changes and say what you would block. That is a direct test of the skill that matters most now that generating configuration is nearly free.
The design round is where level is decided. The prompts are operational rather than abstract: design the deployment pipeline for sixty services and three hundred engineers; take a release process that runs fortnightly over four hours and get it to daily; design multi-region failover for a stateful service and say where the data goes and what you lose; design secret distribution for two hundred services with no long-lived credentials; the Kubernetes bill doubled last quarter, find out why. Strong answers start by asking for numbers (request rate, data volume, growth, recovery objectives, team size, compliance constraints), do arithmetic out loud, choose the smallest thing that meets the requirement, and say explicitly what they are not building yet and what signal would change that.
The incident round is a behavioural interview with a technical spine. You will be asked to walk through a real outage. The structure that works: what the symptom was and how you learned about it (alert, customer, dashboard), what your first three actions were and why, what the cause turned out to be, how you mitigated as distinct from how you fixed, how long each phase took, what the customer impact was in real units, and what changed structurally so it cannot recur in the same form. Interviewers listen specifically for whether you separate mitigation from diagnosis, because the correct instinct under pressure is to restore service first and understand later. They also listen for whether you blame people. A postmortem narrative with a named culprit reads as someone who will make the team hide failures.
Two further stages appear often enough to prepare for. A scripting round, usually Python, Go or Bash, which is practical rather than algorithmic: parse this log, call this API and reconcile the output, write a tool that finds unattached volumes, write a small reconciliation loop. And a hiring manager conversation about collaboration, since a large share of this job is persuading application teams to adopt something they did not ask for. Have one story about a migration you drove to completion across teams you did not manage, including how you handled the team that refused.
- Ask the recruiter for the exact round list, what the practical stage is, how long it takes, whether it is live or take-home, and whether an AI assistant is permitted in each round. Policies vary by company and sometimes by round; getting it wrong in either direction costs the offer.
- In the live troubleshooting round, narrate constantly. Silence reads as stuck. Say what you are checking, what you expect, and what the result eliminated. Interviewers score the diagnostic path, not the time to the answer.
- Gather evidence before forming a theory. The failure mode they are watching for is the candidate who decides in the first thirty seconds that it is DNS and spends the rest of the round proving themselves right.
- In the design round, ask for the numbers before you draw anything, then do the arithmetic aloud. Five hundred requests per second at 4KB each is roughly 2MB per second, which is about 5TB a month before replication. That sentence alone separates you from most candidates.
- Prepare one outage story to full depth, with timestamps, impact in real units, and the structural fix. Prepare a second one where you caused the incident; being asked for that is common, and a credible answer is worth more than a flawless record.
- Bring the sanitised artefacts: a module structure, a pipeline diagram, a runbook, an SLO definition, a postmortem with names removed. Showing a real document beats describing one, and almost nobody does it.
The resume: scale, outcomes and the artefacts that survive screening
DevOps resumes fail in a predictable way. They open with a wall of forty tool names, then list responsibilities ("responsible for maintaining CI/CD pipelines and Kubernetes clusters"), and close with certifications. Nothing in that document distinguishes a person who ran a four-hundred-node fleet across three regions from a person who followed a tutorial. Screeners know this, which is why they look past the tool list for scale and outcomes.
Attach scale to every tool you name, because scale is what makes a claim informative. "Kubernetes" tells a reader nothing. "Four EKS clusters, about 2,000 pods, 90 services, upgraded the fleet across three minor versions with no customer-facing downtime" describes a different candidate. The same applies elsewhere: how many Terraform modules and how many state files, how many pipelines and how many builds a day, what the monthly cloud bill was, how many engineers depended on what you built. If your numbers are small, give them anyway with the context. Being the only infrastructure engineer for a team of fifteen is a real and respected shape of experience.
Then attach outcomes, and prefer the four delivery measures engineering leaders already use, from the DORA research programme now published by Google Cloud: deployment frequency, lead time for changes, change failure rate, and failed deployment recovery time. These are the terms your interviewer thinks in, so using them makes your resume legible. "Cut lead time from merge to production from 6 days to 4 hours by replacing a manual release checklist with GitOps deployment and a merge queue" is a complete bullet: before, after, mechanism. If you never measured, state the shape honestly rather than inventing a figure. "From fortnightly releases to multiple deploys a day" is credible and survives a follow-up question; a fabricated percentage does not.
Cost is the second outcome family and it is underused. Cloud spend is a board-level line item, and an engineer who can show they removed money from the bill without degrading service is immediately interesting. The shapes that work, with your own real figures filled in: moved development clusters to scale-to-zero outside working hours and cut the non-production account from one monthly figure to another; found that a large share of the data transfer bill was cross-availability-zone traffic between two chatty services and removed it by co-scheduling them; replaced always-on capacity with spot and an interruption handler with no change to the SLO. Never reach for a number you cannot defend. Name the mechanism and give the before and after you actually had.
Reliability is the third. State the SLO you defined, the measurement window, the attainment, and what you did when the error budget was spent. "Defined the first SLO for the checkout path at 99.9 percent availability over 28 days, instrumented it in Prometheus, and agreed an error budget policy that paused feature deploys when the budget was exhausted; the policy triggered twice and both times the team shipped reliability work instead" demonstrates something no certification can.
Finally, the parts of a DevOps resume that get ignored. A skills section with sixty comma-separated items is skimmed and discounted, so group it by category and cut anything you would not want to be interviewed on. "Agile", "Scrum" and "stakeholder management" as skills add nothing. Responsibility statements with no result add nothing. Certifications belong in a short line near the end, not at the top. A career objective paragraph is pure cost. What you want instead is a two-line summary stating the shape of your experience in numbers, followed by roles where every bullet carries either a scale figure or a before-and-after.
- Summary format that works (substitute your own figures): "Infrastructure engineer, 6 years. AWS and Kubernetes across 4 clusters and 90 services supporting 120 engineers. Took deploys from fortnightly to 30 a day, cut cloud spend over two quarters, on call for production throughout."
- Name the IaC maturity, not just the tool: remote state with locking, versioned reusable modules, drift detection, policy checks in the pipeline, plan posted to the pull request, no manual console changes. Those phrases tell a reviewer you have operated at scale.
- Name the deployment model: rolling, blue/green, canary with automated analysis, GitOps with Argo CD or Flux, feature flags to decouple release from deploy. State which you ran and why you chose it.
- Include a supply-chain line if you have one: SBOM generation, image signing with cosign, build provenance, base images pinned by digest, third-party CI actions pinned to a commit SHA, admission control rejecting unsigned images. This is one of the fastest-growing requirements in postings.
- Include a security and identity line: workload identity federation instead of static cloud keys, short-lived credentials from OIDC in CI, a real secret manager rather than CI variables, least-privilege IAM you actually tightened.
- Keep a one-page public portfolio README rather than a long document. For infrastructure work, a diagram plus a cost line plus a runbook communicates more than any prose.
- Never let your resume and LinkedIn disagree on figures. Mismatched numbers are an easy and common rejection.
The take-home and the live exercise: what reviewers are actually marking
Most candidates treat the practical stage as a demonstration that they can make something work. Reviewers are marking something else: whether what you produced is safe to run in a company where you would be trusted with production credentials. The difference between a pass and a reject is usually in the parts that are not the happy path.
On a Terraform or OpenTofu exercise, reviewers look first at state and structure. Remote state with locking, not local state committed to the repository. Provider and module versions pinned, not floating. Reusable modules with inputs, outputs and variable validation, not one two-thousand-line root module. No hardcoded account identifiers, regions or secrets. Sensible resource naming and tagging, because tagging is how cost allocation and ownership work later. If you use a secret, it comes from a secret manager or an environment variable, and you note in the README that state contains secret values in plaintext and therefore the backend must be encrypted and tightly access-controlled. That single sentence is one of the strongest signals you can send in a take-home, because it shows you understand the tool's actual risk rather than its syntax.
On the container part, the marks are in the Dockerfile: a multi-stage build so build tooling does not ship to production, a small base image pinned by digest where possible, a non-root user, no secrets baked into layers (they persist in the image even if a later layer deletes them), a .dockerignore, layer ordering that lets the dependency install cache, and an explicit platform target if the team builds on Apple Silicon and runs on amd64. Add a scan step, and say what you do when it reports a high-severity vulnerability in a transitive dependency with no fix available, because that is the real situation and "block the build" is not always the right answer.
On the pipeline, reviewers check permissions and provenance before they check cleverness. Minimal token permissions. Authentication to the cloud via OIDC federation rather than a long-lived access key in a repository secret. Third-party actions pinned to a full commit SHA rather than a moving tag, which is the lesson from the compromises of widely used CI actions where injected code printed secrets into build logs. Separate build and deploy with an approval or an environment gate. Caching that actually hits. A path from a git commit to a deployed artefact that you can trace in one direction.
The README decides more outcomes than the code. Write what you built, how to run it, how to destroy it, what you deliberately left out and why, what you would do differently with a week instead of four hours, and what it costs to run. A reviewer reading "I did not implement multi-region because the stated recovery objective did not require it; here is what I would add if it did" learns more about your judgement than any amount of working infrastructure. Include a teardown command. Leaving the reviewer with resources that cost money is a bad first impression.
For the live exercise, technique matters more than knowledge. Form a hypothesis, state it, test it, say what the result eliminated, and move on. Start from the outside and work in: is the deployment rolled out, are pods running, are they ready, does the Service have endpoints, does DNS resolve, does the network policy allow it, does the application log an error. Check recent change first, because the overwhelming majority of incidents follow a change: a deploy, a config edit, a certificate expiry, a quota limit, a scaling event. If you get stuck, say what you would check next if you had access to something you do not have. Interviewers routinely pass candidates who did not finish but reasoned well, and reject candidates who stumbled into the answer by changing things at random.
- Take-home scoring, roughly in reviewer order: safety and secrets handling, structure and reuse, the README, correctness, then polish. Candidates optimise the last item and lose on the first.
- Always include: a teardown path, a cost note, an explicit list of what you left out, and the assumptions you made. Never include: a committed state file, a real credential, a public object storage bucket, or an IAM policy with Action "*" and Resource "*".
- Time-box honestly and say what you spent. If the instructions say four hours, a twenty-hour submission is a negative signal about judgement, not a positive one about effort.
- Kubernetes failures to be able to diagnose cold: CrashLoopBackOff, ImagePullBackOff, OOMKilled with exit code 137, Pending from insufficient resources or an unsatisfiable node selector or taint, a Service with no endpoints because the label selector does not match, a liveness probe that kills a healthy but slow-starting pod before it initialises (which is what a startup probe is for), a readiness probe so strict that pods never join the endpoints list, CPU throttling caused by a low limit, a PodDisruptionBudget blocking a node drain, a PersistentVolumeClaim stuck Pending for want of a StorageClass, and DNS latency caused by search-domain expansion.
- Know what happens to running workloads during a managed-Kubernetes upgrade, and be able to describe your own procedure: read the deprecated API report, test in a lower environment, surge new nodes, drain with disruption budgets respected, verify, and keep separate rollback stories for the control plane and the node pool.
- If a take-home needs more than about six hours of real work, it is reasonable to ask whether a scoped-down or paid version is possible. A good employer will say yes; the answer tells you something either way.
The questions that decide the technical rounds, and how to answer them
A handful of questions come up repeatedly across employers, and each has a shallow answer that fails and a specific answer that passes. Rehearse these out loud, because the gap between knowing something and explaining it under time pressure is where interviews are lost.
"Walk me through what happens when a user request fails." The shallow answer lists layers. The strong answer picks a real system you ran and traces it: client, DNS, CDN or edge, load balancer and TLS termination, ingress or gateway, Service, pod, application, database and downstream dependencies, naming at each hop what the failure would look like from the outside and which signal distinguishes it. End with the observability you would use to tell the hops apart, which is where distributed tracing earns its keep.
"How do you deploy with no downtime?" Shallow: "rolling updates". Strong: separate the release of code from the release of behaviour with feature flags; use a rollout strategy appropriate to blast radius (canary with automated analysis for risky changes, blue/green when you need instant rollback and can afford double capacity); make every deployment reversible, which means database migrations are expand-and-contract so old and new code can run simultaneously; drain connections properly on shutdown and handle SIGTERM; and have a rollback that is one command and is practised. Say that the hard case is the database, because it is.
"What is your alerting philosophy?" Shallow: thresholds on CPU and memory. Strong: alert on symptoms users feel, derived from SLOs, with an error budget driving whether you page at all; use burn-rate alerts so a fast burn pages immediately and a slow burn opens a ticket; everything else goes to a dashboard, not a pager. Be able to say what you deleted. The strongest version of this answer names the before and after count of paging alerts and what happened to out-of-hours page volume afterwards. Alert fatigue is a universal and genuine problem, and candidates who have measurably fixed it are rare.
"How do you manage secrets?" Shallow: "we use Vault" or "they are in the CI secrets". Strong: no long-lived credentials anywhere, workload identity federation from CI and from pods to the cloud provider, dynamic short-lived database credentials where the platform supports it, secrets synced into the cluster by an operator rather than committed, rotation that actually runs and is tested, encryption at rest of anything that touches state including IaC state, and an answer for the day a credential leaks into a build log (revoke first, rotate, then establish blast radius from the audit log).
"How would you reduce our cloud bill?" Shallow: buy reserved instances. Strong: get visibility first with allocation tags and a per-team or per-service breakdown, because you cannot cut what you cannot attribute. Then work in order of size: idle and oversized compute, non-production running outside working hours, storage with no lifecycle policy, old snapshots and unattached volumes, inter-availability-zone and egress data transfer, over-provisioned managed databases, log and metric volume (cardinality is a silent and large cost), and only then commitment-based discounts, which lock in whatever waste you did not remove first. Say explicitly that you would not trade away an SLO without the owning team agreeing.
"Tell me about a time you broke production." This is not a trap, and a clean record is not the answer they want. Give a real one: what you changed, why it looked safe, what broke, how long it took to notice (usually the interesting number), how you restored service, and what you changed about the system so that the same class of mistake becomes impossible rather than merely discouraged. Guardrails beat good intentions, and saying so is the point of the question.
- Be able to read a Terraform plan out loud and say which lines are dangerous. Create and update-in-place are usually fine; destroy and replace are the ones that end careers, and "forces replacement" on a stateful resource is the line to stop on.
- Know the difference between mitigation and fix, and use the words correctly. Restoring service by rolling back is a mitigation. The fix comes later and is often not technical.
- Have a real opinion on at least one tradeoff, with the reasoning: managed Kubernetes versus serverless containers for a small team, a monorepo pipeline versus per-service pipelines, push deployment versus GitOps pull, one big cluster with namespaces versus many clusters, Terraform versus a cloud-native tool. Opinions with reasons read as experience; neutrality reads as absence.
- When you do not know something, say so and say how you would find out. Infrastructure is too large for anyone to know all of it, and interviewers at this level know that. Bluffing a half-remembered Kubernetes internal is far worse than "I have not run that; here is the closest thing I have run and what I would check first."
- Prepare a short answer on compliance even if you have never worked under it: what change evidence looks like, why separation of duties means the author cannot approve their own production change, and how you keep daily deployment inside that constraint with automated approvals and audit trails.
Getting in with no DevOps title yet
Be clear about the structure of this market, because misunderstanding it wastes months. Genuinely junior DevOps roles are scarce, and the reason is not gatekeeping. The job carries production credentials and a pager, so employers are reluctant to make it anyone's first job. Most DevOps engineers arrived from somewhere else, and the fastest routes are the ones where you are already near production.
The highest-probability route is usually internal. If you work anywhere that has infrastructure, the automation nobody wants to own is available to you: the flaky deployment script, the manual release checklist, the monitoring nobody trusts, the staging environment that is always broken, the cloud bill nobody has looked at. Take one, fix it properly in the open, measure the before and after, and write it up. Internal moves into platform teams happen on the strength of exactly this, and often without a full external interview loop. If your current employer has no infrastructure worth automating, that is itself a reason to move to one that does, even laterally.
The external routes that work, in rough order of how often they do: systems administrator or cloud operations into DevOps (closest skill overlap, usually needs the IaC and container gap filled); technical support or help desk at a software company into DevOps (you already debug production, you need Linux depth and automation); QA automation into build and release into DevOps (you already write code and own pipelines); backend developer into platform engineering (the easiest move of all, because the coding is already there and platform teams prize engineers who have used the platform in anger); network engineer into cloud networking and then wider infrastructure; and in the US, the military or defence-contracting route, where a clearance plus a baseline certification opens doors that are closed to better engineers without one.
If you have none of those and you are building evidence from scratch, build one system, not five tutorials. A credible portfolio project for this role provisions a small but real environment with code: a network, a managed Kubernetes cluster or container service, a managed database, and a modest application. It deploys through a pipeline that builds, tests, scans, signs and promotes an image, and it releases through GitOps rather than a manual apply. It has one SLO you defined, one alert rule that fires on burn rate, a dashboard, and a runbook for that alert. It has a README explaining the decisions and stating the monthly cost. It has a teardown script so it costs nothing when idle. Then do the thing nobody does: break it on purpose, write the postmortem, and commit it. Deliberately killing a node, exhausting a disk, expiring a certificate or revoking a permission teaches you the diagnostic reflexes the live round tests, and the postmortem is the single artefact that most distinguishes a career changer from a tutorial follower.
Two warnings about the switching path. First, the "DevOps bootcamp" market is full of courses selling a tool tour. A course that does not have you on call for anything, debugging anything broken, or accountable for a cost will not get you hired, and the certifications it bundles are the cheap tier. Second, do not skip Linux and networking. Candidates who struggle in the technical screen struggle there, not in Kubernetes. Reading a process table, following a packet, reasoning about file descriptors, understanding what a routing table and MTU do and why TLS fails the way it does is the foundation the rest of the stack sits on, and it is what separates people who can diagnose from people who can configure.
- Ask for the automation work at your current job, and attach a measured outcome. "Reduced the release checklist from 23 manual steps to a pipeline, cutting release time from 3 hours to 12 minutes" is a hiring bullet whatever your current title says.
- Target the adjacent title when your current one does not match. Cloud Engineer, Build and Release Engineer, Infrastructure Engineer and Platform Operations roles hire the same skills and are often less contested than postings using the DevOps label.
- Consultancies, managed service providers and the public sector hire from non-traditional backgrounds more readily than product companies, and they will frequently pay for your cloud certification because their partner status depends on headcount holding them.
- For the portfolio, choose one cloud and go deep. Breadth across three clouds at a shallow level is a weaker signal than one account where you can explain the network design, the identity model and the bill.
- Keep the lab cheap. Scale to zero, use spot where it is safe, set a hard budget alert, and tear down when not in use. Running up a four-figure bill in a personal account is a lesson, but there are cheaper ways to learn it.
- Contribute to the documentation of a tool you use. It is the lowest-friction open source contribution, maintainers welcome it, and it is a public, checkable artefact with your name on it.
Offers, on-call and the questions that tell you what the job really is
DevOps offers have more variables than a straight engineering offer, and the ones that determine whether the job is survivable are rarely in the written offer at all. Ask them before you accept, because the answers are easy to get at offer stage and impossible afterwards.
On-call is the first. Rota size, frequency, whether it is compensated as a stipend, as overtime or not at all, typical and worst-case out-of-hours pages per rotation, whether you are expected to work a normal day after a bad night, whether application teams carry their own pagers, and who decides what gets to page. A team with a one-in-four rota, very few out-of-hours pages per rotation and a culture of deleting noisy alerts is a healthy team. A team where the infrastructure engineer carries the pager for every service in the company has pushed an ownership problem onto one person, and no amount of tooling solves that.
The second is ownership and authority. Can you merge a change to production infrastructure, or does every change need a ticket and an approval from someone on another continent? Who decides architecture? Is there a budget you control, or do you have to ask for every service you want to use? A platform role with accountability for reliability but no authority over what application teams deploy is a well-known trap, and it is worth asking about directly: what happens when a team wants to ship something you believe is unsafe.
The third is the state of the estate. Ask what the oldest unupgradable thing in production is, what the last serious incident was and what changed afterwards, how many manual steps exist in the current release process, whether there is a staging environment anyone trusts, and whether anybody can describe what the cloud bill is spent on. Ask what share of the team's week goes to unplanned work, which is the clearest single indicator of whether you will build anything or spend two years firefighting.
Fourth, compensation. Cite the sources rather than a remembered band: the BLS OES tables for the relevant occupation codes in your metro area, postings in pay-transparency jurisdictions, ITJobsWatch for UK skill-level medians, and self-reported total compensation data for large technology employers, each with its own bias. In many jurisdictions you are not required to disclose salary history and employers may be prohibited from asking; answer with your target and the posted range instead. If there is an on-call stipend, get the amount in writing, and if there is none, treat on-call frequency as part of the compensation calculation rather than a detail.
Finally, a practical note on leaving. Before your last day at any job, extract your numbers and sanitise your artefacts. Deployment frequency and lead time from the pipeline, change failure rate and recovery time from the incident record, SLO attainment from the dashboard, cluster and service counts, spend from the billing console, pipeline durations before and after your work. Write a redacted version of one runbook, one postmortem and one architecture diagram for yourself. None of this is confidential once names, identifiers and customer details are removed, and all of it is the difference between a resume that lists tools and one that proves outcomes.
- Ask: how many people are on the rota, how many out-of-hours pages in a typical week, and what the worst week in the last six months was. The hesitation before the answer is itself informative.
- Ask: what share of the team's time went to unplanned work last quarter. Teams that have measured this and know the number are teams that manage it.
- Ask: what happened to the person who held this role before me. The question works in every discipline and it is more informative here than in most.
- Get in writing: on-call compensation, the rota commitment, whether a cloud certification is paid for and whether exam time is work time, and whether there is a training budget with an actual figure.
- If the role is advertised as platform engineering, ask how they measure adoption of the platform. A platform team with no adoption measure is an infrastructure team with a new name, and the work will be ticket-driven.
- Before your last day: numbers, sanitised artefacts, and a former manager who has agreed to be a reference and knows which figures you will both be quoting.
What a DevOps engineer has to know about AI in 2026-27
Start with the honest calibration, because the hype and the observable change point in different directions. The core of this job has not been automated and is not close to it. Capacity planning, incident command, the decision to roll back, the design of a network and an identity model, the judgement about what is safe to change on a Friday, and the accountability at three in the morning are all unchanged. Fully autonomous remediation in production, which vendors have been promising under the AIOps label for years, is still rare in practice and is not what employers are buying. What has changed is substantial but specific, and it falls into three buckets: a genuinely new body of infrastructure work created by AI workloads, a shift in where the difficulty of the existing work sits, and a new governance problem that lands squarely on this role.
The new work is the biggest story and it is where demand grew. Serving models is an infrastructure problem, and a harder one than serving web applications: accelerator capacity is scarce and expensive, so scheduling it well is a direct financial outcome; model artefacts are large, so image sizes, registry caching and pull storms become real failure modes; inference is bursty and latency-sensitive in ways that break naive autoscaling; and an idle GPU costs money every minute. If your employer has any AI product at all, somebody has to run GPU node pools, handle quota and availability across regions, decide between dedicated, reserved and spot capacity, keep a gateway in front of vendor APIs with per-tenant rate limits and cost attribution, and answer the question of what happens when the model provider has an outage. That person is a DevOps or platform engineer, and the postings now say so.
The shift in existing work is subtler, and interviewers probe it directly. Producing configuration is now nearly free: an assistant writes plausible Terraform, Kubernetes manifests, pipeline YAML and Bash faster than you can. What it does not have is your account's state, your traffic shape, your compliance constraints or your invariants. The characteristic failure now is not generated code that does not work; it is generated code that works and is quietly destructive. A plan that replaces a database instead of updating it. A security group that resolves a connectivity problem by opening a port to the world. An IAM policy with a wildcard action because the narrow one failed. A Helm values change that removes a resource limit to stop an OOM kill. Because writing is cheap, reviewing is where the value moved, and the pull-request review round in the interview loop exists precisely to test whether you catch these.
The governance problem is the newest and the one most often raised in hiring manager rounds. Teams are wiring assistants and agents into real infrastructure: tools that read dashboards, query logs, open pull requests, and in some places execute commands against clusters. "How would you let an agent operate in our environment safely" is now a live interview question, and the strong answer is not philosophical. It is the same answer as for a human operator with less context: its own identity rather than a shared one, short-lived credentials, least privilege scoped to named resources, read-only by default with write paths that go through the normal change process, dry-run and plan output before any apply, blast-radius limits, a full audit trail of what it did and under whose authority, and a hard boundary where prompt-injectable input cannot reach a tool that holds production credentials. Say plainly that the last one is a real attack surface: content an agent reads is untrusted input, and treating it as instructions is the vulnerability.
One more honest note, because it is useful in interviews. The DORA research programme, now published by Google Cloud, has studied AI adoption alongside its delivery measures and has found that individual developers report feeling more productive while team-level throughput and stability do not automatically follow, and can get worse where delivery practices are weak. The implication for this role is flattering and worth saying out loud: when code volume goes up, the constraint moves downstream of writing it, to the pipeline, the review, the testing, the deployment path and the ability to recover. That is your job. Read the current year's report before an interview and cite its findings accurately rather than quoting a number from memory.
Running GPU and inference infrastructure
This is the one clearly new area of DevOps demand, and most candidates have no exposure to it, so even modest real experience is a differentiator. The problems are concrete: scheduling accelerators on Kubernetes, node pools that cost an order of magnitude more than CPU nodes, scarce regional capacity and quota, multi-gigabyte images and model weights, cold starts measured in minutes, and autoscaling behaviour that does not match request patterns. Cost control here is not housekeeping; an idle fleet can be the largest line on the bill.
Show it: Be able to discuss how accelerators are exposed to pods and scheduled, the tradeoff between sharing a device across workloads and dedicating it, why you would keep model weights in object storage with a warm cache rather than baking them into the image, how you scale an inference service that cannot start in under a minute, and how you attribute spend per model and per tenant. If you have none of this at work, run a small quantised model behind a serving stack in your lab, record the cold start time and the cost per hour, and write up what you would change at a hundred times the traffic.
Reviewing generated infrastructure changes for the defects they reliably produce
Writing configuration stopped being the bottleneck; approving it safely became the bottleneck. Infrastructure changes have a far worse blast radius than application code, because one approved plan can delete a database, open a network, or widen a permission across an entire account, and no test suite catches it. Hiring managers are explicitly screening for the engineer who stops these before they merge.
Show it: Read a Terraform plan out loud in the interview and name the dangerous lines: destroy and replace, "forces replacement" on anything stateful, a count or for_each change that re-indexes resources, a security group rule to 0.0.0.0/0, an IAM policy with wildcard actions, a removed lifecycle protection. Have one real story of a generated or copied change you blocked, with the specific failure it would have caused. On the resume, add a bullet about the guardrail you put in so the class of mistake became impossible: policy as code in the pipeline, a required plan review, a deny rule on destructive actions in production.
Governing agent and assistant access to production
Most organisations are currently deciding what an automated agent is allowed to do in their infrastructure, and many have no policy at all. The person who can propose a workable one, rather than either banning everything or shrugging, is valuable immediately. It is also a question with a defensible right answer that draws on skills you already have: identity, least privilege, change control and auditability.
Show it: Have a position you can defend in two minutes: a distinct machine identity per agent, credentials that are short-lived and federated rather than static, read-only by default, write access only through the same pull request and approval path a human uses, dry-run output reviewed before apply, environment separation so production requires an explicit human step, rate and blast-radius limits, full audit logging, and an explicit boundary preventing untrusted content an agent reads from reaching a credentialed tool. Add that you would start in non-production and measure before widening scope, because that is what a cautious employer wants to hear.
Keeping CI capacity and cost under control as change volume rises
When assistants help engineers open more pull requests, continuous integration absorbs the impact. Queue times, runner costs, flaky tests and merge conflicts all scale with change volume, and a pipeline that was comfortable at a few hundred builds a day becomes the company's main complaint when that volume multiplies. This is one of the most common real problems a new platform hire is handed.
Show it: Numbers. Pipeline duration before and after, queue wait at peak, cost per build, cache hit rate, flaky test rate and what you did about it. Techniques to be able to discuss: autoscaled ephemeral runners, aggressive and correct caching, test selection and parallelism, merge queues so main stays green without serialising everything, splitting long pipelines into fast required checks and slower asynchronous ones, and killing redundant builds with concurrency groups. Add the security side, since scaling runners usually means self-hosted ones: never run untrusted pull requests on a self-hosted runner with network access to anything that matters.
Software supply chain integrity, including what models and dependencies you ship
Supply chain requirements moved from best practice to procurement questionnaire. Enterprise and public sector buyers now ask vendors for a software bill of materials and evidence of build provenance, and regulation in several jurisdictions pushes in the same direction for products with digital elements; the phased application dates are exactly the kind of detail that shifts, so check the current position rather than quoting one in an interview. Real incidents keep proving the point, from the xz-utils backdoor to compromises of widely used CI actions that printed secrets into build logs.
Show it: Describe a pipeline that generates an SBOM, signs images, produces build provenance, pins third-party CI actions to a commit SHA and base images to a digest, and enforces signature verification at admission so an unsigned image cannot run. Say what you do about a high-severity finding with no available fix, since blocking is not always the correct answer. If your employer ships anything AI-related, extend the same thinking to model and dataset provenance: where the weights came from, how they are verified, and what you record about which model version served which request.
Using assistants in your own work without losing the ability to operate unaided
Interview policies differ by company and sometimes by round. Some teams now require you to use an assistant in a technical round and watch how you prompt, verify and correct it; others disable it to see whether you can reason from first principles under pressure. More importantly, incidents happen when the tooling is degraded, and the engineer who can only operate with an assistant is a liability at exactly the moment they are needed.
Show it: Ask in writing which rounds permit assistance, and prepare for both. When you use one in front of an interviewer, narrate the verification: what you checked, what you rejected and why. In a take-home, say plainly what you generated and what you changed; reviewers can usually tell, honesty costs nothing and being caught costs the offer. Keep your unaided fundamentals sharp enough to debug a cluster, read a packet capture and write a correct shell pipeline with no help, because that is the round that separates candidates.
What a screen is looking for
These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.
- DevOps Engineer
- Platform Engineer
- Site Reliability Engineer
- Cloud Engineer
- Infrastructure Engineer
- Build and Release Engineer
- Production Engineer
- CI/CD
- Continuous integration
- Continuous delivery
- Continuous deployment
- Deployment pipeline
- GitHub Actions
- GitLab CI
- Jenkins
- CircleCI
- Buildkite
- Azure DevOps
- Argo CD
- Flux
- GitOps
- Merge queue
- Infrastructure as code
- Terraform
- OpenTofu
- HCL
- Terraform modules
- Terraform state
- State locking
- Pulumi
- CloudFormation
- AWS CDK
- Ansible
- Packer
- Policy as code
- Open Policy Agent
- Rego
- Kyverno
- Kubernetes
- EKS
- AKS
- GKE
- OpenShift
- kubectl
- Helm
- Kustomize
- Operators
- Custom resource definitions
- Admission control
- Pod Security Admission
- Network policy
- Ingress controller
- Gateway API
- Service mesh
- Istio
- Cilium
- eBPF
- Karpenter
- Cluster autoscaler
- Horizontal pod autoscaler
- Node draining
- PodDisruptionBudget
- Startup probe
- Liveness probe
- Readiness probe
- Cluster upgrade
- Docker
- containerd
- Dockerfile
- Multi-stage build
- Distroless image
- Image registry
- Amazon ECR
- Trivy
- Grype
- Cosign
- Sigstore
- SBOM
- SLSA
- Build provenance
- AWS
- Amazon Web Services
- Azure
- Google Cloud Platform
- VPC
- Subnetting
- Route 53
- DNS
- TLS
- Certificate rotation
- Load balancing
- CDN
- IAM
- Least privilege
- OIDC federation
- Workload identity
- IRSA
- EKS Pod Identity
- HashiCorp Vault
- External Secrets Operator
- Secrets management
- Linux
- Bash
- systemd
- Python
- Go
- Scripting
- Observability
- Prometheus
- Grafana
- OpenTelemetry
- Distributed tracing
- Structured logging
- Loki
- Datadog
- New Relic
- PagerDuty
- Alerting
- Burn rate alerts
- SLO
- SLI
- Error budget
- Availability
- p99 latency
- Incident response
- Incident commander
- Blameless postmortem
- On-call
- Runbook
- Chaos engineering
- Disaster recovery
- RTO
- RPO
- Backup and restore
- Blue/green deployment
- Canary release
- Rolling update
- Feature flags
- Zero-downtime migration
- Expand and contract migration
- Capacity planning
- Autoscaling
- Spot instances
- FinOps
- Cloud cost optimization
- OpenCost
- Kubecost
- Cost allocation tagging
- DORA metrics
- Deployment frequency
- Lead time for changes
- Change failure rate
- Time to restore service
- Toil reduction
- Internal developer platform
- Backstage
- Developer experience
- Platform engineering
- GPU scheduling
- Inference serving
- Model serving
- vLLM
- Ray
- KServe
- MLOps
- LLM gateway
- Rate limiting
- CKA
- Certified Kubernetes Administrator
- CKAD
- CKS
- AWS Certified Solutions Architect
- AWS Certified DevOps Engineer Professional
- AZ-104
- AZ-400
- HashiCorp Terraform Associate
- CompTIA Security+
- RHCSA
- FedRAMP
- STIG
- SOC 2
- PCI DSS
- Change management
- Separation of duties
Mistakes that cost people this job
Opening the resume with a forty-item tool list and no scale or outcome anywhere.
Attach a number to every tool. "Kubernetes" becomes "4 EKS clusters, 90 services, 2,000 pods, upgraded across three minor versions with no customer-facing downtime". Group the skills section by category, cut anything you would not want to be interviewed on, and make sure every role bullet carries either a scale figure or a before-and-after.
Claiming Kubernetes experience that turns out to be running kubectl against a cluster someone else built and upgraded.
Be precise about your layer and own it. "I deployed to a managed cluster that the platform team ran; I owned the manifests, the Helm charts, the resource limits and the rollout strategy, and I have not run a control plane" is credible and respected. Interviewers probe this with upgrade, networking and etcd questions within two minutes, and an inflated claim collapses immediately.
Guessing in the live troubleshooting round instead of gathering evidence, then changing things at random when the first guess is wrong.
Narrate a diagnostic path. State a hypothesis, test it, say what the result eliminated, and work outside in (deployment, pods, readiness, endpoints, DNS, network policy, application). Check what changed recently first, because most incidents follow a change. Candidates who do not finish but reason cleanly are routinely passed; candidates who stumble onto the answer by mutation are not.
Over-engineering the design round: proposing a service mesh, multi-region active-active and an event bus for a system nobody sized.
Ask for the numbers first (request rate, payload size, growth, recovery objectives, team size, compliance constraints) and do the arithmetic out loud. Then choose the smallest architecture that meets the requirement, and say explicitly what you are not building yet and what signal would change that. Interviewers score chosen tradeoffs, not component inventories.
Treating the take-home as a demo: local state committed to the repo, a hardcoded secret, a single enormous main.tf, no README and no teardown.
Spend the last quarter of your time on safety and the README. Remote state with locking, pinned provider and module versions, no secrets in code, a non-root container built in multiple stages, minimal pipeline permissions with OIDC rather than a static key, a teardown command, a cost note, and an explicit list of what you left out and why. Reviewers mark safety and judgement before they mark cleverness.
Having no numbers because nobody measured, and then rounding up to a figure that sounds right.
Use the shape when you lack the figure: "from fortnightly releases to several deploys a day" is credible and survives a follow-up question; an invented percentage does not. Then fix the root cause. Before you leave your current job, pull deployment frequency, lead time, change failure rate, time to restore, SLO attainment, spend and pipeline durations out of the systems while you still have access.
Ignoring cost entirely, in the resume and in every design answer.
Put a cost sentence in every architecture answer and at least one cost outcome on the resume. Cloud spend is a board-level number, AI workloads made it larger, and an engineer who can attribute spend and remove waste without degrading the SLO is immediately interesting. Start with allocation and visibility, not with reserved capacity, because commitments lock in the waste you did not remove first.
Blaming a person in the incident story, or presenting a flawless record with no outage you ever caused.
Tell a real one: the symptom, how you learned about it, your first three actions, how long until you noticed, the mitigation, the actual cause, and the structural change that makes that class of mistake impossible rather than merely discouraged. A named culprit in your narrative reads as someone who will make the team hide failures, which is disqualifying for a role that depends on honest postmortems.
Searching only for the exact title "DevOps Engineer" and missing much of the market.
Search Platform Engineer, Site Reliability Engineer, Cloud Engineer, Infrastructure Engineer, Build and Release Engineer, Production Engineer and Systems Engineer as well, and read the responsibilities rather than the title. Several of those labels attract fewer applicants for identical work.
Skipping Linux and networking fundamentals because the stack is managed now.
Drill them deliberately, because this is where candidates actually fail the screen. Reading a process table, following a packet, DNS resolution and its failure modes, TLS handshakes and certificate chains, file descriptors, disk and memory pressure, what uninterruptible sleep means, how a connection pool exhausts. Everything above that layer is configuration; this is the part that lets you diagnose rather than guess.
Accepting the offer without asking what on-call actually looks like.
Before signing, get rota size, frequency, compensation, typical and worst-case out-of-hours pages, whether application teams carry their own pagers, and what share of the team's week is unplanned work. These determine whether you build anything in the next two years or spend them firefighting, and they are impossible to renegotiate after you start.
Questions people ask
Do I need a degree or a certification to become a DevOps engineer?
No. There is no licence, no legally required degree and no mandatory certification for DevOps work, and many working DevOps engineers came from help desk, systems administration, networking, QA, the military or a trade. Degrees still matter in graduate programmes at large employers, in some government and defence contracts, and in skilled-worker visa routes where a degree or assessed equivalent counts toward eligibility. After two or three years of production experience, what replaces any credential is a system you personally operated and can describe in detail: its failure modes, its cost, its SLO and the incidents it caused.
Which DevOps certification is actually worth getting?
The Certified Kubernetes Administrator (CKA) carries the most weight, because it is a performance-based exam sat in a live terminal against running clusters rather than a multiple-choice test, so it cannot be passed by memorising a question bank. After that, a cloud certification matching the employer's cloud (AWS Solutions Architect Associate, AWS Certified DevOps Engineer Professional, Microsoft AZ-104 or AZ-400, Google Professional Cloud DevOps Engineer) is mainly useful for passing recruiter and automated filters, and the HashiCorp Terraform Associate is quick and cheap. Generic Agile, Scrum Master and ITIL Foundation certificates do not move a DevOps hiring decision. If you have no DevOps job yet and three months to spend, build and document one real system rather than take a second certification.
Can I get a DevOps job with no experience, and how long does the switch take?
DevOps engineer is rarely anyone's first job, and the reason is structural rather than gatekeeping: the role carries production credentials and a pager, so employers prefer someone who has already been responsible for something running. The realistic paths are to enter an adjacent role first (technical support, systems administration, cloud operations, QA automation, network engineering or backend development) and then take the automation work nobody wants, or to move internally at an employer that already has infrastructure. From systems administration or cloud operations, the switch typically takes six to twelve months of deliberate work to close the infrastructure-as-code, container and pipeline gaps, often without changing employer; from help desk it usually takes longer, because Linux depth and automation both have to be built; from backend development into platform engineering it can be immediate. The limiting factor is almost never study time, it is getting access to something real to operate.
What does a DevOps engineer interview consist of in 2026?
A DevOps engineer loop is typically five stages over two to five weeks: a recruiter stack-match screen (cloud, orchestrator, IaC tool, on-call expectation, salary), a technical screen on Linux, networking and troubleshooting, a practical stage that is either a take-home repository or a live exercise in a deliberately broken environment, an infrastructure design round (design a pipeline for sixty services, get a fortnightly release to daily, design multi-region failover, diagnose a doubled cloud bill), and an incident round where you walk through a real outage you handled. Some employers add a scripting round in Python, Go or Bash, and a growing number add a pull-request review round where you say what you would block. Ask the recruiter for the exact round list and whether an AI assistant is permitted in each one.
What should a DevOps engineer put on a resume?
A DevOps engineer resume needs scale attached to every tool, and outcomes attached to every role. Scale means cluster and node counts, number of services, deploys per day, builds per day, monthly cloud spend, and how many engineers depended on what you built. Outcomes are best expressed in the four delivery measures hiring managers already use from the DORA research (deployment frequency, lead time for changes, change failure rate, failed deployment recovery time) plus cost reductions and SLO attainment, each as a before-and-after with the mechanism named. What gets ignored: a sixty-item skills list, responsibility statements with no result, "Agile" and "Scrum" as skills, and a career objective paragraph.
Do DevOps engineers need to code, and in which language?
Yes, a DevOps engineer writes code, though not at application-developer depth in most roles. Bash and Python are the practical baseline, Go matters if you will write Kubernetes operators, controllers or internal tooling, and fluency in HCL for Terraform or OpenTofu is assumed. Coding rounds for this role are practical rather than algorithmic: parse a log, call an API and reconcile the output, find unattached volumes across accounts, write a small reconciliation loop. Platform engineering roles sit further toward software engineering and often include a genuine coding interview, which is one reason backend developers move into platform teams easily.
What is the difference between a DevOps engineer, an SRE, a platform engineer and a cloud engineer?
The titles overlap heavily and the responsibilities matter more than the label, but the centres of gravity differ. DevOps engineer usually means pipelines, infrastructure as code and deployment, often embedded with product teams. Site reliability engineer centres on production reliability as an engineering discipline: SLOs, error budgets, incident response and eliminating toil, frequently with a stronger coding bar. Platform engineer builds an internal developer platform as a product, with engineers as customers and adoption as a success measure. Cloud engineer centres on the cloud estate itself: accounts, networking, identity, landing zones and cost. Search all four titles, because the same work is advertised under each.
Is AI replacing DevOps engineers?
No, and the observable change runs the other way for the DevOps engineer. Assistants write configuration faster than a human can, so the time spent producing Terraform, manifests and pipeline YAML has fallen, but nothing on the market plans capacity, commands an incident, decides to roll back, or is accountable when production is down, and fully autonomous remediation in production remains rare. Two things grew instead: a new body of infrastructure work around serving models, GPU capacity and inference cost, which is one of the few clearly expanding areas of demand; and the downstream constraint created when more code gets written, which lands on the pipeline, the review, the deployment path and the ability to recover. The DORA research programme has found that AI adoption improves how productive individuals feel without automatically improving team throughput or stability, which is an argument for more delivery engineering, not less.
How much does a DevOps engineer earn?
Cite the source rather than a band, because the US Bureau of Labor Statistics has no DevOps engineer occupation code and aggregator figures for the title vary widely. Employers usually code the work to SOC 15-1252 Software Developers, 15-1244 Network and Computer Systems Administrators, or 15-1299 Computer Occupations All Other, so read the OES tables for all three in your metro area. For live ranges, read postings in pay-transparency jurisdictions such as Colorado, California, Washington, New York and Illinois, and check which rules apply where you are looking, because the list keeps changing. In the UK, ITJobsWatch publishes permanent and contract medians by individual skill, which is useful for pricing a Kubernetes or Terraform specialism. Pay varies more by on-call load, cloud estate size and industry than by title, and an active US security clearance carries a substantial premium.
Do I need Kubernetes to get a DevOps job?
Not universally, but Kubernetes is the single most requested technology in DevOps engineer postings, and not having it narrows your options sharply. Plenty of good jobs run on serverless containers, managed platform services or virtual machines, and in those environments deep Kubernetes knowledge is not the differentiator. If you are building a target skill set from scratch, learn it anyway and learn it operationally rather than superficially: you should be able to explain what happens to running workloads during a cluster upgrade, why a pod is Pending, why a Service has no endpoints, and what a PodDisruptionBudget does to a node drain. Those are the questions the technical screen actually asks.
Put this on a resume in about a minute
Paste your history once and point it at the DevOps Engineer posting you are looking at. No account, no card.
Build my resume free More roles