| What the role is in 2026 | Engineering applied to the problem of keeping a service available, fast and affordable while other people keep changing it. The unit of work is a running production system and the humans who operate it: service level objectives that mean something, alerts that are worth waking up for, incident response that is practised rather than improvised, capacity that is planned rather than discovered, and the automation that removes the manual work the job would otherwise drown in. Writing code is part of it. Being the only person who understands the system is a failure, not a qualification. |
|---|---|
| Titles the same job is posted under | Production Engineer, Platform Engineer, Infrastructure Engineer, Systems Engineer, Reliability Engineer, Cloud Infrastructure Engineer, Observability Engineer, DevOps Engineer, Compute or Capacity Engineer, and at some employers simply Software Engineer with an infrastructure team name attached. Searching only the string SRE misses a large part of the market, and some postings titled SRE are really a renamed systems administrator job. Read the responsibilities, not the title. |
| Licence or credential gate | None. No licence, no registration, no mandatory certification and no required degree. Because nothing gates entry, the screen is the filter, which is why a resume made of tool names loses to one made of measured outcomes. |
| Certifications that actually move a screen | CKA (Certified Kubernetes Administrator) is the one most infrastructure leads respect, because it is performance based in a live cluster rather than multiple choice. CKS is the security variant and requires an active CKA; KCNA is the entry-level option. Google Professional Cloud DevOps Engineer is the certification written most explicitly around SRE practice, including SLOs and error budgets, so it aligns unusually well with what interviewers ask. AWS Certified SysOps Administrator Associate or Solutions Architect Associate, HashiCorp Terraform Associate, and the CNCF Prometheus Certified Associate round out the list. Check current fees, exam length, prerequisites and renewal terms on the issuing body's own page rather than any figure quoted second-hand, including this one. None of them substitutes for one production system you have operated, and a stack of certifications above thin experience reads as a substitute for it. |
| Realistic prep time from an adjacent job | From backend software engineering with production ownership: roughly three to six months of evenings to add Linux depth, networking, observability and the SLO vocabulary. From systems administration or a network operations centre: six to twelve months, because the missing pieces are writing real code and distributed systems reasoning, and both take time to make convincing. From support or technical account management: twelve to eighteen months, usually via a platform or tooling role first. The fastest route in every case is an internal transfer: take the pager, run an incident, write a postmortem that gets read, then apply as someone who already does the job. |
| Typical loop | Recruiter screen, hiring manager call, a practical coding round in Python, Go or Bash (sometimes a standard medium-difficulty algorithms round at large technology companies, so ask which), a systems design round of 45 to 60 minutes with capacity arithmetic in it, a live troubleshooting round on a deliberately broken system, and a behavioural round built around a real incident you handled. Four to six rounds, commonly three to six weeks end to end. Some employers add a short take-home or an infrastructure-as-code review. |
| Where to find real pay numbers | There is no dedicated BLS occupation code for site reliability engineer. Employers file these roles under BLS OES 15-1252 (Software Developers), 15-1244 (Network and Computer Systems Administrators) or occasionally 15-1241 (Computer Network Architects), so the metro-level OES figures for those codes are a lagging floor rather than a market read. For a current read, use posted ranges in pay-transparency jurisdictions (including California, Colorado, Washington, New York, Illinois, Minnesota, Maryland, Massachusetts, New Jersey and Vermont), Levels.fyi for public-company bands by level, and the free US Department of Labor OFLC disclosure data, which publishes offered base salaries by employer, job title and location for visa-sponsored roles. |
| Work shape | On-call is the job, not an edge case, and you should assume a rotation unless a posting explicitly says otherwise. Ask three questions before accepting: how many engineers are in the rotation, how many pages a typical shift carries, and whether on-call is compensated. Compensation varies by employer and by country, and in some jurisdictions standby arrangements are governed by employment law or a collective agreement, so check locally rather than assuming the US norm. Mostly hybrid or remote. Plenty of advertised SRE roles are the first reliability hire at the company, which is the fastest learning available and the heaviest pager. |
What a site reliability engineer owns in 2026, and the five kinds of SRE job
Site reliability engineering started at Google as a simple reframing: treat operations as a software problem, give the team that runs a service the power to refuse changes when it is too broken, and make the trade-off between speed and stability explicit instead of political. That reframing is still the core of the job in 2026, and understanding it is the difference between an interview where you sound like an operator and one where you sound like an engineer. You are not hired to keep everything up. You are hired to decide, with numbers, how much down is acceptable, make the system behave that way at the lowest reasonable cost, and give everyone else a credible answer when they ask whether they can ship.
In practice the work has five recurring pieces. You define what reliable means for a specific service, as a service level indicator measured at a particular point with a particular window, and a target that someone senior has agreed to. You build the observability that makes that number real, and the alerting that pages a human only when a human is required. You run incidents, which means detection, mitigation, communication and a postmortem that produces changes someone actually makes. You plan capacity and manage cost, because an unavailable service and an unaffordable one fail the same business in different ways. And you remove toil, which is the manual, repetitive, automatable work that scales linearly with the size of the service and leaves nothing behind when it is done.
What you do not own matters just as much, because candidates lose credibility by claiming too much. You usually do not own the product roadmap, the feature code, or the decision to accept a risk. You own the quality of the information that decision is made with, and in a mature organisation you own the error budget policy that fires when the number goes bad. The clearest signal that someone has really done the job is that they talk about reliability as a budget that is spent, not as an absolute to be defended.
The title is unreliable and that costs people weeks. Set saved searches on Production Engineer, Platform Engineer, Infrastructure Engineer, Systems Engineer, Observability Engineer, Cloud Infrastructure Engineer and DevOps Engineer alongside SRE, and read the responsibilities every time. Two postings with the same title can be completely different jobs, and one of them will not have a pager at all.
The five kinds of SRE job below hire for noticeably different things. Decide which one you are applying to before you write a single bullet of your resume, because the evidence that wins one of them is close to irrelevant in another. Two further patterns follow the five, and neither is a variant to aim at so much as a thing to recognise before you accept an offer.
- Embedded or product SRE. You are attached to one product team and own their services with them. Heavy on SLOs, incident response, release safety and arguing well with product managers. Hires people who can write code and hold a position. This is the variant where the error budget conversation is a real part of the week.
- Infrastructure or platform SRE. You own the layer everyone else builds on: Kubernetes clusters, CI/CD, service mesh, the observability stack, the paved road. Measured by what other teams can ship without talking to you. Deepest Kubernetes and Linux expectations, and the variant where internal customer empathy is tested hardest.
- Core infrastructure SRE at scale. Storage, networking, databases, load balancing, compute fleets at companies large enough to build their own. The highest bar on distributed systems fundamentals and the slowest loop. Expect deep questions on consensus, replication, partial failure and the specific failure modes of one system you claim.
- The solo SRE, or first reliability hire. A company with real traffic, no reliability practice and a team that is tired. You will define the first SLO, build the first on-call rotation, write the first runbooks and clean up a monitoring estate nobody owns. Fastest learning in the field and the hardest pager. Ask hard questions about executive support before accepting, because without it you become a one-person incident queue.
- Reliability engineering sold as a consulting engagement, at consultancies and managed service providers. The work is real and the exposure to many estates is genuinely educational. The trade is that you rarely stay long enough to live with your own decisions, which is the exact experience interviewers probe for when you move back in-house.
- Not a variant, a warning: SRE in name only. A posting that is really systems administration, ticket queue work or a NOC seat with a modern title. Not a bad job, and sometimes a legitimate door in, but know what you are accepting. The tells are a responsibilities list with no mention of code, no mention of SLOs, and the phrase monitor and escalate.
- Also not a variant: the SRE job that is actually a one-person DevOps function, where the pager is real but so is sole ownership of CI, laptops, the VPN and the AWS bill. Ask what fraction of the week is reliability work and what fraction is everything else nobody else wanted.
The doors in, and exactly what each background is missing
Almost nobody starts as an SRE. The role assumes you have already been responsible for something running, which is why the junior market for it is thin and why most hires are lateral moves made between two and eight years into a technical career. The useful question is not whether you are qualified but which specific gap your background leaves, because interviewers are experienced at finding it and you are better off closing it deliberately than being surprised by it in round four.
The honest ordering of doors, by how quickly they convert, is: internal transfer first, then backend engineering, then platform or DevOps work, then systems administration, then network operations, then support. The internal transfer wins by a distance because it skips the hardest part of the screen, which is proving you have operated something. If your current employer has an SRE or platform team and you can get on their on-call rotation, do that before you apply anywhere else.
- From backend software engineering. You can code, which is the piece others struggle to fake. What is missing is usually depth below the runtime: how Linux actually schedules and allocates, what the kernel is telling you, how DNS and TLS and TCP fail in ways that look like application bugs, and how to read a system you did not write. Fix it by taking production ownership where you are, joining the rotation and spending time with strace, tcpdump, perf and your cloud provider's networking documentation.
- From systems administration or IT infrastructure. You know machines, networks and failure, and you have usually been on call for years, which counts. What is missing is writing software that other engineers would accept in review, and distributed systems reasoning beyond a single host. Fix it by shipping something real in Go or Python with tests, reviews and a deployment pipeline, then rewriting one of your recurring manual tasks as a service rather than a script.
- From DevOps or platform engineering. The closest adjacency, often the same job already. What is missing is usually the reliability discipline itself: pipelines and infrastructure as code are not the same as SLOs, error budgets and incident command. Fix it by defining one real SLO against a service your organisation already runs, instrumenting it, publishing the burn rate and writing the policy for what happens when the budget is exhausted.
- From a network operations centre or a monitoring team. You have seen more incidents than most candidates and you understand escalation and comms. What is missing is the engineering half, and the perception problem that comes with it. Fix it by owning the alerting estate as code rather than as a console, deleting alerts that never led to action and documenting the reduction, then writing automation that closes a class of ticket permanently.
- From technical support or a technical account manager role. You can debug under pressure with a customer watching, which is a real skill and transfers well into incident communications. What is missing is everything a hiring manager cannot see from a ticket queue. Fix it with a deliberate intermediate step: a tooling, platform or support engineering role with code in it, then SRE. Expect twelve to eighteen months and plan it as two moves, not one.
- From a degree with no production experience. This is the hardest door and it is honest to say so. Target junior production engineer, cloud operations or infrastructure engineer roles at companies large enough to run a graduate programme, and build the one thing that substitutes best for experience: a small service you run yourself, with an SLO, alerting, a load test, a documented incident you caused on purpose and a postmortem you wrote.
The three artefacts every SRE interview asks for: SLO, incident, on-call
This is the section that decides whether you get hired, so read it slowly. Interviewers for this role are testing one thing above all others: whether your reliability vocabulary is attached to anything you have actually done. The test is not whether you can define a service level objective. It is whether, when asked for one you defined, you can name the indicator, the measurement point, the window, the target, who agreed it and what it changed. Candidates who have done the work answer immediately and in specifics. Candidates who have read about it produce a definition and then slow down.
Prepare three stories before you apply, write them down, and rehearse them until the numbers come out without hesitation. One SLO, one incident, one piece of on-call or toil work. Every behavioural question in an SRE loop is a different angle on these three, and having them ready is worth more preparation time than another week of algorithm practice.
The SLO story has a required shape. Name the service and who depends on it. Name the indicator precisely: not availability, but the proportion of HTTP requests to a specific endpoint that returned a non-5xx status within a stated latency threshold, measured at the load balancer rather than in the application, because the application cannot report the requests it never received. State the window and whether it rolls. State the target and, critically, say why that number rather than a rounder one, which is where interviewers learn whether you chose it or inherited it. Then say what the error budget policy was: what happened when the budget was burning too fast, who was allowed to override, and whether a release was ever actually stopped. If no release was ever stopped, say so and say why, because an honest answer about an SLO that was decorative beats a fictional one about a policy that was enforced.
The incident story has a different required shape, and the most common mistake is telling it as a mystery novel. Interviewers want the timeline split cleanly: when it started, when it was detected and by what, who was paged, what the first mitigation was, when customer impact ended, and only then what the cause turned out to be. Separate mitigation from fix explicitly, because the single most telling sentence in an SRE interview is some version of we stopped the bleeding first and found the cause afterwards. Name your role honestly. If you were the incident commander, say what you delegated and what you decided. If you were a responder, say what you contributed and who ran it. Then give the postmortem: what the action items were, which ones actually shipped, and what has not happened since. An incident with no durable change is half a story.
The on-call and toil story is the one candidates most often skip, and it is the one that separates people who have lived with a system from people who have visited it. The useful numbers are concrete and you should know yours: how many people were in the rotation, how long a shift was, how many pages a typical shift carried, what proportion of those pages required a human action rather than acknowledging and waiting, and what you did about the ones that did not. Deleting an alert that has paged thirty times and been actionable twice is real reliability work and most candidates never mention it. The same applies to toil: name the manual task, the hours a week it consumed, what you automated or eliminated, and what the hours are now.
A warning about numbers. Do not invent them. Interviewers for this role ask follow-up questions that collapse fabricated figures quickly, usually by asking how it was measured. If you do not remember a number precisely, give the shape and say it is approximate: roughly a dozen pages a week, of which about half were one noisy check. That reads as honest. A precise number you cannot defend reads as a lie and it ends the interview quietly.
- SLI, not metric. Be able to say the difference out loud: an indicator is a carefully defined measure of service behaviour from the user's point of view, chosen because a human cares about it. CPU utilisation is a metric and almost never an indicator.
- Know the difference between SLI, SLO and SLA cold, because it gets asked as a warm-up and a fumbled answer colours the hour. The indicator is the measurement, the objective is the internal target you run to, and the agreement is the contractual promise with a financial consequence attached. The objective is set tighter than the agreement so you find out before the customer does.
- Be able to convert a target into time, in your head, as a sanity check: a month is roughly 43,200 minutes, so 99.9 percent leaves about 43 minutes of budget and 99.99 percent leaves about four. That arithmetic is how you show an executive what another nine actually costs.
- Measurement point matters more than candidates expect. Client side, edge, load balancer, service, downstream dependency: each sees a different failure. Say where yours was measured and what it was therefore blind to.
- Know the four golden signals (latency, traffic, errors, saturation), the RED method for request-driven services (rate, errors, duration) and the USE method for resources (utilisation, saturation, errors), and say which you used and why. Reciting all three without a preference is a tell.
- Burn rate alerting is the current expectation, not static threshold alerting on an SLO. Be able to explain a multi-window, multi-burn-rate alert in plain words: a fast burn pages, a slow burn opens a ticket.
- Separate detection time from mitigation time from resolution time, and know roughly what yours were. Time to detect is the number most teams are quietly worst at and the one an interviewer will dig into.
- Name the incident roles you have used: incident commander, operations lead, communications lead, scribe. If your organisation used different names, say so. If it had no roles, say that too and describe what went wrong as a result, because that is a better answer than inventing a structure.
- Blameless postmortems are expected vocabulary, but the credible version includes the hard part: how action items were tracked, who owned them, and what proportion actually shipped. Most organisations are bad at this and saying so honestly is a strength.
- Toil has a definition worth using precisely: manual, repetitive, automatable, tactical, without enduring value, and growing at least linearly with service size. Using it loosely to mean work I did not enjoy is noticed.
- Have one story about a change you refused or delayed, and one about a risk you accepted that you would accept again. Both show judgement rather than reflex.
- Have one story about an incident you made worse, or a fix that caused a second outage. Senior interviewers specifically ask, and a candidate with no such story is either inexperienced or not being straight.
- If you have never had a formal SLO, say so and describe the implicit one: what level of failure actually triggered a response, what people complained about, what you watched. That is a real answer and it beats a textbook one.
- Numbers that belong in these stories, when you have them honestly: rotation size, pages per shift, alert count before and after, toil hours per week before and after, time to detect, time to mitigate, change failure rate, and infrastructure cost per unit of traffic.
How the hiring loop runs in 2026-27, round by round
Who screens you depends on the company size and it changes what the first conversation rewards. At a large technology employer a recruiter screens against a checklist and your job is to say the words on it clearly: Kubernetes, Linux, Terraform, Go or Python, on-call, SLO, incident. At a company under about three hundred people you are more likely to speak to the hiring manager first, usually the person who owns the pager today, and that conversation is won by specificity rather than keywords. Ask them in the first five minutes what the current on-call load looks like and what the last serious incident was. Their answer tells you more about the job than the posting does, and asking reads as someone who has done it.
The loop below is the common shape in 2026-27. It is not universal, and a good recruiter will tell you the exact stages if you ask. Ask explicitly whether the coding round is algorithms or practical, because preparing for the wrong one is the single most common avoidable failure in this loop.
Timelines have been compressed at mid-sized companies and remain slow at the largest ones. Three to six weeks is normal. If a process stretches past eight weeks without a clear reason, keep applying elsewhere rather than waiting, because a stalled loop usually means an internal problem with the role rather than hesitation about you.
- Recruiter screen, 20 to 30 minutes. Checklist matching, salary expectation, location and on-call willingness. Say clearly and early that you expect to be on call and have been. Hesitation here is sometimes fatal because it is the single hardest thing to staff.
- Hiring manager call, 30 to 60 minutes. The real screen. Expect to be asked for your incident story and your SLO story within the first fifteen minutes. This is also where you should establish which of the five kinds of SRE job this is.
- Coding round, 45 to 60 minutes. At most companies this is practical: parse these logs, write a script that reconciles two inventories, implement a retry with exponential backoff and jitter, write a small concurrent worker in Go. At the largest employers it is often a standard medium-difficulty data structures and algorithms round instead. Python is the safest default language unless the team writes Go, in which case write Go.
- Linux and networking round, sometimes folded into the troubleshooting round. Process states, memory accounting and what the out-of-memory killer does, file descriptors, cgroups and namespaces, what happens between typing a URL and a response arriving, why DNS caching surprised you once, TLS handshake failure modes, TCP retransmission and connection queue limits. Depth here is the clearest separator between candidates from a software background and candidates from an infrastructure background.
- Systems design round, 45 to 60 minutes. See the next section. This is not the generic design interview from a product engineering loop and preparing with only that material is a mistake.
- Troubleshooting round, 45 to 60 minutes. Often the deciding round. See the next section.
- Incident behavioural round, 45 to 60 minutes. Entirely your three artefacts, probed hard. Expect follow-ups that test whether the story is yours: how was that measured, who else was in the room, what did you consider and reject, what would you do differently.
- Optional extras that appear in a minority of loops: a short take-home of four to eight hours, usually instrumenting or fixing a small service; an infrastructure-as-code review where you critique someone's Terraform; a chaos or game-day exercise; and a values or cross-functional round where a product manager tests whether you can have the error budget conversation without either collapsing or picking a fight.
The troubleshooting round and the design round, where offers are decided
The troubleshooting round is the most distinctive thing about an SRE loop and the round most candidates underprepare. The format varies: a shared terminal into a deliberately broken environment, a set of dashboards and logs with a described symptom, or a purely verbal exercise where the interviewer plays the system and answers your questions. The content varies too. The rubric does not. Interviewers are grading a sequence: do you establish the symptom and its blast radius before touching anything, do you form a hypothesis before running a command, does each command actually discriminate between hypotheses, do you narrate what you expected and what you saw, do you ask about recent changes early, and do you mitigate before you investigate when users are affected.
Three specific behaviours separate a strong performance from a mediocre one. First, ask what changed, in the first two minutes, every time, because deployments, configuration changes and certificate rotations are behind the great majority of real incidents and a candidate who starts by reading kernel parameters is signalling the wrong instinct. Second, bound the problem before narrowing it: how many users, which regions, which endpoints, when did it start, is it still happening. Third, say out loud what you would do to stop customer impact right now, even if you then continue investigating, because the interviewer is partly checking whether you know the difference between curiosity and duty. Candidates who find the elegant cause while the imaginary service stays down for forty minutes do not get offers.
A quiet failure mode worth naming: going silent while you think. The interviewer cannot grade silence, and in a real incident silence is a defect. Narrate. If you are stuck, say what you have ruled out and what you would ask a colleague. That is exactly what a good responder does at three in the morning, and it is scored as a strength rather than a weakness.
The design round is also not the one you may have prepared for. A product engineering design interview asks you to sketch a system. An SRE design interview asks you to sketch a system and then do arithmetic about it, a practice Google's SRE material calls non-abstract large system design. Expect to be pushed on numbers: given this request rate and this response size, how much bandwidth, how many machines, how much memory for that cache, how long to refill it after a cold start, what happens at double the traffic, what happens when the largest failure domain disappears. You are allowed to round aggressively. You are not allowed to wave.
The reliability content of that round is predictable and you should rehearse it: where the single points of failure are, what the blast radius of each component is, how a deployment is rolled out and rolled back, what degrades gracefully and what fails hard, what you shed when overloaded and in what order, how retries are bounded so they do not amplify an outage, and what the SLO for the thing you just drew should be. If you draw a retry without a budget, a jitter and a circuit breaker, expect to be asked about retry storms and cascading failure, because that is the classic trap and it is set deliberately.
- Rehearse the first ninety seconds of a troubleshooting round until it is automatic: establish symptom, establish scope, establish start time, ask what changed, state the mitigation you would reach for. Everything after that is easier.
- Practise against real surfaces, not quizzes. A single node kind, k3s or minikube cluster on a laptop is enough. Break it yourself, then debug it: fill a disk, exhaust file descriptors, misconfigure a readiness probe so it always passes, add 200 milliseconds of latency to a dependency with tc, let a certificate expire, set a memory limit too low and watch the out-of-memory killer work.
- Know a small set of commands cold and know what each one would tell you: dmesg, journalctl, ss, dig, curl with timing output, tcpdump, strace, top and vmstat, df and du, kubectl describe and kubectl logs with previous, and the equivalent query in whichever metrics system you claim.
- For the design round, memorise a handful of rough magnitudes rather than precise constants: a disk seek against a memory read, a cross-region round trip against a same-rack one, what a single commodity machine can plausibly serve. Being roughly right in public beats being precisely silent.
- Expect at least one question about cost. Reliability has a price and the senior version of this role is able to say what a nine costs and recommend against buying it. A candidate who treats availability as free reads as junior.
- Ask what the on-call rotation is called upon to do during the design round, and design for it. Saying that you would not page a human for this because the system can shed load and recover is often the strongest sentence in the hour.
The stack employers list in 2026, and where depth actually pays
A posting's requirements list is a wish list written by a committee and nobody expects all of it. What interviewers actually test is narrower and more stable than the list suggests: Linux below the surface, networking, one cloud in depth, Kubernetes, infrastructure as code, one real programming language, and an observability stack you can reason about rather than click through. Depth in a few of these beats familiarity with all of them, and the interview finds out which you have within about ten minutes.
The one strategic choice worth making deliberately is a specialism. Generalist SREs are plentiful. The people who get the strongest offers have one area where they are genuinely deep and can teach it: networking, databases and storage, Kubernetes internals, performance analysis, or now inference serving. Pick one that interests you, go properly deep, and keep enough breadth to be useful everywhere else.
Some things on these lists are near-universally expected in 2026 and some are nice to have. The bullets below separate them, with a note on where depth actually pays rather than merely appearing on the resume.
- Linux, deep. Assumed everywhere, tested everywhere, and the most common place candidates from a pure application background fall down. Process and memory accounting, cgroups and namespaces, file descriptors, the page cache, signals, systemd, and reading kernel messages without panic. Depth pays here more than anywhere else on this list.
- Networking. DNS, TCP behaviour under loss, TLS handshakes and certificate chains, HTTP/2 and HTTP/3 differences that matter operationally, load balancer layer 4 against layer 7, NAT, MTU, and enough cloud VPC design to debug a connectivity problem. Depth pays, because almost every confusing incident eventually becomes a networking question.
- Kubernetes. Effectively mandatory. Not just kubectl: the control plane components and what happens when one is unhealthy, scheduling and why a pod is Pending, probes and what a wrong readiness probe does during a rollout, resource requests against limits, horizontal and vertical autoscaling, Cluster API or Karpenter-style node provisioning, and etcd as a failure domain. CKA is the credential people respect here.
- One cloud, in depth, rather than three shallowly. AWS carries the largest share of postings, Azure is strong in enterprise and regulated sectors, and Google Cloud is heavily represented in companies that run the SRE model closest to its original form. Know its identity model, networking, managed databases, load balancing, quotas and the specific ways it fails.
- Infrastructure as code and GitOps. Terraform or OpenTofu is the default expectation, Pulumi appears, and Argo CD or Flux covers the delivery half. The interview question is almost always about state, drift, blast radius of a plan, and how you review a change that would delete something.
- One programming language you would be allowed to ship in. Python and Go cover the overwhelming majority of postings, with Go dominant in platform and infrastructure teams. Bash is assumed and is not a substitute. Rust appears in a small number of performance-focused infrastructure teams.
- Observability. Prometheus and its query language, Grafana, OpenTelemetry as the collection standard, and at least one of the commercial platforms (Datadog, Honeycomb, New Relic, Splunk, Dynatrace, Elastic, Grafana Cloud). Be able to talk about cardinality and cost, because every observability estate at scale eventually has a bill problem and interviewers like candidates who have met it.
- SLO tooling as a named thing: Sloth, Pyrra, OpenSLO, Nobl9, or the SLO features inside your monitoring platform. Having implemented burn-rate alerting with any of them is a strong, specific signal.
- Incident tooling and process. PagerDuty remains the default pager, with incident.io, FireHydrant and Rootly common alongside it and most monitoring vendors now selling their own on-call product. This corner of the market consolidates and rebrands often, so ask what the team actually runs rather than assuming. What is graded is not the tool but whether you have run the process it encodes.
- Databases and stateful systems. The most under-prepared area among candidates from an application background and a frequent source of real incidents. Replication, failover, connection pool exhaustion, lock contention, the specific recovery behaviour of PostgreSQL or MySQL, and why your cache is now a dependency with no failover. Depth pays.
- Performance analysis and eBPF. The current generation of tooling (bpftrace, Cilium and Hubble, Parca and Grafana Pyroscope for continuous profiling) is a genuine differentiator rather than a requirement. Candidates who can profile a running system without restarting it stand out, especially in core infrastructure roles.
- Chaos and resilience testing. Game days, Wheel of Misfortune exercises, AWS Fault Injection Service, Gremlin, or your own fault injection. Having run one with real participants is worth more in interview than any tool name, because it proves the organisational part, which is the hard part.
What SREs are paid, and where to look it up instead of trusting a band
Do not trust a single salary number for this role, including any you read in an article. The spread is enormous and driven by factors that have nothing to do with your skill: company stage, whether equity is liquid, whether the employer treats SRE as a software engineering level ladder or an operations ladder, and how badly they need someone willing to carry a pager. Two engineers with the same experience can be paid very differently at companies in the same city.
The single most useful structural fact to know before negotiating: at companies that put SREs on the software engineering ladder with the same levels and the same bands, pay is at or near parity with backend engineers. At companies that run a separate operations or infrastructure ladder, it is usually below. Ask the recruiter directly which it is, in those words, early. It is a normal question, it signals that you know the market, and the answer predicts your ceiling at that employer better than the opening number does.
Where to look, in order of reliability. First, posted ranges in pay-transparency jurisdictions, which are legally required disclosures rather than survey estimates and are the closest thing to ground truth. Several US states and cities require a range in the posting, including California, Colorado, Washington, New York, Illinois, Minnesota, Maryland, Massachusetts, New Jersey and Vermont, and the set changes, so search current postings in those places for the same title rather than relying on a list. Second, the US Department of Labor OFLC disclosure data, which publishes offered base salaries by employer, job title and location for visa-sponsored roles, is free, and is real offers rather than self-reports. Third, Levels.fyi for public-company bands by level, which is self-reported but large and well-structured for exactly this kind of role. Fourth, BLS OES figures under 15-1252, 15-1244 or 15-1241 as a conservative floor, remembering that OES lags the market and that none of those codes is specifically this job.
Two components people forget to negotiate. On-call compensation is a real and separately negotiable item at some employers, whether as a stipend per shift, time off in lieu, or a differential; at others it is considered part of base. Ask, because an unpaid heavy rotation is a meaningful pay cut that does not appear in the offer letter. And level matters more than base: moving from a mid level to a senior level usually moves total compensation more than any negotiation within a level, so if you are being levelled below where your evidence supports, argue about the level first and the number second.
- Ask the recruiter in the first call: is SRE on the software engineering ladder here, and is the band the same? The answer predicts your ceiling.
- Ask how many people are in the rotation and how many pages a typical shift carries. A four-person rotation with nightly pages is a different job from a twelve-person rotation that is quiet, at the same salary.
- Ask whether on-call is compensated and how. The answer varies by employer and by country, and in some jurisdictions standby arrangements are governed by employment law or a collective agreement, so check the rules where you will actually be employed.
- Specialisation moves pay. Deep networking, deep database, deep Kubernetes internals and now inference-serving reliability all command more than generalist infrastructure work, because the candidate pool is smaller and visible.
- Industry changes the shape of the offer more than the title does. Finance and trading pay the most in cash for latency-sensitive reliability work and expect the most rigid process. Early-stage startups pay in equity and autonomy. Public sector and regulated industry pay less and are more likely to offer a genuinely humane rotation.
The resume when the work is invisible, and running the search
The structural problem with an SRE resume is that the job's successes are things that did not happen. Nobody notices the outage you prevented. The fix is to convert absence into measured change, which means every bullet should contain a number that existed before and a number that exists now, or the thing you built and what it made possible. A resume made of tool names reads as a list of things you have been in the same room as. A resume made of outcomes reads as someone who has owned something.
What gets read, roughly in order: the top third of page one, the most recent role, and any line with a number in it. What gets ignored: a skills section listing forty technologies, a summary paragraph describing you as passionate about automation and cloud-native technologies, certifications stacked above experience, and every bullet starting with responsible for. Two pages is fine for this role once you are past five years. One page is fine before that.
Write the bullets as before and after. Reduced paging volume from roughly forty alerts a week to under ten by deleting non-actionable checks and moving the rest to burn-rate alerting, with no increase in time to detect. Defined the first SLOs for three customer-facing services and the error budget policy that stopped two releases. Cut the manual release process from ninety minutes of hands-on work to a fifteen-minute automated rollout with automatic rollback. Reduced infrastructure cost per thousand requests by a third by right-sizing requests and limits and moving batch work to spot capacity. Those are the shapes. Use your real numbers, and only numbers you could defend under questioning.
On keywords: applicant tracking systems in 2026 are mostly doing semantic matching rather than naive string matching, but the practical advice has not changed much, because a human still skims and because recruiters still search for exact strings. Use the posting's own vocabulary where it is true of you. If they say service level objectives, do not write SLAs. If they say Kubernetes, do not write K8s only. Spell out an acronym once with the acronym in brackets after it, which covers both the search and the reader.
Run the search deliberately rather than by volume. Twenty well-targeted applications with a tailored first bullet beat two hundred generic ones, and for this role the referral route is unusually effective because the community is small and visible. The people hiring SREs read incident writeups, go to SREcon and local meetups, and notice who answers well in public. Publishing one genuinely good postmortem, or one honest writeup of something you broke and fixed, is worth more than a portfolio of tutorials.
If you have no production experience to point at, build the one project that substitutes best. Not a tutorial deployment: a small service you actually run, with a defined SLI and SLO, Prometheus and Grafana wired up, burn-rate alerts that reach your phone, a load test with k6 or Vegeta that establishes its breaking point, infrastructure as code for all of it, and a documented incident you caused deliberately with the postmortem you wrote afterwards. That is a weekend project repeated over about six weekends, and it gives you honest answers to the three artefact questions that would otherwise sink the interview. Say plainly that it is a personal system rather than production traffic. Interviewers respect that framing and resent the opposite.
- Lead each role with scale and ownership in one line: what the service did, how much traffic, how many services you were on call for, how big the team was. Context makes every bullet below it legible.
- Put the three artefacts on the resume explicitly, not just in your head. One SLO line, one incident line, one toil or alerting line, in the most recent role.
- Quantify the pager. Rotation size, page volume before and after, time to detect before and after. Almost nobody does this and it stands out immediately.
- Name the stack once, inside the role where you used it, rather than in a skills wall. Terraform in the sentence about what you built is worth more than Terraform in a list.
- Drop certifications below experience once you have experience. Above it, for a career changer, CKA plus one cloud certification is enough. Six certifications signal compensation for something missing.
- If you have a public artefact, link it near the top: a conference talk, a postmortem you wrote, a tool you maintain, a substantial contribution to an infrastructure project. For this role, one of these is worth several resume bullets.
- Apply to the kind of SRE job you are actually ready for. A first reliability hire at a fifty-person company is a realistic target from a strong systems administration background and a core infrastructure role at a hyperscaler is usually not, and applying to the wrong one repeatedly produces rejections that teach you nothing.
What a site reliability engineer must know about AI in 2026-27
Start with the honest part, because this role attracts more vendor noise than most and interviewers can tell within one answer whether you have believed it. At the core of the job, less has changed than the marketing suggests. Linux still behaves the way it behaves. TCP still retransmits. A full disk still takes a service down. Capacity planning is still arithmetic. And nobody has automated the part where a human decides at three in the morning whether to roll back a release that might be the cause, with incomplete information and a product director on the call. A strong SRE from 2022 is still a strong SRE. If your plan is to sound current by saying AI a lot, you will interview worse than someone who says plainly that the fundamentals did not move.
What genuinely changed is the environment around the job, in four ways that are worth naming precisely because each one shows up in interviews. The first is change volume. Assistants now produce first drafts of application code, Terraform, Helm charts, Dockerfiles, CI configuration, dashboards and runbooks at most employers, and more change is reaching production, authored faster, by people with less context about the systems they are touching. Change has always been the leading cause of incidents, so the practical consequence for SREs is pressure on everything that makes change safe: progressive rollout, automatic rollback, feature flags, canary analysis, and a review culture that catches the configuration which is syntactically perfect and operationally wrong. Being able to say that you tightened the release path specifically because change volume rose is a current, credible answer to a question about AI.
The second is that a growing number of SREs are being handed inference services to run, and they fail differently from the services you already know. Latency is not one number but two that users feel separately, time to first token and tokens per second thereafter. Capacity is bounded by accelerators with long procurement lead times and provider quotas rather than by autoscaling a stateless fleet, so the classic answer of scaling out is often simply unavailable and the real levers become queueing, admission control and degradation. Correctness is not a status code: a model can return 200 with an answer that is wrong, which means output quality needs its own monitoring and the familiar error-rate SLI covers less of the real risk than it used to. Cost per request becomes something you are expected to watch continuously, because it is far higher than a conventional request and can move by multiples with a configuration change. And graceful degradation acquires a new form specific to this workload: falling back to a smaller or cheaper model, shortening context, or serving a cached answer, rather than simply shedding load.
The third is automation that holds credentials. Agents, assistants and automated remediation systems increasingly have the ability to call internal APIs, open pull requests, scale infrastructure and execute runbook steps. The reliability questions this raises are old questions pointed at a new actor, and SREs are the ones being asked them: what is the blast radius if it is wrong, what is the rate limit, is the action idempotent, is there a kill switch and who can reach it, what does it have permission to do that it does not need, and what happens when it retries ten thousand times because it misread a transient failure. Being able to answer those calmly, with the same vocabulary you would use for any other client of your systems, is a current and valuable skill. Being able to say that you refused to give an automated system write access to something until it had a dry-run mode and an audit trail is better.
The fourth is the tooling pointed at incidents themselves, and here precision matters because it is overclaimed. Monitoring and incident platforms now do alert correlation and grouping, draft incident summaries and timelines, surface recent deploys and changed configuration alongside an alert, and propose likely causes. The useful parts are genuinely useful: drafting a timeline and a first summary saves real time during an incident, and correlation reduces the number of separate pages for one underlying fault. The limits are equally real: these systems are good at surfacing correlation and poor at distinguishing cause, they are only as good as the ownership metadata and runbooks you have maintained, and they produce confident narratives that are sometimes wrong, which is dangerous precisely when a tired responder is relying on them. The professional position, and the one interviewers respond to, is that these tools draft and you decide. Treat a proposed cause as a hypothesis to be tested with evidence, never as a finding. If you have used any of this in anger, say specifically what it did well and what it got wrong, because that answer is almost impossible to fake.
A practical note on the interview itself. Some employers now hand you an assistant during the coding or troubleshooting round and grade how you direct it: whether you ask it for the right thing, whether you verify its output, whether you catch the readiness probe that always returns healthy or the autoscaler configuration that will oscillate. Others still forbid it. Ask the recruiter which, and if assistance is allowed, practise the specific skill of using one while narrating your own reasoning, because candidates who go quiet and paste are marked down even when the answer is right. One more thing worth saying plainly if you are asked whether AI will replace this role: the mechanical parts of SRE have been getting automated continuously for two decades, mostly by SREs, and that has consistently raised the number of systems one engineer can be responsible for rather than reduced the need for engineers. What is not automatable is accountability, and the pager is accountability made concrete.
Keeping the change path safe when change volume rises
Assistance has increased how much code and configuration reaches production and how fast, without increasing the context the authors have about the systems they are changing. Change remains the leading cause of incidents, so the safety of the release path is now doing more work than it was designed for. This is the most concrete, least hyped way AI has affected reliability engineering, and a hiring manager who is living it will recognise immediately whether you are.
Show it: One release path you made safer, described mechanically: progressive rollout with a stated canary percentage and bake time, the automated analysis that decides to promote or roll back, what the rollback does about database migrations, and the change failure rate or recovery time before and after. If you have reviewed generated infrastructure code, name a specific class of error you now look for: an always-passing probe, an over-broad IAM or RBAC grant, a resource limit copied from an example, a retry with no jitter or budget.
Running inference services as a reliability problem
This is the fastest-growing new workload landing on SRE teams, and the failure shapes are unfamiliar enough that experienced engineers get it wrong on first contact. The constraints are different: capacity is bounded by accelerators and quota rather than by autoscaling, latency has two user-visible components, correctness is not a status code, and cost per request is high enough to be a first-class operational concern rather than a quarterly finance conversation.
Show it: Be able to propose the SLIs for a model-backed endpoint without prompting: availability, time to first token, tokens per second, queue wait time, and a quality or refusal-rate indicator measured separately from HTTP status. Describe one degradation strategy you would actually implement, such as falling back to a smaller model, truncating context, or serving from a semantic cache, and say what the user sees when it fires. If you have run one, name the serving stack (vLLM, NVIDIA Triton, KServe, Ray Serve or a provider API behind a gateway) and the saturation signal you alerted on.
Capacity planning under quota and lead time
The classic capacity answer, add instances, is frequently unavailable for accelerated workloads. Supply is constrained by provider quota, reservation commitments and procurement lead times measured in weeks or months, which pushes the problem back into forecasting, admission control, prioritisation between workloads, and arguing for commitments before they are needed. This is recognisably the oldest part of SRE applied to a new constraint, which makes it a good interview topic for showing range.
Show it: Describe a forecast you made and whether you were right, including the demand signal you used and the headroom you held. For accelerated capacity specifically, be able to discuss the trade between reserved, on-demand and spot capacity, queueing low-priority batch work onto idle capacity, and what you would shed first under contention. Knowing that headline accelerator utilisation can look high while the hardware is doing very little useful work, and naming how you would see the difference, is a strong signal.
Blast radius for non-human actors
Automated systems with credentials are now ordinary clients of production, and they behave differently from humans: they are faster, they retry harder, they do not get tired and they do not hesitate. SREs are the people being asked to make that safe, and the organisations deploying them often have not thought about it. This is where reliability and security overlap, and candidates who can hold both frames at once are scarce.
Show it: Have a position, with reasons, on what an automated agent should be allowed to do in your production environment without a human approving it. Be specific about the controls: dry-run mode, scoped and short-lived credentials, per-action rate limits, idempotency keys, an audit trail that names the actor, a kill switch that does not depend on the system being disabled to be healthy, and a circuit breaker on repeated failure. One real story about an automation that misbehaved, and the guardrail you added afterwards, is worth more than any framework.
Alert quality, with and without correlation tooling
Alert fatigue is still the most common chronic illness of an on-call rotation, and correlation and grouping features have reduced the symptom without curing the cause. Every hiring manager has a noisy estate and wants to know whether you will reduce it or add to it. The engineers who actually fix this are the ones willing to delete alerts, which is socially harder than adding them.
Show it: Numbers: alerts before and after, the proportion that required a human action, and what you deleted rather than tuned. Be able to explain multi-window multi-burn-rate alerting in plain words and say why you would page on symptoms rather than causes. If you have used correlation or summarisation features in an incident platform, say concretely what they saved and where they misled you, and make clear that you treated a suggested cause as a hypothesis rather than a conclusion.
Documentation and runbooks that a machine can also use
Runbooks acquired a second reader. Incident tooling and assistants now pull from your runbooks, service catalogue and ownership metadata to draft summaries and suggest next steps, which means vague documentation now degrades automated help as well as human help. The quality of your service catalogue, your ownership records and your runbook structure has become an operational input rather than a hygiene chore, and the organisations that neglected it are discovering that their new tooling is unhelpful for reasons that predate the tooling.
Show it: Describe runbooks you wrote that are structured rather than narrative: a named symptom, how to confirm it, the mitigation with the exact command, the escalation path, and when to stop. Say whether any of them became executable, through automated remediation or a scripted step, and what you insisted on before letting that run unattended. Mentioning that you keep ownership and dependency metadata current because tooling depends on it is a modern, concrete answer.
Incident command, which is the part that did not change
The scarcest skill in this field remains a person who can take command of a confusing, high-pressure situation, coordinate several specialists, communicate clearly to people who are frightened, and decide with incomplete information. No tool does this, and the organisations buying the most incident tooling are frequently the ones whose incidents still go badly, because the problem was never the tooling. Saying so plainly in an interview is a strong signal precisely because it is unfashionable.
Show it: One incident you commanded, told as a timeline with roles named and decisions attributed, including a decision you made with insufficient information and what you did to limit the damage if you were wrong. Describe how you communicated to non-technical stakeholders during it and how often. If you have run a game day or a Wheel of Misfortune exercise, say how you designed it and what it revealed, because practising incident response is itself the evidence that you take it seriously.
What a screen is looking for
These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.
- Site reliability engineering
- SRE
- Production engineering
- Platform engineering
- Infrastructure engineering
- Service level objective (SLO)
- Service level indicator (SLI)
- Service level agreement (SLA)
- Error budget
- Error budget policy
- Burn rate alerting
- Four golden signals
- RED method
- USE method
- Toil reduction
- On-call
- On-call rotation
- Incident response
- Incident commander
- Incident management
- Blameless postmortem
- Mean time to detect
- Mean time to recovery
- Change failure rate
- DORA metrics
- High availability
- Disaster recovery
- RTO
- RPO
- Capacity planning
- Load testing
- Chaos engineering
- Game day
- Fault injection
- Linux
- Linux kernel
- systemd
- cgroups
- namespaces
- strace
- tcpdump
- eBPF
- bpftrace
- Continuous profiling
- TCP/IP
- DNS
- TLS
- HTTP/2
- Load balancing
- CDN
- Anycast
- BGP
- VPC
- Kubernetes
- CKA
- CKS
- Helm
- Kustomize
- etcd
- Karpenter
- Cluster autoscaling
- Service mesh
- Istio
- Cilium
- Envoy
- Docker
- containerd
- AWS
- Google Cloud Platform
- Microsoft Azure
- Terraform
- OpenTofu
- Pulumi
- Ansible
- GitOps
- Argo CD
- Flux
- CI/CD
- GitHub Actions
- Progressive delivery
- Canary release
- Blue-green deployment
- Feature flags
- Rollback
- Python
- Go
- Bash
- Prometheus
- PromQL
- Grafana
- OpenTelemetry
- Distributed tracing
- Datadog
- Honeycomb
- Splunk
- Elastic
- Loki
- Thanos
- Mimir
- Alertmanager
- PagerDuty
- incident.io
- Observability
- Cardinality
- Log aggregation
- PostgreSQL
- MySQL
- Redis
- Kafka
- Replication
- Failover
- Connection pooling
- Caching
- Rate limiting
- Circuit breaker
- Exponential backoff
- Retry budget
- Load shedding
- Graceful degradation
- Cascading failure
- Distributed systems
- Performance tuning
- FinOps
- Cloud cost optimization
- Secrets management
- RBAC
- Least privilege
- Runbook
- Infrastructure as code
- Model serving
- vLLM
- NVIDIA Triton Inference Server
- KServe
- GPU capacity
- Time to first token
- Inference latency
- AIOps
- Automated remediation
Mistakes that cost people this job
Talking about SLOs in the abstract, with no SLO you personally defined.
Prepare one real example and be able to state the indicator, the measurement point, the window, the target, who agreed it and what the error budget policy actually stopped. If your organisation never had a formal SLO, say so and describe the implicit one: what level of failure actually triggered a response and what people complained about. An honest answer about an informal practice beats a textbook definition, and interviewers can tell the difference inside two follow-up questions.
Using SLA, SLO and SLI interchangeably.
Keep them distinct in every sentence. The indicator is the measurement, the objective is the internal target you run the service to, and the agreement is the contractual promise with a financial consequence. Set the objective tighter than the agreement so the budget warns you before a customer does. This gets asked as a warm-up question and fumbling it colours the rest of the hour.
Telling the incident story as a mystery, building to the root cause.
Tell it as a timeline: when it started, when it was detected and by what, who was paged, what the first mitigation was, when customer impact ended, and then the cause. Separate mitigation from fix explicitly. The sentence we stopped the bleeding first and investigated afterwards is one of the highest-value things you can say in an SRE loop, and the dramatic structure buries it.
Investigating happily while the imaginary service stays down in the troubleshooting round.
State the mitigation you would reach for within the first few minutes, even if you then continue investigating. Interviewers are checking whether you know the difference between curiosity and duty. Candidates who find an elegant cause after forty minutes of unmitigated impact routinely fail this round while believing they did well.
Going silent while you think during a live debugging exercise.
Narrate continuously: here is my hypothesis, here is the command that would discriminate between it and the alternative, here is what I expected and here is what I got. The interviewer cannot grade silence, and in a real incident silence is itself a defect. Saying what you have ruled out and what you would ask a colleague is scored as strength, not weakness.
Preparing for the product engineering systems design interview and assuming it transfers.
Prepare for a design round with arithmetic in it. Expect to be pushed on request rates, bandwidth, machine counts, memory for a cache, cold-start refill time, and behaviour at double traffic and with the largest failure domain gone. Rehearse single points of failure, blast radius, rollout and rollback, load shedding order, and bounded retries with jitter and a circuit breaker, because the retry storm trap is set deliberately.
Presenting a wall of tools as the resume, with no numbers anywhere.
Write bullets with a before and an after. Alert volume before and after, toil hours before and after, release time before and after, cost per thousand requests before and after, time to detect before and after. The SRE resume problem is that your successes are events that did not happen, and numbers are the only way to make absence visible to a skim reader.
Inventing or rounding numbers to sound impressive.
Give the shape and flag it as approximate: roughly a dozen pages a week, about half of them one noisy check. Interviewers for this role test numbers by asking how they were measured, which collapses fabrications quickly and quietly. An approximate number you can explain is worth more than a precise one you cannot.
Hesitating on the on-call question in the recruiter screen.
Say early and clearly that you expect to be on call, have been on call, and have questions about the rotation. Then ask them: how many people, how many pages a shift, is it compensated. Willingness is the hardest thing for these teams to staff, and ambivalence about it is one of the few things that ends a screen immediately. Asking informed questions about the rotation is not a red flag, it is the opposite.
Searching only for the title SRE, and concluding the market is small or saturated.
Run saved searches on Production Engineer, Platform Engineer, Infrastructure Engineer, Systems Engineer, Observability Engineer, Cloud Infrastructure Engineer and DevOps Engineer as well, and read the responsibilities rather than the title. Some of the best versions of this job are posted under a different name, and some postings titled SRE are a renamed ticket queue.
Leaning on AI vocabulary to sound current, or dismissing the subject entirely.
Say the narrow true thing: the core craft did not change, and what changed is the environment. Higher change volume stressing the release path, inference services with unfamiliar failure shapes arriving on your rotation, automation that holds credentials and needs a blast radius, and incident tooling that drafts well and concludes badly. One specific example of each beats any amount of enthusiasm, and a candidate who says a tool got a cause wrong and explains how they caught it is more convincing than one who praises it.
Questions people ask
What does a site reliability engineer actually do?
A site reliability engineer applies software engineering to the problem of operating production systems. The work has five recurring parts: defining what reliable means for a service as a measurable service level indicator and objective, building the observability and alerting that make it real and that page a human only when a human is needed, running incidents including detection, mitigation, communication and the postmortem, planning capacity and controlling cost, and removing toil by automating the manual work that would otherwise grow with the service. The job is not to keep everything up at all costs. It is to make the trade-off between change velocity and stability explicit with numbers, so the organisation can decide it deliberately.
Do I need a degree or certification to become an SRE?
No. There is no licence, no registration and no mandatory certification for site reliability engineering, and no specific degree is required, although most people hired into these roles have a technical degree or equivalent production experience. Because nothing gates entry, the resume screen and the interview are the real filters. Two or three certifications help a career changer get read: CKA (Certified Kubernetes Administrator), because it is performance based in a live cluster, Google Professional Cloud DevOps Engineer, because it is the certification written most explicitly around SRE practice including SLOs and error budgets, and one cloud associate certification matching your target employers. None of them substitutes for one production system you have operated and been paged for.
How do I move into SRE from a software engineering job?
Software engineering is the shortest bridge into a site reliability engineer role, because the coding half is already there. The missing piece is almost always depth below the runtime: Linux process and memory behaviour, cgroups and namespaces, networking failure modes in DNS, TCP and TLS, and the ability to debug a system you did not write. Close it from inside your current job first by taking production ownership of your team's services, joining the on-call rotation, defining one real SLO and instrumenting it, and leading one incident with a postmortem you write. Three to six months of deliberate effort usually makes the move credible, and an internal transfer converts faster than any external application because it proves the one thing a resume cannot.
What is the difference between SRE and DevOps?
DevOps is a set of cultural and practice goals about collapsing the wall between development and operations. SRE is one specific, opinionated implementation of those goals that originated at Google, with concrete mechanisms attached: service level objectives, error budgets as a governing mechanism, a cap on how much time the team spends on toil, and blameless postmortems. In the job market the titles overlap heavily and many companies use them interchangeably, so read the responsibilities rather than the title. The practical tell is that a genuine SRE posting mentions SLOs, error budgets and on-call ownership of running services, while a DevOps posting more often centres on pipelines, infrastructure as code and developer tooling.
What is the difference between an SLI, an SLO and an SLA?
An SLI, a service level indicator, is the measurement itself: for example the proportion of requests to a given endpoint that returned a non-5xx status within a stated latency threshold, measured at the load balancer. An SLO, a service level objective, is the internal target for that indicator over a stated window, such as 99.9 percent over a rolling 28 days, and the remaining 0.1 percent is the error budget the team is allowed to spend on change. An SLA, a service level agreement, is the contractual promise made to a customer with a financial or legal consequence if it is missed. Objectives are deliberately set tighter than agreements so the team finds out it is in trouble before the customer does. Expect to be asked this as a warm-up question in an SRE interview.
What does an SRE interview actually test?
A site reliability engineer interview tests six things, in roughly this order of weight: whether you can debug an unfamiliar system in front of an audience, whether your reliability vocabulary is attached to work you have really done, whether you can write usable code in Python or Go, whether you understand Linux and networking below the application layer, whether you can design a system and do arithmetic about its capacity and failure domains, and whether you behave well under pressure with other people depending on you. The behavioural questions are all angles on three artefacts, so prepare them in specifics: one SLO you defined, with the indicator, the measurement point, the window, the target and what the error budget policy stopped; one incident, as a timeline separating detection, mitigation and durable fix, with your real role in it and which action items shipped; and one on-call or toil story, with rotation size, pages per shift and what you deleted or automated. The troubleshooting round and the incident behavioural round decide most offers, because they are the only rounds where an interviewer sees whether you reason from evidence or recite from memory.
How much on-call should I expect as a site reliability engineer?
A site reliability engineer should assume a rotation unless a posting explicitly says otherwise, because owning running services is the job. The load varies enormously and the salary does not tell you which you are getting, so ask three questions before accepting: how many engineers are in the rotation, how many pages a typical shift carries, and whether on-call is compensated. A four-person rotation with nightly pages and a twelve-person rotation that is mostly quiet are completely different jobs at the same pay. Compensation for standby varies by employer and by country, and in some jurisdictions it is governed by employment law or a collective agreement, so check the rules where you will actually be employed rather than assuming.
What programming language should I learn for SRE?
Python and Go cover the overwhelming majority of site reliability engineer postings. Python is the safest single choice: it is used everywhere for automation, tooling and glue, and most practical coding rounds accept it. Go is the dominant language in platform and infrastructure teams and in the cloud native ecosystem itself, so it is worth learning if you are targeting that kind of team, and it is often what the interviewer writes. Bash is assumed and is not a substitute for either. Learn one properly rather than three shallowly, because the round tests whether you can write code a colleague would accept in review, not whether you can recognise syntax.
Can you become an SRE with no production experience?
Becoming a site reliability engineer with no production experience is the hardest door in, and it is honest to say so, because the role assumes you have already been responsible for something running. Two routes work. The first is a deliberate intermediate step: a junior production engineer, cloud operations, infrastructure or platform role at a company large enough to train people, then SRE within a couple of years. The second is building the substitute evidence yourself: a small service you actually run, with a defined SLI and SLO, Prometheus and Grafana wired up, burn-rate alerts that reach your phone, a load test that establishes its breaking point, infrastructure as code for all of it, and a deliberate incident you caused with the postmortem you wrote. Present it honestly as a personal system rather than production traffic. It will not match real experience but it gives you real answers to the three questions that otherwise sink the interview.
Is AI going to replace site reliability engineers?
No, and the honest version of the answer is more useful in an interview than either the enthusiastic or the dismissive one. The mechanical parts of operations have been automated continuously for two decades, mostly by SREs themselves, and the consistent effect has been to raise the number of systems one engineer can be responsible for rather than to reduce the need for engineers. Current tooling drafts incident summaries and timelines, correlates alerts and proposes likely causes, which saves real time, but it is good at correlation and poor at distinguishing cause, and it produces confident narratives that are sometimes wrong. What is not automatable is accountability: deciding to roll back with incomplete information, committing capacity, setting a target, and holding command during an incident. Meanwhile the role gained work, because inference services, higher change volume and automation that holds credentials all land on the reliability team.
Put this on a resume in about a minute
Paste your history once and point it at the Site Reliability Engineer posting you are looking at. No account, no card.
Build my resume free More roles