IT, Cloud & Infrastructure

How to get hired as a platform engineer in 2026-27

The short answer

Platform engineer is an unlicensed job with no required degree and no mandatory certification, so hiring turns on the one thing a DevOps loop does not grade: evidence that you built something other engineers chose to use. The title covers four different jobs in 2026 (rebadged DevOps or cloud operations, genuine internal developer platform product work, developer productivity and build engineering, and domain platforms for data, ML or observability), so classify the posting before you write a line of your resume, because one resume cannot pass all four. For the internal developer platform shape, put adoption and time next to every component you built: how many engineers and services used it, how long provisioning or a service's first production deploy took before and after, how many teams you migrated off the old path, whether you then deleted the old path, and the platform's own availability and support load. The loop is usually a recruiter stack-match screen, a hiring manager conversation about scope and internal customers, a real software round (often Go) rather than a scripting exercise, a platform design round marked largely on what you refuse to build and on your escape hatch and migration plan, and a round with an engineer from a team that would have to use what you ship.

Licence required: noneNo licence, no protected title, no legally required degree and no mandatory certification exists for platform engineering. Degree filters still appear in graduate programmes at large employers, in some government and defence contracts, and in skilled-worker visa routes where a degree or an assessed equivalent counts toward eligibility. Everywhere else the gate is demonstrated production experience plus evidence that internal engineers adopted something you built.
Platform hiring vs DevOps hiring, in one lineDevOps hiring asks whether you can run the system. Platform hiring asks whether you can build something other engineering teams choose to use without being ordered to. Both loops test Linux, Kubernetes, infrastructure as code, pipelines and incident behaviour. Only the platform loop grades adoption, interface design, migration of internal consumers, and what you decided not to build.
Four different jobs share this titleOne: rebadged DevOps, SRE or cloud operations, which is the shape you will see most often on job boards and which hires like DevOps. Two: internal developer platform product work (self-service, golden paths, developer portal, modules and control planes treated as products), which has the highest coding bar. Three: developer productivity and build engineering (monorepo, build caching, test selection, CI duration). Four: domain platforms for data, ML, observability or security, where the domain is the gate. Read the responsibilities for the noun that follows the word build.
Certifications that move a screenHands-on and respected: Certified Kubernetes Administrator (CKA), plus CKAD for application-platform work and CKS for security-leaning roles, because all three are sat in a live terminal against running clusters rather than answered as multiple choice. Useful mainly for recruiter and automated filters: a cloud certification matching the employer's cloud, and the HashiCorp Terraform Associate. No certification demonstrates the part of this job that actually gets you hired, which is building an internal product other teams adopt. Generic Agile, Scrum Master and ITIL certificates do not move a platform hiring decision.
Time to each certification, and the real prerequisitesTerraform Associate: a short multiple-choice exam, realistically a few weeks of evenings if you already write Terraform at work. Cloud associate level (AWS Solutions Architect Associate, Microsoft AZ-104, Google Associate Cloud Engineer): a few weeks to a couple of months. CKA: a timed performance-based exam, typically two to three months of consistent practice from a working Kubernetes baseline and considerably longer from nothing, because it is scored on completing tasks against real clusters under time pressure. CKS requires an active CKA before you can sit it. All of these expire and have to be renewed, and the exam format, duration and curriculum get revised, so confirm the current terms when you book rather than trusting a blog post.
What the platform loop adds that a DevOps loop does notA real software round in a typed language, most often Go for Kubernetes-adjacent work, where you write a small service, a CLI, an API with a schema or a reconcile loop with tests rather than a shell script. An internal product conversation: how you learned what developers needed, what you shipped first, what you refused, how you measured it. An interface design conversation: versioning, backward compatibility, deprecating a module that hundreds of repositories depend on, and the escape hatch. And a cross-functional round with a senior engineer from a team that would have to use your platform, who is marking whether you would become their bottleneck.
Numbers to take with you before you leave a jobWrite these down while you still have access: time from an empty repository to a service running in production, before and after; provisioning lead time for an environment, a database or a queue; how many services and teams moved onto the path you built and out of how many; whether the old path was retired; the platform's own availability against a stated SLO; pipeline success rate and p95 pipeline wall time; interrupts or support tickets per week before and after; CI spend and cloud spend. These are the figures you can never reconstruct once your account is disabled.
Pay: name the source, not an averageThe US Bureau of Labor Statistics has no platform engineer occupation code, which is why aggregator figures for this title vary wildly and are not checkable. Triangulate instead from BLS Occupational Employment and Wage Statistics for 15-1252 Software Developers and 15-1244 Network and Computer Systems Administrators, from employers' own posted ranges under state pay-transparency laws (Colorado, California, Washington, New York and Illinois among them), and from levelling data for large technology employers. Pay for this title tracks the software engineering ladder at product-shaped platform teams and the infrastructure ladder at operations-shaped ones, which is a second reason to classify the posting before you negotiate.

Platform engineer is four different jobs, and the posting tells you which

This title is unusually unreliable, and that is the first thing to understand about the job market for it. Many employers renamed an existing DevOps, SRE or cloud infrastructure team "platform" without changing what the team does, because the word attracts candidates and reads as modern. A smaller set built genuine internal product teams whose users are their own engineers. Both post a requisition headed Platform Engineer, often with overlapping tool lists. Your first task is classification, not resume writing, because the resume, the preparation and the questions you ask are different for each.

Shape one is rebadged operations. The posting leads with cloud, Kubernetes, Terraform, CI/CD, monitoring, on-call, incident response and cost. Responsibilities are phrased around maintaining, supporting and ensuring. The work is running shared infrastructure for the company, and it is a good job. This is the shape you will see most often, and it is what "platform" almost always means at a company with a single infrastructure team and no separate internal customer base to serve. Hiring here is DevOps hiring: scale, reliability numbers, troubleshooting, a design round about a system rather than about an interface.

Shape two is internal developer platform product work. The posting uses product vocabulary: self-service, golden path or paved road, internal customers, developer experience, adoption, abstraction, platform as a product. It names a portal (Backstage, or a commercial one such as Port, Cortex, OpsLevel or Atlassian Compass), a composition or orchestration layer (Crossplane, Kratix, Radius, the Score specification, or a commercial platform orchestrator), reusable Terraform or OpenTofu modules treated as versioned products, Argo CD or Flux with templated application definitions, and custom Kubernetes controllers or custom resource definitions. A programming language usually appears as a first-class requirement rather than as scripting. This shape has the highest coding bar and the most interesting design round, and it concentrates at employers large enough that an internal platform pays for itself.

Shape three is developer productivity and build engineering, sometimes posted as Developer Productivity Engineer, Build Engineer or Developer Experience Engineer. The subject matter is the inner loop and the pipeline rather than the runtime: a monorepo, a build system such as Bazel, Nx, Turborepo, Gradle, Pants or Buck2, remote caching and remote execution, test selection and test impact analysis, flaky test detection and quarantine, merge queues, local and ephemeral development environments, and the wall-clock time of the slowest pipeline. The interview is noticeably different: profiling, caching correctness, hermetic builds, dependency graphs and the discipline of measuring before changing. If you like performance work, this is the shape to look for, and it is less crowded.

Shape four is a domain platform. Data platform, ML or AI platform, observability platform, security platform. Here the domain knowledge is the real gate and the platform skills are assumed. A data platform engineer posting is closer to data engineering than to Kubernetes administration, and an ML platform posting expects you to understand training and serving workloads rather than just the accelerators underneath them. Applying to these with a generic infrastructure resume fails at the screen, because the interviewer wants to hear the domain's failure modes.

You can classify a posting in about ninety seconds. Read the responsibilities and find the noun that follows the verb build: if it is a pipeline or a cluster, you are looking at shape one, and if it is an interface, an abstraction, a portal, a module or a self-service workflow, you are looking at shape two. Count how many times an internal audience is named. Check the required skills for a language stated as a language rather than as scripting. Look at whether on-call is described as production infrastructure on-call or as support for a platform other teams depend on. Notice whose satisfaction is named as the measure of success. Then open the employer's other postings on the same careers page, because a company with a real platform organisation posts several of these and names the stream-aligned teams they serve.

Getting this wrong is expensive in both directions. A strong operator who prepares a product narrative for an operations job sounds evasive about the pager. A strong builder who prepares scale numbers for a product-shaped team has nothing to say in the round that decides it. Classify first, then prepare one of two different candidacies.

How platform engineering hiring differs from DevOps hiring

The difference fits in one sentence, and it is worth memorising because interviewers ask you to articulate it. DevOps hiring asks whether you can run the system. Platform engineering hiring asks whether you can build something other engineers choose to use. Both care deeply about production. Only one of them grades adoption, and the parts of a platform loop that surprise strong DevOps candidates all follow from that single difference.

The first consequence is that your evidence changes shape. An operator's proudest resume line is a fleet upgraded across three Kubernetes minor versions with no customer-facing downtime, and that line is excellent. A platform engineer's proudest line is that most of the company's services moved onto the new deployment path within two quarters and the old path was then deleted. The first proves competence under risk. The second proves that engineers with other priorities, who could have ignored you, did not. Keep the reliability evidence, because the platform is production for internal customers and an unreliable platform is worse than no platform, but lead with the adoption evidence when the posting is product-shaped.

The second consequence is that the coding round is real. Platform loops at product-shaped teams include a software interview in a typed language, most often Go because the Kubernetes ecosystem is written in it, sometimes Java or Python to match the internal customers. You may be asked to implement a reconcile loop that converges on a desired state, a CLI with sensible flags and failure messages, a small HTTP API with a schema and tests, or a concurrency exercise with contexts, cancellation and errors. Candidates who are fluent in Bash and Python glue but have never written and tested a program with an interface fail here, and it is the most common single reason a good DevOps engineer does not convert on a platform loop.

The third consequence is an internal product conversation, whether or not it is on the schedule as one. Expect: how did you decide what to build, how did you find out what developers actually needed, what did you ship first and why the smallest thing, what did you refuse to build, how did you know it worked, and how did you get teams to migrate without a mandate. Candidates whose infrastructure career has been ticket-driven often have nothing to say here, because nobody ever asked them to choose. If that is you, the fix is to go and do one small piece of platform work as a product before you interview, which the section on entry routes covers step by step.

The fourth consequence is that interface and compatibility questions appear, and they are library design questions wearing infrastructure clothes. How do you version a Terraform module that two hundred repositories consume. What is a breaking change to a custom resource definition and how do you ship one. How do you deprecate v1 without stopping anyone's release. What leaks through your abstraction, and what happens to the team whose requirement it cannot express. Treat these as API design, because that is what they are.

The fifth consequence is who is in the room. A DevOps loop is usually infrastructure people throughout. A platform loop generally adds a staff or principal engineer from a stream-aligned product team, and sometimes an engineering manager whose team's delivery speed depends on the platform. That person is not assessing your Kubernetes depth. They are assessing whether you would become the bottleneck between their team and production, whether you would listen when their requirement does not fit your model, and whether you would say no clearly rather than disappearing for a quarter. Answer them in their terms: delivery speed, autonomy, and what they would stop having to care about.

What does not change is the part candidates hope will change. Platform engineering is not an escape from operations. You will still be asked to debug a pod stuck in CrashLoopBackOff, explain why a DNS lookup inside a cluster failed, read a Terraform plan and spot the destructive line, describe how you would cut a cloud bill, and tell an incident story where something you owned broke. A candidate who cannot do those does not pass either loop, and a candidate who presents platform engineering as the senior version of DevOps rather than a different problem sounds like they are avoiding the pager.

What gates this job: no licence, a short credential list, and a coding bar people underestimate

Nothing legal gates platform engineering. There is no licence, no protected title, no required degree and no mandatory certification, and plenty of working platform engineers arrived from systems administration, support, backend development, QA, the military or a trade. Degrees still function as filters in graduate programmes at large employers, in parts of government and defence contracting, and in skilled-worker visa routes where a degree or an assessed equivalent counts toward eligibility. Outside those, after two or three years of production experience the credential that replaces any certificate is a system you personally built, operated and can describe in detail: who used it, what it cost, how it failed, and what you changed as a result.

On certifications, be clear-eyed about what each one buys. The Certified Kubernetes Administrator is the most portable because it is performance-based, sat in a live terminal against running clusters, so it cannot be passed by memorising a question bank, and a hiring manager reads it as evidence you have actually typed kubectl under time pressure. CKAD is the better fit if your platform work is application-facing, and CKS if the role leans toward cluster security and admission control, though CKS requires an active CKA first. A cloud certification matching the employer's cloud is mostly a filter pass rather than a signal to the engineers interviewing you. The HashiCorp Terraform Associate is cheap and quick and worth having if infrastructure as code is central to the posting. All of these have renewal windows, so check what you hold is still current before you put it at the top of a resume.

What no certification covers is the thing this job is actually about. There is no exam that demonstrates you designed an internal interface other teams adopted, ran a deprecation without breaking anyone, or decided what not to build. The developer portal ecosystem has no credential that carries hiring weight, and no hiring manager screens on one. That is why a candidate with three months to spend is almost always better off building and documenting one real piece of platform work with a before and after measurement than collecting a second certificate. The CNCF Platforms Working Group has published a platforms white paper and a platform engineering maturity model, and reading them is worth more in an interview than a certification, because they give you the vocabulary hiring managers now use for maturity, funding and measurement.

There are three real gates, and none of them is paperwork. First, cleared work: in US defence and federal contracting a security clearance genuinely gates the job, sponsorship requires an employer, and the timeline is measured in months, so if you have an active clearance, say so near the top of the resume. Second, an implicit experience gate: platform engineer is rarely anyone's first job, because the role assumes you have been an internal customer of someone else's platform and know which parts of it hurt. Third, a headcount gate: a genuine internal developer platform team only exists where there are enough internal engineers to justify it, so if you want shape two specifically, target larger employers and accept that at a small company the same work exists but is called something else and is one third of a broader job.

The coding bar is where good candidates are most often surprised, so be concrete about what to prepare. For Kubernetes-adjacent platform roles, be able to write Go without an assistant: structs and interfaces, errors as values and wrapping, goroutines with context cancellation, a channel used for something that is not a tutorial, table-driven tests, and a program that reads configuration, does work and exits with a useful message. Understand the controller pattern well enough to explain level-triggered reconciliation against desired state, why a controller must be idempotent and tolerate being restarted mid-operation, and what a finalizer is for. Be able to read Python, TypeScript and Java well enough to debug a build for teams who write them, because you cannot build a golden path for a language you cannot read. For portal work, enough TypeScript and React to write a Backstage plugin rather than only to configure one.

How the loop actually runs, and who screens you

A platform engineering process in 2026 typically runs four to six conversations over three to six weeks. The set of things being tested is stable even when the number of calls is not: a recruiter screen, a hiring manager conversation, a technical screen, a practical stage, a platform design round, and a cross-functional round with somebody who would use what you build. Smaller employers compress that into two or three conversations rather than dropping stages, which usually means the design round carries the technical screen inside it. Ask the recruiter for the stage list and whether there is a take-home, because that answer changes how you spend the week.

The recruiter screen is a stack match plus logistics, and it is the cheapest place to find out what the job is. They will check your cloud, your orchestrator, your infrastructure-as-code tool, your language, your on-call expectations and your compensation range. Use your turn to ask whether the team builds the platform or runs the infrastructure, how many engineers are on the team, how many engineers it serves, and whether adoption is mandated or chosen. Those four answers determine everything about how you prepare, and recruiters answer them readily.

The hiring manager conversation is the scope round and it matters more here than in a DevOps loop. They are establishing whether you think in terms of internal customers at all. Expect to be asked what you own today, who depends on it, how you decide what to work on next, how you handle a team that wants something your platform does not do, and what you have deliberately not built. Have two prepared stories: one where you shipped the smallest useful thing and grew it, and one where you said no to a request and what you offered instead. Ask them what the team deleted most recently, because a platform team that has never retired anything is accumulating surface area it cannot support.

The technical screen is conventional and you should not lose it. Linux process, memory and networking fundamentals; why a pod will not start and how you would find out; DNS and service discovery inside a cluster; TLS and certificate rotation; how a rolling update actually proceeds and what a readiness probe does to it; what you look at first when deploys suddenly get slower. Many teams add a code reading exercise here instead of a code writing one, handing you a hundred lines and asking what is wrong with it. Read it out loud and name the assumptions, which is exactly what the job is.

The practical stage takes one of three forms, and the posting or the recruiter will usually tell you which. A take-home: write a reusable Terraform or OpenTofu module with variables, outputs, a pinned provider, remote state guidance, a test, a README with a working example and a stated upgrade path; or build a small Backstage software template or plugin; or write a controller that reconciles a toy custom resource; or write a CLI that provisions something and is pleasant to be wrong with. A live coding round: Go, usually around ninety minutes, with tests expected. A live debugging round: a deliberately broken environment where you are marked on method rather than speed, so narrate your hypotheses, change one thing at a time, and say what evidence would disconfirm you.

The platform design round is the round that decides product-shaped loops, and the next section covers it on its own because it is where most rejections happen. The cross-functional round that follows is often misread as a culture chat. It is not. A staff engineer from a stream-aligned team is deciding whether working with you would be faster or slower than working around you. Ask what their last three pieces of infrastructure friction were, listen properly, and resist the urge to design a solution in the room before you understand their constraints. The behaviour they are watching for is curiosity about their problem rather than enthusiasm about your model.

Two practical notes on the process. First, take-homes in this discipline are frequently over-scoped by the employer and under-timeboxed by the candidate: ask what the expected time is, hold to it, and write in the README what you would do next and what you deliberately left out, because that note is often read more carefully than the code. Second, if a loop contains no round where a potential internal customer talks to you, that is informative. It usually means the platform is funded as infrastructure and measured on uptime, which is a fine job but a different one from the posting's language.

The resume: which internal developer platform work to describe, and how

Platform resumes fail in a specific and predictable way. They contain the sentence "built an internal developer platform" with no indication of what was in it, who used it, or what got faster. A reviewer reads that as "we wrote some Terraform and a Confluence page", because that is what it most often means. The fix is mechanical: every platform claim gets a component, an internal customer count, a before and after, and where possible a thing that was retired as a result.

Use this bullet shape and fill in your own real figures, because the ones here are only a worked example. Component, who used it, what it replaced, the measured change, and the retirement. "Built a self-service environment provisioning workflow (Terraform modules behind a Go CLI and a Backstage template) used by 14 teams and 180 engineers; environment provisioning went from a ticket with a two to five day wait to a self-service action completing in about twelve minutes; retired the manual request queue and the runbook it depended on." Every element of that sentence survives a follow-up question, and each one invites the interviewer to ask about something you know well. Compare it to "responsible for internal developer platform and self-service tooling", which invites nothing.

The time metrics that reviewers actually respond to are narrow, so learn their names. Time from an empty repository to that service serving traffic in production, which some teams call time to first deploy and others call time to hello world. Provisioning lead time for an environment, a database, a queue, a bucket or a DNS record. Onboarding time to a new engineer's first merged change in production. Lead time for changes and deployment frequency from the DORA delivery measures, which your interviewer already thinks in. Pipeline p95 wall-clock time and pipeline success rate. Interrupt load, meaning how many support requests per week the platform team absorbed before and after you made something self-service. If nobody measured, state the shape honestly rather than inventing a figure: "from a ticket and a multi-day wait to the same day, self-service" is credible and survives questioning, and a fabricated percentage does not.

Adoption is the evidence that separates this resume from a DevOps resume, and retirement is the strongest form of it. Say how many services onboarded out of how many exist, how many teams migrated and over what period, how many repositories consume the current major version of your module, and what share of production deployments now go through the paved road. Then say what you deleted: the old pipeline, the hand-maintained cluster, the wiki runbook, the bespoke per-team Helm charts. Deleting the old path proves adoption was real rather than announced, and it is the single line most likely to get you asked to talk for five minutes.

Do not drop the operational evidence, because your platform is production for people who cannot choose another supplier. State the platform's own service level objective and attainment, not just the applications' SLOs: the availability of the control plane or provisioning API, the success rate of the deployment path, the availability of CI, and what happened when the error budget was spent. Include cost in two directions: the cloud and CI spend you removed, and the spend you made visible and attributable per team, because cost attribution is a platform capability that engineering leaders ask for constantly and few candidates mention.

The artefacts that survive screening are the ones a reviewer can picture. A versioned module with a documented interface, tests and an upgrade guide. A Backstage software template that scaffolds a service with CI, ownership metadata, observability wiring and a deployable manifest already correct. A custom resource definition and controller you wrote, with the reconciliation described in one sentence. A service catalog with ownership data accurate enough to page the right team. A production readiness scorecard that teams actually look at. A design document or internal RFC you wrote and got agreement on across teams. A deprecation plan you executed to completion. Each of these is a five-minute interview answer waiting to happen.

Now the parts that get ignored or actively cost you. A sixty-item comma-separated skills list is skimmed and discounted: group it by category and cut anything you would not want to be interviewed on. "Agile", "Scrum" and "stakeholder management" as skills add nothing. Responsibility statements with no result add nothing. A career objective paragraph is pure cost. Certifications belong in a short line near the end. Listing a developer portal you only ran the demo application of is a trap, because the first question is which plugin you wrote or which entity model you designed. And never let the resume and the LinkedIn profile disagree on a figure, which is an easy and common rejection.

The platform design round: what reviewers are actually marking

The prompts repeat across employers. Design the path from an empty repository to production for three hundred engineers. A product team needs a Postgres database in an hour rather than a week, design that. Design multi-tenant Kubernetes for fifty teams. We have eight hundred services and no reliable ownership data, what do you do first. Design the deployment experience so a team can release on their own on a Friday afternoon. Each of these is deliberately under-specified, and the first thing being marked is whether you ask who the users are and what they do today before you start drawing.

Six things are being scored, and in most loops the architecture is not the heaviest of them. First, scoping: can you identify the thinnest viable platform that solves the stated pain, rather than proposing a portal, a control plane, a service mesh and a multi-region active-active topology for a company with forty engineers. Second, the interface: what exactly does a developer type or click, what do they get back, what do they never have to know. Third, the escape hatch. Fourth, migration and deprecation: how the existing eight hundred services get there, in what order, and who pays for the work. Fifth, the operating model: who is on call, how support requests arrive, what the platform's own SLO is. Sixth, cost, both the cloud bill and the team size required to sustain what you just proposed.

The escape hatch deserves its own paragraph because missing it is the most common single fail. An abstraction with no documented way out gets abandoned by the first team with an unusual requirement, and that team builds a shadow platform, and now you support two. Say how the layer underneath stays reachable: the generated manifests are readable and reviewable rather than hidden, the module exposes the raw resource as an output, there is a documented path for a team that needs something the abstraction cannot express, and the platform does not forbid what it cannot yet model. Then say how you harvest those exceptions back into the product, because a well-run platform treats every escape as a feature request with evidence attached. Candidates who describe the platform as mandatory and complete sound like they have never had a real internal customer.

Build versus buy comes up in almost every one of these rounds, usually as a portal question: Backstage, a commercial portal, or nothing. The answer that scores is about team size and sustained maintenance, not product preference. A self-hosted Backstage deployment is a TypeScript application your team now owns, upgrades and extends, which is reasonable with dedicated engineers and unreasonable as a side project for two people. A commercial portal removes that maintenance and adds a procurement and lock-in conversation. And state the thing most candidates miss: the hard part of a developer portal is not the user interface, it is the catalog data, because ownership metadata goes stale the moment it is not generated from something that must be correct for another reason. A portal with wrong ownership data is worse than no portal, because people stop trusting it and you cannot page the right team.

Multi-tenancy comes up almost as often, so have a position. Namespace per team in a shared cluster is cheap and leaks: noisy neighbours, shared ingress blast radius, cluster-scoped resources like custom resource definitions and webhooks that one team can break for everyone, and a shared upgrade that must suit all tenants. Cluster per team is isolated and expensive in both money and upgrade toil, and it pushes you toward fleet management. Virtual clusters sit in between and buy control plane isolation without full duplication. Whatever you pick, be able to state the defaults you would set on day one: resource requests and limits, quotas, network policy denying cross-namespace traffic by default, pod security standards, who may create a cluster-scoped object, and how you upgrade without asking fifty teams for permission.

A workable structure for forty-five minutes. Spend the first five on clarifying questions: how many engineers, how many services, what languages, what they do today and how long it takes, what hurts most, who funds this. Spend five stating the one or two outcomes you will optimise and how you would measure them, which immediately separates you from candidates who optimise nothing in particular. Spend twenty on the interface and the end-to-end flow, drawing what the developer does and what happens behind it. Spend ten on migration, support model and cost. Keep the last five for what you would not build, what you would kill, and what you would measure after three months to decide whether this was worth it.

The questions that decide it, and how to answer them

"What is the difference between platform engineering and DevOps?" This is a screening question for whether you understand the job you applied to. The answer that works is short and then specific: DevOps is a set of practices for getting software to production safely and the title usually means operating shared infrastructure, while platform engineering means building and running an internal product whose users are your own engineers, so it is measured on adoption and on developer time saved as well as on uptime. Then add the honest caveat, which marks you as experienced: at many companies the two titles are the same job, so you would want to know which this team is. Avoid the answer that platform engineering is the evolved or senior form of DevOps, which sounds like a status claim rather than a distinction.

"Your platform has low adoption after a year. What do you do?" They are testing whether your instinct is to mandate or to investigate. Start with measurement: who is not using it and what they use instead, because the alternative teams chose tells you what your product is missing. Then go and sit with two of those teams and watch them deploy, which finds things no survey does. Then pick the single biggest blocker and fix it rather than adding features. Then reduce the switching cost, usually by doing the first migration for a team yourself. Mandates are the last resort and only work with executive sponsorship and a deadline for deleting the old path, and even then they produce compliance rather than advocacy. Have a real story here if you possibly can, including one where the answer was that you had built the wrong thing.

"A team needs something your abstraction cannot express. What do you do?" The answer is a sequence, not a principle. Unblock them now through the escape hatch, with a review rather than a refusal. Write down the requirement as evidence. If a second team asks for the same thing, it belongs in the product. If no second team asks, leave it as an exception and do not grow the interface for one consumer. Say the quiet part: every option you add to a platform interface is permanent, because somebody will depend on it, so the bar for adding one is higher than the bar for allowing an exception.

"How would you deprecate v1 of a module two hundred repositories depend on?" This is the interface design question and a weak answer costs the round. Ship v2 alongside v1 rather than in place of it. Announce with a timeline that is realistically a quarter or more, not a sprint. Make the migration mechanical: a documented upgrade path, an automated change where possible, and a pull request raised by you against the busiest consumers. Add deprecation warnings into the output engineers already read, which for Terraform means plan output rather than a wiki page. Measure remaining v1 consumers weekly and publish the number. Stop new adoption of v1 first, which is cheap and immediate. Keep a hard removal date and move it only once, publicly. Do not break anyone's release to make a point about hygiene.

"Who gets paged when a service built on your platform breaks?" The expected answer is the owning team for their own service and the platform team for the platform, with the boundary written down in advance and a defined path for the cases where nobody can tell which it is. Then add the thing that distinguishes a candidate who has lived it: the platform team's real failure mode is absorbing other teams' debugging because the platform makes failures hard to attribute, so you invest in error messages, logs and dashboards that tell a developer it was their code rather than your platform. That investment is the difference between a sustainable platform team and one that becomes a help desk.

"How do you know whether the platform is working?" Give three families and one caution. Delivery outcomes: deployment frequency, lead time for changes, change failure rate, time to restore service. Platform-specific measures: adoption and the share of production deployments through the paved road, time from repository to production, provisioning lead time, self-service ratio, and the platform's own availability against its SLO. Experience measures: a short periodic developer survey plus interrupt and support volume, which the DevEx framework organises as feedback loops, cognitive load and flow state. The caution that earns credit: none of these should become an individual performance metric, because the moment deployment frequency is used to rank engineers it stops measuring anything.

Expect a review exercise too, because reviewing is now the scarce skill. You will be handed a Terraform plan or a pull request and asked what concerns you. Name the dangerous lines out loud: destroy and replace on anything stateful, "forces replacement" on a database or a volume, a count or for_each change that re-indexes existing resources, a security group rule opened to the whole internet, an IAM policy with a wildcard action, a removed lifecycle protection, a resource limit deleted to stop an out-of-memory kill, a pinned dependency silently floated to a tag. Then say what guardrail you would add so that class of mistake becomes impossible rather than merely noticed, which is the platform engineer's answer rather than the reviewer's.

Getting in without a platform title yet, and reading the offer

Platform engineer is rarely a first job, and the reason is structural rather than gatekeeping: the role assumes you have been somebody else's internal customer and know which parts of that experience hurt. The realistic routes in are four. From DevOps, SRE or cloud operations, which is the shortest hop and where the gap to close is the coding bar and product thinking rather than infrastructure. From backend development, which is the strongest route into product-shaped platform teams because the coding bar is already met and the gap is production operations, and the move is often immediate. From QA, test infrastructure or release engineering into developer productivity and build engineering, which is a natural and under-contested path. From data engineering into data platform work, where your domain knowledge is the scarce part.

The highest-yield move is internal, at an employer that already has a platform team. You already know the friction, you already have credibility with the teams the platform serves, and platform teams preferentially hire people who have felt the pain they are solving. The practical approach is to do platform work before you have the title: find the piece of friction your own team complains about most, and fix it as a product rather than as a favour.

That last sentence is the whole strategy, so here is the concrete version. Pick one friction, for example that creating a new service takes two weeks of copying from another repository and asking three people for help. Interview three colleagues about what they actually do, and write down the steps and the waiting. Measure the current time properly. Ship the smallest thing that removes most of the waiting, which is usually a template plus a module plus a page of documentation, not a portal. Measure again. Get one other team to adopt it and watch them use it without helping, which is where you learn what is wrong. Write it up in one page with the before and after. You now have the exact artefact the design round and the product conversation are asking for, and you can say the words adoption, escape hatch and deprecation with a real story behind each. This is worth considerably more than a second certification, and it is usually two to six weeks of work.

Public evidence helps if it is small and finished. A Terraform or OpenTofu module published with a versioned interface, tests, a README and an upgrade note. A Backstage plugin or software template. A Crossplane composition or a Kratix promise that provisions something real. A controller built with kubebuilder that reconciles one resource properly, including idempotency and status conditions. A contribution to an infrastructure project, which carries extra weight because it proves you can work inside someone else's code review standards. One finished small thing beats three abandoned ambitious ones, and a README that explains the design decision is read more often than the code.

On the offer, ask the questions that reveal what the job actually is, because this title hides more variation than most. Is adoption of the platform mandated or chosen, and what happens when a team refuses. How many engineers are on the team and how many engineers does it serve. Who is on call for what, is there a support rota or do interrupts land on whoever is free, and how many interrupts does a typical week contain. Does the team control its own roadmap or is the roadmap a ticket queue. What has the team deleted in the last year. How is the platform funded, and whose budget. What fraction of the team's time goes to running existing infrastructure versus building. A clear answer to that last question is the single best predictor of whether this is shape one or shape two.

The warning signs are consistent. A platform team of three serving four hundred engineers with a mandate and no support rota is a burnout posting whatever the compensation. A team whose roadmap is entirely a ticket queue is an internal help desk with a better title. A platform with no evidence of internal customer research will have adoption problems you will be blamed for. A posting that lists every technology in the cloud native landscape is either copied from somewhere or describing three jobs. And a team that has never retired anything is carrying surface area it cannot support, which becomes your on-call.

One narrow point about levelling and pay. Because product-shaped platform teams sit on the software engineering ladder at many employers and operations-shaped teams sit on an infrastructure ladder, the same person with the same skills can be banded differently by two companies using the same job title. Ask which ladder and which level the role maps to before you discuss numbers, and verify the posted range where pay-transparency law requires one. That question is routine and nobody will think less of you for asking it.

Working with AI in this role

What a platform engineer has to know about AI in 2026-27

Start with an honest calibration, because for this role the hype and the observable change point in nearly opposite directions. Nothing has automated the core of platform engineering, and the effect of cheap code generation on this job is the reverse of what you would expect: when producing configuration and application code becomes nearly free, the volume of change rises, the variance in quality rises with it, and the value of a guarded default path rises fastest of all. The judgement this job is paid for (deciding what to abstract, what to expose, how to move hundreds of consumers off an old interface, how to earn adoption without authority) is architectural and social work that no assistant does for you. Autonomous remediation in production, sold under the AIOps label for years, is still rare in practice and is not what employers are buying. What has genuinely changed is specific, and it falls into four areas: a new class of platform user, a new family of platform capabilities, a machine audience for your golden paths, and a capacity and cost problem in continuous integration.

The new user is the agent, and it is the newest requirement showing up in platform postings and interviews. Coding assistants and agents now open pull requests, run pipelines, call internal APIs and want environments and credentials. Every question that follows lands on the platform team, because the platform is where the control points already are. What identity does an agent have, and is it distinguishable from the human who started it. What can that identity do, for how long, and how is it scoped. Which actions require a human approval before they take effect, and is that enforced by policy rather than by convention. Where does an agent get a sandbox that is not production, and how quickly can one be created and destroyed. What does the audit trail show afterwards when somebody asks who changed this. And how do you stop a convenience credential, created to make an agent work on a Friday, from becoming the widest permission in the account. A candidate who can discuss short-lived scoped credentials, workload identity, policy gates on destructive actions and attributable audit logs is answering a question the hiring manager is currently living with.

Golden paths now have a machine audience as well as a human one, and that changes what a paved road should look like. A generator reproduces the most common pattern it can see, so repository conventions, scaffolded templates, machine-readable schemas and a policy layer that rejects the known-bad shape do more work than documentation ever did. Three practical consequences. The paved road should be the easiest thing to generate correctly, which means the template and the module interface matter more than the wiki. Internal documentation needs to be retrievable and unambiguous rather than beautiful, because it is now read by retrieval as often as by people. And several teams now expose internal context to assistants through a tool interface so an assistant can query the service catalog, the deployment state or the runbook rather than inventing an answer. Building that interface is platform work, and saying so in an interview lands well because it reframes an AI question as an abstraction question.

The second new capability family is AI workloads offered as a platform service. If your employer ships anything with a model in it, this work has to land on somebody, and it lands on the platform team. Somebody has to make accelerators available without a ticket, with quotas and cost attribution per team, because an idle GPU costs money every minute and scarce regional capacity has to be scheduled rather than hoped for. Somebody has to put a gateway in front of model providers with per-team keys, rate limits, spend caps, caching and a story for what happens when a provider has an outage. Somebody has to version model artefacts and prompts like any other release, wire evaluation into the pipeline so a model change is gated the way a code change is, and offer a retrieval or vector store as a managed internal service rather than letting six teams each run their own. Postings now say this explicitly, and most candidates have no exposure to it, which is exactly why a small amount of real experience differentiates.

The fourth change is the least glamorous and the most interview-relevant: continuous integration got more expensive. More generated code means more pull requests, more test runs, longer queues, more flakes surfacing and more review load, and the bill arrives at the platform team. Build caching, remote execution, test selection, merge queues, flake quarantine and per-team CI cost attribution stopped being hobby projects and became roadmap items somebody is measured on. A candidate who can say what happened to their pipeline volume over the last year, what it cost, and what they did about it is answering a problem the interviewer has on their own board this quarter. The companion shift is in review: writing changes is cheap now, approving them safely is the bottleneck, and infrastructure changes have a far worse blast radius than application code because one approved plan can delete a database or widen a permission across an account. That is why the pull request review exercise has become a standard round.

Giving coding agents an identity, a sandbox and a limit

This is the newest thing on platform postings and few candidates have prepared for it. Agents now act inside systems that were designed on the assumption that a human with a named account made every change, which breaks attribution, breaks least privilege, and breaks the review model. The platform team owns the control points where this is solved: identity, credential lifetime, policy admission, environment creation and the audit trail. Done badly it produces the widest permission in the account and no way to say afterwards what changed production.

Show it: Be able to describe a concrete design: a distinct machine identity per agent or per workflow rather than a shared service account, credentials minted short-lived from OIDC federation rather than stored, scopes narrowed to the specific repositories and resources needed, a policy layer that refuses destructive actions outright and requires a human approval for a listed set, ephemeral environments an agent can create and destroy without touching production, and logs that distinguish agent actions from human ones. Have one real story, even a small one: the static token you replaced with federation, the deny rule you added, the sandbox you built. On the resume, write the guardrail and who it protected, not the tool name.

Reviewing generated infrastructure changes for the defects they reliably produce

Producing Terraform, manifests and pipeline configuration is now nearly free, so the bottleneck moved to approving it safely. The characteristic failure is not generated code that fails; it is generated code that works and is quietly destructive, because the generator does not have your account state, your traffic shape or your invariants. One approved plan can replace a database rather than update it, open a port to the internet to fix a connectivity problem, widen an IAM action to a wildcard because the narrow one failed, or remove a resource limit to stop an out-of-memory kill. Platform teams now screen for the engineer who catches these before they merge.

Show it: Practise reading a Terraform plan out loud and naming the dangerous lines: destroy and replace on stateful resources, "forces replacement", a for_each or count change that re-indexes resources, a rule opened to 0.0.0.0/0, a wildcard action, a removed lifecycle protection, a floated dependency. Bring one story of a generated or copied change you blocked and the specific outage it would have caused. Then go one level up, which is what distinguishes a platform answer: say which guardrail you added so the whole class became impossible, such as policy as code in the pipeline, a required plan review, a deny rule on destructive actions in production, or an admission policy that rejects the shape entirely.

Designing golden paths and templates that a generator follows correctly

A paved road used to compete with copy and paste from an old repository. It now competes with generation, which reproduces whatever pattern is most visible in your codebase. That makes conventions, scaffolds and machine-checkable schemas a stronger lever than documentation, and it means the consequences of leaving a bad pattern lying around in a popular repository are larger than they used to be. This is a reframe interviewers respond to, because it turns an AI question into an abstraction and defaults question, which is the job.

Show it: Describe a template that scaffolds a service already correct: CI, ownership metadata, observability wiring, secrets handling, a starter SLO and a deployable manifest, with nothing left for a developer to guess at. Say how you made the wrong shape fail fast, with a lint or policy check in the pipeline rather than a review comment. Say how you removed the stale pattern from the repositories people copy from, because leaving it there is now an active cost. If you have exposed internal context to assistants through a tool interface so they can read the catalog or the deployment state, describe it, including what you deliberately did not expose.

Offering AI workloads as a self-service platform capability

This is where a lot of new platform work is being funded, and most candidates have no exposure, so even modest real experience differentiates. The problems are concrete and financial: accelerators are scarce and expensive so scheduling them well is a direct cost outcome, model artefacts are large so image size and registry pull behaviour become failure modes, inference is bursty and latency-sensitive in ways that defeat naive autoscaling, cold starts can run into minutes, and spend without attribution becomes an unanswerable question from finance.

Show it: Be able to discuss how accelerators are exposed to pods and scheduled, when to share a device across workloads versus dedicate it, why model weights usually belong in object storage with a warm cache rather than baked into an image, how you scale a service that cannot start in under a minute, and how you attribute spend per team and per model. For the gateway layer, describe per-team keys, rate limits, spend caps, caching and provider failover. If you have none of this at work, run a small quantised model behind a serving stack yourself, record the cold start and the hourly cost, and write up what you would change at a hundred times the traffic.

Keeping continuous integration and test infrastructure solvent as change volume rises

Pull request and test volume rose when generation made writing cheap, and the platform team pays for it in queue time, flakes and cloud bills. Slow or unreliable CI is also one of the most common complaints in internal developer experience surveys, which means it directly affects the measures your platform team is judged on. Unlike most AI-adjacent work, this is measurable in minutes and currency, which makes it ideal interview material.

Show it: Carry numbers: p95 pipeline wall-clock time before and after, pipeline success rate, queue wait at peak, CI spend per month, and the flake rate you measured rather than estimated. Name the mechanism you used: remote build caching, remote execution, test selection driven by the dependency graph, parallelism with a sane shard strategy, a merge queue, quarantining flaky tests with an owner and an expiry rather than a permanent skip, and cost attribution per team so the conversation about the bill has an owner. State the tradeoff you accepted, because test selection trades a small risk of a missed regression for a large amount of time and interviewers want to hear you say that out loud.

Rolling out AI tooling internally and measuring whether it actually helped

Platform and developer experience teams are frequently the ones handed the internal rollout of assistants and agents: procurement questions, which repositories are in scope, what data leaves the building, licence allocation, and then the inevitable question from leadership about whether it was worth the money. Answering that honestly, with a method rather than a vendor slide, is a differentiating skill right now, because plenty of organisations have spent the money without capturing a baseline they could compare against.

Show it: Describe a measurement approach you would actually defend: the DORA delivery measures plus experience measures from a short periodic survey, compared against a baseline you captured before the rollout, with the caveats stated plainly (confounders, self-selection among early adopters, and the fact that more code merged is not the same as more value delivered). Say explicitly that none of it should be used to rank individual engineers, which is both correct and a signal you have thought about it. If you ran a rollout, name the policy you set about what context may be sent to a third party, and what you did about review load once pull request volume rose.

What a screen is looking for

These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.

Mistakes that cost people this job

Applying to every posting with this title using one resume, without working out which of the four jobs it is.

Classify the posting in ninety seconds: find the noun after the verb build, count how many times an internal audience is named, and check whether a programming language appears as a requirement or only as scripting. Then send either an operations candidacy (scale, reliability, troubleshooting) or a builder candidacy (adoption, interface design, migrations). Ask the recruiter directly whether the team builds the platform or runs the infrastructure.

Writing "built an internal developer platform" with no component, no user count and no before-and-after time.

Name the component, the internal customers, the measured change and what you retired. The shape, with your own figures: "Self-service environment provisioning (Terraform modules behind a Go CLI and a Backstage template) used by 14 teams and 180 engineers; provisioning went from a two to five day ticket to about twelve minutes self-service; retired the manual request queue." If nobody measured, describe the shape honestly rather than inventing a percentage.

Preparing a DevOps candidacy for a platform loop and then being unable to answer how you decided what to build.

Prepare two stories before the loop: one where you shipped the smallest useful thing and grew it from evidence, and one where you refused a request and said what you offered instead. If you genuinely have neither, do one small piece of platform work as a product first. It takes two to six weeks and it supplies the whole product conversation.

Underestimating the coding round because the job is infrastructure.

For a product-shaped platform role, prepare a real software interview. Write Go unaided with tests, contexts and error wrapping; be able to explain level-triggered reconciliation and why a controller must be idempotent; be able to read the languages your internal customers write. This is the most common reason a strong DevOps engineer does not convert on a platform loop.

Designing an abstraction in the design round with no escape hatch.

State the escape hatch explicitly every time: the generated output stays readable and reviewable, the module exposes the raw resource, and there is a documented path for a team whose requirement the abstraction cannot express. Then say how exceptions feed back into the roadmap. Omitting this is the single most common fail in the platform design round.

Reaching for a mandate when asked what to do about low adoption.

Answer in this order: measure who is not using it and what they chose instead, go and watch two of those teams deploy, fix the single biggest blocker before adding any feature, and lower the switching cost by doing the first migration yourself. Mandates come last, need executive sponsorship and a deletion date, and produce compliance rather than advocacy.

Over-building in the design round: a portal, a service mesh, a control plane and a bespoke framework for a company with forty engineers.

Scope to the thinnest viable platform that removes the stated pain, usually a template plus a module plus sensible defaults plus one page of documentation. Say what you would not build and why. Judgement about scope is the scarcest quality in this round, and over-building reads as a risk to the hiring team's roadmap.

Claiming a developer portal on the resume when you deployed the demo application and configured a catalog file.

Be precise about your layer and own it. "Ran a self-hosted Backstage deployment and wrote two plugins and four software templates; the entity model and the ownership data pipeline were mine" is credible, and so is "we use a commercial portal and I own the catalog data quality". The first follow-up question is always which plugin you wrote or where the ownership data comes from.

Treating the cross-functional round with a product engineer as a culture chat.

Treat it as the round that often decides the loop. Ask what their last three pieces of infrastructure friction were, listen without designing a solution in the room, and answer in their terms: delivery speed, autonomy, and what they would stop having to care about. They are deciding whether you would be a bottleneck.

Having no story about something you built that nobody used.

Prepare one honestly, with the reason and the correction. "Three of eleven teams adopted it; the blocker was that it could not express batch workloads, so the next version exposed the underlying job spec and adoption reached nine" is a stronger answer than an unblemished record, which interviewers read as either inexperience or evasion.

Ignoring cost entirely, in the resume and in every design answer.

Carry one cost story with a mechanism and a real before and after, and name the sustaining cost of anything you design: how many engineers it needs to run, what the support rota looks like, and what you would drop if the team lost a person. Cloud and CI spend are board-level lines, and cost attribution per team is a platform capability leaders ask for and few candidates mention.

Accepting the offer without establishing the support model and the ladder.

Get in writing: team size against the number of engineers served, whether adoption is mandated or chosen, who is paged for what, whether there is a support rota or interrupts land on whoever is free, how many interrupts a typical week has, and what fraction of the team's time goes to running existing infrastructure versus building. Also ask which ladder and level the role maps to, because product-shaped and operations-shaped platform teams are often banded differently for the same title.

Questions people ask

What is the difference between a platform engineer and a DevOps engineer, in hiring?

A platform engineer is hired to build an internal product whose users are the company's own engineers, so the loop grades adoption, interface design and migration of internal consumers on top of the infrastructure skills a DevOps loop tests. A DevOps engineer is hired primarily to run shared infrastructure safely, so that loop grades scale, reliability, troubleshooting and incident behaviour. In practice the platform loop adds three things: a real software round in a typed language rather than a scripting exercise, a conversation about how you decided what to build and what you refused, and a round with a senior engineer from a team that would have to use your platform. The honest caveat worth saying out loud in the interview is that at a company with a single infrastructure team the two titles usually describe the same job with different branding.

Do I need a degree or a certification to become a platform engineer?

No. Platform engineering has no licence, no protected title, no legally required degree and no mandatory certification, and many working platform engineers arrived from systems administration, support, backend development, QA, the military or a trade. Degrees still act as filters in graduate programmes at large employers, in parts of government and defence contracting, and in skilled-worker visa routes where a degree or an assessed equivalent counts toward eligibility. After two or three years of production experience what replaces a credential for a platform engineer is a system you built and operated that other teams adopted, described with its user count, its cost, its failure modes and the thing it allowed you to delete.

Do platform engineers write real code, and in which language?

Yes. A platform engineer on a product-shaped team is expected to write and test software in a typed language, most often Go because the Kubernetes ecosystem is written in it, and sometimes Java, Python or TypeScript to match the internal customers or a Backstage deployment. Typical exercises are a reconcile loop that converges on a desired state, a CLI with good failure messages, a small HTTP API with a schema and tests, or a concurrency problem using contexts and cancellation. Underestimating this round is the most common reason a strong DevOps candidate fails a platform loop. On operations-shaped platform teams strong Bash and Python glue is still accepted, which is another reason to classify the posting before preparing.

What internal developer platform work should I put on a platform engineer resume?

A platform engineer should describe components with users and measured time, not responsibilities. The lines that work name the component, the internal customers, the before and after, and what was retired: self-service environment or database provisioning and how long it now takes, a service scaffolding template and the time from empty repository to production, a versioned Terraform or OpenTofu module and how many repositories consume it, a custom resource and controller and what it reconciles, a developer portal with catalog and ownership data accurate enough to page the right team, a scorecard teams actually look at, and a deprecation you ran to completion. Add the platform's own availability against its SLO, pipeline success rate and p95 duration, interrupt volume before and after, and cost removed or made attributable per team. Deleting the old path is the strongest single line on this resume because it proves adoption was real.

What does a platform engineer interview consist of in 2026?

A platform engineer loop usually runs four to six conversations over three to six weeks: a recruiter stack-match screen, a hiring manager conversation about scope and internal customers, a technical screen on Linux, Kubernetes and troubleshooting that often includes reading code rather than writing it, a practical stage, a platform design round, and a cross-functional round with a senior engineer from a team that would use the platform. The practical stage is usually one of three things: a take-home (a reusable module with tests and a README, a Backstage template or plugin, a small controller, or a CLI), a live coding round in Go of around ninety minutes, or live debugging of a deliberately broken environment where method is marked above speed. The design round decides most product-shaped loops, and it is scored heavily on scoping, the escape hatch, the migration plan and the support model rather than on architecture alone. Ask the recruiter for the stage list and whether there is a take-home, because employers vary and the answer changes how you prepare.

Can I move into platform engineering from DevOps, SRE or backend development, and how long does it take?

Yes, and those are the normal routes, because platform engineer is rarely anyone's first job. From DevOps, SRE or cloud operations the infrastructure is already there and the gap is the coding bar and product thinking, which typically takes a few months to a year of deliberate work and often no change of employer. From backend development into a product-shaped platform team the move can be immediate, because the coding bar is met and the gap is production operations and carrying a pager. From QA, test infrastructure or release engineering the shortest path to a platform engineer role is developer productivity and build work, which is measurable and under-contested. The highest yield move is internal, because platform teams preferentially hire people who have been their own internal customers and know which parts hurt.

Is Backstage worth learning for platform engineering jobs?

Backstage is worth learning well enough to write a plugin and a software template if you are targeting platform engineer roles, and not worth listing after only deploying the demo application. Backstage is an open source developer portal created at Spotify and donated to the CNCF, and it appears in a large share of internal developer platform postings alongside commercial alternatives such as Port, Cortex, OpsLevel and Atlassian Compass. The useful thing to understand, and the point that scores in an interview, is that the hard part of any portal is the catalog data rather than the user interface: ownership metadata goes stale unless it is generated from something that has to be correct for another reason, and a portal with wrong ownership data is worse than no portal because nobody can page the right team. Self-hosting it also means owning and upgrading a TypeScript application, which is reasonable with dedicated engineers and unreasonable as a side project for two people.

Do platform engineers carry a pager and go on call?

Yes, most platform engineers are on call, because a platform engineer's output is production for internal customers who cannot choose another supplier. The defensible boundary, and the one interviewers expect you to state, is that the owning team is paged for their own service and the platform team is paged for the platform itself: the provisioning or control plane API, the deployment path, CI, shared clusters and shared infrastructure. The real failure mode is a platform team that absorbs other teams' debugging because failures are hard to attribute, which turns the team into a help desk, so ask in the interview whether there is a support rota, how many interrupts a typical week contains, and whether the team invests in error messages and dashboards that tell a developer the fault was theirs.

Is AI replacing platform engineers?

No, and for a platform engineer the effect has been closer to the opposite. Cheap code and configuration generation raises the volume and variance of change, which raises the value of a guarded default path, policy enforcement and reviewable abstractions, all of which are this role's output. Autonomous remediation in production, promised for years under the AIOps label, is still rare and is not what employers are buying. What genuinely changed is four specific things, and they are now interview material for a platform engineer: agents have become a new class of platform user needing identity, scoped short-lived credentials, policy gates and sandboxes; golden paths now have a machine audience, so templates and machine-checkable schemas matter more than documentation; AI workloads are a new self-service capability to offer, meaning accelerator quota, a model gateway with cost attribution and evaluation in the pipeline; and continuous integration got more expensive as pull request volume rose.

How much does a platform engineer earn?

There is no authoritative single figure for a platform engineer, because the US Bureau of Labor Statistics has no occupation code for the title, which is why aggregator numbers for it vary widely and are not checkable. Triangulate instead from BLS Occupational Employment and Wage Statistics for 15-1252 Software Developers and 15-1244 Network and Computer Systems Administrators, from employers' own posted ranges under state pay-transparency laws (Colorado, California, Washington, New York and Illinois among them), and from levelling data for large technology employers. The structural point matters more than any band: product-shaped platform teams usually sit on the software engineering ladder while operations-shaped ones sit on an infrastructure ladder, so the same platform engineer can be banded differently by two employers using the same job title. Ask which ladder and level the role maps to before discussing numbers.

Put this on a resume in about a minute

Paste your history once and point it at the Platform Engineer posting you are looking at. No account, no card.

Build my resume free More roles