| What the role is | Building, migrating and running a company's infrastructure on AWS, Microsoft Azure or Google Cloud: networking and DNS, identity and permissions, compute and containers, storage and backup, infrastructure as code, monitoring, patching and cost. In a healthy team the output is code and pipelines that other engineers consume rather than consoles you click, and the job sits between the application teams who want capacity and the finance and security people who want it accounted for. The title is used for at least four different jobs, so read the posting for the verb: build a foundation, migrate off a data centre, run an existing estate, or enable developers. |
|---|---|
| Licence or credential gate: none | There is no licence, no protected title, no legally required degree and no mandatory certification for commercial cloud engineering. The real gates are a recruiter's keyword filter, which a single associate-level certification clears, and a hiring manager who wants evidence you have had write access to a cloud estate that mattered. Two genuine exceptions exist. US defence and federal contract work can make a named certification a contractual condition through the Department of Defense cyber workforce qualification rules, so read the requirement written into that specific contract rather than a summary of it. And partner-channel employers often need certified headcount to hold their partner tier, which is why a managed service provider will pay for your exam. |
| The first certification, and how long it takes | One associate-level certification in your primary cloud. AWS Certified Solutions Architect Associate is the most widely recognised; AWS Certified SysOps Administrator Associate fits the operations shape better. On Azure it is AZ-104 for Microsoft Certified Azure Administrator Associate. On Google Cloud it is Associate Cloud Engineer. Check the exam code, the blueprint and whether the exam still exists on the provider's own certification page before you buy a voucher, because all three vendors revise exams and have retired certifications outright. For someone already working in IT infrastructure, expect six to ten weeks of evening study plus hands-on time in a real account. For a genuine beginner, expect three to five months, because the exams assume you can already read a route table and a shell. |
| The second credential that actually changes a screen | HashiCorp Certified: Terraform Associate. It is a short multiple-choice exam, realistically two to four weeks of evenings if you have written real Terraform, and it signals the thing most applicants holding three cloud certifications cannot do. If the postings you want say OpenTofu rather than Terraform, the certification still reads as the same signal: the language and the workflow are the part being tested. Add Certified Kubernetes Administrator (CKA) if the postings mention Kubernetes, EKS, AKS or GKE: it is performance-based, you work in a live cluster against a clock, and it takes most people two to three months of practice. AWS Certified Advanced Networking Specialty is the strongest differentiator if you came from network engineering, because the skill is genuinely scarce. |
| Certification validity, which candidates get wrong | AWS certifications run three years. Microsoft role-based certifications renew annually through a free online assessment, which people forget and then let lapse. Linux Foundation certifications including CKA run two years, as does HashiCorp Terraform Associate. Google Cloud validity differs by level and Google has changed it before, so read the current certification page rather than a forum post. Put the expiry in your calendar on the day you pass. An expired certification on a resume is worse than no certification, because a recruiter who checks the badge finds it dead. |
| The single highest-signal artefact a junior can show | A CI pipeline that deploys infrastructure using short-lived credentials from OIDC federation, with the trust relationship scoped to one repository and branch, and no long-lived access key stored anywhere. Almost no portfolio has it, it takes an afternoon once you understand it, and it proves you think about credentials rather than about services. The companion artefact is a Terraform repository with remote state and locking, one module with a sane variable interface, separate environments, a plan on pull request and apply on merge, and a destroy you have actually run. |
| Experience actually hired, and the feeder roles | Most postings titled cloud engineer ask for two to five years of infrastructure experience with at least one or two in cloud. Junior and associate versions exist and are concentrated at managed service providers, cloud partners, large enterprises running an operations tier, universities and government contractors. The feeder roles that convert are systems administrator, network engineer (the strongest hand, because cloud networking is where estates break), help desk or desktop support with Linux added, NOC or monitoring analyst, data centre technician, database administrator and backend developer. Moving across inside one company is faster than moving between companies, and a year to two years of doing cloud work at the edges of your current job is a normal run-up. |
| Pay: name the source, not an average | The US Bureau of Labor Statistics has no cloud engineer occupation code, which is why aggregator numbers for this title vary wildly. The nearest OES codes are 15-1244 Network and Computer Systems Administrators, 15-1241 Computer Network Architects and 15-1299 Computer Occupations All Other, and the OES tables give you a defensible local floor by metro area. For an actual range, read live postings from employers covered by US pay-transparency law, which requires a range in the posting, and compare the same title at a product company, an enterprise and a managed service provider, because the spread between those three employer types is larger than the spread between cities. |
Cloud engineer is four different jobs, and the posting tells you which
One resume cannot pass all four versions of this title, so classify the posting before you write a word. Read it for the verb, not the service list. Every posting is mostly asking you to build a foundation, migrate off something, run an existing estate, or enable other engineers, and the interview is built around whichever one it is.
Job one is foundation work. You build the landing zone: the account or subscription factory, the organisational hierarchy, the network topology, identity federation to the corporate directory, the Terraform modules every workload inherits, the baseline logging and the guardrails. The work is greenfield or a re-platform, the users are other engineering teams, and the interview goes deep on networking, identity and infrastructure as code. Signals in the posting: AWS Organizations, Control Tower, management groups, Azure landing zone, Terraform modules, a mention of multi-account or multi-subscription strategy.
Job two is migration. You are moving an estate out of a data centre or off a hypervisor, application by application. The work is assessment, dependency mapping, choosing per workload whether to retire, retain, rehost, replatform, refactor, repurchase or relocate, then writing the cutover runbook and the rollback, then doing the cutover at 2am on a Saturday with the application owner on the call. Hypervisor licensing changes have pushed many organisations into re-evaluating their VMware estates, so VMware exit work is a real and well-funded version of this job right now. The interview tests sequencing and nerve, not breadth: what you moved, in what order, what the downtime was, and what you rolled back.
Job three is running the estate, usually called cloud operations even when the title says cloud engineer. On-call, patching, capacity, backup and restore, incident response, access requests, cost, and a ticket queue. This is where most of the hiring volume is and where most people coming from systems administration land first. It is a real engineering job if the team has invested in automation and a frustrating one if it has not, and the difference is visible in the interview: ask who writes the Terraform, and ask whether anyone has deleted a runbook step this quarter.
Job four overlaps DevOps and platform engineering: cloud-native application infrastructure. Kubernetes, pipelines, service deployment, developer self-service, sometimes a service catalogue. If the posting leads with developer experience, internal platform, golden path or self-service, you are reading a platform engineering posting with a cloud engineer title on it, and the loop will include a real coding round rather than a scripting exercise.
Four fast signals separate them. If the first three requirements name Terraform, networking and identity, it is foundation work. If they name an inventory size (400 virtual machines, 90 applications, two data centres), it is migration. If they name a ticketing system, an on-call rota or a service level agreement, it is operations. If they name Kubernetes and a programming language before they name a cloud, it is application infrastructure.
Two questions on the hiring manager call place the role precisely and make you sound like someone who has done it. Ask who merges the infrastructure code that creates a new environment, and ask what the last change was that someone made without filing a ticket. The first answer tells you whether you will have authority or a queue. The second tells you how much the team trusts its own automation.
- Foundation shape: account or subscription factory, organisational hierarchy and policy, hub-and-spoke or Transit Gateway network, identity federation, reusable Terraform modules, centralised logging, cost allocation tagging applied at creation.
- Migration shape: discovery and dependency mapping, per-application disposition, wave planning, cutover runbook and rollback, data replication and the final sync window, decommissioning and the licence savings that funded the project.
- Operations shape: on-call rota, patch and vulnerability remediation, backup and tested restore, capacity and quota management, access request handling, incident response, cost reporting, and the automation that removes your own toil.
- Application infrastructure shape: Kubernetes clusters and add-ons, ingress and certificates, CI/CD, image build and supply chain, secrets delivery, environment provisioning for developers.
- Read the reporting line: into a director of infrastructure or IT operations means jobs one to three; into a director of engineering or platform means job four.
- A managed service provider or cloud partner posting is usually jobs two and three at once, across many customer tenants, which is the fastest way to accumulate estate variety and the fastest way to burn out.
- If the posting lists all four shapes, it is a small company where you will be the whole cloud function. Good for learning, bad for being on call alone, so ask who covers you when you are ill.
The certifications that pay off, in what order, and the ones that do not
Certifications matter more in cloud engineering than in software engineering, and less than the certification industry wants you to believe. They do exactly one job well: they get a resume with no cloud job title past a keyword filter and in front of a human. They do not get you hired, because the hiring decision is made in the troubleshooting round and no exam prepares you for it.
There is one place where a certification is worth real money rather than attention. Managed service providers and cloud partners have to maintain certified headcount to hold their partner tier with AWS, Microsoft or Google, which means your certification has direct commercial value to them. That is why partners will pay for your exam, will hire on aptitude plus a certification rather than on a prior cloud title, and are the single most reliable outside door into this field. If you are trying to break in, apply to partners and say in your cover note that you hold the certification and will take the next one.
The order that works. First, one associate-level certification in the cloud your target employers actually run, which you determine by reading twenty local postings and counting rather than by reading market-share articles. On AWS that is Certified Solutions Architect Associate, or Certified SysOps Administrator Associate if the postings you want are operations-shaped. On Azure it is AZ-104 for Azure Administrator Associate. On Google Cloud it is Associate Cloud Engineer. Before you buy a voucher, open the provider's own certification page and confirm the exam code, the current blueprint and that the exam is still offered, because all three vendors revise exams regularly and AWS has retired specialty certifications with notice periods that third-party course material ignored for months afterwards.
Second, HashiCorp Certified: Terraform Associate. It is cheap, short and sits on the exact line that separates a console operator from an engineer. A hiring manager reading a resume with one cloud associate certification and Terraform Associate assumes you can write infrastructure code; the same resume with three cloud certifications and no Terraform reads as someone who studies instead of builds. If an employer has moved to OpenTofu, say so in the same breath and note that you have run both, because the migration is a question some teams are still working through.
Third, one of these depending on shape. Certified Kubernetes Administrator if the job runs containers, and it is worth more than any multiple-choice exam because it is performance-based: you solve real tasks in a live cluster under time pressure, and nobody passes it without having used kubectl in anger. AWS Certified Solutions Architect Professional or Azure AZ-305 when you are moving toward design and architecture, not before, because the professional exams assume operational experience and reward people who have made the trade-offs. AWS Certified Advanced Networking Specialty if you came from network engineering, because hybrid connectivity, DNS resolution across accounts and transitive routing are where estates actually break and very few candidates can discuss them.
The foundational certifications are worth less than they are sold as, and they are not worthless. AWS Certified Cloud Practitioner, Azure AZ-900 and Google Cloud Digital Leader are useful for exactly one audience: a career changer who needs to show a recruiter that the intent is real and the vocabulary is in place, before the associate exam is passed. On a resume that already shows cloud work they add nothing, and they cost a little credibility because they signal you listed everything you have.
What has close to zero marginal effect: a fourth and fifth certification, a second cloud's associate certification before you are strong in the first, and anything from a training vendor nobody in the industry has heard of. Past two certifications plus Terraform Associate, every spare week is better spent on an artefact with a cost number in the README.
- Step one: one associate certification in your primary cloud. AWS Solutions Architect Associate or SysOps Administrator Associate, Azure AZ-104, or Google Associate Cloud Engineer. Six to ten weeks for someone already in IT infrastructure.
- Step two: HashiCorp Certified Terraform Associate. Two to four weeks if you have written real Terraform. The highest ratio of signal to study hours in this field.
- Step three, pick one: CKA for container work (performance-based, two to three months of practice), AWS Solutions Architect Professional or Azure AZ-305 for the architecture step, AWS Certified Advanced Networking Specialty to differentiate on a scarce skill.
- Skip or de-emphasise: a third cloud vendor's exams, foundational exams once you have real experience, and anything you cannot connect to something you built.
- Linux and networking are the uncertified prerequisites and they are tested harder than any cloud service. If you cannot read a routing table, use tcpdump, interpret a dig result and explain what a TCP handshake failure looks like versus a DNS failure, the certification will not save the interview.
- Validity: AWS three years, Microsoft role-based annually through a free online assessment, Linux Foundation and Terraform Associate two years, Google Cloud varies by level and has changed, so read the page.
- Budget reality: the exam fee is the small cost. The real cost is a cloud account you spend money in, so set a budget alert before your first resource and a destroy habit before your first week ends.
Getting in without a cloud job title
The dominant mistake is applying outward when the shorter path is sideways. A cloud engineer posting from an unknown applicant with no cloud title competes against dozens of people who already hold one. The same person inside a company that is mid-migration, who has spent six months doing cloud work at the edges of their current job, is having a conversation rather than making an application.
If your employer already runs cloud, there is a standard playbook and it works. Volunteer for the migration wave nobody wants. Take the tagging and cost allocation cleanup, which is unglamorous, almost always outstanding, and makes you the person who can read the bill. Ask to join the on-call rota for the cloud estate, because on-call is how you learn what actually breaks. Pick one small thing and own its infrastructure code end to end, including the pipeline that deploys it. Then write down numbers while you still have access: how many accounts, how many workloads, how long provisioning took before and after, what the monthly spend was and what you cut.
Each feeder role has a specific gap to close, and naming yours is more useful than a generic study plan. A systems administrator usually knows the operating system and the failure modes and needs infrastructure as code plus a CI pipeline. A network engineer has the strongest hand of anyone and often does not know it: cloud networking, hybrid connectivity, DNS resolution across accounts and transitive routing are where estates break, and the gap is usually Linux, identity and Terraform. A help desk or desktop support engineer needs Linux, networking fundamentals and scripting before a cloud certification will convert. A developer needs networking and identity, which they typically have never had to own. A database administrator needs compute, networking and IaC, and arrives with backup and restore discipline that most candidates lack. A NOC or monitoring analyst needs to move from watching to changing.
If your employer has no cloud, the most reliable outside doors are, in order: a managed service provider or cloud partner, because certified headcount has commercial value to them and they hire on aptitude; a large enterprise with a cloud operations tier, which hires at junior level and runs a structured induction; a staffing agency on a contract-to-hire basis, which is the fastest to interview and the least stable; a university or public-sector IT department, which pays below market and gives you a real estate to learn on; and a government contractor, where a clearance you already hold is worth more than a certification.
Structured entry programmes exist and are worth checking rather than assuming. AWS runs re/Start, a free skills programme aimed at people who are unemployed or underemployed, delivered through local partners. Apprenticeship routes into cloud and DevOps exist in several countries, including funded apprenticeship standards in the UK, and some large employers run returnship programmes for people coming back after a career break. Check current availability and eligibility with the programme itself, because they are regional and they change.
On bootcamps, the narrow true thing: a bootcamp can compress the learning and it does not substitute for the gap a hiring manager is actually worried about, which is that you have never been accountable for something in production. A bootcamp graduate with an artefact that has a cost number, a teardown and a postmortem beats one with a certificate of completion by a wide margin. If the bootcamp's own project is the same project every graduate submits, it works against you, because the hiring manager has already seen it from the last applicant.
On degrees, the narrow true thing: no degree is required, many working cloud engineers do not have a computer science degree, and the degree matters most for the application-infrastructure shape and at employers with rigid HR frameworks, including some public sector and defence roles where a qualification is a scored criterion. If you have a degree in anything, list it and move on. If you do not, do not explain it; fill that space with evidence instead.
- Inside your current employer: volunteer for the migration wave, take the tagging and cost cleanup, join the cloud on-call rota, own one service's infrastructure code and its pipeline.
- Write the numbers down before you leave a job. Accounts or subscriptions under a baseline, workloads migrated, cutover downtime, provisioning lead time before and after, monthly spend and what you cut, patch compliance, restore test results. Nobody gives you access to these after you resign.
- Close the specific gap for your starting role rather than studying broadly: IaC and CI for a sysadmin, Linux and identity for a network engineer, Linux and scripting for help desk, networking and identity for a developer.
- Best outside doors, roughly in order of reliability: MSP or cloud partner, enterprise cloud operations tier, contract-to-hire through an agency, university or public sector IT, government contractor.
- Check AWS re/Start and local apprenticeship or returnship routes directly for current eligibility rather than trusting a summary.
- Apply to the junior title, not the one you want. Cloud operations analyst, cloud support engineer, infrastructure engineer and associate cloud engineer all become cloud engineer within a couple of years and have a fraction of the applicant pool.
- Say what you have actually been accountable for, even if it was small. A hiring manager would rather hear that you owned the backup restore test for twelve servers than that you completed a 40-hour course.
What hands-on AWS or Azure work to show, and what gets ignored
A hiring manager looking at your portfolio is trying to answer one question: has this person ever broken something that mattered, and do they know why it broke. Tutorial projects do not answer it. Design every artefact around a decision you made and can defend, because the follow-up question is always why, and three why questions is the whole interview.
What gets ignored, honestly. A three-tier web application on virtual machines from a course, because it is the same architecture diagram as every other applicant's. A static site on object storage behind a CDN, on its own, because it demonstrates nothing about identity, networking or state. Console screenshots, because they prove you clicked. A wall of certification badge images, which also breaks resume parsing. A repository with one commit titled initial commit and no README. The capstone project your bootcamp assigned to everyone in your cohort.
What counts, in order of signal. First, a Terraform repository that looks like it belongs to a company: remote state with locking, at least one module with a deliberate variable interface and sensible defaults, separate environments that are not copy-paste of each other, version pins on the provider and the modules, and a README that opens with the decision rather than the tech list. On AWS, know that recent Terraform versions can lock state in the S3 backend itself with a lock file, so the separate DynamoDB lock table is no longer the only answer; being able to say which one you used and why is the point. The Azure equivalent is a storage account using blob lease locking.
Second, a pipeline that deploys that code using short-lived credentials from OIDC federation, with the trust relationship scoped to one repository and one branch, plan on pull request and apply on merge. This is the single highest-signal thing a candidate without a cloud title can show. It takes an afternoon once it clicks, almost no portfolio has it, and anyone who has run a pipeline in a serious environment recognises it instantly.
Third, a landing zone at toy scale, which is still a landing zone. On AWS: an organisation with a small organisational unit structure, one service control policy that actually denies something, CloudTrail aggregated into a separate logging account, and a tagging standard enforced rather than documented. On Azure: a management group hierarchy, one Azure Policy assignment with a real effect, a Log Analytics workspace collecting diagnostic settings, and a naming and tagging convention applied through policy. The scale does not matter. The structure is the point, because it shows you understand that an estate is governed, not just built.
Fourth, networking you can defend line by line. A virtual network with public and private subnets across two availability zones, egress through a NAT gateway, private connectivity to the object store through a gateway or private endpoint so that traffic does not leave through the NAT and get charged on the way, flow logs on, and an explanation of why you used a security group here and a network access control list there. The Azure equivalent is a hub-and-spoke topology with network security groups, a user-defined route sending egress through a firewall, and a private endpoint with its private DNS zone wired up, because the DNS half is where everyone fails.
Fifth, a cost number, which is the artefact almost nobody brings and the one that lands hardest. Build it, run it for a month, and put the bill in the README with what you changed to cut it. A line of the form 'ran at X dollars a month, down from Y after replacing the NAT gateway path to object storage with a gateway endpoint and moving the workload to an ARM-based instance type' tells a hiring manager you understand that cloud engineering is an economic discipline. Use your own real figures, because the first follow-up question is how you worked them out.
Sixth, a teardown and a drift note. Prove that destroy works, that the state is clean afterwards, and say what is not in the code. Every real estate has drift, and a candidate who volunteers theirs is more credible than one who implies perfection. Seventh, one migration with real users, however small: a club's website, a family business's file server, a side project with paying customers. Write the runbook, the rollback and the actual cutover window, then say what went wrong. A migration you reversed and retried is a better story than one that went perfectly, because it shows you had a rollback and used it.
Keep the lab from bankrupting you, because an unexpected bill is the most common reason people stop learning cloud. Set a budget alert and a spend notification on the day you open the account, before the first resource. The three bills that shock beginners are the NAT gateway, which charges by the hour whether or not anything flows through it, a load balancer left running behind a deleted application, and a managed database or accelerator instance someone forgot. Free tier terms are per-account, they change, and they expire, so read the current page rather than a tutorial from two years ago. Build a habit of destroying at the end of a session and rebuilding from code at the start of the next one, which is also the best Terraform practice you can get.
- Terraform repository: remote state with locking, one real module, pinned provider versions, separate environments, README that opens with the decision.
- CI with OIDC federation and no stored access key, trust scoped to repository and branch, plan on pull request and apply on merge. The highest-signal artefact available to someone without a cloud job title.
- A governed estate at small scale: organisational units and one enforcing service control policy, or a management group hierarchy and one Azure Policy assignment, plus centralised logging.
- Networking you can defend: two availability zones, private subnets, NAT egress, a private or gateway endpoint to object storage, flow logs, and a reason for every security group and route.
- A cost figure with a before and after, using your own bill, and the specific mechanism that moved it. This is the artefact that separates you from everyone holding the same certification.
- A working destroy, a stated drift, a runbook, and one postmortem of something you broke yourself with the mechanism named.
- One migration with real users, including the rollback plan and the window. Small is fine. Real is the requirement.
- Budget alert on day one, destroy at the end of every session, and never leave a NAT gateway, load balancer, managed database or accelerator instance running overnight by accident.
The resume: what a cloud hiring manager reads and what gets skipped
Your resume is read twice: first by a parser matching literal strings, then by an engineer who gives it a few seconds before deciding whether to read the bullets. Both readings reward the same thing, which is specificity, so you do not have to write two documents. Keep it to one page under eight years of experience and two pages beyond that, in a single-column layout with no text boxes, no columns, no images and no icons, because every one of those breaks parsing in some applicant tracking system.
Name the cloud and the services as the vendor names them, at least once each, in a bullet rather than only in a skills block. A parser matching 'Amazon VPC' does not match 'virtual networking', and a human reading 'Azure Policy' knows something different about you than one reading 'governance'. This is not keyword stuffing, which is a list of forty service names with no context and is obvious and counter-productive. It is naming the entity you actually worked with while describing the work.
The top third decides whether the rest gets read. Two lines of summary that name the cloud, the scale and the shape: the number of accounts or subscriptions, the number of workloads, the monthly spend if you can share it, and whether you build, migrate, run or enable. Then a short technical block that is ranked rather than flat. 'AWS in production four years. Azure to a support and read level. Terraform, Python, Linux, Kubernetes to CKA level. Google Cloud not run in production.' An honest ranking buys credibility that a list of thirty logos spends.
Every bullet is a change plus a consequence with a unit on it. Cloud engineering is unusually rich in available numbers, and most candidates use none of them. Provisioning lead time, before and after. Workloads migrated and the cutover downtime. Accounts brought under a baseline. Monthly spend reduced, as both a percentage and a figure if your employer allows it. Patch compliance percentage and how it was measured. Mean time to recovery, change failure rate, number of manual runbook steps removed, number of tickets eliminated by self-service. Recovery point and recovery time objectives proven by a restore test you actually ran, which separates you from the many who have never tested one.
A bullet that works looks like this. 'Replaced 19 hand-built VPCs with one Terraform module and migrated all 19 without downtime; new environment provisioning went from 4 days of tickets to 25 minutes self-service, and cross-account DNS resolution stopped being an escalation.' A bullet that does not work looks like this. 'Responsible for AWS infrastructure including EC2, S3, RDS, VPC, IAM, CloudWatch, Route 53, Lambda and CloudFormation.' The second one is a service list wearing a sentence.
Translate your title, do not inflate it. If you were a systems administrator who did cloud work, keep the real title and add a scope line underneath: 'Systems Administrator (cloud estate: 3 AWS accounts, 120 instances, Terraform-managed from 2025).' Enterprises and government contractors run employment verification, and a title that does not match the employer's record ends the process after the offer, which is the worst possible moment. The same applies to dates, which a parser compares against your application form.
Two optional lines that are filters rather than decoration. If you hold a US security clearance, say the level and the status plainly near the top, because for defence-adjacent roles it is the first thing read. If you need visa sponsorship or you are restricted to certain locations, say so rather than discovering it at offer stage. Being filtered out in week one costs you nothing; being filtered out in week six costs you six weeks.
What gets skipped entirely: objective statements, 'passionate about cloud technologies', soft-skill adjectives with no evidence, a full list of every course you completed, certification badge images, a photograph in a market where photographs are not customary, and references available on request. Replace all of it with one more bullet that has a number in it.
- Single column, no images, no text boxes, one page under eight years. The first reader is a parser.
- Name services as the vendor names them, inside bullets, once each. Amazon VPC, AWS IAM, Azure Policy, Microsoft Entra ID, Terraform, Kubernetes.
- Two-line summary naming cloud, scale and shape. Then a ranked technical block that admits what you have not run in production.
- Numbers cloud engineers actually have: provisioning lead time, workloads migrated, cutover downtime, accounts under baseline, spend reduced, patch compliance, mean time to recovery, change failure rate, runbook steps removed, restore test result.
- Keep your real job title and add a scope line underneath. Employment verification catches inflated titles after the offer.
- State clearance level, sponsorship need or location restriction early. Early filtering is cheap; late filtering is expensive.
- Cut objective statements, badge images, course lists and passion language. Every line should be a change with a consequence.
- A short link list at the top beats a portfolio section at the bottom: the Terraform repository, the pipeline, the cost writeup, each with a one-line description of the decision it contains.
How the loop actually runs, and who screens you
The process depends more on who is hiring than on the seniority of the role, so work out the employer type before you prepare. The same candidate needs different preparation for a product company, a managed service provider, an enterprise and an agency placement.
At a product or technology company, expect five stages over three to six weeks. A recruiter screen of 20 to 30 minutes that is a keyword and salary conversation, where the only mistakes available are being vague about what you owned and naming a number before you have information. A hiring manager call of 45 to 60 minutes about scope, on-call, what you were accountable for and why you are leaving. A live technical screen of 45 to 60 minutes on Linux, networking, identity and troubleshooting, which eliminates more candidates than every other stage combined. A practical exercise, either a two to four hour take-home writing Terraform or a script, or a paired session doing the same thing live. Then a design or scenario round, and often a behavioural round.
At a managed service provider or cloud partner, expect two or three stages over one to two weeks. Often a written or multiple-choice technical test, which feels dated and is a real filter. A technical interview that weights certification heavily, because certified headcount has commercial value to them. And a conversation about customers, because you will be on calls with them: expect to be asked how you would tell a customer that an outage was caused by their own developer's change, and expect a question about working across multiple tenants without mixing them up. Partners move fast and will often decide in the same week.
At an enterprise, a bank, a utility or the public sector, expect a panel and a scoring matrix built from the job description, four to ten weeks end to end, and competency questions answered in situation, task, action, result form. The panel may include someone from security and someone from service management who will ask about change control, not about Terraform. There may be a background or security check that adds weeks. The failure mode here is being too informal and too technical: answer the competency question that was asked, with a structure, before you go deep.
Through a staffing agency, expect the fastest and least informative process. A rate conversation first, a checklist-style phone screen from a recruiter who is matching strings, then one or two interviews with the client. Contract-to-hire is common in cloud operations and is a legitimate way in if you treat the contract as six months of evidence building rather than as a job.
Across all of them, understand who is actually screening you at each point. The recruiter is matching your resume to a requisition, so give them literal matches and a clear one-sentence answer to what you own. The hiring manager is assessing scope and risk, so talk about accountability and about what you did when it went wrong. The engineer on the technical screen is assessing whether they would want you on call with them, which is a question about method under uncertainty rather than about recall. The design round assesses judgement, mostly measured by what you decline to build.
Two practical notes on timing. Follow up once, three to five business days after a stage, with one specific thing you thought about afterwards rather than a reminder of your enthusiasm. And in migration-heavy organisations, hiring clusters around the start of a wave, so a role that was not open in March can be open in June. Keep the name of every hiring manager who got to a late stage with you, and send a short note when you have something new to show.
- Product company: recruiter screen, hiring manager call, live troubleshooting round, practical Terraform or scripting exercise, design round, sometimes a values round. Three to six weeks.
- MSP or cloud partner: written technical test, technical interview weighting certification, customer-facing conversation. One to two weeks, decision often in the same week.
- Enterprise or public sector: panel with a scoring matrix, competency answers in situation-task-action-result form, change control questions, possible background or security check. Four to ten weeks.
- Agency or contract: rate first, checklist phone screen, one or two client interviews. Days. Contract-to-hire is a legitimate entry route.
- The live troubleshooting round eliminates the most candidates. Prepare it harder than the design round, however senior you are.
- Questions worth asking in every loop: the on-call paging volume last month, how many accounts or subscriptions exist, and what changed after the last incident.
- Follow up once with a specific thought, not a reminder. Then keep the hiring manager's name, because migration hiring comes in waves.
The technical round: the questions that decide it
Cloud engineering interviews are unusually predictable, because the failure modes of a cloud estate are unusually consistent. The round is almost always a troubleshooting conversation rather than a quiz, and the interviewer is listening for method: do you check in a sensible order, do you say what you expect to see before you look, and do you notice when your own hypothesis has been disproved. A candidate who reaches the right answer by guessing scores worse than one who reaches a wrong answer by a clean process and then corrects.
The single most common question, in some form, is reachability. 'An instance in a private subnet cannot reach the internet. Walk me through it.' A strong answer is ordered and cheap first: confirm what 'cannot reach' means (name resolution failing, connection timing out, or connection refused, because each points somewhere different), then the route table associated with that subnet and whether its default route points at a NAT gateway, then whether the NAT gateway sits in a public subnet with a route to an internet gateway, then the security group egress rules, then the network access control lists in both directions because they are stateless and the return traffic needs its own rule, then DNS resolution, then whether the instance needs the internet at all or should be reaching a private endpoint instead. A weak answer names security groups and stops. The Azure version substitutes user-defined routes, a NAT gateway or firewall as the egress path, network security groups, and the detail everyone misses, which is that a private endpoint whose private DNS zone is not linked to the virtual network resolves to the public address and fails quietly.
Expect the stateful versus stateless question, because it separates people who have configured these from people who have read about them. Security groups are stateful, attach to an elastic network interface, and can only allow. Network access control lists are stateless, attach to a subnet, evaluate rules in numbered order, and can explicitly deny. You reach for a network access control list when you need a blunt subnet-wide block, typically a deny of a specific source range, and almost never otherwise. Saying that out loud, including the almost never, is what a good answer sounds like.
Identity is the other reliable question. On AWS the expected answer is the evaluation logic: an explicit deny anywhere wins immediately; otherwise an allow must exist in an identity-based or resource-based policy; and organisation-level policies, permission boundaries and session policies can only narrow what is already allowed, never grant. Being able to say why a user with an administrator policy still cannot perform an action, because a service control policy at the organisational unit denies it, is the answer that marks experience, and knowing that AWS also has resource control policies for restricting access to resources organisation-wide marks someone who has kept up. On Azure the equivalent is that role-based access control is allow-based with deny assignments as a rare exception, that scope inherits from management group to subscription to resource group to resource, and that Azure Policy is a separate mechanism which governs the shape of a resource rather than who may act on it. Candidates who conflate Azure RBAC and Azure Policy lose the round.
Then the credentials question, and there is one correct answer. 'How do you give a CI pipeline permission to deploy to your cloud account?' Say OIDC federation: the pipeline presents a token from its identity provider, the cloud assumes a role whose trust policy is scoped to that specific repository and branch or environment, and the credentials it gets are short-lived. No access key is stored anywhere. If you answer 'store the access key as a repository secret and rotate it', you have told a competent interviewer that your estate has long-lived credentials in it. Have the follow-up ready too: what the trust policy condition actually looks like, and why scoping it to the organisation alone is not enough.
The cost question has become standard and most candidates are unprepared for it. 'The bill went up 40 percent this month. Find out why.' Work outward: group spend by service to find the mover, then by usage type inside that service because data transfer and request charges hide there, then by tag or by account to find the owner, then look for the usual culprits (a new region someone opened, an environment nobody shut down, unattached volumes and snapshots, cross-availability-zone transfer from a workload that got rescheduled, a log destination collecting far more than anyone intended, an accelerator instance left running). Finish with the real answer, which is that the question is hard because tagging coverage is incomplete, and the durable fix is enforcing tags at creation rather than reporting on them afterwards.
The design or scenario round is marked on judgement and on what you refuse. A typical prompt is to design the network and account structure for three environments plus a connection to an existing data centre. Reviewers are looking for: non-overlapping address ranges chosen with room to grow and documented somewhere; separate accounts or subscriptions per environment rather than separate tags; private by default with deliberate egress; a Transit Gateway or virtual WAN hub rather than a mesh of peering connections that becomes unmanageable; a dedicated connection with a VPN as backup rather than either alone; an explicit answer about where DNS resolution happens for on-premises names and cloud names in both directions; and a statement of what you are deliberately not building yet. Candidates who name every available service lose. Candidates who say 'I would not add a firewall appliance in phase one, and here is what I would need to see first' win.
- Reachability, ordered: define the symptom, route table, NAT or firewall path, security group egress, stateless access control lists both ways, DNS, then ask whether it should use a private endpoint instead.
- Security group versus network access control list: stateful, interface-scoped and allow-only, versus stateless, subnet-scoped and able to deny. Say when you would actually use the second one.
- AWS policy evaluation: explicit deny wins, an allow is required, and organisation policies, permission boundaries and session policies only narrow. Azure: RBAC allows and inherits by scope, Azure Policy governs resource shape, and they are not the same mechanism.
- CI credentials: OIDC federation, short-lived credentials, trust scoped to repository and branch. Never a stored access key.
- Bill investigation: service, then usage type, then tag or account, then the usual culprits, then the tagging enforcement that makes next month's question answerable.
- Restore, not backup. Expect 'your production database is gone, what happens' and answer with a recovery point objective, a recovery time objective, and the date of the last restore you actually performed. A backup you have never restored is not a backup.
- The scripting exercise is usually inventory, tagging or cleaning up orphaned resources against the provider SDK in Python, Bash or PowerShell. The trap is pagination: list calls return pages, and a candidate who processes only the first page produces a confidently wrong report.
- Expect 'tell me about something you broke'. Name the mechanism, the blast radius, how you found out, and the control that now prevents it. A candidate with no story is either inexperienced or not telling the truth.
Pay, on-call, scope, and reading the offer
Do not trust an aggregator average for this title. The US Bureau of Labor Statistics has no cloud engineer occupation code, so every published average is assembled from self-reported data across four different jobs, three employer types and every city in the country. The nearest official codes are 15-1244 Network and Computer Systems Administrators, 15-1241 Computer Network Architects and 15-1299 Computer Occupations All Other, and the Occupational Employment and Wage Statistics tables give you a defensible local floor by metropolitan area rather than a national guess.
For an actual range, read live postings. A number of US states and cities require a pay range in the job posting, and many employers covered by those rules publish a range on postings advertised more widely, which means you can often read real ranges for your target companies from outside those places. Check the current rule where you are applying rather than a list from a few years ago, because the coverage has been expanding. Do the same search three times, filtered to a product company, an enterprise and a managed service provider, because the spread between employer types is larger than the spread between cities. Add union scale if you are looking at public sector roles covered by an agreement, and add the published pay bands that many public employers and universities disclose.
The pattern across employer types is consistent even though the numbers are not. Product and technology companies pay the most and ask for the most, often including an equity component that you should value at zero until you understand the liquidity. Enterprises pay in the middle with better pension and leave and far more process. Managed service providers pay least in cash and most in experience density, because in eighteen months you will have touched many estates, which is worth more at your next move than the salary difference. Contract rates look higher per hour and carry no benefits, no paid leave and no notice protection, so convert them properly before comparing.
On-call is the term of the offer that candidates investigate least and regret most. Ask four questions and insist on numbers. How many people are in the rota, because a rota of three is a different life from a rota of eight. How many pages did the rota receive last month, at what hours, and how many were actionable. Is on-call compensated with pay or time off, and is that written down. And the one that matters most: do you have the authority to fix the thing that pages you, or do you have to wake someone who does. A rota where you are paged for something you are not permitted to change is the worst job in infrastructure.
Several scope questions reveal what the job really is, and asking them is also how you demonstrate seniority. Who approves a production change, and how long does it take. Do you have write access to the infrastructure code, or do you file a ticket with a platform team. How many accounts or subscriptions exist, and is there a baseline they all inherit. What is the monthly cloud spend, because an estate costing tens of thousands a month and one costing millions are not the same job and do not teach the same things. Is there a migration in flight, and who owns the decision to move an application. What happened in the last incident, and did anything change afterwards.
Red flags, stated plainly. No infrastructure as code and no plan to introduce it, which means the job is clicking consoles and will not develop you. One person who holds all the knowledge and is described admiringly as the person who knows where everything is, which means a bus factor of one and a political problem you will inherit. Cost ownership without purchasing or architectural authority, which makes you the person who reports a number that nobody acts on. A role described as cloud engineering that is actually hypervisor administration with a cloud title and a vague roadmap. And an interviewer who cannot say what the on-call paging volume was last month, which usually means nobody measures it.
On negotiating, the leverage in this field is specific. A competing offer is the strongest lever and the only one that reliably moves a band. Below that, the levers that work are a scarce specialism the role needs (hybrid networking, a regulated-industry background, a migration of the exact kind they are starting, accelerator capacity experience), a clearance you already hold, and a willingness to take the on-call rota as written. Certifications do not move an offer, because they were already priced when they got you the interview. If you are transferring internally into a cloud role, know that internal moves usually lag the external market, so take the scope and the title, then plan to test the market in a year or two with a resume full of numbers you could only get from the inside.
- No BLS code exists for cloud engineer. Use OES codes 15-1244, 15-1241 and 15-1299 for a local floor, then read live postings covered by pay-transparency rules for a real range.
- Compare the same title across a product company, an enterprise and an MSP. The employer-type spread is wider than the city spread.
- On-call questions that need numeric answers: rota size, pages last month and at what hours, compensation in pay or time, and whether you may fix what pages you.
- Scope questions: who approves a production change, do you have write access to the infrastructure code, how many accounts, what is the monthly spend, is a migration in flight, what changed after the last incident.
- Red flags: no infrastructure as code with no plan, a single hero holding all knowledge, cost responsibility without authority, hypervisor administration wearing a cloud title, and no measured paging volume.
- Levers that move an offer: a competing offer, a scarce specialism, an existing clearance, accepting the rota as written. Certifications do not, because they were already priced in.
- Internal transfers into cloud usually pay below market. Take the scope, collect the numbers, and test the market in a year or two.
What a cloud engineer has to know about AI in 2026-27
The honest headline first: AI has not changed the core of cloud engineering, and it has changed a great deal around it. Subnets, routes, DNS, identity, state, quotas, backups and bills work exactly as they did three years ago, and no tool reasons about them reliably on your behalf. The 2am diagnosis, the migration sequencing decision, the blast-radius judgement and the argument with an application team that does not want to move are all still yours. Anyone telling you the fundamentals are obsolete is selling a course.
What has genuinely changed falls into five areas, and all five sit squarely in a cloud engineer's lane: a new and scarce class of hardware to procure and schedule, a new class of managed service to make private and observable, a new fast-growing line item on the bill, far more infrastructure code arriving per week from people who cannot defend its defaults, and a new kind of identity taking actions inside your accounts. None of that requires you to train a model. All of it requires you to be good at the job you already have.
This matters in the interview in a specific way. Candidates who claim AI has transformed cloud engineering get one follow-up question and run out of material. Candidates who say 'the fundamentals are the same, and here are the five things in my lane that are actually new' sound like someone who has been doing the work. Be the second one, and have one concrete mechanism ready for each of the five.
There is also a new question you should expect directly: whether you use AI tooling in your own work and where you draw the line. The answer that lands is specific about the boundary rather than enthusiastic or dismissive. Generate the boilerplate module, let an assistant read logs and propose a hypothesis, read every plan line by line, and never let anything apply to production without a human approving the diff. Add the honest observation that an engineer who cannot read a Terraform plan carefully is more exposed now than before, not less, because they can produce far more code than they can defend.
One piece of good news, stated narrowly: AI capacity work is one of the reasons cloud infrastructure headcount is being funded. Teams that struggled to get approval for a second cloud engineer are now hiring one because somebody has to own accelerator quota, the inference endpoint's network path and the bill. If you can talk credibly about those three things, you are applying into demand rather than against it.
Accelerator capacity as a procurement and scheduling problem, not a Terraform problem
The question 'can we have thirty-two GPUs next week' is one a cloud engineer now gets asked, and the answer is almost never a line of infrastructure code. It is a quota request, a region availability check, a reservation decision and a conversation with an account team. Accelerator capacity is genuinely scarce, availability differs sharply by region and by instance family, default service quotas for GPU instance families are often zero until you ask, and a request can take days to be approved. The surrounding engineering matters as much as the chips: multi-node training needs low-latency interconnect and tight placement, the storage layer has to sustain enough throughput to keep the accelerators busy rather than idling expensively, spot or pre-emptible capacity will interrupt a long training run unless checkpointing is set up, and an idle accelerator fleet is one of the most expensive mistakes available in a cloud account.
Show it: Be concrete on one platform rather than vague on three. On AWS: how you raise a service quota for a specific instance family in a specific region, EC2 Capacity Blocks for ML for a bounded training window, on-demand capacity reservations, cluster placement groups, and Elastic Fabric Adapter for the interconnect. On Azure: per-region per-family quota requests, the InfiniBand-enabled virtual machine families for multi-node training, and spot versus reserved capacity. On Google Cloud: reservations and the Dynamic Workload Scheduler for queued capacity. Then say the sentence that marks experience: you check quota and regional availability before you promise a date, because the lead time on capacity is a procurement lead time and no amount of automation shortens it.
Managed inference services as cloud infrastructure, including the quota and throttling model nobody warns you about
Amazon Bedrock, Azure OpenAI within Azure AI Foundry and Google Vertex AI are now ordinary line items in estates you run, and the questions they raise are questions you already know how to answer: what network path reaches this endpoint, what identity calls it, what data can it read, does it egress to the internet, and is the call logged. The part that surprises teams, and that the cloud engineer ends up explaining, is the capacity model. These services are rate-limited in tokens per minute as well as requests per minute; provisioned throughput is a separate commercial purchase rather than a setting; model availability differs by region; and a product team's load test will hit a throttling response that looks exactly like an outage in their dashboard. Being the person who can say 'that is a quota, here is the current limit, here is the request, and here is the retry behaviour your client needs' is immediate value.
Show it: Name the private path on your platform. On AWS: Bedrock reached through a VPC endpoint rather than the public internet, IAM conditions restricting which models a role may invoke, and model invocation logging enabled so there is a record of what was asked. On Azure: Azure OpenAI behind a private endpoint with public network access disabled, managed identity instead of API keys, the private DNS zone actually linked (this is where it breaks), and diagnostic settings shipping to a workspace you can query. On Google Cloud: Vertex AI inside a VPC Service Controls perimeter, which is the control that stops data leaving the project through a managed API. Then add the operational detail: client-side retry with backoff and a circuit breaker, a dashboard on throttling responses rather than only on errors, and regional model availability and data-handling terms checked on the provider's current page rather than remembered, because both move.
The bill, because AI made cost the loudest question in the room
In many organisations the AI line is now the fastest-growing part of the cloud bill, and the person asked to explain it is the cloud engineer. That has promoted cost from a quarterly report into a design constraint. The useful shift is from total spend to unit economics: cost per thousand tokens, cost per accelerator hour, cost per thousand requests, cost per customer. The specific things that blow a budget up are learnable and repeatable: an accelerator fleet idling between jobs, a retrieval pipeline re-embedding an entire corpus on a schedule when it only needed the changed documents, a vector store on expensive provisioned storage, cross-region data transfer feeding a training job from a bucket in the wrong region, an inference endpoint provisioned for peak and running at peak prices all night, and every prompt and response logged into a high-cost log tier where the logging bill outgrows the inference bill.
Show it: Bring one number you moved and the mechanism that moved it, at any scale. Then show the plumbing that makes attribution possible: tags or labels enforced at creation rather than reported on afterwards, a cost allocation scheme that lets you tell a team what their AI workload cost last month, budget alerts wired to a channel someone reads, and an idle-detection job that shuts down what nobody is using. Mention FOCUS, the FinOps Foundation's open specification for billing data, as the direction multi-cloud spend reporting is heading, and say whether your provider's cost export supports it. If a hiring manager asks one cost question and you answer with a mechanism rather than a tool name, you have separated yourself from most of the field.
Reviewing AI-generated infrastructure code, which is now a large share of what reaches your pipeline
The practical effect of coding assistants on this job is volume and provenance. Far more Terraform, YAML and pipeline configuration arrives each week, and a larger share of it was written by somebody who did not read the provider documentation and cannot say why a default is what it is. Generated infrastructure code is usually syntactically clean and wrong in boring, repeatable ways: ingress open to 0.0.0.0/0 because the example was, a storage container left publicly accessible, encryption left at a provider default that is off, an identity policy with a wildcard action and a wildcard resource, no tags so the resource is invisible to cost allocation, an unpinned module or provider version, a deprecated argument that works until the next provider release, and worst of all a plan that quietly destroys and recreates something stateful. You cannot review your way out of this at volume, and that is the whole point.
Show it: Show the pipeline rather than the vigilance. Static analysis in continuous integration with Checkov or Trivy (tfsec was folded into Trivy, so name Trivy if you are choosing today and mention tfsec only as what you inherited), policy as code with Open Policy Agent and Conftest or with Sentinel, a rule set that fails the build rather than warning, and a hard gate requiring explicit human approval for any plan that destroys or replaces a stateful resource. Then bring one specific rule you wrote because a generated change tried something, which is the detail that proves the story is yours. The sentence that lands in an interview: the review bottleneck moved from writing the code to proving the code is safe, so the leverage is in the checks, not in reading faster.
Agent and non-human identity, because something is now taking actions in your accounts
This is the genuinely new infrastructure problem of 2026 and 2027. Agents, and the tool servers they connect to, hold credentials and take actions in response to text they did not author. That makes them the worst variety of non-human identity: broad permissions, long-lived in practice, and driven by untrusted input. The moment an agent that reads a ticket queue also holds a role that can read a bucket or restart a service, prompt injection stops being somebody else's chatbot problem and becomes a permissions question on your estate. Most estates already hold far more credentials issued to automation than to people, and nobody can tell you the ratio in yours without an audit, which is itself a piece of work worth volunteering for. The same applies to the operations assistants teams now point at their own infrastructure, which is a question you should expect to be asked about directly.
Show it: Answer with a control set rather than with alarm. One identity per agent rather than a shared service account, scoped to specific actions on specific resources; short-lived credentials through federation rather than a stored key; egress restrictions so an agent cannot reach arbitrary endpoints; an explicit approval boundary for anything that writes to production, grants a permission or moves data; and an audit trail that attributes the action to the agent rather than to whoever was logged in. Have a worked example ready: an operations assistant that can read logs and metrics, query the configuration inventory, and open a pull request against the infrastructure repository, but that holds no apply permission, cannot modify identity policies, and cannot touch a production database. Being able to produce that scoping on the spot is a question that now actually gets asked.
What has not changed, and saying so without hedging
The parts of this job that are hardest to hire for are the parts no tool has absorbed. Diagnosing a partial failure across a network path, an identity boundary and a DNS resolver. Deciding the order to migrate ninety applications when four of them share a database nobody documented. Judging blast radius before a change and choosing the smaller reversible step. Holding a conversation with an application owner who has no incentive to move. Knowing which 3am page is an outage and which is a noisy alert that should be deleted. An interviewer who has been burned by a confident candidate with no judgement is specifically probing for these, and claiming AI has made them easier is the fastest way to fail the round.
Show it: Say the narrow true thing and then prove it with a story. 'AI has not changed how I diagnose a reachability failure or how I sequence a migration; it has changed how much code crosses my desk and what is running in my accounts.' Then give the diagnosis story with the mechanism named, and the migration story with the dependency you found late and what you did about it. Pair it with a precise account of how you do use the tooling: what you let it generate, what you always read yourself, and the one thing you will not let it do. Specificity about the boundary reads as judgement. Enthusiasm reads as inexperience, and dismissal reads as someone who has not looked.
What a screen is looking for
These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.
- Cloud engineer
- Cloud engineering
- AWS
- Amazon Web Services
- Microsoft Azure
- Google Cloud Platform
- GCP
- Cloud migration
- Landing zone
- AWS Organizations
- AWS Control Tower
- Service control policy (SCP)
- Resource control policy (RCP)
- Azure management groups
- Azure landing zone
- Azure Policy
- Google Cloud organization policy
- Infrastructure as code (IaC)
- Terraform
- OpenTofu
- Terraform modules
- Remote state
- State locking
- AWS CloudFormation
- Azure Resource Manager (ARM)
- Bicep
- Pulumi
- Ansible
- Amazon VPC
- Azure Virtual Network
- Hub-and-spoke network
- AWS Transit Gateway
- Azure Virtual WAN
- VPC peering
- AWS Direct Connect
- Azure ExpressRoute
- Site-to-site VPN
- BGP
- Subnetting
- CIDR planning
- NAT gateway
- Internet gateway
- Route table
- Security group
- Network ACL
- Network security group (NSG)
- VPC endpoint
- Private endpoint
- Private DNS zone
- Amazon Route 53
- Azure DNS
- VPC flow logs
- AWS IAM
- IAM policy
- IAM role
- Permission boundary
- Microsoft Entra ID
- Azure RBAC
- Managed identity
- Workload identity federation
- OIDC federation
- Least privilege
- Short-lived credentials
- Amazon EC2
- Auto Scaling
- Elastic Load Balancing
- Azure Virtual Machines
- Azure App Service
- AWS Lambda
- Azure Functions
- Serverless
- Containers
- Docker
- Kubernetes
- Amazon EKS
- Azure AKS
- Google GKE
- Helm
- Amazon S3
- Azure Blob Storage
- Amazon EBS
- Amazon RDS
- Azure SQL Database
- Backup and restore
- Disaster recovery
- RPO
- RTO
- High availability
- Multi-AZ
- CI/CD
- GitHub Actions
- GitLab CI
- Azure DevOps Pipelines
- Jenkins
- Git
- GitOps
- Amazon CloudWatch
- Azure Monitor
- Log Analytics
- Prometheus
- Grafana
- Observability
- Linux
- Bash
- PowerShell
- Python
- boto3
- Azure CLI
- AWS CLI
- SDK pagination
- Patch management
- AWS Systems Manager
- Azure Update Manager
- Configuration management
- Cloud cost optimization
- FinOps
- Cost allocation tagging
- AWS Cost Explorer
- Azure Cost Management
- Reserved instances
- Savings Plans
- Spot instances
- Rightsizing
- FOCUS billing specification
- Policy as code
- Open Policy Agent (OPA)
- Checkov
- Trivy
- Drift detection
- Well-Architected Framework
- Cloud Adoption Framework
- VMware migration
- Data centre exit
- Cutover runbook
- Rollback plan
- Change management
- On-call
- Incident response
- Mean time to recovery
- GPU capacity
- Service quotas
- Capacity reservations
- Amazon Bedrock
- Azure OpenAI
- Vertex AI
- AWS Certified Solutions Architect Associate
- AWS Certified SysOps Administrator Associate
- AWS Certified Solutions Architect Professional
- AWS Certified Advanced Networking Specialty
- Microsoft Certified Azure Administrator Associate (AZ-104)
- Azure Solutions Architect Expert (AZ-305)
- Google Associate Cloud Engineer
- HashiCorp Certified Terraform Associate
- Certified Kubernetes Administrator (CKA)
Mistakes that cost people this job
Collecting a third and fourth certification instead of getting production access to a cloud account. Cloud Practitioner, then Solutions Architect Associate, then SysOps, then Developer, with nothing built.
One associate certification in your primary cloud plus Terraform Associate clears the filter. After that, every spare week goes into one artefact with a decision in the README, a cost figure and a working destroy. The certification gets you read; the artefact gets you hired.
Claiming AWS, Azure and Google Cloud at equal depth. The resume reads broad, then the interview goes three questions deep into whichever one the employer runs and the answers stop.
Pick the cloud your target employers actually run, determined by reading twenty local postings and counting, and go deep enough to discuss specific policy evaluation and specific networking behaviour. Then rank honestly in writing: in production, support level, not run in production. The ranking buys credibility a flat list spends.
Being a console-only engineer. You can build it by clicking and you cannot write the Terraform, so every change you propose becomes somebody else's ticket and you never get the senior title.
Learn Terraform to the point of writing a reusable module with a deliberate variable interface, remote state you treat as sensitive, and a pipeline that fails a plan. Then say it in interview terms: 'I shipped the module, not the ticket.' That single sentence separates you from most applicants holding the same certifications.
A portfolio of tutorial projects: a three-tier application on virtual machines, a static site behind a CDN, a serverless to-do list. No decision, no cost figure, no teardown, and identical to every other applicant's.
Build fewer things and document the decision in each. Why two availability zones and not three, why a gateway endpoint instead of NAT egress, what the month cost and what you changed to cut it, what is still drifted. One artefact with a defensible decision outperforms six without.
Telling an interviewer you store a long-lived cloud access key as a pipeline secret and rotate it quarterly. It is said casually and it ends the round at anywhere competent.
Use OIDC federation and short-lived credentials, with the trust relationship scoped to one repository and branch, and be able to describe the trust policy condition. If your current employer still uses stored keys, say that, say you know why it is wrong, and say what you would change first.
Resume bullets that list services instead of describing changes. 'Responsible for AWS infrastructure including EC2, S3, RDS, VPC, IAM, CloudWatch and Lambda.'
Every bullet gets a change and a consequence with a unit on it: provisioning lead time before and after, workloads migrated and the cutover downtime, accounts brought under a baseline, monthly spend reduced, patch compliance, a restore you actually tested. Collect those numbers before you leave the job, because you cannot get them afterwards.
Applying outward exclusively, hundreds of applications to cloud engineer postings from a help desk or sysadmin resume, while the migration project at your own employer goes unstaffed.
Do the cloud work where you already have access. Volunteer for the migration wave, take the tagging and cost cleanup, join the cloud on-call rota, own one service's infrastructure code. A year or two of that converts far more reliably than cold applications, and it gives you numbers no outside candidate has.
Treating Linux and networking as legacy because the job title says cloud. The candidate knows service names and cannot read a route table, interpret a dig result or tell a DNS failure from a TCP timeout.
Drill fundamentals harder than services, because the live troubleshooting round is built on them and it eliminates more candidates than every other stage combined. Practise explaining a reachability failure out loud in a fixed order, and practise saying what you expect to see before you look.
Having no answer to the cost question. The bill is somebody else's problem, so the candidate cannot say how a NAT gateway is charged or how to find out why spend rose.
Learn the investigation path (service, then usage type, then tag or account, then the usual culprits) and the enforcement fix (tags applied at creation, not reported afterwards). Cloud engineers who cannot read a bill get capped, because cost is now how the business measures the function.
Backups nobody has restored. The candidate describes a backup policy and retention schedule, then cannot name the date of the last successful restore test.
Run a restore, time it, and quote the recovery point and recovery time objectives you actually proved rather than the ones in the policy document. A sentence of the form 'we restored the production database into an isolated account in under an hour in August, which is how I know the stated four-hour objective was real' is worth more than any certification.
Accepting an on-call rota you never asked about, then discovering it is three people, uncompensated, and pages for things you have no authority to change.
Ask for numbers before the offer: rota size, pages last month and at what hours, how many were actionable, whether compensation is written down, and whether you may fix what pages you. Asking these also reads as experience, because only people who have been burned ask them.
Talking about AI in the abstract, either claiming it has transformed cloud engineering or dismissing it entirely. Both answers die on the first follow-up question.
Say the narrow true thing: the fundamentals are unchanged, and five things in your lane are genuinely new (accelerator capacity and quota, managed inference endpoints and their throttling model, the AI line on the bill, the volume of generated infrastructure code, and agent identity). Have one concrete mechanism ready for each.
Questions people ask
What does a cloud engineer do?
A cloud engineer builds, migrates and runs a company's infrastructure on AWS, Microsoft Azure or Google Cloud. The work covers networking and DNS, identity and permissions, compute and containers, storage and backup, infrastructure as code, monitoring, patching and cost. In a well-run team a cloud engineer's output is code and pipelines that other engineers consume rather than consoles they click, and the day involves as much negotiation with application teams and finance as it does configuration. The title is used for at least four distinct jobs: building a landing zone and foundation, migrating an estate off a data centre or hypervisor, running an existing estate on an on-call rota, and providing cloud-native infrastructure for developers. Reading which of those a posting means is the first thing to do before applying.
Which cloud certification should a cloud engineer get first?
A cloud engineer should get exactly one associate-level certification in the cloud their target employers actually run, which you determine by reading twenty local postings and counting rather than by reading market-share articles. On AWS that is Certified Solutions Architect Associate, or Certified SysOps Administrator Associate if the roles you want are operations-shaped. On Azure it is AZ-104 for Azure Administrator Associate. On Google Cloud it is Associate Cloud Engineer. Check the current exam code, the blueprint and that the exam is still offered on the provider's own certification page before buying a voucher, because the vendors revise and sometimes retire exams. Then add HashiCorp Certified Terraform Associate, which is short, cheap and signals the exact skill that separates a console operator from an engineer. Beyond those two, a cloud engineer gains very little from a third or fourth certification and should spend the time on an artefact instead.
Can you become a cloud engineer with no experience?
Becoming a cloud engineer from a standing start is possible but it is rarely a direct jump, and planning for one is the common mistake. Most postings with this title ask for two to five years of infrastructure experience with at least one or two in cloud, because the job means changing a production estate where a wrong move causes an outage or a bill. The reliable route is to take an adjacent job first, such as help desk with Linux added, systems administrator, network engineer, NOC analyst, data centre technician or backend developer, then move across inside the same company by volunteering for the migration, the cost and tagging cleanup and the cloud on-call rota, which commonly takes a year to two years. From outside, managed service providers and cloud partners are the most reliable door for an aspiring cloud engineer, because certified headcount has direct commercial value to them and they hire on aptitude plus a certification rather than on a prior cloud title.
Do you need a degree to be a cloud engineer?
No degree is required to work as a cloud engineer, there is no licence and no protected title, and many working cloud engineers do not hold a computer science degree. A degree matters most in two situations: the application-infrastructure shape of the role, which overlaps software engineering and sometimes screens on it, and employers with rigid human resources frameworks including parts of the public sector and defence contracting, where a qualification can be a scored criterion. If you have a degree in anything, list it and move on. If you do not, do not explain it anywhere on the resume and use that space for a cost figure, a migration you ran, or a restore you tested instead.
Should a cloud engineer learn AWS or Azure?
A cloud engineer should learn whichever one their target employers run, and the way to find out is to read twenty job postings within commuting distance or within the remote market being targeted, then count. In practice AWS is strongest at technology companies, startups and much of the digital-native market, while Azure is strongest at organisations already standardised on Microsoft products, which includes a very large share of enterprises, government and healthcare, often through an existing Microsoft agreement. Google Cloud is strongest in data-heavy and machine-learning-heavy organisations and in specific industries. Depth in one beats shallow coverage of three for a cloud engineer, because the interview goes three questions deep into one platform. Once you are genuinely strong in one, the second takes weeks rather than months, since the concepts transfer and mainly the names and the sharp edges change.
How long does it take to become a cloud engineer?
For someone already working in IT infrastructure, a realistic plan to become a cloud engineer is nine to eighteen months: six to ten weeks of evening study for the first associate certification, a few months of building real artefacts and taking on cloud work inside the current job, then an internal transfer or an external move. For someone starting from outside IT entirely, plan on two to three years, usually routed through a first job in support or operations, because what employers are buying is accountability for production systems and that cannot be studied. The fastest version of the path is being inside an organisation that is mid-migration and volunteering for the work nobody else wants, which can compress it to under a year.
What does a cloud engineer interview test?
A cloud engineer interview tests troubleshooting method far more than recall. The round that eliminates the most candidates is a live conversation about Linux, networking and identity: a reachability failure traced in a sensible order through route tables, egress paths, security group rules, stateless network access control lists and DNS; the difference between a stateful security group and a stateless network access control list and when you would actually use the second; how policy evaluation resolves when an explicit deny exists; and how a pipeline should get credentials, where the only good answer is OIDC federation with short-lived credentials rather than a stored access key. Expect a practical Terraform or scripting exercise, where pagination in provider SDK list calls is the standard trap. Expect a cost investigation question. Expect a design round marked largely on what you decline to build. And expect to be asked what you broke, which a cloud engineer should answer with a mechanism, a blast radius and the control that now prevents it.
How much does a cloud engineer earn?
There is no official occupational wage figure for a cloud engineer, because the US Bureau of Labor Statistics has no code for this title, which is why published averages for it vary so widely. The nearest official codes are 15-1244 Network and Computer Systems Administrators, 15-1241 Computer Network Architects and 15-1299 Computer Occupations All Other, and the Occupational Employment and Wage Statistics tables for your metropolitan area give a defensible local floor. For a real range, read live postings from employers covered by pay-transparency rules, and run the same search three times filtered to a product company, an enterprise and a managed service provider: for a cloud engineer the spread between those employer types is wider than the spread between cities. Product companies pay most, enterprises sit in the middle with better benefits and more process, and managed service providers pay least in cash while delivering the fastest accumulation of estate experience.
Is cloud engineering being automated away by AI?
No, and a cloud engineer who claims otherwise in an interview loses the round. The core of the job is unchanged: diagnosing a partial failure across a network path, an identity boundary and a DNS resolver, sequencing a migration of applications that share undocumented dependencies, judging blast radius before a change, and deciding which page at 3am is a real outage. What has changed sits around that core, and it is work being added to a cloud engineer's plate rather than taken away: accelerator capacity and quota to procure and schedule, managed inference endpoints to make private and observable including their tokens-per-minute throttling model, an AI line that is often the fastest-growing part of the cloud bill, a much larger volume of generated infrastructure code that has to be gated by policy as code rather than reviewed by hand, and agent identities taking actions in accounts you are accountable for. In several organisations that last list is the reason a cloud engineer headcount got approved at all.
What is the difference between a cloud engineer, a DevOps engineer, an SRE and a platform engineer?
A cloud engineer owns the cloud estate itself: accounts and subscriptions, networking, identity, compute, storage, backup, patching and cost, usually expressed as infrastructure as code. A DevOps engineer is more often focused on the delivery path, meaning pipelines, build and release, and the automation between a commit and production. A site reliability engineer owns the reliability of running services through service level objectives, error budgets, incident response and the engineering work that reduces toil, and is measured on availability rather than on infrastructure delivered. A platform engineer builds an internal product that other engineers choose to use, and is measured on adoption. The titles overlap heavily and employers use them loosely, so for a cloud engineer the practical test is to read the posting for the verb and for the reporting line: into infrastructure or IT operations means cloud engineering, into engineering or platform means one of the other three.
Put this on a resume in about a minute
Paste your history once and point it at the Cloud Engineer posting you are looking at. No account, no card.
Build my resume free More roles