| Licence required: none | No licence, no mandatory certification and no legally required degree gates cloud architecture work. Two real exceptions sit at the edges. In US federal and defence contracting an active security clearance genuinely gates the job, sponsorship needs an employer, and the timeline runs in months, so put a clearance in the top third of the resume. And in a few jurisdictions the unqualified word architect is reserved by statute for licensed building architects, which is part of why postings read Cloud Solutions Architect or Cloud Infrastructure Architect rather than Architect; if you are job hunting outside the US, check how the title is used locally rather than assuming. |
|---|---|
| Four different jobs share this title | One: internal enterprise cloud architect, owning a landing zone, standards, an architecture review forum and usually a migration programme. Two: customer-facing architect at a cloud provider, a systems integrator or a managed service partner, where the work is discovery, whiteboarding, proposals and reference architectures across many accounts. Three: product or platform architect at a software company, designing the company's own system, which is often titled principal engineer instead and hires like one. Four: specialist architect for FinOps and cost, security, network, data or AI infrastructure, where the specialism is the gate. Read the posting for whose system you would be designing: your employer's internal estate, a customer's, or the product you sell. |
| Certifications: they gate one of the four jobs, not all four | At consulting partners and managed service providers a professional-level cloud certification is close to mandatory, because the employer's partner tier with the cloud provider is counted in certified staff, so the certification is a commercial requirement and the firm will usually pay for it. At a product company a certification is a filter pass at most, and a wall of badges above thin production experience actively hurts, because the panel reads it as compensation for missing scars. The useful professional-level credentials are AWS Certified Solutions Architect Professional, Google Cloud Professional Cloud Architect and Microsoft Certified Azure Solutions Architect Expert, each in the cloud the employer actually runs. |
| Time to each certification, and what each really assumes | Associate level (AWS Solutions Architect Associate, Microsoft AZ-104, Google Associate Cloud Engineer): a few weeks to a couple of months of evenings from a working baseline. Professional level (AWS SAP, Google Professional Cloud Architect, Azure AZ-305): realistically two to four months of evenings on top of real production exposure, because the questions are long scenarios with several defensible answers and the exam is scored on picking the one that fits the stated constraint. FinOps Certified Practitioner from the FinOps Foundation is short and cheap and reads well on cost-leaning postings. TOGAF matters mainly to large enterprises and to recruiter filters in Europe and India, and almost nowhere else. All cloud certifications expire and the curricula are revised, so confirm the current exam code, format and renewal window when you book rather than trusting a blog post. |
| The artefacts that prove the level | Five, and a panel will ask for at least two of them: an account or subscription topology (how many accounts, the organisational unit structure, the preventive controls, the network pattern, centralised logging, break-glass access); a migration plan with a disposition per application across the seven options of retire, retain, rehost, relocate, replatform, repurchase and refactor, plus the wave results that followed; a cost model with unit economics (cost per tenant, per transaction, per thousand requests, per gigabyte ingested) and one reduction attributed to a specific method; resilience targets per service tier with a failover you actually tested; and architecture decision records that name the rejected option and the condition that would make you revisit the choice. |
| What the loop adds beyond a senior engineer loop | A 60 to 90 minute open architecture design round, usually a migration, a multi-tenant SaaS with data residency, a landing zone for a named number of teams, or a resilience target you have to price. A past-decision deep dive in which being unable to name a design you got wrong is close to disqualifying. A money round, either standalone or folded into the design: what drives your current bill by line item, and how you would cut it without a rewrite. An influence round, because an architect usually has no direct reports and has to get standards adopted anyway. And at partner firms a mock customer session in front of people playing a client, marked on listening more than on drawing. |
| Numbers to write down before you leave a job | You cannot reconstruct these once your account is disabled: number of accounts or subscriptions and organisational units under your design; monthly cloud spend under your remit and the trend; applications migrated, in how many waves, over how many months, with downtime per cutover and whether a rollback was used; commitment coverage and utilisation for reserved capacity and savings plans; tagging or cost allocation coverage as a percentage; RTO and RPO per tier and the date of the last tested failover; number of teams that adopted the reference architecture and how many exceptions you granted; the before and after on any cost reduction with the method named. |
| Pay: name the source, not an average | The US Bureau of Labor Statistics has no cloud architect occupation code, which is why aggregator figures for this title are wide and uncheckable. Triangulate instead from BLS Occupational Employment and Wage Statistics for 15-1241 Computer Network Architects, 15-1252 Software Developers and 11-3021 Computer and Information Systems Managers, from employers' own posted ranges under state pay-transparency laws (Colorado, California, Washington, New York and Illinois among them), and from levelling data for named technology employers. One structural thing moves the number more than the title does: at a product company the architect is usually levelled as a staff or principal individual contributor with equity, while at a consulting partner the package often carries a utilisation or sales-attached component and little equity. Ask which ladder in the first conversation, because that answer sets your range before any negotiation does. |
Cloud architect is four different jobs, and the posting tells you which
Before you write a line of your resume, work out which of four jobs the posting describes, because the preparation diverges almost immediately. All four say cloud architect. They differ on whose system you design, who you have to persuade, whether a certification is a commercial requirement or a nice to have, and whether the interview contains a coding round or a customer presentation.
Shape one is the internal enterprise cloud architect. The employer is a bank, insurer, retailer, health system, manufacturer or government body with an existing estate and a migration programme somewhere in it. The posting names a landing zone, governance, standards, an architecture review board, a cloud centre of excellence, cost management, and compliance regimes such as PCI DSS, HIPAA, SOC 2 or a sector regulator. You design for internal teams who do not report to you, and the hardest part of the job is not technical: it is getting forty delivery teams to use the thing you specified while granting exceptions carefully enough that the standard survives. The screen cares about estate scale, how many applications you moved, and whether you have run governance without authority.
Shape two is the customer-facing architect, at a cloud provider itself, a global systems integrator, a regional consultancy or a managed service provider. The posting mentions pre-sales, discovery, workshops, statements of work, proof of concept, a partner practice, or billable utilisation. You see many estates rather than one, you write a lot of proposals, and the hiring process will test how you behave in front of a customer who has already decided on the wrong answer. This is the one shape where certification is close to mandatory, because partner tier with the cloud provider is counted in certified heads, and it is also the fastest way to accumulate migration experience if you do not have any.
Shape three is the product architect at a software company. The subject matter is the company's own system: multi-tenancy model, region strategy, data residency, the control plane, tenant isolation, the per-customer cost of serving, and the resilience the contract promises. At a lot of these employers there is no architect title at all and the equivalent is staff or principal engineer, which matters for your job search because searching the word architect alone hides the jobs you would most want. The loop here looks like a senior software loop with an infrastructure accent: real code or infrastructure-as-code review, a design round that goes deep on failure modes, and no tolerance for someone who has stopped touching a terminal.
Shape four is the specialist. FinOps or cloud cost architect, cloud security architect, network architect, data platform architect, AI infrastructure architect. The specialism is the gate and general cloud breadth is assumed. A cloud security architect posting is read by security people and wants identity boundary design, segmentation, key management and threat modelling. A FinOps-leaning posting wants a cost model, commitment strategy and unit economics. Applying to these with a generalist resume fails at the screen, because the interviewer is listening for the specialism's failure modes.
There is a fifth pattern worth recognising so you can avoid it: the posting that is really a senior cloud engineer job with the word architect used as a retention device, or an architect job with no decision authority, where the role produces diagrams that delivery teams ignore. Two questions in the screen separate it. Who signs off a design in this company, and what happens when a team declines to follow a standard. If nobody can answer the first and the second answer is nothing, the title is decorative.
Classification takes about two minutes. Read the responsibilities for the possessive: our platform, our customers' environments, the client, the product. Count how many compliance regimes are named, because that tells you how much of the job is paperwork. Check whether a programming language appears as a requirement or only as a nice to have. Look for the word utilisation or billable, which only ever appears in shape two. Then open the employer's other postings on the same careers page, because an organisation with a real cloud architecture function posts several related roles and names the teams they serve.
- Shape one tells: landing zone, governance, standards, architecture review board, cloud centre of excellence, migration programme, named compliance regimes, stakeholder management. Prepare estate scale, migration counts and a governance story.
- Shape two tells: pre-sales, discovery, workshop, proposal, statement of work, proof of concept, partner practice, utilisation, travel. Prepare a customer whiteboard and expect the certification requirement to be real.
- Shape three tells: our product, multi-tenancy, tenant isolation, region strategy, control plane, latency targets, cost to serve, and a language in the required list. Prepare like a staff engineer interview and search for the staff and principal titles too.
- Shape four tells: the specialism appears before or instead of the word cloud, and the tool list is narrow and deep (identity providers and policy engines for security; commitment and allocation tooling for cost; routing, peering and private connectivity for network; accelerators, schedulers and model serving for AI infrastructure).
- A warning sign: one posting asking for landing zone design, hands-on Terraform delivery, a 24/7 on-call rotation, pre-sales support and FinOps ownership is describing an understaffed function, not a broad role. Ask how many people are in the architecture function and how many engineers it serves.
- One question resolves most of it with a recruiter: is this role designing our internal estate, designing for customers, or designing the product we sell? Any decent recruiter knows, and the answer picks your preparation.
What actually gates it: experience at decision scope, not paperwork
Nothing legal gates cloud architecture. There is no licence, no protected title in IT, no required degree and no mandatory exam, and plenty of working cloud architects arrived from systems administration, networking, backend development, data engineering, support or the military. Degrees still act as filters in graduate schemes, in parts of government and defence contracting, and in skilled-worker visa routes where a degree or an assessed equivalent counts toward eligibility. Outside those, the gate is an experience gate with a specific shape: have you owned a decision that would have cost a migration to reverse, and can you defend it in detail two years later.
That is the sentence to internalise, because it explains most rejections. A candidate with eight years of excellent cloud engineering can be turned down for an architect role while a candidate with five years gets it, and the difference is usually that the second one owned the choice rather than implemented it. Panels probe for this with questions that sound simple. Who decided that? What else did you consider? What were you told to do and what did you argue for? If every answer puts the decision somewhere above you, the panel concludes you have been a very good implementer, which is a different job.
On certifications, be clear about what each one buys, because candidates waste months here. At a consulting partner or a managed service provider a professional-level certification in the employer's primary cloud is close to a hard requirement and the firm will fund it, because the partner tier they sell against is counted in certified staff. At a product company the same certificate is a filter pass and nothing more, and the engineers on the panel will not mention it. Nowhere does a certification substitute for the production scar tissue the deep dive is looking for, and a resume that leads with six badges above two years of experience invites exactly the sceptical questioning you do not want.
If you are choosing one credential, choose the professional-level architecture certification in the cloud your target employers actually run, and run it alongside real work rather than instead of it. AWS Certified Solutions Architect Professional, Google Cloud Professional Cloud Architect and Microsoft Certified Azure Solutions Architect Expert all test long constrained scenarios rather than service trivia, which makes studying for them a decent rehearsal for the design round: the skill being examined is picking the option that fits the stated constraint rather than the most impressive option. A second cloud at associate level is worth having if your employers are genuinely hybrid; a second cloud at professional level is usually a worse investment than one more real migration.
Two specialist credentials carry weight where the posting leans their way. FinOps Certified Practitioner from the FinOps Foundation is short, inexpensive and directly relevant on cost-focused postings, and it gives you the vocabulary (allocation, showback, chargeback, commitment coverage, unit economics) that hiring managers in that area use. For security-leaning architecture, CISSP still opens doors in regulated enterprises and government contracting even though it is broad rather than cloud-specific, and the cloud providers' own security specialty exams are a better technical signal. TOGAF is worth having only if you are targeting large enterprises or markets where recruiters screen on it, and it will not impress a product company.
What no certification covers is the half of the job that decides promotions: writing the design down so other people can act on it, running a review in which someone senior disagrees with you, pricing an option before it exists, sequencing a migration so the risky cutover is not the first one, and saying no to a team in a way that leaves them still willing to talk to you. Those are assessed in the interview directly and they are the reason a candidate with three months to invest is usually better off producing one real artefact than collecting a second certificate.
One more real gate, often missed: hands-on currency. The most common quiet rejection for architect candidates is the panel concluding that the person has not touched the platform in three years. Interviewers test it fast, with a question that only a practitioner answers crisply. What does a NAT gateway cost you and how would you avoid it. What breaks first when you put a managed database behind a private endpoint. Read this Terraform plan and tell me which line is destructive. Keep enough hands in the system to answer those without hedging, even if you no longer build day to day.
- The gate in one sentence: you must have owned at least one decision whose reversal would require a migration, and be able to defend it two years later with the numbers attached.
- Worth doing if you target partners and consultancies: the professional-level architecture certification in their primary cloud, because it is a commercial requirement there rather than a signal.
- Worth doing if the posting leans at cost: FinOps Certified Practitioner, measured in days rather than months, mainly for the shared vocabulary.
- Worth skipping: a second professional-level cloud certification, generic Agile and Scrum certificates, and vendor badges for tools the employer does not use. None of these appear in an architect hiring decision at a product company.
- Worth reading instead, because interviewers use the vocabulary and expect you to know it: the AWS Well-Architected Framework pillars, the Microsoft Azure Well-Architected Framework and Cloud Adoption Framework, the Google Cloud Architecture Framework, the FinOps Foundation's framework, and the four DORA delivery measures. These are free and they are the language of the review conversation you will be asked to run.
- Put an active security clearance in the top third of the resume with the level and the year of last investigation, because for cleared roles it is the first filter applied and the absence of it ends the conversation regardless of your design skill.
- Keep one honest line about your hands-on currency, such as the last thing you personally built or debugged and when. Volunteering it is far better than being caught by a probe.
The five artefacts that prove architect level
If you want to step up to architecture, stop thinking about years and start thinking about artefacts. There is a short, stable list of documents that a panel treats as proof. Each one is something you can produce in your current job without permission, and each one answers a question an interviewer will ask you anyway. Produce them, keep sanitised copies, and the deep dive stops being a memory test and becomes a walkthrough.
Artefact one is the account and network topology. Not a pretty diagram: the decisions. How many accounts or subscriptions and on what boundary (environment, business unit, blast radius, billing), the organisational unit structure, which preventive controls are enforced centrally and which are advisory, the network pattern and why (hub and spoke, transit gateway, cloud WAN, or deliberately flat because the estate is small), the IP address plan and who owns it, how private connectivity to on-premises works and what happens when it fails, where logs and audit trails land and who can read them, how break-glass access works and who signs for it, and how a new team gets a working account in a day rather than a month. The strongest version of this artefact includes the thing you chose not to centralise and why.
Artefact two is the migration plan with the results next to it. An application inventory with a count, a disposition per application across the seven standard options (retire, retain, rehost, relocate, replatform, repurchase, refactor), the dependency mapping method you used, the wave sequence with the reasoning for what went first, the cutover runbook with its rollback point and the go/no-go criteria, and the data migration method per class of system. Then the part most candidates omit: what actually happened. How many applications moved, over how many months, how many waves slipped and why, downtime per cutover against what you promised, whether you ever used the rollback, what you retired instead of moving, and the run cost before and after. A plan with no outcome reads as a slide. A plan with outcomes, including the wave that went badly, reads as experience.
Artefact three is the cost model. Two halves. The first is a model of what a design costs before it is built, broken into the lines that actually move: compute with a commitment assumption, storage by tier and growth rate, data transfer including cross-zone and egress, managed service minimums and idle cost, logging and observability ingest, licensing, and support. The second is unit economics, which is the thing that marks you as an architect rather than a cost cutter: cost per tenant, per active user, per transaction, per thousand requests, per gigabyte ingested, per model call. Then one reduction you can attribute to a method rather than to a season: a rightsizing exercise with the measurement that justified it, commitment coverage raised from one figure to another, a storage lifecycle policy, an architectural change that removed cross-zone chatter or an egress path, idle environments shut down on a schedule, or a managed service swapped because the minimum was larger than the workload. Always say the base, the method and the period, because a saving without a base is not a number.
Artefact four is the resilience specification with a test behind it. RTO and RPO stated per service tier rather than per company, the failure domains you designed against (instance, zone, region, provider, dependency, human error, corrupted data), the recovery mechanism per tier, and the result of the last time you tested it with the date. The most persuasive version includes a region strategy you argued against: you looked at active-active across regions, priced it, found the business tolerance was two hours rather than two minutes, and built multi-zone with tested restore instead. Panels mark that higher than a multi-region design nobody asked for, because it shows you let the requirement pick the architecture.
Artefact five is a set of architecture decision records. One page each, in a consistent form: the context and the constraint, the options considered, the decision, the consequences accepted, and the condition that would make you revisit it. Write them for decisions your team is making right now, even small ones, even if nobody asked. Six months of this gives you exactly the material a past-decision deep dive wants, in your own words, with the rejected options already written down. It is the single highest-return habit for someone trying to move into architecture, and it costs an hour a week.
A sixth artefact is worth having if your target is shape one: the adoption record for a standard you wrote. A reference architecture or a set of approved modules, how many teams adopted it, how long adoption took, how many exceptions you granted and on what grounds, and whether you ever retired the old pattern. Governance that exists only as a document is not evidence. Governance with an adoption number and an exception log is.
- Sanitise, do not skip. Replace customer names, internal hostnames, IP ranges and absolute spend with ratios and orders of magnitude, then say out loud that you have sanitised it. Panels respect that; they do not respect a candidate who brings a client's live diagram.
- Each artefact needs a count, a unit and a date. Twelve accounts, four organisational units, 180 applications, nine waves, eleven months, 40 minutes of downtime per cutover, RPO of five minutes for tier one, last tested in March.
- Keep one artefact per shape. A landing zone and an adoption record for enterprise roles, a migration plan and a customer-facing design for partner roles, a multi-tenancy and cost-to-serve design for product roles.
- Write the decision you got wrong into the artefact rather than hiding it. Every panel asks, and an answer with the correction and the cost attached outperforms a flawless record that sounds rehearsed.
- If you do not have these yet, pick the one your current job already contains. Almost every team has an undocumented cost problem, an untested restore, and a decision nobody wrote down. Those are three artefacts sitting in plain sight.
- Do not bring a slide deck to a technical round unless asked. Bring the numbers in your head and draw on the whiteboard. Architects who can only present are a recognised failure mode and interviewers screen for it.
How the loop actually runs, and who screens you
The process for an experienced architect role runs two to five weeks at a product company or a partner, and routinely six to twelve weeks at a large enterprise or in government, where an architecture appointment goes through a panel, sometimes a board, and occasionally a written submission. Expect four to six conversations. The variation between the four shapes is mostly in who is in the room, not in the stages themselves.
The first screen is a recruiter and it is a matching exercise, not an assessment. They are checking the cloud (an estate on Azure will usually not shortlist a pure AWS architect for an internal role, however transferable the thinking is), the scale, the industry if it is regulated, the certification if the employer is a partner, location and clearance, and the salary band. Give them crisp numbers in the first two minutes: the cloud, the spend under your remit, the number of applications or accounts, the regulated regime if any, and the size of the biggest migration you have run. Recruiters pass on candidates they can summarise.
The second conversation is the hiring manager, and it is really about scope and authority. They want to know what decisions you have made rather than what systems you have touched, and they are deciding whether your scope matches theirs. This is the right place to ask your own scope questions: who signs off a design here, is there an architecture review forum and does the architect chair it, does the role carry delivery accountability or advisory only, how many teams does the function serve, and is the estate mid-migration or post-migration. The answers change the job more than the title does.
The third stage is the architecture design round and it is the one that decides most outcomes. Sixty to ninety minutes, open prompt, a whiteboard or a shared drawing tool. The prompts repeat across employers: migrate an estate of a named size within a named deadline; design a multi-tenant SaaS for customers who require data residency in three regions; design the landing zone for an organisation of forty engineering teams; take a system and meet a stated RTO, then tell me what it costs; or our bill tripled in a year, find it. The next section covers what reviewers are marking, because this round is scored on behaviour at least as much as on the design that ends up on the board.
The fourth stage is a deep dive on your own past work, often with the most senior person you will meet. It goes from the diagram down to a specific failure: why that database, what you did when the cutover window closed early, what the design cost, what you would do differently. Being unable to name a design you got wrong is close to disqualifying at this stage, because every architect with real mileage has one, and a candidate without one is either junior or not candid. Prepare one properly: the decision, the constraint you misread, the symptom that revealed it, the correction, the cost of the correction, and the rule you adopted afterwards.
The fifth stage varies by shape and is where candidates are least prepared. At a product company it is an engineering round with code or infrastructure as code: review a Terraform plan and name the destructive line, write a small script, read someone else's module and say what is wrong with it. At a consulting partner it is a customer simulation, where two or three people play a client with a budget, a deadline and a preferred wrong answer, and you are marked on whether you listened before drawing. In regulated enterprises it is a compliance and security round about identity boundaries, data classification, encryption and key custody, logging retention and evidence for auditors. In cost-leaning roles it is an explicit money round. Ask the recruiter what the stages are; almost all of them will tell you, and the answer changes what you rehearse.
The final stage is usually a stakeholder or executive conversation, with an engineering director, a CIO or CTO, or a business owner. It is not technical. They are testing whether you can explain a trade-off to someone who controls the budget without jargon and without pretending the risk away. Prepare three sentences that explain your biggest design in business terms: what it let the company do, what it cost, and what you accepted as a risk. If you cannot do that, this round goes badly however well the design round went.
- Typical stage list: recruiter screen, hiring manager scope conversation, architecture design round (60 to 90 minutes), past-design deep dive with a principal or head of architecture, a shape-specific round (code and IaC, customer simulation, security and compliance, or cost), then an executive conversation.
- Enterprise and public sector additions: a panel with set questions and scoring, sometimes a written design submission or a presentation to a review board, and references taken seriously. Timelines run in weeks rather than days and silence between stages is normal rather than a signal.
- Partner and consultancy additions: a mock customer whiteboard, a proposal or statement of work exercise, and questions about utilisation, travel and how you behave when the client is wrong.
- What the recruiter screen is really filtering: the cloud, the regulated industry, the certification if it is a partner, the clearance, the location, and whether your scale is within an order of magnitude of theirs.
- Ask for the design round format in advance: whiteboard or shared document, whether you may ask questions throughout, and whether cost is in scope. All three change your preparation and asking reads as professional rather than nervous.
- Bring your own questions to every stage and make at least one of them about authority. Who signs off a design, what happens when a team declines a standard, and what the last exception granted was. Candidates who ask these sound like architects; candidates who ask only about technology sound like engineers.
- If a process has no design round at all, be curious rather than relieved. An architect role with no design assessment often means the function has no decision authority, and you should ask directly what the last significant decision the architecture team made was.
The resume: scale, decisions and money, in that order
An architect resume fails differently from an engineer resume. The engineer version fails on missing tools. The architect version fails because the reader cannot tell what you decided or at what scale, so there is nothing to calibrate against. Every bullet should let a reader answer two questions without asking: how big was this, and what did you personally choose.
Open with a four or five line summary that is pure calibration, not adjectives. The cloud or clouds and the depth in each, the scale under your remit (accounts or subscriptions, applications, monthly spend as an order of magnitude, users or tenants, regions), the industries and any regulated regime you have delivered under, the biggest migration you have run stated as a count and a duration, and the specialism if you have one. A reader should be able to place you in a band in fifteen seconds. If you have a clearance, it goes here.
Then write bullets around decisions with numbers in front. The pattern that works: the decision, the scale, the constraint that drove it, the measured outcome. Designed the account topology for a 40-team organisation on a blast-radius boundary, 14 accounts and 4 organisational units with preventive controls enforced centrally, cutting new-team provisioning from six weeks to two days. Dispositioned 180 applications across the seven migration options, retiring 31 of them, and moved the remainder in nine waves over eleven months with a maximum cutover downtime of 40 minutes against a promised two hours. Rejected active-active multi-region after pricing it against a two-hour business tolerance and delivered multi-zone with quarterly tested restore, keeping the resilience budget inside the agreed figure. Each of those tells a reader what you own.
Money belongs on an architect resume and most candidates leave it off. Monthly or annual spend under your design is a scale signal, and reductions are an outcome signal, but a reduction without a method and a base is not credible. Write the method. Raised savings plan and reserved capacity coverage from a low fraction to most of the steady-state fleet. Removed a cross-zone data path that accounted for a large share of transfer cost. Introduced cost allocation tagging and took allocation coverage from partial to near complete, which is what made the rest possible. If you cannot share absolute figures, use ratios and say explicitly that the figures are ratios because the employer was confidential. Panels accept that; they do not accept vagueness.
Know what gets ignored or actively counts against you. A thirty-service logo list proves nothing, because every cloud architect has touched most of them; keep one compact skills block at the bottom for the keyword screen and nothing more. Claiming expert across AWS, Azure and Google Cloud reads as shallow and invites a brutal depth probe in whichever one you know least. A certification wall above a thin experience section reads as compensation. Phrases like designed scalable, highly available, secure solutions are filler, because nobody designs unscalable insecure ones on purpose. And a single lift and shift from years ago described as a cloud transformation dates you badly; if your most recent architecture work is old, lead with what is current.
Tailor by shape, because the same career supports different documents. For enterprise roles, lead with estate scale, migration counts, governance and adoption, and name the compliance regimes. For partner roles, lead with the number and variety of client environments, the deals or proposals you supported, the certifications, and the industries. For product roles, lead with the system you designed, the tenancy and region model, the latency and availability targets you actually hit, the cost to serve, and the fact that you still write code. For specialist roles, lead with the specialism's own metrics and let the general cloud work sit underneath.
Two practical mechanics. Applicant tracking systems still parse plainly, so keep it to a single column with real headings, no text inside images, and put the literal job title you are applying for somewhere in the document if it honestly describes you, because Cloud Architect, Cloud Solutions Architect and Cloud Infrastructure Architect are searched as different strings. And keep a one-page architecture portfolio separate from the resume: three sanitised artefacts, one page each, that you can send when a hiring manager asks for an example. Very few candidates have one, and it ends the question of whether you have really done the work.
- Lead the summary with calibration: clouds and depth, accounts or subscriptions, applications, spend order of magnitude, tenants or users, regions, regulated regimes, biggest migration as a count and a duration.
- Bullet pattern: decision, scale, driving constraint, measured outcome. If a bullet has no number and no decision, it is a duty, not evidence, and it should be cut.
- Always include at least one rejected option. Chose X over Y because of Z constraint is the single most architect-sounding sentence on a resume, and it survives every follow-up because you already know the reasoning.
- State cost work with a method and a base, or as ratios if the employer was confidential. Say which it is.
- Cut: thirty-service logo lists, expert in three clouds, designed scalable and secure solutions, responsible for, and any certificate older than its renewal window presented as current.
- Name the compliance regimes you have actually delivered under (PCI DSS, HIPAA, SOC 2, ISO 27001, FedRAMP, a sector regulator) rather than claiming compliance experience generally, because the screen for regulated roles searches for the specific regime.
- Keep a separate one-page architecture portfolio of three sanitised artefacts, ready to send the moment a hiring manager asks for an example of your work.
The architecture design round: what reviewers are actually marking
The design round is scored on behaviour as much as on the diagram, and the marking is more consistent across employers than candidates expect. Interviewers are watching for a sequence: did you establish what is actually required before designing, did you name the constraint that drives the shape, did you consider and reject a cheaper option out loud, did you reason about failure and blast radius, did you price it at least roughly, did you say what you would not build, and did you finish with a sequence for getting there from the current state. Candidates who jump to a diagram in the first three minutes lose most of those marks regardless of how good the diagram is.
Spend the first five to ten minutes on requirements and write them on the board where everyone can see them. Who uses this and how many of them. Request rate at peak and the shape of the peak. Data volume now and growth per year. Latency target and at which percentile. RTO and RPO, per tier if the tiers differ. Compliance regime and data residency constraints. Budget, or at least whether cost or speed is the binding constraint. The deadline and what happens if it slips. And the team: how many engineers, what they already run, what they are on call for. That last one is a legitimate design input and most candidates never ask it, which is exactly why asking it stands out. A design that requires a Kubernetes platform team at a company with four engineers is a wrong answer however elegant.
Then say the constraint out loud before you draw. This design is driven by the residency requirement, so I am starting from the region topology. Or this is driven by the eleven-month deadline, so the first wave has to be rehost and the refactoring comes later. Naming the driver tells the interviewer you are optimising for something real, and it gives the rest of the hour a spine. Interviewers write that sentence down.
Reject something cheaper, explicitly. This is the most reliable single way to look senior. The cheaper option is almost always one of: a single region instead of several, multi-zone instead of multi-region, a managed service instead of a self-run one, one shared database instead of per-tenant isolation, a scheduled batch instead of streaming, a queue instead of an event platform, rehost now and replatform later, or buying instead of building. Price it, say what it fails to deliver against the requirements you wrote on the board, and only then take the more expensive path. A candidate who arrives at the expensive design without visibly considering the cheap one reads as someone who will overspend the company's money.
Price the design in line items, not in a total. Say which lines dominate: compute and whether you assume commitments, storage by tier and growth, data transfer across zones and out to the internet, managed service minimums and anything that charges while idle, logging and observability ingest (which surprises people more often than compute does), licensing, and support plan. Then say the order of magnitude and that you would verify against current pricing rather than quoting a figure from memory. Saying I would check current rates is a strength here, not a weakness, because inventing a price is a tell that you have never owned a bill. Interviewers mainly want to see that you know which lines move.
Reason about failure concretely rather than saying highly available. Pick a component and walk it: this zone disappears, what breaks, what is the user impact, what recovers automatically and what needs a human, how long. Then do the awkward ones that candidates skip: the shared dependency that takes everything with it, a bad deploy, a credential expiring, a quota being hit, a region-level provider outage, and corrupted or deleted data, which is the failure a multi-region design does not help with at all. Say where the blast radius boundary sits and which account, subscription, network or key boundary enforces it. Then say what you would monitor to know it is working and what alert would page a human.
Finish with the two things most candidates run out of time for, which means you should budget for them: what you are deliberately not building, and how you get there from where the company is today. The refusal list is short and specific: not multi-cloud, not multi-region in phase one, not a service mesh at this team size, not per-tenant infrastructure until a customer pays for it, not a custom control plane when a managed one exists. The migration sequence is a phase plan with a first wave chosen because it is low risk and informative, a point at which you would stop and re-measure, a rollback position, and the first metric you would look at after go-live. Ending there, rather than on the diagram, is what separates a design from an architecture.
- Common prompts: migrate an estate of N applications in M months; design a multi-tenant SaaS with residency in three regions; design a landing zone for forty teams; meet a stated RTO and price it; the bill tripled, find it; connect three clouds and a datacentre; design for a regulated workload under a named regime.
- First ten minutes, out loud: users, peak rate, data volume and growth, latency target and percentile, RTO and RPO per tier, compliance and residency, budget or the binding constraint, deadline, and the size and skills of the team that will run it.
- Say the driving constraint before you draw, then design from it. Interviewers record that sentence as evidence of senior reasoning.
- Always reject a cheaper option explicitly and say what it failed to deliver. This is the single highest-value behaviour in the round.
- Price in line items and name the dominant ones, then say you would verify current rates. Do not invent per-unit prices.
- Walk at least three concrete failures including data corruption or deletion, which multi-region does not solve, and name the boundary that contains each blast radius.
- Reserve the last ten minutes for the refusal list and the migration sequence, with a rollback position and the first metric you would check after go-live.
- Keep the whiteboard legible and labelled, and restate the design in four sentences at the end. A reviewer who has to reconstruct your design from an unlabelled diagram scores communication down, and communication is half the job.
The questions that decide it, and how to answer them
A handful of questions recur in almost every cloud architect loop, and the strong answers have a shape in common: a number, a decision, a trade-off named honestly, and a consequence you lived with. Rehearse these out loud, because the failure mode is not ignorance, it is rambling.
Walk me through an architecture you own, end to end. Give the shape in four sentences before any detail: what the system does, the scale, the two or three decisions that define it, and the one thing you would change. Then let them steer. The mistake is starting at a component and narrating upward, which uses twenty minutes and never reaches a decision. Have the numbers ready: requests, data volume, spend, availability achieved, team size.
Tell me about a design decision you got wrong. Answer with the constraint you misread rather than with bad luck. I chose a single shared database for tenant data because the first eight customers were small, and I misjudged how fast one would grow; the symptom was noisy-neighbour latency on the largest tenant within a year; we extracted that tenant to dedicated capacity in a quarter, which cost roughly one engineer quarter of unplanned work; since then I ask for the expected distribution of customer size, not the average, before choosing a tenancy model. Constraint, symptom, correction, cost, rule. Never answer that you cannot think of one.
What drives your current cloud bill? This question separates architects from diagram producers in about forty seconds. You should be able to name the top three or four lines in rough proportion, say which of them is growing fastest, name the cheapest thing you have not yet fixed and why, and say what your commitment coverage is. If you genuinely do not know, say that you do not own the bill and then say what you would look at first, in order: allocation coverage, the largest line item, idle and orphaned resources, data transfer paths, logging ingest, and commitment coverage against the steady-state fleet.
When would you not go multi-region, and when is multi-cloud right? Both of these are maturity tests disguised as architecture questions. The honest answer on regions: multi-region costs you duplicated capacity, data replication, replication lag you have to design around, a much harder consistency story and a traffic management layer, so it is justified by a business tolerance measured in minutes, by residency requirements, or by latency to distant users, and not by a general wish to be resilient. Most organisations get better availability per pound from multi-zone with tested restore. On multi-cloud: portability costs you the managed services that make a cloud worth using and doubles the identity, networking, observability and skills surface, so the defensible versions are specific rather than general. An acquisition you have to run, a regulator or customer contract that requires it, a single workload placed deliberately where it runs best, or a commercial lever in a negotiation. Choosing multi-cloud as an abstract hedge is the answer that costs people this round.
How do you get teams to follow a standard when they do not report to you? Answer with a mechanism rather than with influence as a personality trait. Make the compliant path the easiest path, so an approved module or a provisioned account does more than a document ever will. Enforce the small number of things that are genuinely non-negotiable with preventive controls rather than review, because a control that cannot be bypassed does not need persuading. Make everything else advisory with a visible exception process that has an owner, an expiry date and a recorded reason, because an exception with an expiry is a negotiation and an exception without one is a permanent fork. Then show the adoption number, which is what tells the interviewer this actually happened.
What would you do in your first ninety days? Resist the urge to propose a target architecture. The answer that lands: read the bill and the account structure, find out what is already decided and by whom, list the top risks by what would hurt most if it failed, test one restore, talk to the delivery teams about what currently slows them down, and only then publish a short set of decisions with the reasoning. Then name the one thing you would fix immediately regardless, usually an identity or blast-radius problem or an untested backup, because an architect who changes nothing in a quarter is not useful and one who redesigns everything in a quarter is dangerous.
- Have ready, with numbers: one system you own end to end, one decision you got wrong, one cost reduction with a method, one failover you tested, one standard you got adopted, and one time you said no to a senior stakeholder and what happened next.
- Rehearse the four-sentence version of your biggest design, and a non-technical version of the same thing for the executive round. Both get asked.
- Expect at least one hands-on probe: read a Terraform plan and name the destructive line, explain why a private endpoint broke name resolution, say what a NAT gateway costs you and how to avoid it, or explain the difference between a security group and a network access control list and when each one bites.
- Expect one identity question with real depth: how you would grant a workload access across accounts without a long-lived key, where you draw the permission boundary, how you handle break-glass, and how you would detect a role that has quietly become too wide.
- Ask your own authority questions, and ask them in every loop: who signs off a design, does the architecture function chair its own review, what was the last exception granted, and is the estate mid-migration or post-migration.
- If you are asked a question about a service you have not used, say so in one clause and reason from the primitives it is built on. Interviewers mark bluffing far harder than gaps, and architects are expected to meet unfamiliar services weekly.
Stepping up without the title yet, and reading the offer
Most cloud architects were cloud, platform, infrastructure, network or backend engineers first, and the move is rarely a clean external jump. The gate is having owned a decision at architect scope, and the cheapest place to acquire that is the job you already have, which is why an internal move is usually easier than an external one: your employer can already see the decisions you made, while an external panel can only hear you describe them.
Start by producing architect output without the title. Write architecture decision records for the choices your team is making this quarter and circulate them, because the act of writing options and consequences down changes how people treat you within a few weeks. Volunteer for the work nobody wants that is secretly architecture: the cost review, the disaster recovery test, the vendor evaluation, the security exception nobody wants to own, the production readiness review, the capacity forecast, the cross-team design review. Each one produces an artefact and a story. Then ask the current architect to co-author the next reference architecture with you, which is both a mentorship and a credential.
Pick the entry route that matches your situation honestly. If you are at a large organisation with an architecture function, the internal route is: artefacts, then a visible design someone senior sponsored, then the first vacancy. If you are at a small company with no architect title, you are probably already doing parts of the job and the issue is evidence, so document what you decide and target scale-ups rather than enterprises, because an enterprise panel wants scar tissue you cannot have yet. If you have narrow depth and need breadth fast, a consulting partner or managed service provider is the most efficient accelerator available: several migrations a year, many estates, funded certifications, and a two-year resume that an internal role takes six years to build. The cost is real and you should go in knowing it: utilisation pressure, travel or at least client hours, less depth per system, and the risk of becoming fluent in proposals and rusty in production.
If you are coming from outside infrastructure, there are two viable bridges and one dead end. From software engineering, the bridge is owning the deployment, cost and resilience of a system you wrote, then generalising it. From networking or systems administration, the bridge is identity and the landing zone, because the organisational boundary work is closer to what you already do than application design is. The dead end is certifications with no production exposure: it produces candidates who can describe services and cannot answer what happened the one time it broke, and panels identify them in the first ten minutes.
On timing, be realistic about the market you are applying into. Cloud migration is no longer new, which changes the mix of jobs: there are fewer greenfield programmes and more work in optimisation, consolidation, cost reduction, repatriating workloads that were wrongly placed, regulated and sovereign requirements, and building the infrastructure for AI workloads. That is good news for anyone with operating and cost evidence and bad news for anyone whose only story is a migration finished years ago. If your experience is mostly the initial move to cloud, add one current artefact before you apply: a cost model, a tested restore, a consolidation, or an accelerator placement decision.
When an offer arrives, read the role rather than the title. Ask whether there is an architecture review forum and whether this role sits on it. Ask who can overrule a design and how often that has happened. Ask whether the role carries delivery accountability or is advisory, because advisory architects in organisations that do not want advice have short tenures. Ask the spend under the function's remit and the number of engineers it serves, which together tell you the real scope. Ask whether the estate is mid-migration or post-migration, because a post-migration architect job is mostly governance and cost, and some people love that and some people leave within a year. At a partner, ask the utilisation target and whether architects carry pre-sales expectations, and get the answer as a number. And ask what the last significant decision the architecture function made was, because the answer tells you more about the job than the job description does.
One last check before you sign: find out whether the title sits on the engineering ladder or in a separate architecture band, because that determines your next promotion and often your equity. At a product company an architect is usually a staff or principal individual contributor and the path onward is clear. In an enterprise architecture band the path onward may be management or nothing, and the band tends to pay less than the equivalent product-company level. Neither is wrong, but find out which one you are joining before you negotiate, not after.
- Four weeks of unglamorous work that changes an interview: write three architecture decision records for choices already being made, run one cost review and produce a model with unit economics, test one restore and write the result with a date, and document the account topology you already have including what you would change.
- Volunteer for these specifically: cost reviews, disaster recovery tests, vendor evaluations, security exceptions, production readiness reviews, capacity forecasts, cross-team design reviews. They are architecture work in disguise and nobody competes for them.
- Fastest breadth accelerator: a consulting partner or managed service provider, for the number of estates and the funded certifications. Known costs: utilisation pressure, travel, less depth, and the risk of drifting out of production.
- Search the titles that hide the job: Cloud Solutions Architect, Cloud Infrastructure Architect, Principal Cloud Engineer, Staff Infrastructure Engineer, Platform Architect, Solutions Architect, Enterprise Architect (cloud). At product companies the role you want is often advertised as principal engineer.
- Where the roles are in 2026 and 2027: optimisation and cost, consolidation, workload repatriation, sovereign and regulated estates, and infrastructure for AI workloads. Fewer pure greenfield migrations than three years ago.
- Offer questions worth more than the salary conversation: who signs off a design, what was the last exception granted, is this ladder engineering or a separate architecture band, is the estate mid-migration or post-migration, what is the spend under this function, and at a partner what is the utilisation target as a number.
What a cloud architect has to know about AI in 2026-27
Start with an honest calibration, because for this role the hype and the observable change point in different directions. Nothing has automated the core of cloud architecture. Eliciting requirements from a business that has not agreed with itself, choosing between two defensible designs on cost and risk, sequencing a migration so the dangerous cutover is not the first one, and taking personal accountability for a decision that will be expensive to reverse are not tasks an assistant performs. Autonomous remediation in production, sold under the AIOps label for most of a decade, is still uncommon in practice and is not what employers are buying. If you go into an interview claiming AI has transformed cloud architecture, the panel will hear someone who has read about the job rather than done it.
What has genuinely changed is specific, and a cloud architect is now expected to have an opinion on five things that barely existed as architecture concerns three years ago. The first is the return of capacity as a real constraint. For a decade the operating assumption was that compute is infinite and placement is a cost question. Accelerator capacity broke that: availability is region-specific and quota-gated, the instance types you want may not exist in the region where your data has to stay, lead times for committed capacity are real, and at the extreme end power and cooling at a provider's site is the thing rationing supply. The architectural consequences are concrete. Placement is now driven by where capacity exists as well as by latency and residency, which can conflict directly with a data residency requirement and force an explicit decision. Quotas and reservations become design inputs rather than operational details. Workloads need a queue and a scheduler rather than an assumption of on-demand availability, and batch and training work needs to be interruptible. And committed capacity shifts your cost model from variable to fixed, which changes who must approve it.
The second is that inference cost became a first-class line in the cost model, and it behaves more like data transfer than like compute: driven by volume and by choices made in application code, invisible until someone adds it up, and capable of growing faster than the revenue it supports. The architect is usually the person asked to put structure around it, and the structure is familiar once you see it as a gateway problem. A single path to model providers rather than keys scattered in services. Per-team and per-feature credentials so spend can be attributed. Rate limits and hard spend caps that fail safely rather than silently running up a bill. A caching layer for repeated requests. Routing between a small cheap model and a large expensive one by task rather than by default. A fallback position for a provider outage, which is now a dependency in your availability calculation whether or not it appears on your diagram. And unit economics for the feature: cost per request, per document, per customer, per resolved ticket. Being able to say what your AI features cost per unit, and which design choice moved that number, is currently a strong differentiator in interviews because many candidates have never been asked.
The third is that sending data to a model is a data transfer decision, and it lands squarely in the architect's remit. The questions are the ones you already know how to answer in a different setting: where does the data physically go, who else can see it, how long is it retained, is it used to train anything, what classification of data is permitted on that path, and what is the audit trail afterwards. The design choices are real and they have different cost and control profiles: a provider's public API, a model service running inside your own cloud account and region, an open-weights model you host yourself on your own accelerators, or a mix with routing by data classification. Regulation exists here and it is becoming more specific about documentation, data governance, human oversight and logging for higher-risk uses, alongside existing data protection and sector rules. Treat the dates with care in an interview: obligations and deadlines in this area have been amended since they were first published, so state the obligation and say that the current timing should be confirmed with counsel rather than asserting a month. Saying that is read as competence, not evasion, and getting a deadline wrong in front of a compliance-minded panel is a hard failure.
The fourth is non-human identity, which is the newest thing appearing in cloud architect interviews and the one fewest candidates have prepared for. Coding agents now open pull requests, run pipelines, call internal APIs and ask for environments and credentials, and every cloud control model was designed on the assumption that a named human made each change. The architect owns the answer. A distinct identity per agent or workflow rather than a shared service account. Credentials minted short-lived from federation rather than stored. Permissions scoped to specific resources, with a policy layer that refuses destructive actions outright and requires a human approval for a listed set. Sandbox accounts or subscriptions that an agent can create and destroy without touching production. Audit logs that distinguish an agent's action from a person's, so that after an incident someone can actually say what changed production. And a standing answer to the question of how a convenience credential, created to unblock an agent on a Friday, is prevented from becoming the widest permission in the account. If you can discuss this concretely you are answering a problem a lot of hiring managers are living with right now.
The fifth is that generated infrastructure code changed where the risk sits. Producing Terraform, Bicep and manifests is now nearly free, so the volume of configuration change has gone up, the variance in its quality has gone up with it, and the bottleneck has moved to approving it safely. The characteristic defect is not generated code that fails. It is generated code that works and is quietly wrong, because the generator does not have your account state, your traffic shape or your invariants: a plan that replaces a database rather than updating it, a security group opened to the world to fix a connectivity problem, a wildcard permission because the narrow one did not work first time, a resource re-indexed by a changed loop. The architect's leverage is to make the class impossible rather than to review every instance, which means preventive controls over detective ones: organisation-level policy, approved and versioned modules, policy as code in the pipeline, required plan review on stateful resources, drift detection, and guardrails that reject the wrong shape before it reaches an account. Interviewers ask how you keep quality when volume rises, and that is the answer.
There is a sixth thing that is less about technology and more about the role. Architects are now frequently the person who has to replace a shadow AI stack with a sanctioned path, which is an old job in new clothes: find out what teams are already doing, provide a route that is easier than the one they improvised, enforce the few things that are genuinely non-negotiable, and grant exceptions with expiry dates. If you have done that for any technology, the story transfers directly, and telling it well is better than claiming expertise in a toolchain you used twice.
Designing around accelerator capacity, quota and placement
This is the clearest reversal of a decade-old assumption in cloud architecture. Capacity is no longer effectively infinite for the workloads that matter most to employers right now, and that turns placement, quota and commitment into design inputs rather than operational detail. It also creates genuine conflicts an architect has to resolve explicitly: the region with capacity may not be the region your data is allowed to sit in, and the committed contract that secures capacity turns a variable cost into a fixed one that someone senior has to approve. Hiring managers building anything with a model in it are living this constraint and notice immediately when a candidate has not met it.
Show it: Be able to describe one real placement decision with the conflict in it: the region you wanted, the capacity or quota that was not there, the residency or latency requirement that pulled the other way, and what you chose. Then show the mechanics you used: quota requests and their lead times, reserved or committed capacity and what coverage you held, a queue and scheduler so work waits instead of failing, checkpointing so training can be interrupted, and cost attribution per team for shared accelerators. On the resume, write the decision and the constraint, not the instance type.
Putting inference spend inside a cost model with unit economics
Model spend is a variable cost driven by application choices, it is usually invisible until someone totals it, and it can outgrow the revenue of the feature it powers. Architects are the people asked to make it predictable and attributable, and the pattern is a gateway rather than a novelty: one path out, per-team keys, limits, caching, routing by task, and a fallback for a provider outage. Very few candidates can state the cost per unit of anything they have built, so being able to is a differentiator rather than a baseline.
Show it: Bring one unit figure and the design choice that moved it: cost per request, per document processed, per customer, per resolved ticket, before and after. Describe the controls you put in place, specifically hard spend caps that fail closed, per-team attribution, a cache with a measured hit rate, routing a cheap model for the common case and an expensive one for the hard case, and what happens to your availability number when the provider is down. If you have not owned this, say so and describe what you would measure first: spend by feature, requests by model, cache hit rate, and the share of calls that did not need the large model.
Deciding where data may go when a model is involved
Calling a model is a data transfer with a destination, a retention policy and a training-use question attached, which makes it an architecture decision rather than a procurement one. It is also the area where a candidate can do real damage by sounding confident: regulation here is specific about documentation, data governance, human oversight and logging for higher-risk uses, and the timing of obligations has been amended since first publication. The architect's value is drawing the boundary clearly and keeping the evidence, not reciting a deadline.
Show it: Be able to lay out the options with their trade-offs: a provider's public API, a managed model service inside your own account and region, self-hosted open weights on your own accelerators, and a routing policy that sends data to different paths by classification. Say which classes of data are allowed on which path and what enforces it, how retention and training use are contracted, and where the audit record lives. State obligations without inventing dates and say that current timing should be confirmed with counsel. One concrete artefact beats any amount of theory: a decision record for a model choice that names the data classification and the rejected option.
Giving agents an identity, a sandbox and a limit
Every cloud permission model assumes a named human made the change, and agents break attribution, least privilege and the review model at once. The architect owns the control points where this is fixed: identity, credential lifetime, policy admission, environment creation and the audit trail. Done badly it produces the widest permission in the account and no way to say afterwards what changed production. This is the newest requirement showing up in interviews for the role and the one where prepared candidates stand out most sharply.
Show it: Describe a concrete design: a distinct machine identity per agent or workflow rather than a shared account, short-lived credentials minted from federation rather than stored, scopes narrowed to named resources, a policy layer that denies destructive actions and requires human approval for a listed set, sandbox accounts an agent can create and destroy, and logs that separate agent actions from human ones. Have one real story even if it is small: the static key you replaced with federation, the deny rule you added, the sandbox you stood up. Write the guardrail and who it protected, not the tool name.
Reviewing generated infrastructure changes, then making the class impossible
Generated configuration is now cheap and plentiful, so the risk moved from writing it to approving it. The dangerous output is not the change that fails; it is the change that succeeds and is quietly destructive, because the generator does not know your account state or your invariants. An architect who only reviews harder does not scale. The answer that marks you as senior is turning each class of defect into something a control prevents, which is the same instinct that makes a landing zone work.
Show it: Practise reading a Terraform or Bicep plan out loud and naming the dangerous lines: destroy and replace on a stateful resource, forces replacement, a loop change that re-indexes resources, a rule opened to the whole internet, a wildcard action, a removed deletion protection, a dropped dependency. Bring one story of a change you blocked and the outage it would have caused. Then go one level up and name the preventive control you added so the class could not recur: organisation policy, an approved module, a policy-as-code check in the pipeline, required review on stateful changes, or drift detection with an owner.
Replacing an improvised AI stack with a sanctioned path
In most organisations teams started using models before anyone designed for it, so the architect inherits a scattered estate of keys, accounts, prototypes and data flows nobody approved. The job is not to issue a prohibition, which produces more shadow usage, but to provide a route that is easier than the improvised one and to enforce only the few things that are genuinely non-negotiable. This is governance work an architect has always done, and framing it that way in an interview is far more credible than claiming deep expertise in a new toolchain.
Show it: Tell the story as a governance story with numbers: what you found when you looked, what the sanctioned path offered that the improvised one did not, how many teams moved onto it and over what period, what you enforced centrally versus left advisory, how many exceptions you granted and whether they had expiry dates, and whether you retired the old path. If you have done this for any other technology, a shared logging standard, a secrets system, a deployment pipeline, the same story transfers and you should say so plainly rather than overclaiming on AI specifics.
What a screen is looking for
These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.
- Cloud Architect
- Cloud Solutions Architect
- Cloud Infrastructure Architect
- Solutions Architect
- Enterprise Architect
- Principal Cloud Engineer
- Staff Infrastructure Engineer
- Platform Architect
- Cloud Engineer
- Infrastructure Architect
- Security Architect
- Network Architect
- Data Platform Architect
- AI Infrastructure Architect
- AWS
- Amazon Web Services
- Microsoft Azure
- Google Cloud Platform
- GCP
- Oracle Cloud Infrastructure
- Hybrid cloud
- Multi-cloud
- Private cloud
- Sovereign cloud
- Landing zone
- AWS Control Tower
- AWS Organizations
- Service control policies
- Azure Landing Zone
- Management groups
- Azure Policy
- Google Cloud organization policy
- Resource hierarchy
- Account topology
- Subscription design
- Blast radius
- Guardrails
- Preventive controls
- Detective controls
- Policy as code
- Open Policy Agent
- Rego
- Conftest
- Checkov
- Cloud Custodian
- Well-Architected Framework
- Well-Architected Review
- Cloud Adoption Framework
- Google Cloud Architecture Framework
- Reference architecture
- Architecture decision record
- ADR
- Architecture review board
- Design authority
- Cloud Centre of Excellence
- Technology standards
- Exception process
- Cloud migration
- Migration strategy
- 7 Rs
- Rehost
- Replatform
- Refactor
- Repurchase
- Relocate
- Retire
- Retain
- Lift and shift
- Application portfolio assessment
- Dependency mapping
- Migration wave plan
- Cutover runbook
- Rollback plan
- Go/no-go criteria
- AWS Migration Hub
- Azure Migrate
- Database Migration Service
- Datacentre exit
- Workload repatriation
- Mainframe modernisation
- VMware migration
- Multi-tenancy
- Tenant isolation
- Noisy neighbour
- Control plane
- Data plane
- Region strategy
- Data residency
- Data sovereignty
- Availability zones
- Multi-region
- Active-active
- Active-passive
- Pilot light
- Warm standby
- RTO
- RPO
- Disaster recovery
- DR test
- Failover testing
- Backup and restore
- Immutable backups
- Business continuity
- SLA
- SLO
- SLI
- Error budget
- Chaos engineering
- Game day
- Infrastructure as code
- Terraform
- OpenTofu
- Terraform modules
- Module versioning
- Remote state
- Terragrunt
- Pulumi
- AWS CDK
- CloudFormation
- Bicep
- ARM templates
- Ansible
- Packer
- GitOps
- Argo CD
- Flux
- Kubernetes
- EKS
- AKS
- GKE
- OpenShift
- Helm
- Serverless
- AWS Lambda
- Azure Functions
- Containers
- ECS
- Fargate
- Cloud Run
- Event-driven architecture
- Microservices
- Modular monolith
- API gateway
- Service mesh
- Message queue
- SQS
- Kafka
- Event Hubs
- Pub/Sub
- Caching
- CDN
- CloudFront
- Networking
- VPC
- VNet
- Hub and spoke
- Transit Gateway
- Cloud WAN
- VPC peering
- PrivateLink
- Private endpoints
- Direct Connect
- ExpressRoute
- Cloud Interconnect
- VPN
- BGP
- DNS
- Route 53
- IPAM
- Subnetting
- Network segmentation
- Security groups
- Network ACLs
- Egress control
- Firewall
- WAF
- DDoS protection
- Zero trust
- IAM
- Identity federation
- Single sign-on
- SAML
- OIDC
- OIDC federation
- Workload identity
- Least privilege
- Permission boundaries
- Privileged access
- Break-glass access
- Machine identity
- Non-human identity
- Short-lived credentials
- Secrets management
- HashiCorp Vault
- KMS
- Key management
- Customer managed keys
- BYOK
- Encryption at rest
- Encryption in transit
- Certificate management
- Data classification
- Audit logging
- CloudTrail
- Centralised logging
- SIEM
- Compliance
- PCI DSS
- HIPAA
- SOC 2
- ISO 27001
- FedRAMP
- GDPR
- DORA regulation
- Regulatory reporting
- Evidence collection
- FinOps
- Cloud cost optimisation
- Cost model
- TCO
- Unit economics
- Cost per tenant
- Cost to serve
- Showback
- Chargeback
- Cost allocation tagging
- Reserved instances
- Savings Plans
- Committed use discounts
- Commitment coverage
- Rightsizing
- Autoscaling
- Spot instances
- Karpenter
- Storage lifecycle policies
- Data transfer costs
- Egress charges
- Idle resource elimination
- Budgets and alerts
- AWS Cost Explorer
- Azure Cost Management
- Kubecost
- OpenCost
- Capacity planning
- Quota management
- Observability
- OpenTelemetry
- Prometheus
- Grafana
- CloudWatch
- Azure Monitor
- Distributed tracing
- Incident response
- Postmortem
- Runbook
- On-call
- Change management
- CAB
- DORA metrics
- Deployment frequency
- Lead time for changes
- Change failure rate
- Time to restore service
- GPU
- Accelerators
- Accelerator quota
- Reserved capacity
- GPU scheduling
- Model serving
- Inference
- Inference cost
- LLM gateway
- Model routing
- Prompt caching
- Spend caps
- Rate limiting
- Token spend
- Retrieval augmented generation
- Vector database
- AI governance
- EU AI Act
- Model risk
- Human oversight
- AI coding agents
- Agent identity
- Human approval gate
- Sandbox accounts
- Shadow AI
- Stakeholder management
- Executive communication
- Technical writing
- Whiteboarding
- Pre-sales
- Discovery workshop
- Statement of work
- Proof of concept
- Proposal
- Utilisation
- Partner practice
- TOGAF
- AWS Certified Solutions Architect Professional
- Google Professional Cloud Architect
- Azure Solutions Architect Expert
- AZ-305
- AZ-104
- CKA
- CISSP
- FinOps Certified Practitioner
- Security clearance
Mistakes that cost people this job
Applying to every posting with this title using one resume, without working out which of the four jobs it is.
Classify in two minutes: whose system would you design (your employer's estate, a customer's, or the product you sell), how many compliance regimes are named, whether a language is a requirement, and whether the word utilisation appears. Then send the enterprise resume, the partner resume, the product resume or the specialist resume, not the same document four times.
Leading the resume with a wall of certifications above thin production experience.
Lead with calibration numbers: accounts, applications, spend order of magnitude, regions, tenants, regulated regimes, and the biggest migration as a count and a duration. Put certifications in one line near the bottom unless the employer is a cloud partner, where they are a commercial requirement and belong higher.
Writing duties instead of decisions. Responsible for cloud architecture, designed scalable and secure solutions, owned the AWS environment.
One line per decision: what you chose, what you rejected, the constraint that drove it, and the measured outcome. Chose X over Y because of Z, and the result was N, is the sentence pattern that survives every follow-up because you already know the reasoning.
Jumping straight to a diagram in the design round.
Spend the first five to ten minutes writing requirements on the board: users, peak rate, data volume and growth, latency target and percentile, RTO and RPO per tier, compliance and residency, budget, deadline, and the size and skills of the team who will run it. Then name the driving constraint out loud before you draw a box.
Designing for a scale the business does not have. Active-active multi-region, per-tenant infrastructure and a service mesh for a company with four engineers and a two-hour recovery tolerance.
Let the requirement pick the architecture and say so. Price the cheaper option, state what it fails to deliver against the targets you wrote down, and only then spend. Rejecting multi-region on a tested two-hour tolerance scores higher than building it unasked.
Offering multi-cloud as a general hedge when asked about resilience or lock-in.
Say the narrow true thing: portability costs you the managed services that make a cloud worth using and doubles the identity, network, observability and skills surface. Then name the defensible cases, which are an acquisition you have to run, a contractual or regulatory requirement, one workload placed deliberately where it runs best, or a commercial lever in a negotiation.
Being unable to say what drives the current cloud bill.
Know your top three or four cost lines in rough proportion, which is growing fastest, your commitment coverage, and the cheapest thing you have not fixed yet. If you do not own the bill, say so and give the order you would investigate in: allocation coverage, largest line item, idle and orphaned resources, data transfer paths, logging ingest, commitments.
Having no design you got wrong, or blaming circumstances for the one you name.
Prepare one with five parts: the constraint you misread, the symptom that revealed it, the correction, what the correction cost, and the rule you adopted afterwards. Panels mark the rule highest, because it shows the lesson generalised rather than being survived.
Presenting as someone who hands down diagrams and no longer touches the platform.
Volunteer your hands-on currency before they probe for it: the last thing you personally built or debugged and when. Be ready to read a Terraform plan and name the destructive line, explain a private endpoint breaking name resolution, or say what a NAT gateway costs and how to avoid it.
Claiming expert-level depth across AWS, Azure and Google Cloud.
State one primary cloud with depth, the others at the level you actually have, and say which you would need a ramp in. Panels probe the weakest claim, and an honest ranking survives that probe while a tri-cloud claim does not.
Describing governance as documents written, with no evidence anyone followed them.
Attach adoption to every standard: how many teams adopted it, over what period, how many exceptions you granted and whether they had expiry dates, and whether the old pattern was retired. Make the compliant path the easiest path and enforce only the genuinely non-negotiable items with preventive controls.
Claiming AI has transformed cloud architecture, or asserting a regulatory deadline from memory.
Be precise about what changed: accelerator capacity and placement, inference cost as a line item with unit economics, data boundaries when a model is called, identity for agents, and preventive controls now that generated configuration is cheap. State regulatory obligations without inventing dates and say the current timing should be confirmed with counsel.
Questions people ask
What does a cloud architect do?
A cloud architect owns the design decisions that would be expensive to reverse in a cloud estate: the account or subscription topology and the blast-radius boundaries, the network pattern and private connectivity, the identity and permission model, where data is allowed to live, the resilience targets per service tier and the mechanism that meets them, the migration strategy and its sequencing, and the cost model including unit economics. The output of a cloud architect is designs, decision records, reference architectures and standards more than it is code, although the good ones stay hands-on enough to read a plan and debug a network path. A practical test of the boundary: if a choice can be undone inside a sprint, it is probably not the cloud architect's to make, and if undoing it needs a migration, it is.
How do you become a cloud architect?
Most people reach cloud architect from cloud, platform, infrastructure, network or backend engineering, and the gate is not years but having owned a decision whose reversal would need a migration. The reliable route is to produce architect evidence before you hold the title: write architecture decision records for the choices your team is already making, document the account topology you have and what you would change, run a cost review and produce a model with unit economics, test one restore and write up the result with the date, and volunteer for the cost reviews, disaster recovery tests, vendor evaluations and cross-team design reviews that are architecture work in disguise. Then move internally if your employer has an architecture function, or join a consulting partner or managed service provider if you need breadth fast, because they expose you to several migrations a year and fund the certifications.
Do you need a certification to be a cloud architect?
No certification is legally required to work as a cloud architect, and at a product company a certificate is a filter pass at most, with the panel deciding on production evidence instead. There is one place where it is effectively mandatory: consulting partners, systems integrators and managed service providers, because their partner tier with the cloud provider is counted in certified staff, which makes a professional-level certification a commercial requirement rather than a signal, and those firms usually pay for it. The practical answer for most cloud architect candidates is one professional-level architecture certification in the cloud your target employers actually run, taken alongside real work rather than instead of it, because no exam demonstrates the part of the job that gets you hired.
Which cloud certification is best for a cloud architect?
For a cloud architect, the useful answer is the professional-level architecture certification in the cloud the employer actually runs: AWS Certified Solutions Architect Professional, Google Cloud Professional Cloud Architect, or Microsoft Certified Azure Solutions Architect Expert. All three test long constrained scenarios with several defensible answers rather than service trivia, which makes studying for them a reasonable rehearsal for the design round. Add FinOps Certified Practitioner if the posting leans at cost, a provider security specialty or CISSP if it leans at security, and TOGAF only for large enterprises and markets where recruiters screen on it. A second cloud at professional level is usually a worse investment than one more real migration, and all of these expire, so confirm the current exam code and renewal window when you book.
How long does it take to become a cloud architect?
Reaching cloud architect typically takes five to ten years of technical work, but the variance is explained by decision ownership rather than by time served: engineers who document decisions and volunteer for cost reviews, disaster recovery tests and design reviews get there years earlier than equally skilled engineers who only implement. From a standing start with no infrastructure background, expect two to four years to competent cloud engineer and then two to four more to architect scope. The fastest observable accelerator is a consulting partner or managed service provider, where several migrations a year and funded certifications can build in two years the breadth an internal role takes six to accumulate, at the cost of utilisation pressure and less depth per system.
How much does a cloud architect make?
There is no single band for a cloud architect, and any article quoting one number is averaging across four different jobs. Source it yourself: the US Bureau of Labor Statistics has no cloud architect code, so triangulate from Occupational Employment and Wage Statistics for 15-1241 Computer Network Architects, 15-1252 Software Developers and 11-3021 Computer and Information Systems Managers, add employers' own posted ranges under state pay-transparency laws in Colorado, California, Washington, New York and Illinois among others, and add levelling data for named technology employers. One structural factor moves the figure more than the title: at a product company a cloud architect is usually levelled as a staff or principal individual contributor with equity, while at a consulting partner the package often carries a utilisation or sales-attached component and little equity. Ask which ladder in the first conversation.
What is the difference between a cloud architect and a cloud engineer?
A cloud engineer builds and operates the environment and is measured on whether it works: delivery, reliability, automation and the pager. A cloud architect decides the shape the engineer builds against, meaning the topology, the identity and network boundaries, the resilience targets, the migration sequence, the standards and the cost model, and is measured on whether those decisions were still right two years later. Engineers ship changes most days; a cloud architect writes more prose and diagrams and reads more code than they write, though a cloud architect who has stopped touching the platform is a recognised failure mode and panels probe for it in the first ten minutes. In pay terms the two are often closer than the titles suggest, because a senior cloud engineer band and a cloud architect band frequently overlap.
What is the difference between a cloud architect and a solutions architect?
The titles overlap heavily and the real difference is usually audience rather than skill. A cloud architect most often designs an employer's own estate: landing zone, migration, standards, cost, resilience. A solutions architect title more often means customer-facing work at a cloud provider, a systems integrator or a software vendor, where the job is discovery, whiteboarding, proposals and reference architectures across many customer environments, with pre-sales expectations and a utilisation or quota component attached. Read the posting rather than the title, because plenty of employers use cloud architect for the customer-facing job and solutions architect for the internal one, and the interview differs sharply: the customer-facing version includes a mock client session marked on listening.
What do cloud architect interviews test?
A cloud architect interview centres on a 60 to 90 minute open design round, and it is marked on behaviour as much as on the diagram: whether you elicited requirements before designing, named the driving constraint, rejected a cheaper option out loud, reasoned about concrete failures and blast radius, priced the design in line items, said what you would not build, and finished with a migration sequence and a rollback position. Around that sit a hiring manager conversation about decision authority, a deep dive on a past design including one you got wrong, a money round on what drives your bill and how you would cut it, an identity and security round, and a stakeholder conversation that tests whether you can explain a trade-off to someone holding the budget. At a consulting partner there is also a mock customer whiteboard; at a product company there is usually real code or an infrastructure-as-code review.
Is cloud architect a dying job because of AI?
No, and the honest version of the answer is more useful than either extreme: a cloud architect's core work of eliciting requirements, choosing between defensible designs on cost and risk, sequencing migrations and owning a decision that is expensive to reverse has not been automated, and autonomous remediation in production remains uncommon despite a decade of AIOps marketing. What has changed is the workload mix. A cloud architect in 2026 and 2027 is expected to have an opinion on accelerator capacity and placement, inference cost as a line item with unit economics, where data goes when a model is called, identity and limits for agents acting inside the estate, and preventive controls now that generated infrastructure code is cheap and plentiful. The job pool has also shifted from greenfield migration toward optimisation, consolidation, repatriation, sovereign requirements and AI infrastructure.
Put this on a resume in about a minute
Paste your history once and point it at the Cloud Architect posting you are looking at. No account, no card.
Build my resume free More roles