AI & Machine Learning

How to get hired as an AI solutions architect in 2026-27

The short answer

An AI solutions architect is hired on demonstrated design judgement, not on a credential: no jurisdiction licenses the title, and most offers are decided by one round, a reference architecture you draw live for a scenario you have not seen. Nearly all of these jobs sit in one of four places that interview differently, so work out which one the posting is before you tailor anything: a cloud or model vendor's field organisation, a consultancy or systems integrator, an enterprise's own architecture or AI platform group, or a software vendor's implementation team. If you are arriving from cloud or enterprise architecture most of your skill transfers, and three gaps decide the loop: evaluation (proving quality when the output is not deterministic), unit economics (tokens, context length, caching, GPU memory), and the authority model for agents that call tools. The fastest way in is one written artefact: a reference architecture for a real workload, with a cost model per thousand requests, an evaluation harness over a labelled set, and a decision record naming what you rejected and why.

What the role is in 2026-27Deciding, writing down and defending how an organisation will actually run an AI workload end to end: where the data comes from and who is allowed to see it, what the model is and what replaces it when the vendor retires it, how quality is measured and who signs off on the threshold, what it costs per thousand requests, what happens when it is wrong, and who operates it after the architect leaves. The deliverable is a decision with a trade-off attached, not a diagram.
Read the posting for this before anything elseWhich of four employers it is. Vendor field role (cloud or model provider, pre-sales, quota-adjacent). Consultancy or systems integrator (billable delivery architect, utilisation and scoping). Enterprise in-house (architecture or AI platform group, standards and vendor selection). Software vendor implementation or forward-deployed team (one product, many customers). Same title, four different interviews, four different resumes.
Posting titles that mean this jobAI Solutions Architect, GenAI Solutions Architect, AI Specialist Solutions Architect, AI/ML Solutions Architect, Generative AI Architect, AI Platform Architect, Enterprise AI Architect, Principal Architect (AI), Field CTO (AI), AI Delivery Architect, Applied AI Architect, Forward Deployed Architect. Specialist Solutions Architect with an AI qualifier is the vendor spelling.
Credential gateNo jurisdiction licenses this title and no certification is legally required. Certifications act as a screening filter, and inside cloud vendors and their partners they are closer to structural, because partner competency tiers are counted in certified individuals, which is why those employers often pay for them and expect them within the first few months. The professional or expert tier cloud architect credentials carry the most weight, the AI credentials from the same vendors carry some, and anything from an issuer the hiring manager cannot name carries none. Catalogue names, exam codes and prices in this area change often, so check the issuer's own page before paying.
The round that decides itA live reference architecture, commonly 45 to 75 minutes. A scenario described in two or three minutes, a whiteboard or shared canvas, and an expectation that you ask before you draw, draw the data path before the model, attach real numbers with units, name one residual risk, and say what you would not build. Vendor and consultancy loops usually add a customer-facing simulation where someone plays a sceptical stakeholder.
What gets you screened in from outsideOne public artefact that shows judgement rather than enthusiasm, in rough order of weight: a written reference architecture for a named workload including a cost envelope and an evaluation plan, a conference or meetup talk with the slides online, an open evaluation harness for a task you can describe, a decision record library showing rejected options, a production retrospective saying what broke and what it cost. A list of courses completed is not an artefact.
Level, and realistic time to hired by starting pointThis is a senior individual contributor title in practice. Most requisitions expect several years of architecture, consulting or platform engineering before the AI qualifier, and new graduates are not hired into it. From a cloud solutions architect job, budget 8 to 16 weeks of deliberate work on the three gaps plus a normal search. From enterprise architecture with no hands on keys, budget 4 to 8 months, because you need to have built and measured something recent. From data or machine learning engineering, the gap is the customer-facing muscle rather than the technology, and the fix is reps in front of stakeholders. The most common successful move is internal, from a cloud or platform team into the same employer's AI programme.
Pay, and where to check it rather than trust a bandNo BLS occupation code names this title. The nearest authoritative US baselines are OES 15-1241 Computer Network Architects, 15-1252 Software Developers and 15-1299 Computer Occupations All Other, and all three undercount vendor field pay. For the field variant the package is usually base plus a variable component tied to consumption or quota, and the variable share differs enough between employers that you have to ask for the split in writing before you compare a vendor number to an in-house salary. Check posted ranges under pay-transparency rules in states including California, Colorado, Illinois, New York and Washington, the employer's own posting, published levels data for the large vendors, and H-1B labour condition filings for named employers.

What an AI solutions architect actually does, and which of the four jobs the posting means

The job is not drawing. Diagrams are the artefact, decisions are the product. An AI solutions architect is the person who writes down, in a form a security reviewer and a finance partner can both argue with, how a specific AI workload will run inside a specific organisation: which data it touches, under whose permissions, through which model, at what measured quality, at what cost per thousand requests, degrading how, owned afterwards by whom. Everything that makes AI projects die in month seven lives in that sentence, and almost none of it is model selection.

In practice the role exists because two groups keep talking past each other. Engineers can build a convincing prototype in days, and with a coding assistant often in one. Enterprises still cannot deploy it, because the prototype reads a folder of documents with no regard for who is allowed to see them, has no measurement beyond a product manager's impression, has an unmodelled cost that scales with context length, calls a model version that will be retired, and nobody has agreed who gets paged at 2am. The architect's value is closing that gap before the money is spent, and the ability to say no to a design is a bigger part of the job than the ability to propose one.

It is worth being precise about the neighbours, because interviewers test the boundary. An AI engineer builds and owns the running system. A machine learning engineer trains, tunes and serves models. An MLOps or platform engineer owns the pipelines and the infrastructure. A data architect owns the model of the data itself. An enterprise architect owns the portfolio and the standards across many systems. An AI solutions architect sits across a workload end to end, usually several workloads at once, and is accountable for the parts nobody else wants: the permission model in the retrieval path, the evaluation set, the cost envelope, the exit plan, and the hand-off.

Before you tailor anything, work out which of the four employers the posting comes from. It changes the interview, the resume, the metric you will be judged on, and whether you will write code. The signals are in the text: consumption, workloads, customers and partners mean vendor field. Engagement, client, statement of work, utilisation and practice mean consultancy. Standards, governance, reference architecture, centre of excellence and stakeholders mean in-house. Deployment, onboarding, implementation and a named product mean software vendor.

Coming from cloud or enterprise architecture: what transfers, and the gaps that decide the loop

Most of this role transfers directly from cloud architecture, and the people who assume they are starting over waste months. Identity and authorization, network topology and private connectivity, data residency and sovereignty, tenancy models, storage tiering, queueing and backpressure, observability, disaster recovery with a stated recovery point and recovery time, cost allocation and showback, change management, the discipline of writing a decision down with its rejected alternatives: all of it is the same, and all of it is scarce among people who arrived from the model side. Say so plainly in the interview rather than apologising for not having trained a model.

Four gaps decide the loop, and they are the things a cloud architect has genuinely never had to do. The first is evaluation. Every system you designed before was deterministic enough that acceptance testing meant pass or fail. Here the same input can produce a different output, quality is a distribution, and the only honest answer to does it work is a labelled set of cases, a scoring method, a baseline number and a threshold agreed with whoever signs off. An architect who has no answer to how will we know it is good enough is rejected in the whiteboard round, by every employer type.

The second gap is unit economics and capacity. You are used to instances, cores and gigabytes. The currency here is tokens, and the dominant cost driver in most retrieval systems is not the clever part, it is the length of the context assembled on every single call. You need to be able to do the arithmetic out loud: tokens in plus tokens out, times calls per day, times the contracted rate per million, before adding the index, the embedding passes, the re-ranking and the human review time. If the workload is self-hosted you need the memory arithmetic too: parameters times bytes per parameter at your chosen precision, plus the key-value cache, which grows with sequence length and concurrency and is what actually decides how many requests fit on a GPU.

The third gap is the authority model. The moment the system calls tools rather than just writing text, the design question stops being what does the model say and becomes what is this thing allowed to do, with whose credentials, and what cannot be undone. That is an authorization problem, which is good news for a cloud architect, but it has a twist: the model takes instructions from text the organisation does not control, so you cannot rely on it refusing. Authority has to sit outside the model, in scoped credentials, server-side limits, idempotency, and human confirmation on irreversible actions.

The fourth gap produces the worst incidents: retrieval is an access-control problem before it is a relevance problem. An index built by crawling a document store flattens permissions by default, and the first demo where an assistant cheerfully summarises a document the asker was never entitled to read ends the programme. Permission-aware retrieval, meaning filters applied at query time from the caller's identity rather than scrubbed afterwards, is something interviewers specifically listen for.

Timelines, honestly. From a working cloud solutions architect job, 8 to 16 weeks of deliberate effort on those four things, done as one real artefact rather than as reading, is enough to interview well. From enterprise architecture with no recent hands-on work, 4 to 8 months, because the deep dive round will find out and because you need something you built and measured yourself. From data engineering or machine learning, the technology is not the problem: book yourself into stakeholder-facing situations and get used to being interrupted by someone who does not care how it works.

How hiring works, stage by stage, and how it differs by employer

There is no licence, no board exam and no portfolio review panel. What gates this role is a loop of four to seven conversations, one of which is a live design exercise, and a strong preference for people the hiring team has already watched work. Internal movement and referral through a partner relationship fill a large share of these seats, which means the highest-return action for most people is not applying harder, it is getting visible to the AI programme at the employer or partner they already touch.

The shape below is the vendor loop, which is the most structured and the most commonly asked about. Consultancies compress it and add scoping. Enterprises stretch it and add a written exercise and a committee. Software vendors add a hands-on element because you will be in the code.

Timelines vary, but the usual shapes are: vendor loops 3 to 8 weeks from first screen to offer, longer when a cross-organisation reviewer has to be scheduled; consultancies fastest at 2 to 4 weeks, because a billable seat is open now; enterprises 4 to 10 weeks, often going quiet in the middle, which is usually budget approval rather than rejection.

One structural note worth knowing before you negotiate: in vendor field organisations the architect is often paired with an account executive and sits inside a sales organisation's planning cycle. That affects the questions you get asked (how do you handle a customer who wants something our product does not do), the metric on your scorecard, and the fact that part of your pay moves with someone else's number. None of that is a reason to avoid the job, but walking in without knowing it reads as naive.

The reference architecture whiteboard round, worked end to end

Here is a scenario of the kind actually used. A mid-size property insurer wants its claims adjusters to receive, for each new claim, a drafted summary of the claim file and a coverage assessment that cites the specific policy clause it relies on. The inputs are policy documents as PDFs, adjuster notes and claim history inside a claims system of record, loss photographs, and a policy master on a mainframe that nobody is allowed to change. Peak volume is 12,000 claims a day across roughly 1,800 adjusters. The data cannot leave the country. The regulator expects decisions to be explainable and the business expects this to cut 20 minutes off each claim. You have 50 minutes and a whiteboard.

Spend the first eight to ten minutes asking, and say that is what you are doing so it does not look like stalling. The questions below are not padding: each one changes the drawing. Then draw in a fixed order, and say the order out loud, because a repeatable process is most of what is being scored. Data sources first, then permissions, then ingestion and indexing, then retrieval, then prompt assembly, then the model call, then output handling and citation, then the human step, then logging and evaluation, then the feedback loop. The model box goes on the board fifth or sixth, not first. Interviewers notice.

Mark the trust boundary explicitly. In this system the adjuster notes and the claim correspondence contain text written by claimants and third parties, which means an attacker can contribute to the prompt. Say that out loud, say that you therefore will not let the model hold any authority to act on the claims system, and that drafting is advisory with a human commit step. If a later phase automates anything, say which actions could ever be automated (set a status flag, request a document) and which never could without a second control (issue a payment, deny a claim).

Then do the numbers. Out loud, with assumptions stated as assumptions. Suppose the assembled context runs around 8,000 tokens per draft, because the policy excerpts dominate, and the output is around 700 tokens. At 12,000 drafts a day that is roughly 96 million input tokens and 8.4 million output tokens a day, so multiply each by your contracted rate per million and add them. Now say the three things that follow: input dominates, so the single biggest cost lever is retrieving three clauses instead of thirty and caching the stable part of the prompt; the per-request cost is small but the daily number is a real budget line finance will ask about; and the human review minute is probably the most expensive item in the system, so the business case rests on the time saved per claim rather than on the token price. Do not invent a vendor price in the room. Say you would price it against the contracted rate and show the formula.

Non-functionals, in roughly the order they decide the round. Residency: a region-pinned model endpoint and a region-pinned index, with a written answer on where inference logs go, how long they are kept, and what the provider is contractually allowed to do with the content. Permissions: adjusters see claims in their book of business, so retrieval filters by identity at query time, and the index carries the access metadata rather than relying on a post-filter. Latency: adjusters work a queue, so a 9 second draft is annoying but survivable while a 90 second one is abandoned; state a target such as p95 under 10 seconds, and say you would stream the draft so perceived latency is the first token, not the last. Availability and degradation: if the model endpoint is unavailable the queue must not stop, so the fallback is no draft rather than a blocked workflow, and that needs to be designed, not discovered. Auditability: store the retrieved clause identifiers, the prompt version, the model version and the adjuster's edit, because the explainability requirement is satisfied by the citation trail and the human decision, not by any claim about model interpretability.

Evaluation is where most candidates stop talking, and it is the part the hiring manager cares about most. Say it concretely. Take a few hundred closed claims with known outcomes, have two experienced adjusters label the correct coverage position and the clauses that support it, and score three things separately: is the cited clause the right clause, is the summary free of statements the file does not support, and does the recommended position match the adjuster's. Report them as three numbers, not one. Set the go-live threshold with the business before the build, keep a holdout, and name the online metrics you would watch afterwards: edit distance between the draft and the submitted summary, override rate, time per claim, and complaint volume. Then say who owns the eval set, because an eval set with no owner is dead within two quarters.

Close in the last three minutes with three sentences that separate experienced architects from everyone else. One residual risk you are accepting, with its compensating control (the model will occasionally cite a clause that is real but not the governing one, so the citation renders as a link to the clause text inline and the adjuster must open it to commit). One thing you would not build (no automated claim decisions in phase one, no fine-tuned model until retrieval quality is measured, no custom vector store if the existing data platform already provides one). And the operating answer: who runs this, who gets paged, what the rollback is, and what happens when the model version you built against is retired.

The other rounds: deep dive, cost and sizing, and the customer in the room

The deep dive is a breadth test with two spikes, and which spike you fail is predictable from your background. People from the architecture side lose on AI mechanics, usually on retrieval behaviour, context cost and evaluation. People from the model side lose on identity, private networking, tenancy and operations. Prepare the half you did not come from, and prepare it to the level of explaining a trade-off rather than reciting a definition.

The questions that recur are mechanism questions, not trivia. Why does chunk size change answer quality, and in which direction for a long legal clause. What does a re-ranker buy you and what does it cost in latency. Now that long context windows are cheap enough to abuse, when does stuffing the whole document set in beat retrieval, and when does it just pay for the same tokens on every call. When is fine-tuning the right answer rather than retrieval, and what does it not fix. What changes when you move from a hosted endpoint to self-hosted weights, in money, in operations and in risk. What breaks when you upgrade the model under a prompt that was tuned to the old one. How do you keep a multi-vendor posture without reducing everything to the lowest common denominator. Each of those has a real answer with a trade-off in it, and a candidate with production scars gives it in thirty seconds.

The customer simulation is a separate skill and people with strong technical rounds fail it. The failure is almost always the same: answering the objection you prepared instead of the one that was raised, and defending instead of conceding. The move that works is to restate the objection in the stakeholder's own terms, concede the part that is true, narrow the claim to what you can actually support, and attach a next step with a date and an owner. If you do not know, say you do not know and name when you will come back. Senior people in the room have heard confident nonsense before and they are listening for whether you do it.

The behavioural round is scored on specificity. The architect stories that land are about decisions with consequences: a design you killed after a measurement contradicted you, an estimate you got wrong and what you changed in how you estimate, a time you were overruled and what happened, a customer who had already failed once with a different vendor and what you did differently. Numbers in the story are what separate a real anecdote from a rehearsed one.

Certifications: the few that move a decision, and the many that do not

Certifications matter unevenly across the four employers, and getting this wrong wastes months. At cloud vendors and especially inside their partner ecosystem they are structural: partner competency tiers are counted in certified individuals, so consultancies and systems integrators need badges on the bench, often ask for them explicitly, will pay for them, and sometimes make one a condition of passing probation. At software vendors they matter if the vendor has its own credential. In-house they are close to irrelevant beyond getting past an automated screen, where a recruiter is matching the requisition text.

The order of weight is consistent. A professional or expert tier cloud architect credential on the platform the employer uses or sells is the only one that reliably changes a screening decision, because it is the baseline the role is built on. The AI credential from the same vendor is a useful second, mostly as evidence you took that platform's AI services seriously. A data platform credential matters when the shop is built on that platform. Everything else is a tiebreaker at best.

Names, codes and prices in this area change more often than in any other part of cloud certification, and credentials get renamed, split or retired between one hiring cycle and the next. Check the issuer's own page for the current name, exam code, prerequisites and price before you book, and do not put a retired credential on a resume as though it were current. If a credential has an expiry, show the year.

Two pieces of sequencing advice. First, if you are employed somewhere that will pay for a certification and give you study time, take it there rather than paying yourself. Second, if you are choosing between one more certification and one written reference architecture with a cost model and an evaluation plan, write the document. In the rounds that decide the offer the badge has nothing to say, and the document answers four questions.

The resume and the portfolio: what belongs, what gets ignored

An architect resume is read for decisions and numbers. A tool inventory is read as a tool inventory. The bullet that works has four parts: the decision you made, the constraint that forced it, the alternative you rejected and why, and the measured outcome with a unit attached. Three bullets of that shape beat fifteen lines of responsibilities, and most candidates for this role submit fifteen lines of responsibilities.

Tailor to the employer variant, because the readers want different proof. A vendor hiring manager wants evidence of customer-facing work at volume: how many customers you worked across, how many workloads reached production, the workshops you ran, how many proofs of concept converted, the size of the environments. A consultancy wants engagement shape: scope, duration, the team size you carried technically, estimate accuracy, whether the client extended. An in-house panel wants adoption: which pattern you set, how many teams built on it, what it removed, how many security exceptions you avoided. The same project can be written three ways and should be.

Lead the document with a four line summary that names the variant you are applying to, the platforms you are deep in, the industries you know, and one number. Then the evidence. Then a short technologies block a screener can match against the requisition, kept to names actually in use rather than every model released this year. Certifications get one line each with the year.

Two specific things that get ignored and one that actively hurts. Ignored: a list of every model and framework you have touched, and verbs like spearheaded and drove with nothing measured after them. Also ignored: a screenshot of a diagram with no numbers on it, which reviewers skip because it says nothing a hundred other diagrams do not. Actively harmful: claiming production experience you do not have in a field small enough that the deep dive round will find out in four minutes. Say prototype when it was a prototype and name what you learned from it. That reads as senior.

Where the jobs are, what to search, and a twelve-week plan

These roles concentrate in four pools, and the one most candidates ignore is the largest. First, cloud, model and data platform vendors, hiring field and specialist architects. Second, the partner ecosystem: global systems integrators, regional consultancies and boutique AI delivery shops, which collectively hire more of these people than the vendors do and which fill seats faster because the seat is billable now. Third, enterprises in regulated and data-heavy sectors standing up internal AI platforms: financial services, insurance, healthcare and payers, pharmaceuticals, energy and utilities, telecoms, logistics, and public sector. Fourth, software vendors with an implementation or forward-deployed function.

Search the title variants rather than the single title, and search the vendor spelling. Specialist Solutions Architect with an AI qualifier, GenAI Architect, AI Platform Architect, Enterprise AI Architect, AI Delivery Architect, Principal Architect with AI in the body text. Also search the pattern language rather than the title: postings asking for retrieval augmented generation, a model gateway, an evaluation harness, agentic workflows or responsible AI review are the same job under a different name. Cloud partner directories are a good way to find the regional consultancies you have never heard of, and those are frequently the fastest hires.

If you are already inside a company with a cloud or platform team, the internal route beats the external one by a distance. The AI programme in most large organisations is understaffed on exactly the skill you have and is being asked for architecture decisions it cannot make. Volunteer for the unglamorous part: the evaluation standard, the cost model, the permission model in retrieval. Those three documents make you the obvious internal candidate within a quarter, and they are the portfolio if you later go outside.

Twelve weeks is enough from a working architecture job. The plan below assumes evenings and one weekend day, and it is built so that every week produces an artefact rather than a reading list.

Working with AI in this role

What an AI solutions architect must know about AI in 2026-27

The role is an AI role, so being vague here is fatal in a way it is not in most jobs. The thing to understand about 2026 is where the difficulty has moved. Building a working prototype is now cheap, often a day's work with an assistant, and that has not made the architect less necessary, it has changed what the architect is for. Nobody needs help getting a demo. They need help answering the four questions a demo cannot answer: how do we know it is good enough, who is allowed to see what it reads, what does it cost at volume, and what happens when the model underneath it changes. An architect fluent in those four and ordinary at prompt design will out-interview the reverse every time.

It is worth saying plainly what has changed less than the coverage suggests. The reasons enterprise AI projects fail are mostly the reasons enterprise projects have always failed: the data is not accessible or not trustworthy, identity and entitlement are modelled inconsistently across the systems involved, no single person owns the outcome, the integration with the system of record is harder than the clever part, and nobody planned the change management for the people whose job the thing alters. The model is rarely the bottleneck. Interviewers who have lived through a failed programme listen specifically for a candidate who knows that, because the candidate who thinks the hard part is model selection will spend the first six months on the wrong problem.

Two genuine shifts are worth naming. The build-versus-buy line has moved: a large amount of what organisations hand-built in 2023 and 2024, custom retrieval stacks, bespoke orchestration, homegrown guardrails and evaluation scaffolding, has since become a feature of platforms they already license, so an increasing share of an architect's value is knowing what not to build and being able to say so to a team that has already started. And the default workload has moved from answering questions to taking actions, which turns a content-safety conversation into an authorization conversation. A resume full of custom infrastructure from 2024 needs one sentence explaining why it was the right call then.

The eight things below are what actually get tested, in roughly the order they decide rounds.

Designing the evaluation before the architecture is approved

This is the discriminating question in the role. The output is non-deterministic, so every classical acceptance gate the organisation owns (a test suite, a sign-off, a service level with a pass or fail) silently stops working, and the vacuum gets filled by someone's impression of the demo. Programmes die here: they cannot prove improvement, so they cannot justify the next quarter's budget, so they are cancelled while still working. An architect who does not design the measurement has shipped a demo with a deployment plan attached.

Show it: In the whiteboard round, draw the evaluation harness as a box on the diagram, not as a remark. Specify the task, the labelled set and who labels it, how many cases and how you chose them, the scoring method (exact match where the answer is checkable, a rubric scored by a model with a measured agreement rate against human spot checks where it is not, task completion and trajectory checks for agentic flows), the baseline, the threshold required to go live, the holdout you keep, and the online metrics that replace the offline ones after launch. Name the owner of the set. Then give a number from your own work: the size of the set you held and the score you refused to ship below.

Retrieval as an access-control problem before it is a relevance problem

The fastest way to end an internal AI programme is for one person to receive a well-written summary of a document they were never entitled to read. Indexing a document store flattens its permissions by default, and the common shortcuts (filter after retrieval, scrub the corpus, trust the prompt) fail in ways that are hard to detect and impossible to explain to a risk committee. This is also where cloud and enterprise architects have a real advantage, because it is ordinary authorization work in unfamiliar clothes.

Show it: Say the words permission-aware retrieval and then be concrete: access metadata carried into the index at ingestion, filters applied at query time from the authenticated caller's identity rather than from a parameter the client supplies, a re-check against the source system before content is shown where the stakes justify the latency, and a defined behaviour when an access list changes after indexing. Explain how you would prove tenancy isolation in a test rather than assert it. For agentic flows, add that tool calls run under scoped credentials derived from the caller, never under a service account holding the union of everyone's permissions.

Unit economics and capacity planning in tokens, cache behaviour and GPU memory

Finance will ask, and the answer decides whether the workload scales or is quietly capped. The counter-intuitive fact most candidates miss is that in retrieval systems the dominant cost driver is the length of the context assembled on every call, not the output and not the model tier, which makes retrieval precision a cost lever as much as a quality lever. Self-hosting adds a second arithmetic people get wrong in the other direction, assuming parameter count sets the hardware when concurrency and sequence length are what actually fill the memory.

Show it: Do the arithmetic out loud with assumptions labelled as assumptions: tokens in plus tokens out, per request, times requests per day, times the contracted rate per million, plus embedding passes, index storage, re-ranking and the human review minute. Then name the levers in order: retrieve less and better, cache the stable prefix, route easy traffic to a smaller model behind a quality check, batch anything asynchronous. For self-hosting, give the shape (weights as parameters times bytes per parameter at your precision, plus a key-value cache that grows with context length and concurrency, plus headroom) and then name the only four reasons to do it: residency, unit cost at sustained volume, a latency floor, or a data policy that forbids a third party.

Choosing between retrieval, long context and fine-tuning, and saying which one the problem needs

Cheap long context has made a new wrong answer available: paste everything in on every call, which works in a demo and prices badly at volume while making citation and permission filtering harder. Fine-tuning is the other reflex, usually proposed to fix accuracy when the real problem is that the right document was never retrieved. Interviewers use this to separate people who have run the comparison from people repeating a vendor blog post, because the three options fail differently and the choice drives most of the rest of the design.

Show it: Give the discriminators rather than a preference. Retrieval when the corpus changes, is large, or has per-document permissions, and when you need a citation. Long context when the relevant material is small, bounded and genuinely needed whole (one contract, one patient record), and say what it costs per call. Fine-tuning for format, tone, a narrow classification task or latency and cost reduction on a stable task, and say plainly that it does not install facts and does not fix retrieval. Then say how you would decide: run the cheap option against the evaluation set first and only move when the numbers justify the cost, and treat distillation into a smaller model as the optimisation you reach for after quality is proven, not before.

Authority and blast radius for agents that call tools

As soon as the system acts rather than drafts, the design question changes from what does it output to what is it allowed to do with whose credentials, and which of those actions cannot be undone. The complication specific to AI is that the model reads text the organisation does not control (a claimant's letter, a supplier's invoice, a web page, a calendar invitation), so an attacker can contribute to its instructions and the model cannot reliably tell the difference. Controls that live inside the prompt are therefore not controls. This is the subject most likely to come up in a security partner's round, and the answer had better not be a filter.

Show it: Produce a tool authority table: per tool, the credential and its scope, whether the action is reversible, the server-side limit (amount, rate, recipient allowlist), whether human confirmation is required, and what is logged. Say explicitly that you cannot eliminate injected instructions, so you bound what the system can do and close the paths by which data leaves. Phase the autonomy: advisory first with a human commit, then automate only the reversible low-value actions, with irreversible ones behind a second control permanently. Name the detection you would add at the tool-call layer, and say who reviews those logs.

Model lifecycle, vendor churn and a credible exit plan

The component at the centre of your architecture is a third party's product on a retirement schedule you do not control, and prompts tuned against one version regress against the next. Organisations have been surprised by this more than once, and procurement now asks about it. An architect who cannot answer how the system survives a model change is proposing a system with an undated expiry.

Show it: Describe the gateway pattern honestly: a single internal entry point that owns authentication, routing, quota, logging, cost attribution and version pinning, with applications depending on your interface rather than on a vendor SDK. Then be honest about its limit, that a thin abstraction over several providers drifts toward the lowest common denominator, so you abstract the call and the plumbing while accepting that a provider-specific capability is a deliberate coupling you document. Add the regression gate: the evaluation set is the upgrade test, you run the new version against it before any switch, and you keep the old version pinned until the numbers clear. Say what migration window you would ask the provider for in the contract rather than assuming one.

Governance and regulated-sector obligations, stated without inventing a date

A large share of these jobs are in financial services, insurance, healthcare and the public sector, where the architecture has to survive a risk committee rather than a technical review. You need to speak the vocabulary and, just as important, not overstate it. The specific trap is compliance dates: obligations in this area have been amended and deferred, and a candidate who states a tranche date as settled fact in front of a legal or risk partner loses credibility in the one room where credibility was the point.

Show it: Speak to the structure rather than the calendar: the EU AI Act works by risk tier with the heaviest obligations on high-risk uses, covering governance of training and validation data, documentation and record-keeping, human oversight, accuracy and robustness, transparency where people are interacting with a system, and conformity processes, with separate obligations for general-purpose models. Name the NIST AI Risk Management Framework as the voluntary structure most US enterprises map onto, and ISO/IEC 42001 as the management system standard an organisation can be certified against. Add the sector layer that applies: protected health information handling and software-as-a-medical-device questions in healthcare, model risk management expectations in banking, adverse action and explainability duties in consumer credit, records retention in the public sector. Then say the sentence that earns trust: these obligations and their timing have been amended more than once, so I would confirm the current position with counsel before we commit to a date in a plan.

Knowing what not to build, and being able to say it to a team already building it

The commodity line has moved fast enough that the most expensive mistakes now are custom components nobody needed: a bespoke vector store beside a data platform that ships one, a homegrown orchestration layer, a hand-rolled evaluation tool, a guardrail service duplicating a licensed one. Every undifferentiated component is a multi-year staffing commitment. Interviewers for this role are hiring someone who will delete scope, and they ask about it deliberately.

Show it: Have a named test and use it unprompted: build where the differentiation is and the business logic is specific, buy where the capability is commodity, and never own an undifferentiated component you will have to staff. In the whiteboard round, say explicitly what you are not building and why, including one thing the interviewer probably expected you to build. In the behavioural round, tell the story of a project you reduced or killed, with the number that justified it, and what you did to keep the team on side while doing it.

What a screen is looking for

These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.

Mistakes that cost people this job

Drawing the model box first and building the architecture around it.

Draw the data path first: sources, who is allowed to see what, ingestion, index, retrieval with identity filters, prompt assembly, and only then the model. Interviewers score the order because it predicts how you will spend the first month of the job. The model is the most replaceable component in the diagram and the one the customer cares least about by month three.

A diagram with no numbers on it.

Produce five numbers unprompted even when the scenario omitted them, labelled as assumptions: requests per day and at peak, tokens in and out per request, cost per thousand requests, a p95 latency target, and the quality threshold required to go live. Say the arithmetic out loud. An architecture without quantities is a slide, and every interviewer in this field has seen a thousand slides.

Designing the greenfield system instead of the one that has to integrate with what exists.

Ask what the systems of record are, whether you can read them live or only through a nightly extract, what cannot be written to, and which data quality problem predates the project. Then design around those constraints explicitly and say which you would try to change and which you would accept. The mainframe, the nightly batch and the inconsistent entitlement model are the project. The retrieval pipeline is the easy part.

Treating evaluation as somebody else's problem, or answering a human will check it.

Design the measurement as part of the architecture: the labelled set and who labels it, the scoring method, the baseline, the go-live threshold agreed with the business before build, the holdout, the online metrics after launch, and the named owner of the set. If there is a human reviewer, quantify them: how many minutes per case, what the override rate is, and what you will do when review becomes the bottleneck.

Promising that the system will not hallucinate, or offering a prompt-level fix for an authority problem.

Say that you will measure how often and in which direction it is wrong, bound what a wrong output can cause, make the wrongness visible to the person who can catch it, and review the numbers on a schedule. For anything that acts, put the authority outside the model: scoped credentials, server-side caps, idempotency, reversibility, human confirmation on irreversible actions, and logging at the tool-call layer. A filter or a better prompt as the headline control is scored as a fail.

Pasting the whole corpus into a long context window because the window is now cheap.

Price it per call before you propose it, then say which constraint decides: long context when the material is small, bounded and needed whole, retrieval when the corpus is large or changing or carries per-document permissions, and retrieval whenever you need a citation the reviewer can click. The same tokens on every request is a recurring bill, and a long window does not give you entitlement filtering.

Proposing a custom build of something the customer already licenses.

Check the existing platform first and say you are doing it. Build where the business logic is genuinely specific, buy where the capability is commodity, and name the undifferentiated components you refuse to own because they become a multi-year staffing commitment. If you built custom infrastructure in 2024 that is now a platform feature, say why it was the right call then and what you would do today.

Quoting a regulatory deadline or a compliance date as settled fact.

Describe the obligation and its structure (risk tiers, governance of training and validation data, human oversight, documentation, transparency) and then say that the timing in this area has been amended more than once and that you would confirm the current position with counsel before putting a date in a plan. In a room containing a legal or risk partner, that sentence builds more credibility than a confident date, and a wrong date destroys it.

Having no answer on data handling when a non-technical reviewer asks.

Know, for the platform you are proposing, whether the provider trains on customer data and where that is written down, what the log retention is and whether it can be set to zero, which region inference runs in, and what the subprocessor position is. Say you would verify it against the current contract rather than quoting a marketing page. Deals and internal approvals stall on this more often than on architecture.

Leading with the vendor's product rather than the customer's decision.

Spend the first eight to ten minutes of any design conversation on the user, the decision the output supports, the volume, the constraints and the cost of being wrong. Map to products afterwards. Vendor interview panels reject product-first answers harder than customers do, because an architect who pitches before diagnosing creates the escalation they will have to clean up.

Submitting a resume that is a tool inventory and a list of responsibilities.

Write bullets with four parts: the decision, the constraint that forced it, the alternative you rejected and why, and the measured outcome with a unit. Name the artefacts you produced (reference architecture, decision records, cost model, evaluation plan, rollout plan, runbook) because those are the deliverables of the job. Cut the list of every model you have touched.

Collecting certifications instead of producing one artefact.

Get the professional or expert tier cloud architect credential for the platform your target employers use, because it is the screening filter, and let an employer pay for the rest. Then spend the time on one written reference architecture with a cost model, an evaluation plan and a decision record. The badge has nothing to say in the whiteboard round. The document answers four of its questions.

Failing the customer simulation by defending instead of conceding.

Restate the objection in the stakeholder's own words, concede the part that is true, narrow your claim to what you can actually support, and attach a next step with a date and an owner. If you do not know, say so and say when you will come back. Senior people in that room have heard confident nonsense before and are testing for exactly this.

Never saying what you would not build, and never committing to a recommendation.

Close every design round with three sentences: the residual risk you are accepting and its compensating control, the thing you are deliberately not building in phase one, and who operates this after you leave. Then state a recommendation and the condition under which you would reverse it. Being wrong is survivable in this round. Being vague is not.

Questions people ask

What does an AI solutions architect actually do?

An AI solutions architect decides, writes down and defends how a specific AI workload will actually run inside a specific organisation: which data it touches and under whose permissions, which model and what replaces it when that model is retired, how quality is measured and who signs off on the threshold, what it costs per thousand requests, how it degrades when a dependency fails, and who operates it afterwards. The artefacts an AI solutions architect produces are a reference architecture, a set of architecture decision records, a cost model, an evaluation plan with a labelled set, a phased rollout plan and a hand-off runbook. The job is less about drawing diagrams than about removing scope: deciding what not to build is a larger part of the role than proposing what to build.

Can I move from cloud architect into an AI solutions architect role, and how long does it take?

Moving from cloud architecture into an AI solutions architect role is the most common route in, and most of the skill transfers directly: identity, authorization, private connectivity, data residency, tenancy, disaster recovery with numbers, cost allocation, and the discipline of writing a decision down with its rejected alternatives. Three gaps decide the loop for an AI solutions architect arriving from cloud: evaluation (proving quality when the output is not deterministic), unit economics (tokens, context length, caching, GPU memory), and the authority model for agents that call tools. From a working cloud solutions architect job, 8 to 16 weeks of deliberate work on those gaps, expressed as one real written architecture with a cost model and an evaluation plan, is enough to interview well. From enterprise architecture with no recent hands-on work, budget 4 to 8 months.

Which certifications actually help an AI solutions architect get hired?

For an AI solutions architect the only credential that reliably changes a screening decision is a professional or expert tier cloud architect certification on the platform the employer uses or sells: AWS Certified Solutions Architect Professional, Google Cloud Professional Cloud Architect, or Microsoft Certified Azure Solutions Architect Expert. The vendor AI credentials (Google Cloud Professional Machine Learning Engineer, Microsoft Certified Azure AI Engineer Associate, and the current AWS machine learning exam, whose name has changed recently enough to be worth checking) are a useful second, and Databricks, Snowflake or NVIDIA credentials matter inside shops built on those platforms. Certifications matter most at cloud vendors and their partner consultancies, where certified headcount feeds partner competency tiers, and least for in-house roles. Exam names and prices in this area change often, so an AI solutions architect should confirm the current exam on the issuer's own page, and should choose one written reference architecture with a cost model over a fourth certificate.

What does the reference architecture whiteboard interview test?

The reference architecture whiteboard round tests whether an AI solutions architect has a repeatable process that works out loud on a scenario they have never seen. You get a business situation described in two or three minutes, usually with a system of record, a regulated constraint and a volume, then 45 to 75 minutes to ask before drawing, draw the data path before the model, mark the boundary where untrusted text enters, attach real numbers (requests per day, tokens per request, cost per thousand requests, a p95 latency target, a quality threshold), design the evaluation as a box on the diagram, accept one residual risk with a compensating control, and say what you would not build and who operates it afterwards. The three failures that end this round for an AI solutions architect are a diagram with no numbers, a greenfield design that ignores the mainframe and the nightly batch, and the phrase a human will check it offered in place of a measurement.

Do I need to write code to be an AI solutions architect?

An AI solutions architect is not hired to ship production code, but is expected to be hands-on enough to have built the thing small and to have numbers from doing it: standing up a retrieval path with permission filtering, running an evaluation set, reading a trace, and doing the token and memory arithmetic. The deep dive round finds out quickly, so an AI solutions architect who last touched a terminal four years ago needs to fix that before interviewing rather than after. The amount of code varies by employer: a software vendor's implementation architect writes a good deal, a vendor field architect writes demos and reference implementations, a consultancy architect reviews more than they write, and an in-house platform architect writes reference patterns other teams copy.

What is the difference between an AI solutions architect and an AI engineer?

An AI engineer builds and owns the running system, while an AI solutions architect decides its shape and is accountable for the parts no single team wants: the permission model in the retrieval path, the evaluation set and its threshold, the cost envelope, the exit plan for when the model version is retired, and the hand-off to whoever operates it. An AI solutions architect usually works across several workloads at once and spends a large share of the week with stakeholders, security reviewers and finance rather than in an editor. The practical test in a posting: if the deliverable is shipped software it is an engineering role, and if the deliverable is a decision somebody else implements it is an architecture role.

How much does an AI solutions architect earn, and where should I check?

No BLS occupation code names the AI solutions architect title, so the honest answer is to check sources rather than trust a quoted band. The nearest authoritative US baselines are OES 15-1241 Computer Network Architects, 15-1252 Software Developers and 15-1299 Computer Occupations All Other, all of which undercount vendor field pay. For an AI solutions architect in a cloud or model vendor's field organisation the package is usually base plus a variable component tied to consumption or quota, and the variable share differs enough between employers that you should ask for the split in writing before comparing the number to an in-house salary. Check posted ranges under pay-transparency rules in states including California, Colorado, Illinois, New York and Washington, the employer's own posting, published levels data for the large vendors, and H-1B labour condition filings for named employers.

Is a machine learning background or a PhD required for an AI solutions architect job?

Neither a PhD nor model training experience is required for an AI solutions architect, and most people in the title arrived from cloud architecture, consulting or platform engineering rather than from research. What is required is enough mechanics to avoid being confidently wrong: how retrieval selects documents and how chunking changes answers, why context length drives cost, when a long context window beats retrieval and when it just repeats the bill, what fine-tuning fixes that retrieval does not, what an evaluation harness is and what a rubric scored by a model is worth, what a tool call and an agent loop actually are, and what sets throughput on a GPU. An AI solutions architect candidate from a research background usually has the opposite gap and should expect the loop to probe identity, networking, tenancy and operations hard.

What belongs on an AI solutions architect resume, and what gets ignored?

An AI solutions architect resume should be built from decisions with measured outcomes: the choice made, the constraint that forced it, the alternative rejected and why, and a number with a unit (requests per day, p95 latency before and after, cost per thousand requests, evaluation set size and the threshold held, workloads taken to production and how many survived a year). Name the artefacts produced, because they are the job: reference architecture, decision records, cost model, evaluation plan, rollout plan, runbook. What gets ignored on an AI solutions architect resume is a list of every model and framework touched, verbs like spearheaded with nothing measured after them, a diagram screenshot with no numbers, and more than three certifications. Claiming production experience you do not have is worse than useless in a field small enough for the deep dive round to find out in four minutes.

Is AI solutions architect a durable job, or a title that disappears when the hype settles?

The work an AI solutions architect does is durable even if the title gets relabelled, because the underlying problems are permanent: entitlement-correct access to data, measuring a system whose output is a distribution, modelling cost in a currency that scales with context, integrating with systems of record that cannot be rewritten, and deciding what not to build. What has already shifted is the emphasis. Employers now ask an AI solutions architect what reached production and stayed there rather than what was prototyped, which favours candidates with unglamorous evidence such as an evaluation set and a cost model over candidates with a demo. Expect the title to drift toward AI platform architect, and toward plain solutions architect with AI assumed, in the way that cloud architect stopped being a separate species.

Put this on a resume in about a minute

Paste your history once and point it at the AI Solutions Architect posting you are looking at. No account, no card.

Build my resume free More roles