Data & Analytics

How to get hired as a data architect in 2026-27

The short answer

To get hired as a data architect in 2026-27, show decisions rather than tools: what you chose, what you rejected, why, and what it would have cost to be wrong. The loop is usually a hiring-manager conversation, a 60-90 minute data system design session, a deep dive on a past decision, and a governance interview, and the governance interview is where senior candidates most often lose the offer. Expect at least one AI question: where embeddings live, how document permissions carry into a retrieval index, how you would prove what data a model was trained on. Resume screens match on scale with units and named modeling patterns, not on a list of forty tools.

What the role ownsThe shape of the data: conceptual, logical and physical models; storage, table format and catalog topology; the contracts between producing systems and consuming teams; and the governance layer. Accountable for the decisions that are expensive to reverse.
Closest confusionsA data engineer builds and operates pipelines. An analytics engineer owns transformation and the metric layer. A platform engineer owns the infrastructure underneath. The architect decides the shape all three build against.
Typical loopRecruiter screen, hiring-manager conversation, a 60-90 minute data system design session, a deep dive on a past architecture decision, and a governance conversation. Three to six weeks at a product company; longer where an architecture review board meets on a fixed cadence. Multi-day take-homes are rare at this level.
Who screens youA recruiter matching platform and modeling vocabulary, then a head of data, director of data platform, or, for the AI-program variant, a head of AI or engineering. Regulated employers add a governance, privacy or security reviewer who holds a veto.
PayNo single band. Start from the US Bureau of Labor Statistics OES code 15-1243 (database architects) for the national median and percentile spread, then levels.fyi and the ranges employers post under state pay-transparency laws. Which of the three job types you are in moves pay more than the title does.
Resume lengthTwo pages. One page strips the industry, scale and regulatory context that makes an architect's experience legible to a panel, and leaves a list of verbs.
Evidence that landsSource-system counts, row and byte volumes, dual-run duration and what reconciled exactly versus what carried a stated tolerance, warehouse spend with the lever named, access-approval times, and incident rates before and after data contracts.
What changed by 2026The data platform stopped being only a reporting system. It now feeds retrieval and model training, so lineage, classification and purpose limitation became load-bearing, and unstructured content became an architect's problem.

Data architect vs data engineer vs analytics engineer vs platform engineer

The job is deciding the shape of an organization's data, and being accountable for the decisions that are expensive to reverse. Concretely: the conceptual, logical and physical models; where data lands, in what table format, under which catalog; the boundaries between domains and who owns each; the contracts between systems that produce data and the teams that consume it; and the governance layer (classification, lineage, access, retention) that makes the whole thing defensible to a regulator, an auditor or a customer.

A data architect writes more prose and more diagrams than code, and most of the prose is about options that were rejected. That is the part of the job people underestimate, and it is the part interviews actually test. Anyone can name the thing they built. Far fewer can tell you what else was on the table, what it would have cost, and what would have to be true for the other choice to have been right.

Three adjacent roles get confused with this one constantly, and applying to the wrong one wastes months. The distinction is not seniority. It is what you are measured on.

The three kinds of data architect job: read the posting's verbs

"Data architect" describes at least three genuinely different jobs. They want different evidence, they interview differently, and a resume tuned for one reads as underqualified for another. Work out which you are looking at before you write a word. The fastest test is the verbs in the posting.

How hiring for this role works in 2026-27

Two separate things happened to this market and they are worth keeping apart, because merging them produces a tidy story that is wrong. Data team hiring contracted from 2023 onward on cost grounds, the end of cheap capital and general headcount discipline, well before AI-assisted development was materially adopted for pipeline work. Since then, AI-assisted development has changed what the remaining roles do rather than being the original cause of the contraction. Architect postings have stayed visible through both, though that is a pattern I would describe rather than a measured share: someone has to be accountable for the platform an AI program is now leaning on, and the cost of a bad architectural decision went up once models started being grounded on the output.

Application volume went the other way. AI-assisted applying means a posting now draws far more applicants and nearly all of them look tailored, because they are. Two things cut through: a referral, and specificity a general-purpose rewrite cannot fake. A line naming a source system, a row count and what reconciled exactly reads as lived. A line saying "drove enterprise data strategy" reads as generated, whether or not it was.

The first pass on your resume is increasingly a model summarizing it for a recruiter, followed by a human reading the summary. That rewards plain structure, explicit nouns and numbers with units. It punishes design-led layouts that scramble reading order and skills buried in prose. Do not use a two-column template that interleaves columns when parsed. Put the platform names where they cannot be missed.

Here is the loop, stage by stage.

What data architects are paid, and what actually moves it

There is no single band for this title, and anyone quoting you one number is averaging across three different jobs. Get your own figures rather than trusting a blog: the US Bureau of Labor Statistics publishes database architects separately under OES code 15-1243, with a national median and a 10th-to-90th percentile spread by state and metro area; levels.fyi and Glassdoor carry self-reported bands with company names attached; and employers in several US states and across the EU now publish ranges on the posting itself under pay-transparency rules. Read the posted range for the actual companies you are targeting. It is the only figure that is both current and about you.

Five things move the number more than the title does.

Which of the three job types you are in. The regulated enterprise variant usually pays a defined architecture band with little equity. The scale-up variant is often levelled on the engineering ladder as staff or principal, which means base plus meaningful equity and a wider spread. The AI-program variant has been paying a premium because the supply of people who have actually shipped a permissioned retrieval path is thin.

Which ladder the role sits on. This is the question candidates skip and then regret. In many companies the staff data engineer band and the data architect band are the same band. Ask in the first or second conversation whether the role is on the IC engineering ladder or in a separate architecture function, and what the level maps to. That single answer determines your range more than any negotiation you do later.

Industry, and whether you have its data shapes. Financial services, insurance and health care pay for domain proximity because the ramp on their source systems is long. Location still matters for the enterprise roles, which are more often hybrid, and less for the AI-program roles, which are more often remote or remote-leaning.

Whether you can show a cut-over. Designing a target state and running a migration to completion are priced differently, because one of them has been tested.

The resume: evidence that lands, and what to keep confidential

The resume's job is to make a stranger believe you have owned consequential decisions at a scale near theirs. Almost everything that fails does so by describing scope instead of decisions.

What reliably gets skipped: "responsible for enterprise data architecture"; a forty-item tool list; certifications with no system attached; "led data strategy"; vendor course badges; anything that would be equally true of the other six people on that team. A panel cannot distinguish you from your org chart on that material, so it does not try.

What lands, in rough order of power. Every figure below is an example of the shape, not a benchmark, use your own numbers, because you will be questioned on them.

The design interview: what is actually being graded

Expect 60 to 90 minutes on a shared canvas. Common prompts: design the data platform for a company with these systems and these consumers; we have 40 sources and one team, what do you do first and why; model a subscription business where plans change mid-cycle; design retrieval for an internal assistant over our contracts and support tickets. That last one is now routine and is where otherwise strong candidates come apart.

The grading is less about the final diagram than about the first five minutes. Strong candidates ask about grain, volume, latency tolerance, who consumes it, and what regulatory regime applies, before drawing anything. Candidates who start sketching a medallion architecture in the first thirty seconds are demonstrating a default, not a decision.

Name a trade-off, then pick, then say what would change your mind. "Batch hourly, because the only consumer with a sub-minute requirement is the fraud team and they already have their own stream. If a second real-time consumer shows up, this becomes streaming ingestion with a batch mirror, and here is the cost of that." Choice, reason, reversal condition. That single move separates architects from senior engineers more reliably than anything else in the loop.

Say what you would not build. Most candidates over-architect under pressure: streaming everywhere, a full federated mesh for a sixty-person company, seven layers where three will do. Right-sizing is rarely held against you as long as you say out loud that you are sizing to the organization and name the threshold at which you would revisit; over-building routinely is, and so is silently under-scoping a genuinely large problem. On mesh specifically, size the answer to what survived rather than the branded methodology: "two domains with owned contracts and a shared catalog now; I would push ownership further out as teams get their own engineers, which is the part of mesh that actually held up, the full federated model needs platform investment an organization this size cannot staff."

Be ready to sequence, not just to design. "What do you do in the first 90 days" is asking whether you can order work against risk: what you stabilize first, what you leave alone, what you can migrate without a cut-over, and what needs a dual run. A target-state diagram with no order of operations is a slide, not a plan.

You will very likely be asked about table format and catalog, and this is the question a 2023-vintage answer fails. Have a position on all five of these.

The governance interview, question by question

This is the stage senior candidates most often underprepare, partly because it is easy to mistake governance for paperwork. In an interview it is not. Every question is really asking one thing: do you know the difference between a rule that is written down and a rule that is enforced by the system?

Getting there from data engineering: and where these jobs are actually posted

Most data architects arrive from data engineering, usually after they have owned a domain end to end rather than after a set number of years. The transition fails in a predictable way: the candidate's evidence is still about throughput and uptime, which are engineering outcomes, while the job is about decisions and their reversal cost. At three to five years in, the realistic target is not the architect title, it is owning a domain, with the decision records to prove it, and applying to the scale-up and AI-program variants where that is enough.

Start generating architect evidence before you have the title. Write architecture decision records for the choices your team is already making: one page each covering context, options, decision, consequences, and what would make you revisit. Six months of those is a portfolio, and it is exactly what the past-decision deep dive asks for. It is also the cheapest way to find out whether you enjoy the part of the job that is writing.

Volunteer for the work that is architectural in substance regardless of title: the source-system assessment nobody wants, the cost review, the classification exercise, the migration plan, the data contract negotiation with a producing team that does not report to you. That last one is the real skill, and a panel can tell within minutes whether you have done it, so here is how to do it without seniority. Get the meeting through your manager, framed as a reliability problem rather than a governance one, because "your change broke our finance report on the 14th and cost us a day" gets a calendar invite and "we would like to introduce data contracts" does not. Open with the incident, the date and the cost, not the proposal. Ask for notification first, not enforcement: tell us before the schema changes. Offer to absorb the work, you write the tests, you own the contract file, they just review it. If they say no, instrument the breakage, count it for a quarter, and come back with the number. That sequence is itself an interview story.

Close the two gaps engineers most often carry into an architect loop. The first is modeling vocabulary: be fluent on grain, conformed dimensions, slowly changing dimension types and when each applies, and know the shape of Data Vault and normalized enterprise modeling even if you have only ever built star schemas. The second is money: know what your platform costs, which workloads drive it, and what you would cut first. Architects get asked about cost in almost every loop and engineers frequently cannot answer.

Keep a hands-on proof, and make it the right one. The fastest way to be quietly downgraded in a panel is to be an architect who can no longer build, the engineering peer is specifically checking. The highest-value build in 2026 is a small permissioned retrieval system, because it maps directly onto the design prompt and its governance follow-up. Four properties are what make it count: a corpus of a few thousand documents with per-chunk group identifiers filtered at query time; hybrid retrieval with a reranker so exact identifiers actually match; the embedding model and version stamped on every row with a dual-index path for upgrading; and a measured cost per answer, broken into retrieval and generation. A weekend gets you a working version. That is not a portfolio piece to link on your resume, it is the thing that lets you answer four interview questions from experience instead of from reading.

On sourcing, which most guides skip: the title is unreliable, so search for the shape. Search "data modeler" and "enterprise architect" alongside "data architect" for the enterprise variant; search staff and principal data engineer postings and read the description for architecture language for the scale-up variant; and for the AI-program variant search engineering and AI platform postings, because those roles frequently do not appear under data at all. Referrals matter more here than in most roles because the community is small: talks at data and catalog project meetups, contributions to the open catalog and table-format projects, and the authors of the internal-platform writeups you actually learned from are more realistic routes than cold outreach. Remote is common for the AI-program and scale-up variants and less common for the regulated enterprise one, where review-board culture still runs on rooms.

Finally, be honest with yourself about whether you want it. In many companies the staff data engineer band and the architect band are the same money. What the title changes is the day: more writing, more meetings, more influence over teams you do not manage, and materially less time in an editor. If you want to ship code most days, the staff engineering track pays comparably and costs you less of the keyboard.

Working with AI in this role

What a data architect has to know about AI in 2026-27

The honest version of what changed: the data platform stopped being a reporting system. It now also feeds model training, fine-tuning and, far more commonly, retrieval for assistants and agents running against the company's own content. That moved several things from someone else's problem to yours, and it changed what an interviewer expects you to have an opinion about.

Three shifts drive everything below. Serving requirements changed, because a retrieval path needs chunk-level provenance, permission enforcement at query time and sub-second lookups, none of which a nightly dimensional mart was built for. Governance became load-bearing rather than documentary, because "where did this field come from and may we use it for this purpose" now has legal and contractual weight. And unstructured content (contracts, tickets, transcripts, PDFs) became an architect's responsibility, at a volume nobody inventoried.

What has not changed is the judgment. A model will write DDL, dbt models, mapping documents and migration scripts quickly and mostly plausibly. It will not tell you the grain of a fact table in a business you have not interrogated, which source is authoritative when two disagree, whether a dimension is slowly changing and of which type, or how much reversal cost a decision carries. What has shifted is the ratio: generated transformations now arrive faster than review capacity grew, so the architecture's ability to constrain and check what lands matters more than anyone's ability to author it.

In interviews this shows up less as "do you know about LLMs" and more as a design prompt, design retrieval over our contracts and tickets, plus a governance follow-up about permissions and provenance. Candidates who have only read about it draw an embedding box with no access-control story. Candidates who read about it more recently draw an embedding box with an ACL filter and no lexical retrieval, no reranker and no way to tell whether it works. Both are tells.

Retrieval that carries the source permission model through to query time

The characteristic failure of internal RAG is answering a question using a document the person asking could not have opened in the source system. Permissions live in the source; the index is a copy; nothing propagates by itself. This is the first thing a good interviewer probes and the first thing a real deployment gets wrong.

Show it: Describe a design where access control reached query time: ACLs or group identifiers stored per chunk and applied as a filter at search, re-sync when source permissions change, and what you did about documents whose permissions change faster than the index refreshes. Name the corpus you deliberately excluded because the permission model could not be honored, the exclusion is the part that sounds lived.

Retrieval quality as an architecture concern, and evaluation data as a real asset

Pure dense vector search fails on exactly what people search internal corpora for: contract clause numbers, product codes, ticket IDs, error strings, acronyms. The working default for enterprise document retrieval is hybrid, lexical (BM25) plus dense, with a reranking stage, metadata filtering and chunking tied to document structure rather than a fixed token count. And none of those choices can be defended without evaluation, which in most organizations lives in a spreadsheet and a notebook, unversioned and unowned, so nobody can say whether last quarter's change helped.

Show it: State your retrieval design with the lexical half included and say what the reranker cost you in latency and money. Then show that you treated golden question sets, eval runs and human feedback like any other asset: a schema, versioning tied to the embedding and prompt versions, retention, ownership, and access rules for anything containing customer content. Even a modest version of this is unusual enough to be remembered.

The embedding lifecycle, including the re-index you will definitely have to do

Changing embedding models always requires re-embedding the whole corpus, vectors from different models are not comparable, so there is no clever design that avoids it. The mistake is treating that as a failure rather than a budgeted, recurring cost. What lifecycle design actually buys you is knowing what is in the index, and being able to do the re-embed without an outage.

Show it: Take a position on where vectors live: inside the existing platform (pgvector, a warehouse-native index, OpenSearch) or a dedicated store, and justify it from that workload's scale and latency, not from vendor preference. Then describe four specifics: the embedding model and version stamped on every row so you can tell what needs re-embedding; a durable store of chunk text and metadata so you re-embed from your own data rather than re-crawling sources; change detection so steady-state re-embedding is incremental; and a shadow or dual index so the upgrade cuts over with no outage and you can A/B the new model before switching reads.

Agent access: machine identity, scoped writes, and what MCP exposes

Most discussion of AI and data is about the read path. The live architecture question in 2026 is the write path. An agent that takes actions needs its own identity, an authorization model distinct from the human who invoked it, and an audit trail of what it did to what. "Which service principal did the agent use, and can it write to production?" is being asked in real design reviews, and MCP is now the common way tools and datasets get exposed to agents, which makes the set of exposed tools a governance surface an architect owns.

Show it: Describe the access model: machine identity per agent rather than a shared service account, delegated on-behalf-of authorization so the agent cannot read what the invoking human cannot, read-only by default with a sanctioned and reviewed write path for anything else, rate and blast-radius limits, and agent actions logged at the same fidelity as human access. If you have exposed data through MCP, say which tools and datasets you chose to expose, to whom, and what you refused to expose.

Lineage whose terminal nodes are models and indexes, not dashboards

Traditional lineage stops at the BI layer. When the consumer is a model, the useful question is which datasets contributed to which model version or which grounding index, and that edge is usually missing. Without it you cannot answer an audit question, assess the blast radius of a bad upstream load, or respond when you discover a licensing restriction on data you already trained on.

Show it: Show a lineage graph whose endpoints include a model version or a retrieval index. Say how the edges were captured (orchestration metadata, dbt exposures, catalog integration) and give one concrete question it let you answer that you previously could not, such as which index needed rebuilding after a bad load, or which model versions saw a dataset you later had to withdraw.

Purpose limitation on training and grounding data

Consent, purpose limitation, contractual restrictions on vendor-supplied data and sector rules all now apply to a use case they were never assessed for. "We had the data" is not the same as "we may use it to train." This is the question a legal reviewer asks first, in every jurisdiction, regardless of whether any AI-specific statute applies.

Show it: Describe a classification scheme where the label drove enforcement rather than documentation, and say how purpose was attached to a dataset and checked at use. The strongest version of this answer is a time you said no: a dataset you kept out of a training or grounding corpus, the specific restriction behind it, and the alternative you proposed instead.

Bringing unstructured content under the same governance as tables

Object storage full of PDFs, transcripts and tickets is now a primary input, and it very often contains personal data nobody inventoried. The tooling to govern it exists (Unity Catalog volumes, Purview's file scanning, document assets in the enterprise catalogs) but in most estates the object storage was never onboarded, so the content has no owner, no classification and no retention. That is an organizational gap, not a tooling one, and closing it is what a growing share of architect postings are actually about.

Show it: Give the numbers: how much content, in what formats, and what you did. The metadata model you defined for documents, how classification was applied at that volume (sampling, automated classifiers, whatever you actually used), what retention you set and how it is enforced, and how chunking and metadata were designed so provenance survives into a retrieval answer.

The unit economics of an AI feature, which bill nothing like query workloads

Scan-based query pricing does not describe any of this, and the cost is not where people expect. In almost every production assistant, generation tokens dominate the bill by an order of magnitude or more; retrieval, embedding and index storage are the small terms. A candidate who answers a cost question with vector-index economics and never mentions tokens looks like someone who costed a pilot.

Show it: Give a real per-unit figure you measured: cost per answer, split into input and output tokens, and name the lever you pulled. Retrieved context is what drives input tokens, so chunk size and top-k are direct cost levers; so are prompt caching, routing cheap requests to a smaller model, and reserving the expensive model for the cases that need it. Then the smaller terms: cost of a full re-embed of a corpus of a stated size, and index storage with the memory-versus-disk trade-off for the index type you chose.

Guardrails for a much higher volume of machine-generated SQL

Analysts and engineers now generate transformations with model assistance faster than review capacity grew. The architectural response is not to ban it; it is to make the platform reject what is wrong, so review is not the only control.

Show it: Describe one control you put in place and its measured effect: required tests or contracts failing CI, ownership enforced at model creation, a certified metric layer that cut duplicate and contradictory definitions, or a write path that keeps ungoverned tables out of a production schema. Pair it with the number that moved.

What a screen is looking for

These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.

Mistakes that cost people this job

Applying to data engineer postings with an architect resume, or the reverse. Both read as mismatched to the opposite screen, and the rejection tells you nothing.

Read the posting's verbs. Design, define, standardize, approve, review means architect. Build, ingest, operate, on-call, optimize runtime means engineer. The same career supports both resumes, but not the same document, change which bullets lead.

Listing tools and scope instead of decisions and scale: a forty-item skills block, and bullets like "owned the enterprise data platform" that a panel cannot calibrate against.

One line per decision: what you chose, what you rejected, the reason, the outcome, opened by numbers with units: source systems, terabytes, daily model runs, consuming teams, spend. Keep the tool list to one block at the bottom, for the keyword screen only.

Claiming governance with no mechanism: "established data governance," "defined the data strategy," "implemented a RACI."

Name the mechanism and the number it moved. A classification scheme that drives actual grants. Ownership enforced in CI. A contract that fails a build. Then: approval time, incident rate, ungoverned tables eliminated.

Treating AI as a line item: "familiar with LLMs and vector databases", on a 2026 architect resume.

One concrete architecture decision. Where embeddings live and why. How document permissions carry into retrieval. The cost per answer and the lever that changed it. A dataset you excluded from a training corpus and the restriction behind it.

Over-architecting in the design interview: streaming everywhere, a full federated mesh for a sixty-person company, seven layers where three would do.

Size the answer to the organization and say so out loud. "At this size, two domains with owned contracts and a shared catalog; I would push ownership further out as teams get their own engineers." Right-sizing is rarely held against you; over-building routinely is.

Drawing before asking. Sketching a medallion architecture in the first thirty seconds demonstrates a default, not a decision.

Spend the first five minutes on grain, volume, latency tolerance, consumers and regulatory regime, and write the answers on the canvas. Every later trade-off then has a stated reason, which is what is actually being graded.

Presenting an inherited architecture as one you chose. It collapses the moment an interviewer asks what else you considered.

Be explicit about what you inherited, what you changed and why. "I inherited the warehouse; my decision was to stop expanding it and route the new event domain to object storage, because…" A smaller real decision beats a large borrowed one.

Hiding the migration that went badly, or omitting cut-over mechanics entirely.

Give the mechanics: dual-run duration, what reconciled exactly versus what carried a stated tolerance and why, what broke, what you changed afterward. Panels have run migrations. A clean narrative with no scar tissue reads as a plan someone else executed.

Taking scale-up evidence into a regulated-enterprise loop, or enterprise evidence into a scale-up one. Greenfield speed reads as naive to a review board; standards and committees read as slow to a team of five.

Pick the variant your evidence fits and lead with what it values, cut-overs, master data and audit-surviving lineage for the enterprise; end-to-end ownership and a platform that survived a growth multiple for the scale-up. If you have under six years, the enterprise variant is not yet your market; stop spending applications on it.

Missing the AI-program roles entirely because you only search data job titles.

Search engineering and AI platform postings too. These roles are frequently titled staff engineer, principal engineer or AI platform architect, report into a head of AI or engineering, and never appear under a data org's listings, which is also why they are less crowded.

Publishing your employer's confidential figures, or handing a panel their actual internal design document.

Publish outcomes and orders of magnitude; band anything sensitive ("cut warehouse spend by roughly 40%") and never name a vendor's contract price or a control's weakness. Write a sanitized architecture decision record for interview use, it demonstrates the writing skill the role is measured on and carries no risk.

Questions people ask

What does a data architect do?

A data architect decides the shape of an organization's data and is accountable for the decisions that are expensive to reverse: the conceptual, logical and physical models; where data is stored and in what table format and catalog; the boundaries between domains and who owns each; the contracts between systems that produce data and the teams that consume it; and the governance layer of classification, lineage, access and retention. The output is designs, decision records and standards more than code. A useful working test of the role: if the decision can be undone in a sprint, it is probably not an architect's to make; if undoing it takes a migration, it is.

What is the difference between a data architect and a data engineer?

A data engineer builds and operates pipelines and is measured on whether data arrives (freshness, reliability and throughput. A data architect decides the shape the engineer builds against) the model, the platform topology, the domain boundaries, the contracts and the governance, and is measured on whether those decisions were right two years later. Engineers ship code most days; architects write more prose and diagrams and read more code than they write. In compensation terms the two are often closer than the titles suggest: in many companies the staff data engineer band and the data architect band are the same band, so the move is about what your day looks like more than about pay.

How do you become a data architect?

Most arrive from data engineering, but analytics engineering, database administration, BI development and software architecture are all normal routes, and the common requirement is not years, it is having owned a domain end to end and being able to defend the decisions. Start producing architect evidence before you hold the title: write architecture decision records for the choices your team is already making (context, options, decision, consequences, revisit conditions), and after six months you have exactly the material a past-decision deep dive asks for. Volunteer for the source-system assessments, cost reviews, classification exercises, migration plans and data contract negotiations with teams that do not report to you. Then target the two job variants that hire on demonstrated ownership rather than tenure: the scale-up architect and the AI-program architect. The regulated-enterprise variant is buying a decade of source-system scar tissue and will not look at you until you have it.

How much does a data architect make?

There is no single band, and any article quoting you one number is averaging across three different jobs. Source it yourself from three places: the US Bureau of Labor Statistics publishes database architects separately under OES code 15-1243, with a national median and a 10th-to-90th percentile spread by state and metro; levels.fyi and Glassdoor carry self-reported bands with company names attached; and employers in several US states and across the EU publish the range on the posting itself under pay-transparency rules. What moves the number most is not the title but which ladder the role sits on, at a scale-up the architect is often levelled as a staff or principal IC with equity, while a regulated enterprise usually has a defined architecture band with little equity. Ask which ladder in the first or second conversation, because that answer sets your range before any negotiation does.

What skills does a data architect need?

Five clusters. Modeling: grain, conformed dimensions, slowly changing dimension types, star schemas, and the shape of Data Vault and normalized enterprise models even if you have only built dimensional. Platform: a warehouse or lakehouse, an open table format (Iceberg or Delta) and a defensible position on catalog choice, change data capture, orchestration, dbt, and enough SQL, Python and Terraform to verify what you designed. Governance: classification that drives enforcement, lineage, access models, retention, residency, and the regimes that apply to your industry. AI and retrieval: permission-preserving retrieval, hybrid search and reranking, the embedding lifecycle, and the unit economics of an AI feature. And influence: an architect's authority runs mostly through teams they do not manage, which is why data contract negotiation is the skill panels probe for most directly.

What should a data architect resume include?

Two pages, with a one-line context header for each role giving industry, team size, source-system count, data volume and regulatory regime, so a reader can calibrate every bullet beneath it in a second. Lead each role with scale in real units, then give decisions rather than responsibilities: what you chose, what you rejected, why, and the outcome. Include a migration with genuine cut-over mechanics, a cost result with the lever named, a governance mechanism with the number that moved, and at least one concrete AI-adjacent architecture decision. Add a short Architecture decisions section of three to five lines, very few candidates do, and it is exactly what the interview examines. One caution most advice omits: band or genericize anything your employment agreement covers, and never hand over your employer's actual internal design document; write a sanitized one for interview use instead.

How has AI changed the data architect role?

The data platform stopped being only a reporting system and now also feeds model training and, far more often, retrieval for assistants running over the company's own content. Four things moved onto the architect's desk as a result. First, retrieval design that preserves the source system's permissions at query time, because the index is a copy and nothing propagates by itself. Second, the embedding lifecycle, re-embedding on change, stamping the model version on every row, and a dual-index path for the full re-index that any embedding-model upgrade unavoidably requires. Third, lineage and classification that reach past the warehouse to a model version or a grounding index, because provenance and purpose limitation are now legally material. Fourth, unstructured content (contracts, tickets, transcripts) which the tooling can govern but which most estates never onboarded. Separately, AI-assisted tooling generates transformations faster than review capacity grew, so the platform's ability to reject what is wrong matters more than anyone's ability to author it.

Is the data architect role being automated away?

No, and the pressure runs the other way. Models generate DDL, dbt models and migration code quickly, which absorbs routine build work lower in the stack, but they do not determine the grain of a fact table in a business nobody has interrogated, decide which source is authoritative when two disagree, or weigh how expensive a decision will be to reverse. What changed is the ratio: more generated SQL lands per human reviewer, so accountability for the shape of the platform and for the controls that catch bad work became more valuable. The contraction in data hiring that began in 2023 was driven by cost discipline rather than by AI, and architect postings have stayed visible through it, a pattern worth describing rather than a measured share, but one consistent with what candidates report encountering.

What is the hardest part of a data architect interview?

The governance conversation, which senior candidates consistently underprepare because they mistake governance for paperwork. The questions that fail people are mechanism questions: who owns a table and how that is enforced rather than merely recorded; how a sensitivity label becomes an actual access grant; how a deletion request is satisfied across a lakehouse with time travel, derived marts, BI extracts, backups and a vector index, and how snapshot expiry or VACUUM is the mechanism by which the delete actually completes; how residency works in a global warehouse and what it costs; and how you would prove to an auditor what data a model was trained on. Prepare one real governance failure, told in two minutes, with what you changed afterward. A near-miss or an inherited mess works, the reviewer is testing whether you noticed, not whether you were senior enough to fix it.

What certifications does a data architect need?

None are required and none will win a panel. DAMA International's CDMP and the cloud vendors' architect certifications do one useful thing: get you past a keyword screen at employers that filter on them, which regulated enterprises sometimes do. TOGAF, from The Open Group, appears on enterprise-architecture postings and signals familiarity with review-board culture more than technical depth. All are worth listing and none is worth discussing at length in an interview, where every question is about a decision you made and can defend. If you have limited time, six months of architecture decision records will do more for your candidacy than any certification on this list.

Put this on a resume in about a minute

Paste your history once and point it at the Data Architect posting you are looking at. No account, no card.

Build my resume free More roles