Data & Analytics

How to get hired as a data engineer in 2026-27

The short answer

To get hired as a data engineer in 2026-27, show one pipeline you owned in production and be able to state its volume, its freshness target, what it cost to run, and what broke. The stack that actually gets tested is narrower than the postings suggest: SQL at depth, Python that survives messy input, one warehouse or lakehouse (Snowflake, Databricks or BigQuery), dbt or SQLMesh for transformation, one orchestrator (Airflow or Dagster), and enough Apache Iceberg and catalog literacy to say where a table lives and which engine may write to it. No licence, registry or degree gates this role, and certifications decide an offer in one place only, consultancies and cloud partners that need certified headcount; the loop is a live SQL and Python screen, a take-home or live build, a pipeline design round graded largely on the questions you ask before you design anything, and increasingly a code-review or incident round. Candidates lose on two things above all: a take-home that cannot be safely re-run, and no numbers attached to anything they built.

What the role ownsGetting data from where it is produced to where it is used, reliably and at a known cost: ingestion and change data capture, transformation and modeling, orchestration and scheduling, storage layout and table format, data quality checks, and the serving layer the warehouse exposes to analysts, applications and now retrieval systems. Measured on whether data arrives: freshness against an agreed target, correctness, pipeline reliability, and warehouse spend.
Closest confusionsAn analytics engineer owns transformation and metric definitions, mostly in SQL and dbt, and rarely touches ingestion or infrastructure. A data platform engineer owns the infrastructure the pipelines run on. A data architect decides the shape and is accountable for decisions that take a migration to undo. A data scientist consumes what you build. Below roughly fifty people, one person is all of these and the title is whatever the founder typed.
Credential gateNone. No licence, no registry, no board exam, no required degree. Certifications (Snowflake SnowPro, Databricks Certified Data Engineer Professional, Google Professional Data Engineer, AWS Certified Data Engineer - Associate, Microsoft Fabric Data Engineer DP-700) change outcomes in one place: consultancies and cloud partners that need certified headcount to hold partner status. A product company reads the project section first and the certification line last.
Typical loopRecruiter screen, hiring-manager conversation, a live SQL and Python screen, then a take-home (commonly timeboxed at three to five hours) or an extended live build, a 45 to 60 minute pipeline design round, and a behavioral or cross-functional round. Three to six weeks at a product company. On the panel: a data engineer for the code review and often an analyst or data scientist who consumes your pipelines. Faster at consultancies and in contract hiring, where a strong SQL screen plus references can close in a week.
Time to job-readyFrom data analyst or BI developer, with SQL already in hand, months rather than years: the gap is engineering practice (git, dbt, CI, orchestration, on-call), and it can be closed inside your current job. From software engineering, a shorter gap than most people assume - modeling and warehouse mechanics. From no technical background, plan on a year of deliberate work with one deployed, scheduled pipeline running the whole time, because run history is the evidence.
PayThere is no single band, and the US Bureau of Labor Statistics does not publish data engineer as its own occupation. Triangulate: BLS OES 15-1243 (database architects) and 15-1252 (software developers) bracket it with state and metro percentile spreads; levels.fyi carries company-attributed bands at the technology end; and posted ranges are now mandatory in a growing list of US states (California, Colorado, New York, Washington, Illinois, Massachusetts and others). In the EU, Directive (EU) 2023/970 on pay transparency has a 7 June 2026 transposition deadline, so adverts in your own market are becoming better data than any salary survey.
Resume and evidenceOne page under about six years, two pages after. The screen looks for a named warehouse, a named orchestrator, SQL and Python, and at least one number with a unit on it. What lands: rows or events per day, table sizes in GB or TB, the freshness target and whether you hit it, pipeline count and how many you were on call for, monthly warehouse spend and the lever that moved it, backfill duration, incident count and time to detection, and one thing you deleted.
What changed by 2026Open table formats became the default interchange layer rather than a bet, so Iceberg and catalog questions appear in ordinary interviews. Transformation code is now generated faster than anyone can review it, so employers test review and constraint rather than authoring speed - which is why code-review and incident rounds have appeared in loops that did not have them. And the warehouse feeds retrieval systems and agents as well as dashboards, putting unstructured content, embedding pipelines and chunk-level permissions on the data engineer's plate.

Data engineer vs analytics engineer vs platform engineer: read the posting's verbs

Four jobs share this title's vocabulary and they interview differently. Applying to the wrong one is the most common reason a qualified candidate gets a rejection that explains nothing: the resume was read against a different rubric than the one it was written for.

The verbs tell you which job it is. Ingest, orchestrate, operate, on-call, backfill, optimize runtime means data engineer. Model, transform, test, document, define metrics means analytics engineer. Provision, Terraform, Kubernetes, cluster, cost means data platform engineer. Design, standardize, approve, review means architect. One career supports several of these resumes, but not the same document, and the fix is usually reordering bullets rather than rewriting them.

Company size changes the job more than the title does. At a startup, data engineer means ingestion, dbt, the BI tool, the on-call rota and explaining to the CEO why revenue moved. At a large company it can mean one Kafka topic family and nothing else. Ask in the first conversation which layers you would own, because the answer decides whether the next two years make you broader or deeper. Both are defensible; only one matches what you want after this job.

The stack that actually gets tested in 2026-27

A typical posting lists twenty-five technologies. A typical interview tests six or seven, and two of them carry most of the weight. Depth in one warehouse, one orchestrator, dbt or SQLMesh, Python and SQL beats shallow familiarity with everything, because every question that discriminates between candidates is a depth question: what happens on a re-run, what happens when the schema changes, what happens when yesterday's data arrives today.

Two shifts come up in 2026 interviews that were optional conversation three years ago. The first is the open table format. Apache Iceberg became the default interchange layer, with Delta Lake strong inside the Databricks estate, and the live question is no longer whether to use one but which catalog governs it - the Iceberg REST catalog interface, Unity Catalog, AWS Glue and S3 Tables, Snowflake's open catalog, Apache Polaris, Nessie, Lakekeeper - and therefore which system enforces access and which engines may write. Iceberg v3 added deletion vectors and row lineage, which matter in practice because they change how row-level deletes and updates perform on a large table. If you can say who writes, who only reads, where the metadata lives, and what compaction and snapshot expiry you scheduled, you are ahead of most candidates.

The second is that orchestration moved from task-shaped to asset-shaped thinking. Dagster started there. Airflow 3, released in 2025, pushed the same way: asset and data-aware scheduling, DAG versioning, a task execution API that permits remote execution, and a rewritten UI. This changes what a good design answer sounds like - you talk about the datasets a job produces, their freshness expectations and what is stale, rather than a chain of tasks that ran. Interviewers who have migrated notice immediately which vocabulary you use.

The transformation layer is consolidating, so ask rather than assume. dbt is still the default and what most take-homes expect, but the engine was rewritten (Fusion) and dbt Labs and Fivetran announced a combination in late 2025, so a team may be running dbt Core, dbt Cloud, Fusion, or SQLMesh, and the answer tells you how much of their build and state management is theirs to maintain. Asking which one and why is a better question than naming the tool.

What not to fake: streaming. Kafka, Flink and exactly-once semantics are real requirements at a minority of employers, and at those employers the questions get specific fast - consumer groups and rebalancing, offsets and replay, watermarks and allowed lateness, state backends, why Kafka 4.0 dropped ZooKeeper in favor of KRaft. If you have only run batch, say so and say what you would need to learn. The failure mode is claiming streaming experience and then proposing a streaming architecture for a problem whose freshness requirement was four hours, which tells the panel you reach for the impressive answer rather than the correct one.

How hiring for this role works, stage by stage

There is no licence and no registry. Hiring is entirely evidence-based, which cuts both ways: nothing stops you applying, and nothing vouches for you either, so the burden of proof sits on artefacts you control.

The common product-company loop runs five stages over three to six weeks. A recruiter screen that is mostly keyword matching and salary range. A hiring-manager conversation about what you owned - where most rejections actually happen, because candidates describe tools instead of ownership. A live SQL and Python screen, usually 60 minutes, which is the narrowest gate in most loops. Then a take-home or an extended live exercise. Then a pipeline design round, and a behavioral or cross-functional round with someone who consumes your data.

Four variations are worth planning for. Large technology companies often substitute a standard software engineering coding round for the SQL screen, so data engineers who have not practiced data structures get caught by a medium-difficulty algorithms question. Consultancies and system integrators compress the loop, weight certifications and client-facing communication more heavily, and may ask you to present rather than build. Contract and day-rate hiring (common in the UK, Netherlands, Germany and Australia) can run one technical conversation and a reference, close in days, and care far more about the specific platform in use than general capability; in the UK, ask whether the role is inside or outside IR35 before discussing rate, because the two numbers are not comparable.

The fourth variation catches people off guard: the review round and the incident round. In a review round you are handed a pull request, usually transformation code, and asked what is wrong with it. In an incident round you are told the morning dashboard shows revenue down sharply and the pipeline finished three hours late, and asked what you do, in order. Both exist because employers decided that producing code is no longer the scarce skill and judging it is. Prepare them explicitly: have a diagnosis order you can say out loud - is the data wrong or is the pipeline wrong, what changed, which upstream source, what do row counts look like per day, is this a duplicate or a drop - and resist guessing a cause before you have asked what changed.

One practical note on the take-home. Unsupervised take-homes are getting shorter or being replaced by paired and live builds, because an unsupervised submission no longer proves much about authorship. If a brief is open-ended and unpaid at more than about five hours, it is reasonable to ask to timebox it, or to offer to walk through an existing project of yours instead. Employers who refuse both are telling you something useful.

The SQL and Python screen: what is actually asked

This is the most predictable part of the process. The questions cluster tightly, because the underlying skills are the ones you use every day in a pipeline.

On SQL, the recurring patterns are deduplication (ROW_NUMBER partitioned by a business key, ordered by a timestamp, filter to one), running totals and period comparisons with window functions, gaps and islands, sessionization of event streams with a timeout, type 2 slowly changing dimension lookups where a fact must join to the dimension version valid at the event time, as-of joins, and the question that separates people who have shipped from people who have studied: at what grain does this result set sit, and what happens to your aggregate when a one-to-many join silently duplicates rows. Expect to be asked why a query is slow, and to have something better than 'add an index'. On a columnar warehouse the answer involves bytes scanned, partition pruning, clustering, join order and cardinality, and whether the expensive operation is a shuffle.

On Python, this is not a competitive-programming round at most employers. It is a messy-input round. Read from an API with pagination and rate limits, retry with backoff, validate records against a schema, handle the row with a date in the wrong format or a missing key without taking the whole batch down, write it somewhere, and make the whole thing safe to run twice. Being asked to write a pytest test for the function you just wrote is common, and being unable to is a real signal. Know why a generator matters when the file is larger than memory, and have an opinion about where pandas stops being the right tool and why Polars or DuckDB is often the answer in a pipeline.

The most useful preparation is unglamorous. Work window-function problems until dedup and sessionization are automatic rather than recalled. Install DuckDB locally and practice against a dataset you can inspect, so the feedback loop is seconds. And rehearse narrating: a candidate who says 'I am going to get to one row per customer per day first, then aggregate, because otherwise this join will fan out' reads as senior before the query is finished, while silence followed by a correct query reads as uncertain.

Ask whether you may use an AI assistant, and assume the answer shapes the round rather than removing it. Many employers now allow it and then interrogate the output, because that mirrors the job. If it is allowed, do not let it produce code you cannot defend line by line. The question after the generated query is always some version of 'why did it write it that way, and what breaks if the input has duplicates' - and that question is the actual interview.

The take-home and the design round: the rubric nobody shows you

The take-home is usually some version of: here is a messy source (a CSV dump, a public API, a sample of event JSON), build a pipeline that lands it somewhere queryable, model it, and tell us what you decided. The brief sounds open. The grading is not, and it is remarkably consistent across employers.

What is actually checked, roughly in order of weight: does it run from a clean clone following only the README; can it run twice without duplicating or corrupting data; are there tests, and do they test something that could plausibly fail rather than asserting that 1 equals 1; is the modeling sane and is the grain of each table stated; is bad input handled explicitly rather than crashing or silently dropped; and does the README say what you decided, what you traded away, and what you would do with another week. That last item is the cheapest way to look senior, and most submissions omit it.

What is not graded: cleverness, breadth of tooling, or a dashboard. Adding Kafka, Airflow, Spark and Kubernetes to a problem whose dataset fits in a few hundred megabytes is a negative signal, because the panel reads it as someone who cannot size a solution. Python, DuckDB or Postgres, dbt and a Makefile, with tests and a clear README, routinely beats a sprawling submission. Respect the timebox and say what you left out because of it: 'I spent the stated four hours on correctness and tests, and skipped incremental loading, which I would implement as follows' is a strong answer, not an apology.

The design round is a conversation, not a drawing exercise, and the first two minutes decide the grade. The interviewer says something like 'design a pipeline that ingests our order events and serves analytics'. The candidates who do well do not start designing. They ask: what is the freshness requirement and who is downstream; what volume, in events per day and bytes per day, and what is the peak; is this append-only or do records update and get deleted; what is the business key and can it change; where does it come from - an API, a database, a queue, a file drop; what is the acceptable cost; and what happens if the data is wrong, who notices and how fast.

Then the design is largely implied by the answers, and the remaining work is naming failure modes before you are asked. Late and out-of-order data, and the watermark or lookback window you use. Duplicates, and why you prefer at-least-once delivery with an idempotent upsert over a claim of exactly-once. Schema evolution, and whether an added column passes silently while a changed type fails the build. Backfill, and whether reprocessing ninety days runs through the same code path as the daily job or a separate one that will drift from it. Monitoring, and which three checks tell you it is broken before an analyst does: freshness, row-count variance against an expectation, and uniqueness or referential checks on the key. Cost, with a number attached.

Say the batch answer when batch is the answer. Most real analytics requirements are satisfied by hourly or daily batch, and the engineer who proposes a scheduled incremental load and explains that streaming would add operational burden for no benefit at this freshness target is demonstrating the judgment the round exists to test. Keep the streaming design in your pocket for when the requirement is seconds, and when it is, be able to explain state, watermarks and replay.

The resume, and the portfolio project that substitutes for experience you do not have

The screen is fast and mechanical: a named warehouse or lakehouse, a named orchestrator, SQL and Python, scale with units, recent work. Everything else is read only if those survive. The most common self-inflicted wound is a list of verbs with no magnitudes, which is indistinguishable from the bullets of someone who finished a bootcamp project.

Write bullets that open with the number, name the mechanism, and end with the consequence. Not 'built ETL pipelines using Airflow and dbt' but a bullet in this shape, with your own figures: 'owned 40 Airflow DAGs loading 12 sources into Snowflake, tens of millions of rows a day, 6 a.m. freshness target met in all but one month last year; cut monthly spend by about a third by clustering the two largest fact tables and killing three unused hourly refreshes'. If confidentiality blocks real figures, use the order of magnitude. An order of magnitude with a unit is credible; an adjective is not.

Keep the tool list to one block near the bottom, grouped by layer, written with the names an automated screen expects, and honest. Inflating it is a bad trade: the screen may pass you and then a panel asks a specific Kafka or Spark question and the conversation ends there. Two pages after six years, one before, and no skills matrix with star ratings.

If you do not yet have production experience, the project section is the whole resume, and most candidate projects fail the same way: a Kaggle CSV read into pandas in a notebook demonstrates nothing a data engineer is hired for. The project that works has six properties and can be built over a few weekends. A source that changes over time, so pull from a live public API on a schedule - transit, weather, government filings, GitHub events, an exchange - rather than downloading a static file once. A real schedule: GitHub Actions on a cron, or a small Airflow or Dagster instance, with run history visible. A warehouse or lakehouse target: DuckDB with Parquet or Iceberg on object storage, Postgres, or a free tier. dbt models with tests and a stated grain. An explicit idempotency story, so re-running yesterday neither duplicates nor loses rows. And monitoring, even if monitoring is one freshness check that opens an issue.

Then write the thing almost nobody writes: a README that reads like an engineering document. What the pipeline does, the volumes it actually handles, the freshness it targets, what it costs per month even if that is four dollars, three decisions you made and what you rejected, the failure you hit and how you found it, and what you would change at a hundred times the volume. One such project with many months of run history and a real commit log outperforms a stack of certifications and a tutorial portfolio, and it gives you something to go three levels deep on in the hiring-manager conversation - the round most people fail with generalities.

Pay, getting in, and where these jobs are actually posted

Any article quoting you one data engineer salary is averaging across at least four different jobs in several labor markets. Source it yourself, because you can, and because walking into the conversation with the posted ranges for your market and level changes the negotiation more than any technique does. Three sources, in order of reliability. Posted ranges: mandatory in a growing set of US states (California, Colorado, Illinois, Maryland, Minnesota, New Jersey, New York, Vermont, Washington and Massachusetts among them), and in the EU, Directive (EU) 2023/970 carries a 7 June 2026 transposition deadline. BLS OES tables: no data engineer code exists, so bracket with 15-1243 (database architects) and 15-1252 (software developers), which is the right instrument for geography rather than company. And levels.fyi, skewed toward large technology companies and inflated at the top end, but the only source attaching company names and levels to numbers.

What moves the number, in rough order of effect: the employer's industry and funding, far more than your skill; the level you are hired at, which is set by how much scope you can evidence rather than by years; whether the company puts data engineers on the software engineering ladder or a separate and usually shorter data ladder, worth asking in the first or second conversation; on-call and production ownership, which tends to be paid; and scarce specialisms - real-time systems, regulated data in finance and healthcare, and large-scale cost ownership. What moves it less than candidates expect: the number of tools you know, certifications outside the consultancy world, and the brand of your bootcamp or degree. Note also that the same work is sold three ways - permanent, contract or day rate, and consultancy - and in the UK the inside or outside IR35 determination changes take-home enough that comparing rates without it is meaningless. Ask which shape the role is before discussing money.

On entry level, the honest version: it is harder than it was at the start of the decade, for mixed reasons. Some is market and budget. Some is that the simplest layer of the work - writing a straightforward transformation or a connector by hand - is now partly generated, so the tasks that used to season a junior engineer are thinner on the ground. Anyone telling you AI has not touched the bottom rung is not watching; anyone telling you the role is disappearing is also wrong, because the operational, correctness and cost ownership at the center of the job has not been automated. The practical consequence is that the bar for a first role is production evidence rather than coursework, and the fastest route to production evidence is often a job adjacent to the one you want.

The routes that work, roughly in order of reliability. From data analyst or BI developer: you have SQL and business context, the gap is engineering practice, so volunteer for the dbt repository, the orchestration, the CI and the on-call, then apply as an analytics engineer and move across. From software engineering: you have the practice and the gap is modeling and warehouse mechanics, so take the integration and pipeline work nobody wants. From database administration or older ETL platforms (Informatica, SSIS, Talend): the modeling instincts transfer well, and the task is translating them into the current stack in public, because those resumes are screened out on vocabulary rather than ability. From support, operations or QA: internal transfer is the realistic path, and the lever is building the thing your own team needs and letting it be used.

Where the roles are. Engineering-led and data-specific job boards; the careers pages of companies that have recently bought Snowflake or Databricks, since they hire for months afterwards and partner directories and case studies are a legitimate way to find them; consultancies and system integrators, who hire continuously and will train platform skills; and the sector most people overlook - companies where data engineering is cost or compliance rather than product, meaning insurance, utilities, healthcare systems, logistics, local government and universities. Those hire steadily, interview less theatrically, state ranges more often, and are frequently the better first job because the data is genuinely messy and you will own something quickly.

Two tactics with an unusually high return. Contribute something small and real to a tool in the stack - a dbt package, an Airflow or Dagster issue, an Airbyte or Meltano connector, a docs fix in an Iceberg client. It is public evidence read by exactly the people hiring, and the bar for a useful first contribution is lower than it looks. Then apply where your project overlaps the employer's problem and say so in two sentences in the first message: the pipeline you built against a public transit API is a better opening with a logistics company than any cover letter. Volume applying performs badly here, because the screen is looking for a stack match and a generic resume matches nothing exactly. Fifteen applications with the bullets reordered to lead with the warehouse and orchestrator named in the posting beat two hundred identical ones. Keep one master document and generate the targeted version per posting; it is a ten-minute job and the highest-return habit in a data engineering job search.

Working with AI in this role

What a data engineer has to know about AI in 2026-27

Start with the honest version, because overclaiming here is the fastest way to lose a technical panel. The core of data engineering has changed less than the discourse suggests. Grain and modeling still decide whether numbers are right. Idempotency, backfills, late data and schema evolution are the same problems with the same answers. Batch did not die. On-call did not get easier. Nobody has automated the 7 a.m. question of why the finance dashboard is wrong.

What did change is real and sits in three places. First, the authoring ratio: SQL, dbt models, DAGs and connector code are now generated quickly and usually plausibly, so the scarce skill shifted from writing transformations to reviewing them and constraining what is allowed to land. That is why review rounds and 'what is wrong with this pull request' rounds have appeared in loops that did not have them. Second, the consumer set changed: the warehouse now feeds retrieval systems and agents as well as dashboards, which brought unstructured content, embedding pipelines, chunk-level permissions and evaluation datasets into scope at many employers. Third, metadata became load-bearing, because an agent writing SQL reads your column descriptions and your tests and will cheerfully join on the wrong key if the semantics live only in an analyst's head.

The failure modes of generated transformation code are specific, and knowing them by name is what a review round tests. Generated SQL is fluent about syntax and careless about grain: the classic output is a join that fans out and inflates an aggregate, a dedup that keeps an arbitrary row rather than the latest, an incremental filter that is not deterministic across runs, a window frame that quietly differs from the stated intent, and a cast that truncates. None of these fail loudly. They produce a number that is wrong by a plausible amount, which is the most expensive kind of bug in this job. The engineer who says that in an interview, and then names the test that catches each one, is demonstrating exactly what employers are now hiring for.

In interviews this shows up less as 'tell me about LLMs' and more in three concrete shapes. A design prompt that includes unstructured content: build retrieval over our contracts, tickets or support transcripts. A follow-up about permissions and provenance: can the person asking the question see the document the answer came from, and can you say what a model was trained on. And a question about the warehouse as an agent surface: if we point a text-to-SQL agent at this, what breaks. Candidates who have only read about it draw an embedding step with no access-control story and no evaluation. One concrete decision you made - including something you refused to index - is worth more than fluency in the vocabulary.

Treating an embedding pipeline as a pipeline, not a notebook

The work that lands on data engineers is not model training; it is an ingest, parse, chunk, embed and load pipeline over documents, with all the usual obligations and a few new ones. Most first attempts are a script somebody ran once, so nobody can say what is in the index, what failed to parse, or how stale any of it is. That is the same class of problem as an unmonitored nightly load, and you already know how to solve it.

Show it: Describe it with the ordinary properties: the source and how changes are detected, parse failures captured rather than skipped silently, chunk text and metadata stored durably in your own tables so you can re-embed from your data instead of re-crawling, the document and chunk identity that makes re-runs idempotent, and freshness and volume monitoring. Say what your parse failure rate was and what you did about the PDFs that turned out to be scans.

Embedding model versions and the re-embed backfill

Vectors from different embedding models are not comparable, so changing model means re-embedding the whole corpus. There is no clever design that avoids it. The mistake is treating that as a disaster rather than a budgeted, recurring backfill - which is a thing you already know how to plan.

Show it: Stamp the embedding model and version on every row and say why: it tells you exactly what needs re-embedding. Then describe the cutover as you would any large backfill - a shadow index, writes to both, a comparison on a fixed question set before switching reads, and the cost and duration of the re-embed with a number attached.

Carrying source permissions through to the chunk, and knowing when to refuse

The characteristic failure of an internal retrieval system is answering a question using a document the asker could not have opened. Permissions live in the source system; your index is a copy; nothing propagates by itself. This is the first thing a good interviewer probes and the first thing a real deployment gets wrong.

Show it: Describe access control reaching query time: ACLs or group identifiers stored per chunk and applied as a filter at search, re-sync when source permissions change, and what you did about documents whose permissions change faster than your refresh. Then name the corpus you deliberately excluded because its permission model could not be honored. The exclusion is the part that sounds lived rather than read.

Reviewing generated transformation code at volume

More transformation code is being proposed than any team can carefully read, and review capacity did not grow. Employers are explicitly testing for this, which is why pull-request review rounds have appeared. The value is not suspicion of generated code; it is a fast, repeatable checklist that catches the small number of ways it goes wrong.

Show it: State your review order out loud: what grain does each CTE produce, does any join fan out, is the dedup deterministic, is the incremental filter stable across re-runs, do the window frames match the stated intent, are the casts lossless, and what test would have caught this. In a review round, find the grain bug first. It is usually there, and finding it first is the signal.

Making the warehouse legible to an agent, and deciding what to expose

Text-to-SQL agents and warehouse MCP servers mean something non-human now reads your schema and writes queries against it. It will join on the wrong key if nothing tells it which key is right, and it will average a ratio if nothing defines the metric. Documentation, column descriptions, tests and a semantic or metric layer stopped being hygiene and became the interface. The set of tables and tools you expose is a governance surface somebody owns, and increasingly that somebody is the data engineer.

Show it: Say which curated layer you would expose and which raw tables you would not, how metrics are defined once rather than reinvented per query, where column descriptions and tests live so humans and tools read the same semantics, read-only access by default with any write path named and reviewed, and query cost limits on an agent that can otherwise generate a very expensive full scan on a loop.

LLM calls inside a pipeline, treated as an unreliable external API

Extraction, classification and enrichment steps that call a model are now ordinary pipeline stages: pull the fields out of an invoice, categorize a ticket, normalize a free-text address. They break the assumptions batch pipelines are built on, because they are non-deterministic, rate limited, slow and metered. Engineers who have run one in production talk about it very differently from engineers who have not.

Show it: Describe the controls: output validated against a schema and rejected rather than trusted, retries and a dead-letter path for records that will not parse, responses cached on an input key so a re-run does not re-pay, the model and prompt version stamped on every output row so results are reproducible, a batch endpoint where latency does not matter, and a cost ceiling with an alert. Then give the unit economics: cost per thousand documents, and the accuracy you measured on a sample you labeled by hand.

Evaluation sets and feedback as first-class datasets

Every AI feature that works has an evaluation set behind it, and in most organizations that set lives in a spreadsheet and a notebook, unversioned and unowned, so nobody can say whether last quarter's change helped. The data engineer is the person who knows how to make it a real asset, and that is a visible, high-leverage contribution still unusual enough to be remembered in an interview.

Show it: Describe golden question sets, evaluation runs, human ratings and thumbs-up signals as modeled tables: a schema, a version tied to the embedding and prompt versions in use, retention, ownership, and access rules for anything containing customer content. Say what you learned when the eval run first disagreed with the demo.

Training-data and dataset provenance you can actually prove

'Which data went into this model, and were we allowed to use it for that purpose' now has legal and contractual weight. The EU AI Act's obligations phase in across 2026 and 2027 on a timetable that has been actively debated, so check its current state rather than quoting a date at an interviewer. Separately, a deletion request has to reach the training snapshot and the vector index, not only the production table, and most pipelines were not built with that in mind.

Show it: Describe dataset versioning as a snapshot you can reconstruct: what was included, as of when, from which sources, under which restrictions, and the lineage that proves it. Then say how a deletion or purpose restriction propagates from the source system through derived tables, training snapshots and the retrieval index, and which of those paths you had to build by hand.

Spend on AI workloads, named and controlled

Cost control was always part of this job, and AI workloads added line items that surprise people: embedding a large corpus, per-document extraction, vector storage and index memory, and agents issuing unbounded queries against the warehouse. Being the engineer who can state the unit cost of an AI data feature is as valuable as it was to be the engineer who could state the cost of the nightly load.

Show it: Give unit economics for something you built: cost to embed the corpus once and per month in steady state, cost per thousand documents extracted, cost per thousand retrieval queries, and the lever that changed it - a smaller model for the easy cases, caching, a batch endpoint, different chunking, or not indexing a corpus nobody searched.

Knowing, and saying, where AI changed nothing

A panel of working data engineers is suspicious of candidates who narrate everything as transformed. Saying plainly which parts of the job are unchanged is a credibility signal, and it is also true, which is why it works.

Show it: Say it directly when asked: the grain of a fact table in a business you have not interrogated, which source is authoritative when two disagree, whether a dimension needs type 2 history, what idempotency means for this particular key, how to reconcile a migration, and who gets paged at 3 a.m. are unchanged. What changed is that the code arrives faster, so the constraints - tests, contracts, CI, lineage, access - carry more weight than they used to.

What a screen is looking for

These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.

Mistakes that cost people this job

A resume of verbs with no magnitudes: 'built ETL pipelines using Airflow, dbt and Snowflake' repeated across three jobs.

Open every bullet with a number and a unit, then the mechanism, then the consequence. Rows or events per day, source count, pipeline count, freshness target and whether you met it, monthly spend and the lever that cut it. If the real figures are confidential, give the order of magnitude. An adjective is not evidence.

A take-home that cannot be run twice. The second run duplicates rows, or the load appends blindly, or the only recovery is to drop the table.

Make idempotency the first design decision and say so in the README: the business key, the upsert or merge, and what happens to late-arriving and deleted records. Graders run submissions twice specifically to see what happens. This single property separates the shortlist from the pile.

Starting the design round by drawing boxes. Thirty seconds in there is a Kafka topic on the whiteboard and nobody has said how fresh the data needs to be.

Spend the first three minutes asking: freshness requirement and who is downstream, volume per day and at peak, append-only or updates and deletes, the business key and whether it is stable, the source type, the cost ceiling, and who notices when it is wrong. The questions are a large part of the grade, and the design largely follows from the answers.

Reaching for streaming, Spark and Kubernetes on a problem whose data fits on a laptop, because it looks more senior.

Size the solution to the requirement and say the sizing out loud: 'at this volume and a four-hour freshness target this is a scheduled incremental load; streaming would add operational burden for no benefit, and here is the threshold at which I would change my mind'. Panels read right-sizing as seniority and over-engineering as inexperience.

Claiming exactly-once delivery, or claiming streaming experience you do not have.

Say at-least-once delivery with an idempotent upsert, which is what almost every production system actually does, and name the key that makes it idempotent. If your experience is batch, say so and say what you would need to learn. Overclaiming survives the recruiter screen and dies in the panel, where the follow-up is about consumer rebalancing or watermarks.

Skipping SQL practice because the job description says Spark, or because an assistant can write the query.

Drill the discriminating patterns until they are automatic: deduplication with ROW_NUMBER, sessionization, gaps and islands, point-in-time joins against a type 2 dimension, and detecting fan-out before aggregating. The live SQL screen is the narrowest gate in most loops, and 'the assistant writes it' is not an answer when the interviewer asks why the aggregate is inflated.

Not knowing what your own pipelines cost. Asked about warehouse spend, the candidate has never seen the bill.

Before interviewing, find the real numbers for something you own: monthly warehouse or cluster cost, the three most expensive queries or jobs, and one change you made or would make. Cost ownership is one of the clearest seniority signals in this role and one of the easiest to acquire.

A portfolio of notebooks: a static Kaggle CSV read into pandas, charted, and pushed to GitHub.

Build one pipeline with a source that changes over time, a real schedule with visible run history, dbt models with tests, an explicit idempotency story, one monitoring check, and a README stating volumes, cost, decisions and the failure you hit. One project with many months of run history beats five notebooks and three certifications.

Collecting certifications as the route in, then applying to product companies with no project to show.

Know what a certification buys: it unblocks consultancies and cloud partners that need certified headcount, and it is read last at a product company. If you want one, take the one matching the platform in the postings you are targeting, and spend the remaining time on the project that gives you something to be interviewed about.

Applying to analytics engineer and data engineer postings with the same document, and getting silence from both.

Read the verbs. Model, transform, test, document, define metrics means lead with dbt, modeling and the metric layer. Ingest, orchestrate, operate, on-call, backfill means lead with the orchestrator, CDC, reliability and cost. Keep one master document and reorder it per posting. Ten minutes, and it changes the response rate more than volume does.

Treating AI as a line item on a 2026 resume: 'familiar with LLMs and vector databases'.

Show one concrete decision instead. The embedding pipeline you ran and its parse failure rate. How document permissions reached query time. The cost per thousand documents extracted and the lever that changed it. The corpus you refused to index because its permission model could not be honored. One decision beats a vocabulary list.

Describing pipelines you built and never the ones that broke.

Bring one failure with the full arc: the symptom, how it was detected and how long that took, the diagnosis order you followed, the root cause, the fix, and the check you added so it could not recur silently. Interviewers hiring for an on-call rota are listening for exactly this, and most candidates never offer it.

Questions people ask

What does a data engineer actually do?

A data engineer gets data from where it is produced to where it is used, reliably and at a known cost. In practice that means ingestion and change data capture out of source systems, transformation and modeling in a warehouse or lakehouse, orchestration and scheduling, storage layout and table format decisions, data quality checks, and operating all of it, including being on call when a load fails. The work is judged on whether data arrives: freshness against an agreed target, correctness, reliability, and the monthly bill. A useful test of whether a task belongs to a data engineer: if it is about whether the numbers arrive, it is yours; if it is about what the numbers mean, it is usually an analytics engineer's or an analyst's.

What is the difference between a data engineer and an analytics engineer?

An analytics engineer owns the transformation layer and the metric definitions, working mostly in SQL and dbt on data that has already landed, and typically does not own ingestion, infrastructure or an on-call rota. A data engineer owns getting the data there and keeping it there: connectors and change data capture, orchestration, storage layout, reliability and cost. The skills overlap heavily in SQL and modeling, which is why moving from analytics engineering into data engineering is one of the most reliable routes in. Pay is often similar at the same level, so the real difference is what your day looks like, and whether your phone can go off at 3 a.m.

What stack should I learn to get a data engineering job in 2026-27?

Narrower than the postings suggest, and in depth rather than breadth. SQL at a level where deduplication, window functions, grain and point-in-time joins are automatic. Python for messy input: APIs, retries, validation, generators, pytest, plus one of Polars or DuckDB. One warehouse or lakehouse properly - Snowflake, Databricks or BigQuery - including its partitioning and pricing model. dbt for transformation, with tests and incremental models, or SQLMesh. One orchestrator, Airflow or Dagster, including retries, backfills and freshness checks. Enough Apache Iceberg and catalog literacy to say where a table lives and which engine may write to it. Then git, Docker and a CI pipeline that runs your tests. Add Kafka and Flink only if you are targeting jobs that genuinely stream, because those interviews go deep quickly.

Do I need a degree or a certification to be a data engineer?

No. There is no licence, no registry and no required degree, and people enter from analytics, software engineering, database administration, support and unrelated careers. Certifications - Snowflake SnowPro, Databricks Certified Data Engineer Professional, Google Professional Data Engineer, AWS Certified Data Engineer - Associate, Microsoft Fabric Data Engineer (DP-700) - have one reliable use: consultancies and cloud partners need certified headcount, so a certification can unblock that screen. At a product company it is read after the project section and rarely decides anything. If your time is limited, one deployed project with a scheduled pipeline, tests and a README with real numbers moves more interviews than two certifications.

What does a data engineering interview test?

Usually five things, in this order. A live SQL and Python screen, the narrowest gate in most loops, concentrating on deduplication, window functions, grain and fan-out, incremental and merge logic, and handling messy input in Python. A take-home or live build, graded mainly on whether it runs from a clean clone, whether it can be run twice safely, whether there are tests that could actually fail, and whether the README states your decisions. A pipeline design round, where the questions you ask before designing - freshness, volume, update semantics, keys, cost, who notices when it breaks - carry much of the grade. Increasingly, a code review round on transformation code with a grain bug in it, and an incident round where a dashboard is wrong and you must say what you check, in order. Finally a conversation with someone who consumes your data, testing whether you ask what it is for.

How much does a data engineer make?

There is no single band, and the US Bureau of Labor Statistics does not publish data engineer as an occupation, so any article quoting one figure is averaging across several different jobs. Triangulate from three sources. Posted ranges, now mandatory in a growing list of US states (California, Colorado, Illinois, Maryland, Minnesota, New Jersey, New York, Vermont, Washington and Massachusetts among them), and in the EU, Directive (EU) 2023/970 on pay transparency, which member states must transpose by 7 June 2026. BLS OES codes 15-1243 (database architects) and 15-1252 (software developers), which bracket the role and are the right instrument for comparing geographies. And levels.fyi for company-attributed bands at the technology end. What moves your number most is the employer's industry and funding, the level you are hired at, and whether the company puts data engineers on the software engineering ladder or a separate shorter one. Ask which ladder in the first or second conversation.

Is data engineering being automated by AI?

Partly, and specifically. Writing a straightforward transformation, a connector or a DAG is now substantially generated, which has thinned out the simplest tasks and made the entry level harder than it was at the start of the decade. What has not been automated is the center of the job: deciding the grain, knowing which source is authoritative when two disagree, making a pipeline idempotent and backfillable, reconciling a migration, controlling cost, and diagnosing a late or wrong pipeline under time pressure. The net effect on hiring is a shift in what gets tested. Employers now probe review and judgment, including handing you generated transformation code and asking what is wrong with it, because producing code stopped being the scarce part and judging it did not.

What AI work actually lands on a data engineer now?

Mostly pipelines, which is the reassuring part. Ingesting and parsing unstructured content such as contracts, tickets, transcripts and PDFs. Chunking and embedding as scheduled, monitored pipeline stages rather than one-off scripts. Stamping the embedding model version on every row so you know what needs re-embedding when the model changes. Carrying source permissions down to the chunk so a retrieval answer cannot cite a document the asker could not open. Modeling evaluation sets and human feedback as real versioned tables. Dataset provenance good enough to say what a model was trained on, and deletion that reaches the training snapshot and the vector index rather than only the production table. And keeping the cost of all of it legible. Separately, making the warehouse safe for a text-to-SQL agent - a curated layer, defined metrics, column descriptions, read-only access, query cost limits - is becoming an ordinary data engineering task.

What is the best portfolio project for a data engineering job?

One pipeline with six properties, built over a few weekends. A source that changes over time, so a live public API - transit, weather, government filings, exchange data, GitHub events - rather than a static CSV, because that makes incremental loading a real problem. A real schedule, so GitHub Actions on a cron or a small Airflow or Dagster instance, with visible run history. A warehouse target, so DuckDB with Parquet or Iceberg on object storage, Postgres, or a free tier. dbt models with tests and a stated grain. An explicit idempotency story, so re-running yesterday neither duplicates nor loses rows. And one monitoring check that fires when freshness slips. Then a README that reads like an engineering document: volumes, freshness target, monthly cost even if it is four dollars, three decisions and what you rejected, the failure you hit and how you found it, and what you would change at a hundred times the volume.

How long does it take to become a data engineer?

It depends entirely on what you already have, because the gate is production evidence rather than a credential. From data analyst or BI developer, months rather than years: you have SQL and business context, and the gap is engineering practice, so inside your current job move the reporting SQL into a dbt project with tests, put it in git, add CI, take over the scheduling, and ask to be added to the rota that gets paged when the morning load fails - that sequence produces exactly the evidence a data engineering screen looks for, and analytics engineer is a legitimate intermediate title if a direct move is slow. From software engineering, also months: the gap is modeling and warehouse mechanics, so volunteer for the integration and pipeline work. From no technical background, plan on about a year of deliberate work, with one scheduled pipeline running the whole time so that by the time you apply it has real run history. In every case, say the numbers when you apply: how many models, how many sources, what volume, what freshness target, what the rota was.

Put this on a resume in about a minute

Paste your history once and point it at the Data Engineer posting you are looking at. No account, no card.

Build my resume free More roles