Data & Analytics

How to get hired as an analytics engineer in 2026-27

The short answer

To get hired as an analytics engineer in 2026-27, publish a small transformation project a reviewer can clone and run in under two minutes, with every modelling decision written down: the grain of each mart, which source you treated as authoritative and why, the incremental strategy that survives late-arriving data, and the tests that would actually fail. The role sits between data engineer and data analyst: the data engineer lands the data, the analyst answers the question, and the analytics engineer owns the tested, documented, version-controlled tables and metric definitions in between. There is no licence and no required certification for this job, so the shortlist is made almost entirely on readable evidence, SQL, dimensional modelling judgment, and a repository someone can open. The usual loop is a recruiter screen, a 45-to-60-minute SQL exercise, a modelling case, often a code review of somebody else's dbt project, and a stakeholder conversation; expect at least one question about making the warehouse answerable by an AI assistant, because the semantic layer that makes that work is now part of the role.

What the role ownsThe transformation layer and the definitions on top of it: staging, intermediate and mart models; the grain of each table; the metric or semantic layer; tests, freshness thresholds and contracts; documentation and deprecation. The output is dependable tables and agreed definitions that other people build on, not pipelines, and not dashboards.
Closest confusionsA data engineer lands and moves data and is measured on whether it arrives. A data analyst uses the modelled tables to answer a business question and argue for a decision. A data architect decides the shape all of it is built against. An analytics engineer turns landed raw data into trustworthy modelled data, in version control, with tests.
Credential gateNone. No licence, no registration, no required degree. The dbt Analytics Engineering certification, SnowPro, Databricks and Google data certifications exist and occasionally clear a recruiter filter, but no hiring manager in this role will trade a certificate for a readable repository and a clean SQL screen.
Typical loopScreen; a 45-to-60-minute SQL exercise; a modelling or data case; frequently a code review of an existing dbt project; a stakeholder communication round; the hiring manager. Two to five weeks end to end. Short take-homes (commonly two to five hours, often a dbt project against seed files or DuckDB) are still normal in this role, unlike in much of engineering hiring.
Who screens youA recruiter matching on SQL plus the warehouse plus dbt, then an analytics engineering manager, head of data or analytics manager. At larger companies a data platform engineer runs the technical rounds and an analyst you would serve sits in on one, with a real veto.
Pay, and the lever that moves itThere is no dedicated US BLS occupation code for analytics engineer, which is itself worth knowing, the nearest OES codes, 15-2051 (data scientists) and 15-1243 (database architects), both mismap. Use ranges published under state pay-transparency laws and levels.fyi for companies with a public ladder. The largest single lever is not skill, it is which ladder the role sits on: the same work paid on a platform or software engineering band earns materially more than on an analyst band, and that is decided by the team you join. Ask which ladder and level the posting maps to on the first call.
Evidence that landsModel counts and the direction they moved; test coverage tied to an incident it caught; build time and warehouse spend with the lever named; conflicting metric definitions consolidated into one certified definition with a named sign-off; and migration detail, how many objects, how long you dual-ran, what reconciled exactly and what carried a stated tolerance.
What changed by 2026Writing SQL stopped being the bottleneck and reviewing it became one. The semantic layer turned into the interface AI assistants query, which moved documentation, naming and certified metrics from housekeeping into the part of the job an employer will fund.

Analytics engineer vs data analyst vs data engineer vs data architect

The distinction is about what your output is, not about seniority or how technical you are. A data engineer's output is data arriving: extraction, loading, streaming, orchestration, platform reliability. A data analyst's output is an answer and an argument, this is what the number is, this is what it means, here is what I think we should do. An analytics engineer's output is a dependable thing other people build on: a set of modelled tables with a stated grain, tests that fail when something is wrong, documentation someone else can read, and a metric definition that finance and marketing have both agreed to.

The title was popularised around 2019 and 2020, largely by dbt Labs (then Fishtown Analytics), to name work that already existed without a name. What created the role was the economics of ELT. Once managed connectors landed raw tables cheaply and warehouse compute got elastic, extraction stopped being the hard part. The hard part became: there are four tables that all look like customers, two of them disagree about churn, one has duplicates from a replication retry, and three dashboards compute revenue differently. That is not a pipeline problem and it is not an analysis problem. It is a modelling problem, and it needs software practices, version control, code review, tests, CI, environments, releases.

Titles are an unreliable guide in both directions, and this costs candidates real interviews. Jobs that are genuinely analytics engineering are regularly posted as BI developer, analytics developer, data analyst (with dbt buried in the requirements), decision support engineer, reporting engineer, or data engineer (analytics). Meanwhile some postings titled analytics engineer are a data engineering job: Python, orchestration, streaming, infrastructure. Read the requirements, not the title. The triad that means analytics engineering is SQL plus a cloud warehouse plus a transformation framework. The triad that means data engineering is Python plus orchestration plus infrastructure, with SQL as a supporting skill.

The four kinds of analytics engineer job, and how to tell them apart from the posting

These four jobs share a title and almost nothing else about the day. Applying to all of them with one resume is a common reason a qualified candidate gets no response, because each one screens for a different lead bullet. Work out which one a posting is before you tailor anything.

Read the posting for three signals. How big is the existing stack, a posting that names the warehouse, the orchestrator, the BI tool and the dbt deployment has a stack; a posting that says you will help us choose does not. Who do you report to, head of data and a central team means platform work; a marketing or finance director means embedded. And how much of the job is analysis, if dashboards and stakeholder requests sit in the responsibilities, you are also the analyst, which is fine but changes what to lead with.

How hiring for this role actually works in 2026-27

There is no credential gate, which cuts both ways. Nobody can block you for lacking a licence, a degree or a certificate, and plenty of good analytics engineers came from finance, operations, support or teaching. But because there is no gate, the filter has to come from somewhere, and it comes from evidence a reviewer can read in ten minutes. That is why a small public repository matters more in this role than in almost any other data job, and why the candidate pool divides so cleanly between people who have one and people who have a list of course certificates.

The first screen is usually keyword-shaped and shallow. A recruiter checks for SQL, a named warehouse (Snowflake, BigQuery, Databricks, Redshift), a transformation framework (dbt is still the default; SQLMesh, Dataform and Coalesce appear), a BI tool, and sometimes Python and an orchestrator. Spell those out as nouns, including the flavour, dbt Core versus dbt Cloud, and the warehouse by name. Writing "cloud data warehouse" where the posting says Snowflake fails a match a human would have made.

After that the loop is unusually hands-on compared with other data roles, and short take-homes have not gone away here. A two-to-five-hour dbt exercise against seed files or a local DuckDB database is normal, and is often the most informative round for both sides. Treat scope as negotiable: ask what the time box is, ask what they will grade, and if the answer is open-ended and unpaid at over about five hours, say so and offer a live pairing session instead. Many teams accept, and the ones that refuse have told you something.

One stage has grown fast and catches people out: being handed an existing dbt project or a pull request and asked what is wrong with it. It is an efficient filter, because reading someone else's models is the actual job, and because a model cannot do it for you while you are talking. Prepare for it specifically, with real code. dbt Labs' jaffle shop repositories are the canonical small example; GitLab publishes its own production analytics project; dbt package source code on the package hub is another source of real models. Clone one, read it, and practise narrating what you would change and in what order.

Timelines run two to five weeks at a product company, faster at consultancies and startups, slower in government, healthcare systems and universities where a requisition passes through a committee. Referrals carry more weight in this niche than in most, because the community is small and legible: the dbt Community Slack, Locally Optimistic, and local data meetups and conferences are places where hiring managers genuinely find people.

What analytics engineers are paid, and what actually moves it

Anyone who gives you a single national band for this role is guessing. The honest starting point is that analytics engineer has no dedicated occupation code in the US Bureau of Labor Statistics system, so official wage data does not cleanly describe it. The nearest OES codes are 15-2051 (data scientists) and 15-1243 (database architects), and both mismap, the first skews toward statistical and modelling work, the second toward enterprise architecture. Look them up for the national and metro percentile spread, and read them as weak bounds rather than as the answer.

The better sources are tied to actual offers. Pay-transparency laws now cover a growing list of US states (Colorado, California, Washington, New York, Illinois and Massachusetts among them) so a large share of postings carry a range, and remote postings open to those states usually do too. Collect twenty ranges for your level and location from companies you would actually join and you have a far better picture than any salary article. For companies with a public engineering ladder, levels.fyi shows the band and the equity split. Otherwise ask the recruiter on the first call; in some states, California included, they must give you the range for the role on request. Outside the US there is less posted data: UK, EU and Canadian postings often omit ranges, so lean on recruiter conversations, community salary surveys and the ladder question below.

The single largest lever is not your skill level, it is which ladder the job sits on. The same work, building the same models, pays on an analyst band in one company and on a software or platform engineering band in another. The gap between those two bands at the same company is often large enough to dwarf a two-level promotion inside the analyst band. The determining factor is organisational: whether the data team sits inside engineering, and whether analytics engineer exists as a rung in the engineering ladder or gets slotted into Analyst III with Analyst III's ceiling.

So ask, in the first conversation, in roughly these words: which ladder does this role sit on, what level does it map to, and is that the engineering band or the analyst band. It is a normal question, and the answer predicts both your offer and your next three years of raises. If the title does not exist in their ladder at all, that is worth probing, it often means promotion requires leaving the role.

After the ladder, the things that visibly move pay: whether you also own orchestration, infrastructure and cost, which pushes you toward the data platform band; industry, with fintech, health insurance, adtech and quantitative firms at the top and nonprofits, higher education, local government and agencies at the bottom; whether the company has a usable amount of data, since scale buys both budget and interesting problems; and contract versus permanent, where hourly rates look higher and often are not once you account for benefits and unpaid gaps.

What moves pay less than candidates expect: certifications, the number of tools you can list, the number of dashboards you have built. What moves it more than candidates expect: being the person who can be trusted with a definition the CFO quotes externally, and being able to show you cut warehouse spend. Both are evidenceable and both are rare.

The portfolio project: the highest-leverage thing you can build

In a role with no licence, the shortlist is made on what a reviewer can verify quickly. A small, runnable, well-documented transformation project does that better than anything else, and the striking thing is how few applicants send one. Most send a course certificate or a notebook of charts. You are not competing with a high bar here; you are competing with an absence.

Build it to be read, not to be impressive. Fifteen to twenty-five well-reasoned models beat a hundred and twenty. The reviewer is a working analytics engineer with ten minutes and a browser tab open; they will clone it, try to run it, open the README, open two model files, and look at your tests. Optimise for that sequence exactly. Make the repository public, give it a descriptive name rather than dbt-project-final, pin it on your profile, and put the link in your resume header.

Make it run in under two minutes with no credentials: dbt with DuckDB, data committed as seeds or a small Parquet file, a single command. A project that requires the reviewer to create a Snowflake trial account does not get run, and a project that does not get run does not count. Put the exact command in the first three lines of the README, then clone your own repository into a fresh directory and verify that command works from nothing.

Pick data with real mess in it: duplicates from retried loads, records that arrive late, corrections to past rows, a dimension whose values change, inconsistent keys across two sources. Concrete options that have all of this and are free: NYC taxi trip records, Citi Bike or other bikeshare trip exports, GTFS transit feeds, city building-permit data, SEC EDGAR filings, and open baseball play-by-play data. A single clean CSV is useless for this purpose, because none of the interesting decisions can occur. If your chosen dataset turns out to be clean, introduce the mess deliberately and document that you did.

The README is the most-read file in the repository, and most candidates waste it on a tool list. Make it ten to fifteen lines of decisions, each one naming the alternative you rejected and why: why this grain for this mart, which source you treated as authoritative when two disagreed, why this incremental strategy and what lookback window, why you snapshotted this dimension rather than joining it as-is, what you deliberately left out. A reviewer reading that is reading the thing they interview for, judgment, before they have run a line of your code.

The resume: what lands, and what gets skipped

Two pages once you have three or more years of relevant work; one page before that. Applying the one-page rule too early in this role strips the industry, the data volume and the stack context that make a bullet legible, and leaves a list of verbs that could describe anyone. Put the repository link in the header next to your email, not in a footer nobody reaches.

Every bullet should answer: what did you change, what did it cost or save, and how do you know. Numbers with units, and the lever named. "Reduced spend by forty percent" is unverifiable and reads as decoration. "Replaced a nightly full refresh of a two-billion-row fact table with an incremental merge, cutting the daily warehouse bill by about a third" is a sentence another analytics engineer believes immediately, because they recognise the lever. Use your own real figures; if you do not have a figure, name the lever and the direction rather than inventing a percentage.

Name the stack explicitly and in the employer's words, because the first screen is a keyword match done by someone who cannot infer. Snowflake, BigQuery, Databricks, Redshift, Postgres, dbt Core, dbt Cloud, SQLMesh, Airflow, Dagster, Fivetran, Airbyte, Looker, LookML, Power BI, Tableau, Hex, Mode, Sigma. One compact block near the bottom for the screen, and the same nouns woven into the bullets for the human. Match the posting's spelling too: most US postings say modeling, materialized and optimization, and a keyword filter is literal about it.

The second half of the job is cutting. The items in the second list below are not disqualifying, they are inert, and they crowd out evidence that is not. Two deserve a longer note. Certifications placed above experience actively hurt here, because the reviewer reads them as a signal that there was no experience to lead with, keep them to one line at the bottom, where they can still clear a filter. And course or bootcamp projects described in the grammar of production work ("architected an end-to-end data platform") invite a line of questioning you cannot survive. Describe them accurately as projects, with the decisions in them, and they read as honest and useful instead of inflated.

The SQL, modelling and code review interviews, question by question

The SQL screen in this role is not a LeetCode screen. It is specific to warehouse analytics work, and it mostly tests whether you think about grain and about correctness under messy data, not whether you can invert a tree. The recurring shapes are predictable enough to drill until they are mechanical, because the interviewer's real interest is the commentary you provide while writing, not the final query.

Say the grain out loud, early and often. Before writing, state what one row of your result represents; after a join, state whether the grain changed. It is the habit the job is made of, and it is the clearest separator between candidates who advance to the modelling round and candidates who do not. If you take one thing from this guide into a live interview, take this.

The modelling case gives you a business, a handful of source tables and a vague request. The grading is not whether you produce the architecture they had in mind. It is whether you ask what decision the output serves, declare the grain of each mart, notice that two sources disagree and say which you would treat as authoritative and how you would prove it, handle a fact that arrives late and a dimension that changes, choose tests that would fail for a real reason, and decline to build something. "I would not build a mart for that, let the analyst query the fact table, and we can promote it if three teams ask" is a strong answer, not a dodge.

The code review round is where preparation pays most, because the failures are a finite list. You are graded on taste and on the order you would address things, not on completeness. Lead with what makes numbers wrong, then what makes the project unmaintainable, then style. Saying "the first two are correctness, the rest I would raise as comments and not block the merge" is exactly the judgment they are looking for.

The communication round usually hands you a discrepancy: the finance report says one number, the growth dashboard says another, explain it to a VP. The answer has three beats and most candidates perform only the middle one. Reconcile first, in units, the gap is 412 orders, all of them cancellations counted in one and excluded in the other. Then explain which definition is correct for which purpose, without making anyone wrong. Then say what you will change so it cannot recur: one certified definition, the other model deprecated with a date, the dashboard repointed. The third beat is what gets you the offer.

Finally, expect at least one question about something you deleted, deprecated or killed. In a role whose failure mode is accumulation, willingness to remove things is a real signal, and a candidate whose record is all additions looks like someone who will double the model count in a year.

Getting in without the title, and where these jobs are actually posted

The highest-probability route into an analytics engineer job is inside the company you are already at, especially if you are an analyst, a BI developer, a report writer, or the person in finance or operations who became the SQL person. The move is not subtle: volunteer for the transformation migration nobody wants, put the shared logic into version control, own the metric definitions, start reviewing other people's pull requests, and get the title changed before you look externally. Having held the title for a year, even at a small company, changes your external market far more than another certificate will.

From data engineering, the move is technically easy and gets undersold in interviews. The gap to close is not SQL, it is stakeholder work and semantics, spend your examples on the time you arbitrated a definition between two teams, not on the Airflow DAG. From analysis, the gap is engineering practice: version control, code review, testing, CI, environments, and the discipline of building something other people depend on rather than something you own. Either way, name the gap yourself and say what you did about it. Interviewers trust a candidate who has correctly identified their own weakness far more than one who claims symmetry.

Searching job boards for the exact phrase analytics engineer misses a large part of the market. The best single query is the tool: search dbt as a keyword with no title filter, and the postings that surface are almost all this job regardless of what they are called. Then add the title variants (analytics developer, BI engineer, BI developer, reporting engineer, decision support engineer, data analyst with dbt in the requirements) as saved searches.

Beyond the big boards, three places reliably carry postings that never reach them: community job channels, including the dbt Community Slack and equivalent spaces; the careers pages of analytics consultancies and dbt implementation partners, who hire continuously and are the most forgiving entry point for a switcher; and the engineering blogs of data-heavy companies, which tell you both that they are investing and exactly what to mention in a cold note.

Cold outreach works unusually well in this niche, because the community is small and most inbound applications are undifferentiated. Six lines: one specific thing you read about their stack from their own blog or conference talk, one sentence on what you would want to own, one link to a repository that runs in two minutes, no attachment. Send it to the analytics engineering manager, not to a recruiter. The conversion rate is not high, but it is far better than an application through a portal, and it costs twenty minutes.

Working with AI in this role

What an analytics engineer has to know about AI in 2026-27

The honest summary: AI has changed this job less at its core and more at its edges than the hype suggests, and the direction of the change is mostly favourable. Generating SQL is now easy, and generating SQL was never the hard part. Deciding the grain of a fact table, arbitrating which of two sources is authoritative, agreeing with the CFO what a renewal is, and knowing that a correction will arrive eleven days late are all still the hard part, and none of them have been automated. Do not walk into a 2026 interview claiming the role is being eliminated, and do not claim nothing happened either.

What actually changed, in order of how much it affects your week. First, authoring got fast and reviewing did not, so the ratio flipped: a team now produces more candidate transformations than it can carefully read, and the constraint on a project became review capacity, tests and contracts rather than typing speed. Second, the big one for your career, the semantic layer became the interface AI assistants query. Conversational analytics products (Snowflake Cortex Analyst, Databricks Genie, conversational analytics over Looker, ThoughtSpot, Omni, Hex and others) are only as accurate as the semantic model, naming, descriptions and certified metrics underneath them, and that layer is owned by analytics engineers. Documentation stopped being housekeeping and became the thing that determines whether the company's AI analyst gives correct answers. Third, agents and assistants now issue queries themselves through warehouse and BI connectors, including MCP servers, which makes row-level security, access control and query cost guardrails part of your job rather than someone else's.

What is being automated around the role, stated plainly: boilerplate staging models generated from source schemas, first drafts of documentation and YAML, suggested tests, and a slice of the ad-hoc "can you pull me X" traffic that self-serve assistants now absorb. What is not being automated: anything where the right answer depends on a business agreement or on knowledge of how the data actually behaves. The practical consequence for a candidate is that the most exposed work is the most junior work, writing the query someone else specified, so the evidence you need to show early is modelling judgment, not throughput. On tooling churn, note it and move on: dbt remains what postings name, its newer Fusion engine is appearing in job descriptions, and dbt Labs combining with Fivetran was announced in late 2025. None of that changes what you need to learn, and saying so calmly is a better interview answer than reciting vendor news.

The failure mode employers have now seen for themselves is specific, and being able to describe it is worth more than any tool list: a text-to-SQL or conversational product demoed beautifully, went to the business, and returned confidently wrong numbers. Not because the model wrote bad SQL, but because there were four tables with customer in the name, the status column had undocumented values, two metrics with the same name meant different things, and nothing in the warehouse told the model which one to use. The fix is not a better model. The fix is a curated, documented, certified layer with a small number of blessed entities and metrics, which is exactly the work this role has always done, and now has a visible business case for.

So the interview question is rarely "do you know about LLMs". It is some version of: how would you make our warehouse answerable by an assistant, or we tried a conversational analytics tool and it gave wrong answers, what would you do. The weak answer names a vendor. The strong answer narrows the surface (expose a handful of certified entities and metrics, not 900 tables), fixes the metadata (descriptions, synonyms, value enumerations, deprecation markers, ownership), enforces permissions at query time so the assistant cannot answer from something the asker could not open, and then does the part nobody thinks of: builds an evaluation set of real business questions with known correct answers and measures accuracy before and after. Being the person who proposes measuring it is the whole differentiator.

On your own workflow, be straightforward. Saying you do not use AI reads as incurious; saying you accept generated models without review reads as dangerous. The useful answer is specific about where you trust it and where you do not: fine for boilerplate staging models, YAML, docs drafts, regex, window function syntax and translating a stored procedure; never trusted for the grain decision, the incremental strategy, or whether a join fans out, and every generated model gets the same tests and the same review as a hand-written one. Have one real example of generated SQL you caught being wrong, and say what the error was. Interviewers remember that answer. A growing number of take-homes now explicitly permit AI tools and ask you to say where you used it; if the instructions are silent, ask, and disclose it either way.

Owning the semantic layer as the interface AI tools query

Conversational analytics only works if a curated layer tells the model which entities and metrics exist and what they mean. This is the clearest case in data where AI raised the value of an existing job, and it is the part of the role most likely to appear in your next job description.

Show it: Describe a layer you curated: how many entities and metrics you exposed versus how many tables exist, what you deliberately hid, how a metric got certified and who signed off, and how you handled two metrics that shared a name. With no work example, build it into your portfolio project, a small metric or semantic definition set over your marts, with descriptions and synonyms, plus a note on what you chose not to expose.

Evaluating a natural-language-to-SQL surface with a real question set

Teams deploy these tools and often measure nothing, which is why trust collapses after the first wrong answer. A candidate who talks about measurement rather than vendors is immediately distinguishable, and the skill transfers to any AI feature the data team is asked to support.

Show it: Describe or build an evaluation set: twenty to fifty real business questions with known correct answers, run against the semantic layer, scored for correctness, with failures categorised (wrong table, wrong filter, wrong grain, ambiguous metric). Report the before and after of a metadata fix. Even thirty questions over a toy project makes the point.

Reviewing generated SQL for the specific errors it reliably makes

The project's constraint is now review, not authoring. The valuable reviewer knows where generated transformations break, rather than reading every line with equal suspicion.

Show it: Name the failure classes you check first: a join that silently fans out, an incremental filter with no lookback, a filter that drops NULLs unintentionally, a timezone or day-boundary assumption, a metric that double counts after a union, plausible column names that do not exist. Then describe the guardrails that mean review is not the only defence: contracts on marts, a unique test on every declared grain, CI that compares row counts against production.

Documentation and metadata treated as a product with a consumer

Descriptions, column-value enumerations, ownership and deprecation markers used to be a nicety for humans who could ask a colleague. An assistant cannot ask a colleague, so gaps now become wrong answers given to executives.

Show it: Show a documentation standard you enforced and how you enforced it: a CI check that fails when a mart column lacks a description, a deprecation convention with a removal date, an ownership field that routes questions. Quantify it if you honestly can, the share of mart columns documented when you arrived and when you left.

Access control and cost guardrails for machine query traffic

Assistants and agents query in loops, retry, and scan more than a human would, and they query on behalf of a person whose permissions live in the source system rather than in the index or the semantic layer. Both the cost profile and the access risk are new, and both land on whoever owns the warehouse layer.

Show it: Describe how a person's permissions reached query time rather than being applied after the fact, what you did about the table you refused to expose because its permission model could not be honoured, and the cost controls you put in place: warehouse sizing and timeouts for the assistant's role, result caching, query tagging so spend can be attributed to the AI surface, and an alert when it crossed a threshold.

Saying plainly what AI has not changed about the role

Interviewers at this level are tired of candidates who overclaim. Naming precisely which parts of the job are untouched demonstrates that you understand the job, and it is simply true, which matters.

Show it: Have a crisp list ready: grain decisions, source-of-truth arbitration, agreeing a definition with a stakeholder who has an incentive, late and corrected data, migration and cut-over sequencing, deprecating something people still use, and deciding what not to build. Then name the one thing that did genuinely change in your own work, with an example. A clear no plus a specific yes is the credible answer.

What a screen is looking for

These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.

Mistakes that cost people this job

Applying with an analyst resume that leads with dashboards and insights, to a role that is graded on modelling and engineering practice.

Lead with the tables, tests and definitions you own, and the practices you work inside: version control, pull requests, CI, environments. Keep the analysis work, but reframe it as what the models were for. One line about a decision you influenced is enough; the rest of the space goes to evidence you can be depended on.

Treating dbt fluency as the whole job. Candidates who can explain ref, macros and materialisations in detail but cannot state the grain of a table they built get rejected, and usually never find out why.

Drill grain until stating it is automatic: before writing, after each join, and in every model description. The framework is the easy half of the role and interviewers know it; they are buying modelling judgment, and the framework is just how you express it.

A portfolio project that cannot be run. It needs cloud credentials, or a credentials file that is not in the repository, or it fails on the first command.

dbt plus DuckDB, data committed as seeds or a small Parquet file, and the exact command in the first three lines of the README. Verify it by cloning into a fresh directory yourself. A project that does not run is worth the same as no project.

Describing scale as an achievement with no quality attached: owned a 600-model dbt project. A panel hears a possible mess, especially if nothing was ever removed.

Pair size with health and with subtraction: test coverage on the mart layer, build time, cost, and the number of models you deleted after checking consumption. Growth plus pruning reads as ownership; growth alone reads as accumulation.

Answering a modelling case by naming patterns: "I would use a star schema with slowly changing dimensions", without stating a single grain or asking what decision the output serves.

Ask what the output is for, then state the grain of each table in plain words, then name the two or three source conflicts you expect and how you would resolve them. Patterns are the vocabulary, not the answer. Then say what you would not build yet, and why.

Over-engineering the take-home: forty models, four layers, custom macros and a monorepo structure, for a dataset with six tables. It reads as someone who cannot calibrate.

Match the structure to the data and say that you did. Fifteen models, three layers, two macros, with a README line explaining what you would add if the data were a hundred times bigger. Judgment about proportion is scarce and visible.

Ignoring the stakeholder half of the role, then losing the communication round. Analytics engineering is partly a translation job, and people who dislike that part do not last in it.

Prepare one real story about a definition dispute you resolved: who disagreed, what each side meant, how you reconciled the numbers, what got certified, and what you deprecated. Practise the discrepancy answer in three beats, reconcile in units, explain which definition serves which purpose, state the change that prevents a recurrence.

Not knowing what your own work costs. Asked roughly what a model costs to run, or how to find the most expensive model in a project, many candidates have never looked.

Before interviewing, open the query history and cost views on whatever warehouse you can reach, including a free tier, and learn where spend attribution lives. Have one example of a cost reduction with the lever named. Cost fluency is rare at every level and lands disproportionately well.

Leading the resume with certifications and courses, in a role where no credential is required.

Lead with work, including adjacent work. The finance analyst who rebuilt the monthly close in SQL has a stronger first bullet than anyone with four certificates. Keep certifications to one line at the bottom, where they can still clear a keyword filter without setting the frame.

Mishandling the AI question in either direction: claiming the role is being automated away, or claiming you never use AI tools.

Say the narrow true thing. Authoring got fast, reviewing became the constraint, and the semantic layer became the interface AI tools query, which raised the value of documentation and certified metrics. Grain, source arbitration, late data and stakeholder agreement did not change. Bring one example of generated SQL you caught being wrong and what the error was.

Searching only for the exact title, and so seeing a fraction of the market. A large share of these jobs are posted as BI developer, analytics developer, reporting engineer, or data analyst with dbt in the requirements.

Search for dbt as a keyword with no title filter, and save searches for the title variants. Then add the two channels the boards never show: analytics consultancies and implementation partners, who hire continuously and are the friendliest entry point for a switcher, and community job channels where hiring managers post directly.

Accepting an open-ended unpaid take-home of ten or more hours, or accepting one with no stated rubric.

Ask what the time box is and what they will grade. Deliver to the box and write in the README what you would have done with more time. If the scope is genuinely large and unpaid, say so and offer a live pairing session instead. Many teams accept, and a team that will not discuss it has shown you how it will treat your time later.

Questions people ask

What is an analytics engineer?

An analytics engineer turns raw data that has already landed in a warehouse into trustworthy, tested, documented tables and agreed metric definitions that other people query. The work is mostly SQL and configuration, committed through pull requests into a version-controlled project, with tests, code review, CI and separate environments, which is why the role is described as applying software engineering practice to analytics. The output is a dependable thing others build on, as distinct from a data engineer's output (data arriving reliably) and an analyst's output (an answer and a recommendation). A practical test: if your work lands by merging a pull request someone else reviewed, you are doing analytics engineering; if it lands by saving a dashboard or editing a stored procedure in place, you are not yet, whatever the title says.

What is the difference between an analytics engineer and a data analyst?

A data analyst answers business questions and argues for decisions, and is measured on the quality and influence of those answers. An analytics engineer builds the modelled layer the analyst works from, and is measured on whether those tables can be trusted and whether people can find and understand them. The analyst's deliverable is usually a dashboard, a document or an experiment readout; the analytics engineer's deliverable is a mart, a test suite, a metric definition and documentation. The two roles overlap heavily in skills, both live in SQL, and diverge in accountability, and at small companies one person does both. Analytics engineering is often paid on a higher band, but that depends more on which ladder the company places the role on than on the work itself.

What is the difference between an analytics engineer and a data engineer?

A data engineer gets data into the warehouse and keeps it arriving: connectors, change data capture, streaming, orchestration, platform reliability and infrastructure-level cost, mostly in Python and infrastructure code. An analytics engineer takes it from there and shapes it: staging, modelling, grain, tests, metric definitions and documentation, mostly in SQL and YAML. The clean division is before and after the raw landing zone. In practice the line moves by company, at a small company one person does both, and at a large one the analytics engineer may also own orchestration of the transformation layer. If a posting emphasises Python, streaming, Kubernetes and on-call for pipelines, it is a data engineering job regardless of the title on it.

Do I need to know Python to get hired as an analytics engineer?

Usually not, which surprises people. The core of an analytics engineering job is SQL plus a transformation framework plus the warehouse, and many analytics engineers write very little Python. Python helps in three situations: orchestration work where you touch Airflow or Dagster, writing a small amount of glue or a custom data test, and roles that sit closer to the data platform team. If a posting asks for Python without naming a use for it, it is usually a nice-to-have and a strong SQL and modelling candidate clears it. Conversely, strong Python with weak dimensional modelling fails this loop, so if you are rationing study time, spend it on grain, incrementality and testing rather than on Python.

Do I need a dbt certification to get an analytics engineer job?

No. There is no required credential for analytics engineering: no licence, no registration, no mandatory degree. The dbt Analytics Engineering certification is real and can occasionally help a recruiter justify passing you along, particularly if you are switching careers and have no relevant job title yet, and the same goes for warehouse certifications such as SnowPro or the Databricks and Google data credentials. But no hiring manager will trade a certificate for a readable repository and a clean SQL screen. If you have a fixed number of weekends, build a small runnable project with a decisions README before you sit an exam. The project gets you interviews; the certificate gets you past a filter you can often get past anyway.

What do analytics engineers get paid?

There is no reliable single band, and the reason is worth knowing: analytics engineer has no dedicated US Bureau of Labor Statistics occupation code, so official wage data does not describe it directly. The nearest OES codes are 15-2051 (data scientists) and 15-1243 (database architects), and both mismap. The usable sources are ranges published in postings under state pay-transparency laws (Colorado, California, Washington, New York, Illinois and Massachusetts among the states that require them), levels.fyi for companies with a public ladder, and asking the recruiter directly on the first call, in some states, including California, they must give the range for a role on request. The largest single lever is which ladder the role sits on: the same work paid on a platform or software engineering band earns materially more than on an analyst band, and that is decided by the team you join rather than by your skill. Ask which ladder and level the role maps to before you discuss numbers.

How do I get an analytics engineer job with no analytics engineer experience?

The highest-probability path is to do the work under a different title at your current employer and then rename it. If you are an analyst, a BI or report developer, or the person in finance or operations who became the SQL person, volunteer for the transformation migration, move shared logic into version control, own the metric definitions, and start reviewing other people's pull requests. Holding the title for even a year transforms your external market. If that is not available, the two external routes that reliably open are analytics consultancies and dbt implementation partners (they hire continuously, train deliberately, and give you several stacks in two years) and contract-to-hire roles. In both cases a small runnable portfolio project with a decisions README is what gets you the first conversation, because there is no credential standing in for it.

What does the SQL interview for an analytics engineer actually test?

Warehouse analytics SQL, not algorithm puzzles. The recurring shapes are deduplicating to one row per key with a defensible tiebreak, window functions for running totals and first and last touch, building a date spine and filling gaps, cohort retention, gaps and islands over sessions, spotting and fixing a join that fans out and inflates a measure, NULL behaviour in NOT IN and in outer joins, and timezone and day-boundary handling. Interviewers care at least as much about your narration as your syntax, and the highest-value habit is stating the grain of your result before you write and again after each join. Many screens also ask roughly what the query would cost and how you would make it cheaper.

Is dbt still the standard for analytics engineering in 2026, and should I learn SQLMesh instead?

dbt is still the default and is what postings name, so learn it first. SQLMesh is a credible alternative with real adoption, particularly where teams want column-level change awareness and virtual data environments; Dataform appears in Google-centric shops and Coalesce in some enterprises. Warehouses have also absorbed part of this work directly through declarative features such as dynamic tables and materialized views. The practical answer for a candidate: be genuinely fluent in one framework and able to discuss the trade-offs of the others for ten minutes. Interviewers increasingly ask what you would choose and why, and the good answer is about the team's review practices and environment needs rather than feature lists. The underlying skills (grain, incrementality, testing, lineage) transfer across all of them, which is also the right thing to say.

Will AI replace analytics engineers?

No, and the honest version is more useful than either extreme. Generating SQL is now fast and cheap, which was never the expensive part of analytics engineering. What has not been automated: deciding the grain of a table, arbitrating which of two conflicting sources is authoritative, agreeing with a finance stakeholder what a renewal means, handling data that arrives late or gets corrected, sequencing a migration, and deprecating something people still depend on. What has changed is that review became the constraint rather than authoring, and that the semantic layer turned into the interface AI assistants query, which raised the value of documentation, naming and certified metrics. The most exposed work is the most junior kind, writing a query someone else fully specified, so a candidate should lead with modelling judgment rather than volume of output.

Put this on a resume in about a minute

Paste your history once and point it at the Analytics Engineer posting you are looking at. No account, no card.

Build my resume free More roles