Data & Analytics

How to get hired as a data analyst in 2026-27

The short answer

To get hired as a data analyst in 2026-27 you need three things: SQL you can write correctly while someone watches, one piece of work where you turned a vague business question into a number somebody acted on, and a target list that is not only tech companies. No licence or certification gates this job, the real gate is a live SQL exercise, usually 30 to 60 minutes, where the interviewer watches how you handle join grain, NULLs, date truncation and dedup, and where sanity-checking your own output counts as much as the final query. AI changed the job at the edges rather than the core: assistants now draft a competent query in seconds, which has thinned the pure "pull me this number" work and raised the bar for juniors, while choosing what to measure, knowing which table is the audited one, and standing behind a number in front of a director is still a person's job. The most likely first analyst job is a reporting or operations seat inside a hospital system, insurer, bank, utility, university, logistics firm, government agency or retailer, and it is often filled by internal transfer, referral or contract-to-hire rather than from the public applicant pile.

Licence or credential requiredNone. Data analyst is an unlicensed, uncredentialed occupation in the US: no board exam, no registration, no mandatory certification. The functional gate is a live SQL exercise. Certifications (Microsoft PL-300 for Power BI, Tableau Certified Data Analyst, SnowPro Core, Databricks Certified Data Analyst Associate, the Google Data Analytics Professional Certificate) are tie-breakers and HR checkboxes, never substitutes. In healthcare specifically there is one near-exception worth knowing: Epic's reporting certifications (Clarity, Caboodle, Cogito) normally require sponsorship by an Epic customer, so you earn them after you are hired, not before.
How long it takes to become hireableFrom zero, plan on roughly 6 to 12 months of deliberate work to be interview-ready: SQL through window functions, one BI tool to real depth, Excel properly, enough Python to clean a file, and two finished projects on messy real data. From an adjacent seat you already hold (operations, finance, customer support, marketing, claims, clinical coding, logistics) it is often 3 to 6 months, because domain knowledge is the slow half and you already have it. Many postings list a bachelor's degree as preferred or required; non-tech and public-sector employers enforce that more literally than tech companies do, and the field of the degree matters less than that it exists.
Typical hiring loopTech or tech-adjacent: application through Greenhouse, Ashby, Lever or Workday; recruiter screen of 20 to 30 minutes; live SQL screen of 30 to 60 minutes on CoderPad or similar; a case or short take-home; hiring manager; one or two stakeholder interviews; sometimes a 20-minute presentation of the take-home. Two to five weeks. Large non-tech employer: fewer and less technical stages, but a panel of three to six people scoring against a rubric, often in one block, over three to eight weeks. Public sector: a scored application, a structured panel where every candidate is asked identical questions, then references and a background check, over two to four months.
Who screens youFirst an ATS keyword match or a recruiter filtering on tool names. Then an analytics manager, BI manager or senior analyst who grades the SQL. Then the business stakeholders you would serve (a marketing director, a revenue-cycle manager, a supply-chain lead, a controller) who grade whether you can explain a number and its caveats without jargon. At a non-tech employer the stakeholder panel usually holds the real veto, and the technical bar is lower than the communication bar.
Pay: where to check instead of a quoted bandThere is no single federal occupation code for this title, which is why quoted "average data analyst salary" figures are unreliable. Use the BLS Occupational Employment and Wage Statistics tables for the code that matches the actual job: 15-2051 Data Scientists (business intelligence analysts sit here, as O*NET 15-2051.01), 13-1161 Market Research Analysts, 13-2051 Financial and Investment Analysts, 15-2031 Operations Research Analysts, 15-2041 Statisticians, 43-9111 Statistical Assistants, 13-1111 Management Analysts. BLS publishes 10th, 25th, median, 75th and 90th percentiles by state and metro: read the percentiles, not the mean. Then read live posted ranges under pay-transparency laws (Colorado, California, Washington, New York, Illinois, Minnesota, Maryland, Hawaii, Vermont, New Jersey, Massachusetts, District of Columbia), levels.fyi for tech total compensation, the Robert Half Salary Guide for enterprise and contract rates, and the OPM General Schedule plus locality tables for federal roles in the 0343, 1530/1529, 1560 and 2210 series.
The actual technical barSQL: joins including anti-joins, conditional aggregation, CTEs, window functions (ROW_NUMBER, LAG and LEAD, running totals, ranking with ties, percent of total), date truncation and time zones, dedup, cohort and funnel patterns. One BI tool deep enough to build a semantic model rather than drag fields: DAX in Power BI, LOD expressions in Tableau, LookML in Looker. Excel to pivot-table and lookup fluency. Python or R is commonly listed and rarely tested hard in a reporting job; it is tested properly in a product analyst loop. Legacy stacks are normal and not a red flag: SQL Server with SSRS, SAS in insurance, pharma and government, Alteryx in finance, Epic Clarity and Caboodle in hospitals.
What AI changed, honestlyQuery writing got cheap and the pure ad-hoc-request job thinned. The judgment did not move: choosing the metric, knowing which of four revenue tables is the audited one, catching an answer that is plausible and wrong, and being accountable for it. Net effect on candidates: the junior reporting rung is harder to get onto, and the differentiating skill is now making self-serve answers correct, documented tables, defined metrics, and a semantic layer an assistant can query without inventing an answer.
Related titles that are the same job, or notBusiness Intelligence Analyst, Reporting Analyst, Insights Analyst, Decision Support Analyst and Product Analyst are usually this job under another name. Business Analyst is often requirements and process work with little SQL: read the posting before assuming. Analytics Engineer is the data-modelling half split out, and usually pays more. Data Scientist at a tech company normally means experimentation and modelling; at many non-tech employers the same title is used for work a product analyst does elsewhere.

Four different jobs share this title: read the posting's verbs

"Data analyst" names at least four jobs with different day-to-day work, different interview loops and different pay. Applying to all of them with one resume is the most common reason a qualified person gets no replies. The posting tells you which one it is if you read the verbs and the tool list instead of the title: build, publish, maintain, refresh, support means reporting; test, measure, instrument, diagnose means product; model, transform, document, dbt means analytics engineering; forecast, reconcile, variance, close means a domain finance or operations seat.

The split decides your preparation. A reporting loop tests SQL and a BI tool and will barely mention statistics. A product loop tests experiment design and metric definition and will not care how pretty your dashboard is. A domain loop is mostly about whether you understand the business process: a healthcare revenue-cycle panel will ask what a denial is, and no amount of window-function practice covers that.

One more title trap. A "Business Analyst" posting often describes requirements gathering, process mapping and stakeholder documentation with SQL as a nice-to-have. That is a different career with different interviews. Check whether the posting names a warehouse and a BI tool, or names Visio, Jira and user stories.

What is actually true about the 2026-27 analyst market

Two separate things happened to analyst hiring, and merging them produces the wrong preparation. First, from 2023 onward, data and analytics headcount tightened on cost grounds (the end of cheap capital, flatter org charts, hiring freezes) well before AI-assisted analysis was meaningfully adopted. Second, from roughly 2024, assistants genuinely did change what the remaining analyst does. The second did not cause the first.

The practical consequence: the easiest version of the entry-level job (sit near a stakeholder, take requests, write the query, send the number) is the part most exposed, because a stakeholder with a text-to-SQL tool pointed at a tidy semantic layer can get a plausible number without you. That rung got narrower rather than disappearing; it moved up a step. Employers now expect a junior to arrive able to model data, document it, and notice when a self-serve answer is wrong.

Application volume went the other way. Applying became nearly free, so a well-titled analyst posting now draws a very large pile, almost all of it tailored, because a model tailored it. Two things follow. Referrals and internal transfers are disproportionately effective because they skip the pile. And specificity is the only thing that survives a screen: a bullet naming a system, a row count and a decision reads as lived, while "leveraged data-driven insights to optimise business outcomes" reads as generated whether or not it was.

Where the jobs are is the single most useful correction to most candidates' target lists. The volume of analyst hiring sits outside tech: health systems and payers, insurance, banking and credit unions, state and local government, federal contractors, school districts and universities, utilities, transit agencies, grocery and retail chains, distribution and logistics, manufacturers, non-profits and professional services. These employers hire continuously, often pay less than tech, screen on SQL plus Excel plus domain rather than on a portfolio, and receive far fewer applicants per posting. They are also where contract-to-hire is normal: a six-month contract through a staffing firm is a legitimate way in and converts often enough to take seriously.

Expect their stacks to be unfashionable, and do not treat that as a downgrade. SQL Server with SSRS, a decade of Excel models, SAS in insurance and government, Alteryx in finance, Epic Clarity and Caboodle in hospitals. Learning the specific stack of the employers in your metro beats learning the stack a tech company blog post recommends, because the posting filter is literal: a hospital hiring for Clarity reporting will shortlist the person who already names it.

How hiring works, stage by stage

There is no licence, no board and no portfolio review committee. The loop is short compared with engineering, and the decisive stage is almost always the live SQL exercise. Below is the common tech or tech-adjacent version.

Two variants matter. At a large non-tech employer, expect fewer and less technical stages but a panel: three to six people including the stakeholders you would serve and often an HR representative, scoring against a rubric, sometimes in a single two-hour block. The questions skew behavioural and practical, walk us through a report you built; two departments disagree about a number, what do you do; describe a time you found an error. SQL may be a short written exercise or, surprisingly often, untested, which means these jobs go to the candidates who explain their work clearly to non-technical interviewers.

In the public sector the process is longer and more formal. The first screen is a scored check that your application states each listed minimum qualification in the posting's own words, with dates and hours per week; a resume that describes the same experience in different language gets marked as not meeting minimum qualifications. Federal resumes are expected to be long and specific, not one page. Then a structured panel where every candidate gets identical questions, then references and a background check. Two to four months end to end is normal. Federal analyst work sits mostly in the 0343 (management and program analysis), 1530 and 1529 (statistician and mathematical statistician), 1560 (data science, a series OPM established in 2021) and 2210 series. Some of it requires US citizenship, some requires a clearance, and a clearance is both a delay and a durable advantage once you hold one.

On take-homes: they got shorter and more often supervised, because an unmonitored take-home stopped being informative once a model could do it. A reasonable employer now gives a two-to-three-hour dataset exercise and spends 30 minutes asking you to defend your choices, or runs a live pair-analysis session instead. Treat any take-home demanding more than four hours of real work as information about how the company treats people's time, and say so politely if you decline.

Know what a take-home is actually scored on. Almost nobody scores model accuracy or chart beauty. They score: did you state the question you answered, did you notice what is wrong with the data, did you say what you excluded and why, did you end with a recommendation, and did you name what would change your mind. A two-page write-up with a clear answer beats a forty-cell notebook with no conclusion.

Pay: where to look, and what actually moves it

Do not trust a single quoted "average data analyst salary." The title spans a reporting analyst at a county government and a product analyst at a public software company, and no survey separates them cleanly. There is also no single federal occupation code for the title, which is the root of most of the bad figures circulating online.

Build your own picture from four sources, in this order. First, the BLS Occupational Employment and Wage Statistics tables, which publish 10th, 25th, median, 75th and 90th percentile wages by state and metro. Look up the code that matches the real job: 15-2051 data scientists (where business intelligence analysts sit, as O*NET 15-2051.01), 13-1161 market research analysts, 13-2051 financial and investment analysts, 15-2031 operations research analysts, 15-2041 statisticians, 13-1111 management analysts, and 43-9111 statistical assistants, which is closer to a junior reporting seat than people like to admit. Read the percentile spread for your metro, not the national mean.

Second, live postings with ranges disclosed under pay-transparency law. Colorado, California, Washington, New York, Illinois, Minnesota, Maryland, Hawaii, Vermont, New Jersey, Massachusetts and the District of Columbia require a posted range, and many national employers now post one everywhere rather than maintain two versions of a posting. Twenty current postings for your exact title, level, industry and metro is better evidence than any survey. Read the bottom of the range: an external candidate who does not match the stack exactly rarely lands above the midpoint, and the top is usually reserved for internal promotions.

Third, levels.fyi for tech companies, where total compensation includes equity and the base figure alone misleads. Fourth, the Robert Half Salary Guide for enterprise salaries and contract hourly rates, and the OPM General Schedule plus locality tables for government, where grade and step determine the number exactly, the negotiating lever there is the grade and the step you are brought in at, not the salary.

What actually moves the number, roughly in order of effect: industry (finance, tech and health payers pay well above non-profit, education and local government for identical work); metro, or the remote policy that replaces it; whether you own a metric close to revenue; and stack depth, where someone who can model in dbt and own the semantic layer is paid closer to an analytics engineer than to a report builder. Then domain scarcity: actuarial-adjacent insurance data, clinical and claims data, payments and fraud, pricing, and anything requiring a clearance all carry a premium because the pool is small.

Two practical notes. Contract hourly rates often look better than the equivalent salary and usually come without benefits or paid time off, so compute the annual equivalent including the gap between contracts before comparing. And the largest raise most analysts ever get is the second job, not a promotion in the first. A first analyst job is worth taking slightly underpaid if it gives you production SQL against real data, because that is what makes you a credible candidate eighteen months later.

The SQL screen: what is actually being tested

Almost everyone who fails an analyst loop fails here, and almost nobody fails on exotic syntax. They fail because they write a query that runs, returns a believable number, and is wrong, and they do not notice.

What a competent interviewer is watching: do you ask about the grain of each table before joining, do you check whether a join multiplied rows, do you handle NULLs deliberately, do you reach for a window function instead of a correlated subquery, do you narrate your reasoning, and do you sanity-check the output before announcing it. What gets a candidate through is not a clean query, it is a checked one. A bug you catch and fix in front of the interviewer is evidence of how you work; a wrong number handed over confidently is the most common way an analyst loop ends.

Ask which dialect you are in before you start, and then stay in it. Postgres, T-SQL, Snowflake and BigQuery differ in date functions, string handling and how they treat quoting, using the wrong one and not knowing it is a small error that reads as inexperience. Saying "I usually write Postgres, is this T-SQL?" reads as the opposite.

Practise with a schema and a question, out loud, against a real database. A local Postgres or DuckDB instance with a public dataset is enough. Typing SQL into a text editor without executing it is the main reason people feel ready and are not. DataLemur, StrataScratch, the HackerRank SQL track and LeetCode's database section use the right question format; the point of the practice is to get to the stage where you notice your own wrong answers.

Two habits that read as professional in a live screen. Say the sanity check out loud before you run it: "this should be one row per order, so let me check the count against the orders table." And when you are stuck, say what you are trying to express in English before trying to express it in SQL. Both convert a stuck moment into evidence of how you think, which is what is being graded.

The case, the statistics questions, and the two-minute answer

An analyst case is not a consulting case. It is a diagnosis exercise with a correct shape: something moved, so establish whether it is real, then find the cause, then say what you would do. The failure mode is naming a cause in the first ten seconds and then defending it.

Take the standard version: weekly active users dropped week over week. A strong answer starts by asking what decision depends on the answer and how long you have. Then it separates measurement from reality, was there a release, a tracking change, a bot-filter change, a pipeline failure, a late-arriving load? Then it checks the baseline: what did the same week look like in prior years, is there a holiday, is this inside normal week-to-week variance at all? Only then does it segment by platform, app version, geography, new versus returning, acquisition channel and device, to see whether the drop is broad or concentrated, a concentrated drop almost always has a specific cause, a broad one is usually seasonality, a measurement change or something external. It finishes by naming the first query it would run and the result that would prove the hypothesis wrong.

Say the mix-shift and Simpson's paradox possibilities out loud when they apply: a total can fall while every segment holds steady, if the mix of segments changed. Interviewers notice candidates who think about composition.

The metric-definition question is the other staple: how would you measure whether this feature succeeded, or what should this team's primary metric be. What is being tested is whether you pick something that moves when the thing you care about moves, is hard to game, and the team can actually influence, and whether you volunteer a guardrail metric alongside it. Define it precisely enough to compute: numerator, denominator, time window, inclusion rules, which users are excluded.

The statistics bar is modest for a reporting analyst and real for a product analyst. Know these cold and in plain English rather than as formulas: what a p-value is and specifically what it is not (it is not the probability the hypothesis is false), how to read a confidence interval, why stopping a test the first time it looks significant inflates false positives, what drives required sample size, what to do with a flat result (often it is a real answer, and the test may simply have been underpowered to detect an effect you would have cared about), multiple-comparison inflation when you slice a test ten ways, novelty and primacy effects, and why mean revenue per user is a poor summary of a heavily skewed distribution.

Finally, the two-minute answer. Every stakeholder interview is secretly testing one thing: can you lead with the answer. Practise the pattern, here is what we found, here is what I recommend, here is the one caveat that matters, and here is the evidence if you want it. Analysts who narrate methodology first and reach a conclusion in minute six lose offers to analysts no more skilled who said the answer first.

The resume and the portfolio: what gets read and what gets skipped

Analyst resumes fail predictably: they list outputs instead of consequences. "Built 14 dashboards in Tableau" tells a hiring manager nothing, because dashboard count is not value and some of those dashboards were nobody's. A line in the shape of "replaced nine weekly spreadsheets with one dashboard used by the store-operations team, cutting several analyst-hours a week" tells them what you are, and if you know the real figures, use them.

Write one line per piece of work: what question or problem, what you did, what changed as a result, with units. The numbers that make a line credible are rows or records, source systems integrated, how many people or teams consume the output, hours or dollars saved, error rate reduced, the decision taken. Never estimate upward to make a line sound better; an inflated number is the thing an interviewer probes. If a piece of work genuinely did not change a decision, say what it enabled instead, a reconciliation that closed a gap, a definition that stopped two departments reporting different revenue.

Format plainly and in a single column. The first pass is increasingly a model summarising your resume for a recruiter, and a two-column design that interleaves when parsed turns your experience into nonsense. Keep a tools block near the bottom with exact names (Power BI with DAX, Snowflake, dbt Core, PostgreSQL, SSRS), that block exists for the keyword screen, not for the human.

What gets skipped: objective statements, soft-skill adjectives, "proficient in Microsoft Office," long certification lists, coursework presented as experience, and any claim to Python you could not act on in an interview. Listing a tool you cannot discuss is a bad trade: the question will come, and the answer damages everything else you said.

Now the portfolio question, answered honestly. A portfolio does not get you hired as an analyst; it gets you past the no-experience objection. Hiring managers do not read five projects. One good project, described in a page, with a link to the SQL or the notebook, does the job. A Titanic or Iris notebook does nothing at all, they have seen thousands, and a model can produce one in a minute, so it signals nothing about you.

What makes a project count: real messy data you had to clean, a question somebody could act on, explicit documentation of the judgment calls (what you excluded, how you handled missing values, which of two conflicting fields you trusted and why), and a conclusion with a recommendation and a limitation. The cleaning decisions are the evidence, because they are specific to your dataset and cannot be faked. Good sources: your city or state open data portal, data.gov, CMS provider and hospital data, FRED, BLS, SEC EDGAR, GTFS transit feeds, state crash or inspection records. Pick a domain that matches the jobs you want so the project doubles as domain evidence.

Better than any portfolio, if you can get it: real work with someone else's data. A small non-profit, a church, a youth sports league, a local business, a team at your current employer that runs on spreadsheets. It comes with a stakeholder, a deadline, bad data and an actual decision, which is to say it is the job, and you can talk about it in an interview like a job.

Getting the first one, and where these jobs are actually posted

Ranked by probability of success, these are the routes into a first analyst job.

One: the internal transfer. If you work anywhere that has data, you have an unfair advantage, you know the business, you know whose numbers are wrong, and a hiring manager is comparing you with strangers. Start by solving a real reporting problem for your own team, then ask the analytics manager what they wish someone would fix. People move into analyst seats from customer support, claims processing, pharmacy tech, warehouse supervision, bank branch work, school administration and recruiting coordination constantly. It is the most common origin story in the field and almost nobody writes about it.

Two: contract-to-hire through a staffing firm that places into enterprises. Lower pay, less glamour, real production data on your resume in six months. Three: the employer in your metro that nobody is excited about, a regional insurer, a utility, a county health department, a distributor. Four: a small company where you will be the only analyst, which is a fast education and a weak mentorship situation; take it with your eyes open. Five: a formal analyst rotation or new-grad programme if you are near graduation, these exist mostly in banking, insurance, health systems and large retailers, and they recruit on a calendar, usually in the autumn for the following year.

What does not work in 2026: applying to 300 remote postings at recognisable companies through one-click apply. A remote analyst posting draws a national and often international pool, and the screening has to be brutal. A local, hybrid or in-office role at an unfashionable employer has a fraction of the competition for the same work.

On bootcamps and certificates: the Google Data Analytics Professional Certificate and its equivalents teach a reasonable foundation and are worth doing if you need structure, but so many people hold them that they differentiate nothing. The ones that occasionally matter are vendor certifications in the tool a specific employer uses (PL-300 where the shop is Power BI, Tableau Certified Data Analyst where it is Tableau, SnowPro Core where it is Snowflake) because a non-technical recruiter or a government HR screen can tick the box, and because they prove you learned a product rather than a concept. None of them replaces writing a correct query while someone watches.

Where to look: company career pages directly, which is the highest signal-to-noise channel and the one where you can apply within hours of a posting going live; LinkedIn with the location filter set to your metro rather than Remote; Indeed and Glassdoor for non-tech employers; USAJOBS plus state, county and city portals for government; your state hospital association and local health system career pages for healthcare; and the job boards of the specific staffing firms that place analysts in your city. Then referrals: a short, specific message to someone doing the job you want, asking one answerable question about their work rather than asking for a referral, converts far better than a connection request with a resume attached.

One last thing worth internalising. Nearly every analyst job is, at bottom, somebody needing to make a decision and not trusting their numbers. Everything above is in service of becoming the person they trust with that. The SQL is the gate; the trust is the job.

Working with AI in this role

What a data analyst has to know about AI in 2026-27

The honest version: at the core of this job, less changed than the headlines claim. Writing a query got dramatically faster. Deciding what to measure, knowing which of four revenue tables is the audited one, noticing that an answer is plausible and wrong, and being the person a director trusts with a number, none of that moved. Anyone telling you AI eliminated data analysis is describing a job where the analyst was a human query interface, which was always the most replaceable version of the role.

What did change is worth being precise about. Natural-language querying stopped being a demo and shipped inside tools employers already pay for: Snowflake Cortex Analyst, Databricks AI/BI Genie, Power BI Copilot, Tableau's assistant and Pulse metrics, conversational analytics in Looker, ThoughtSpot Spotter. Business users can now ask a question and get a chart. They can also get a confidently wrong chart and find out in a board meeting. That single fact created most of the new work in this job.

The gap between the demo and your warehouse is not a secret, and you can check it rather than take anyone's word. On tutorial-grade schemas these tools look close to solved; the public benchmarks built on realistic enterprise schemas and workflows, BIRD and Spider 2.0, score far lower, and the difference is largely context: undocumented columns, four overlapping order tables, business rules that exist only in someone's head. Read those benchmark pages before you form an opinion, and before an interviewer asks you for one.

So the value moved one layer down, into the thing those tools read: the semantic layer and the documentation under it. Metric definitions, table and column descriptions, synonyms for how the business actually talks, verified queries, row-level security, deprecated tables genuinely hidden. An assistant pointed at an undocumented warehouse will pick a table and sound certain. An assistant pointed at a modelled one is usually right. The analyst who owns that layer has become the person the AI rollout depends on, and it is now a routine interview topic rather than a specialism.

Two cautions. Do not claim fluency you cannot demonstrate: "I use Copilot" is not an answer, and interviewers have started asking what it got wrong for you. And know your employer's rules about data leaving the building, pasting customer rows, PHI, or anything under a confidentiality term into a consumer chatbot is a fireable mistake in a regulated industry, however convenient. Knowing which tool is the approved one, and saying so unprompted, reads as professional maturity. The same restraint applies to agentic tooling: assistants that connect straight to a warehouse and run scheduled checks are real, early, uneven and heavily over-claimed, so have a modest opinion about where a human approval step belongs rather than a confident one about autonomy.

Validating a query you did not write

An assistant produces syntactically perfect SQL with repeatable semantic failures: joining at the wrong grain so revenue double-counts, missing a soft-delete column it did not know existed, using a deprecated table that still returns rows, silently dropping records with NULL join keys, truncating dates in the wrong time zone. Each yields a number that looks fine. Catching these is now a larger share of the job than writing the query was, and it is precisely what a live SQL screen tests.

Show it: Have a validation routine you can state in one breath, and state it in the interview: check the row count against the base table, reconcile the total to a number someone already trusts, trace three individual records end to end, and look at the distribution rather than only the total. Then tell a specific story: a generated query that returned a believable number, how you caught it, and what the real figure turned out to be.

Owning the semantic layer and the documentation it reads

Text-to-SQL accuracy on a real warehouse depends far more on metadata quality than on which model is behind it. Defined metrics, described columns, documented grain, hidden deprecated tables and business synonyms are what turn a demo into something a finance director will rely on. Documentation used to be the work nobody did and nobody punished; now it is load-bearing, because the description you write on a column decides whether the next hundred self-serve answers are right. Employers are hiring for this explicitly, look for "semantic model", "metric layer", "certified datasets", "dbt documentation" in postings.

Show it: Name the artefact and the tool: dbt models with tests, descriptions and a metrics layer; LookML with defined measures; a Power BI semantic model with certified datasets and documented DAX; Snowflake semantic views; a dated metric definitions page you own. Then give the before and after in human terms, three teams reporting three different revenue numbers, one agreed definition, and the dispute ending.

Using an assistant well on your own work

Analysts who get real speed from these tools work differently from those who do not. The pattern that works is specification first, small steps and verification at each step: describe the grain and the expected output shape, generate one CTE at a time, run it, check the count. The pattern that fails is asking for the whole analysis in one prompt and accepting the answer. Interviewers can tell the difference within two questions.

Show it: Describe your workflow concretely: where you let it draft (boilerplate transformations, a DAX measure you half-remember, regex, chart formatting, a first-pass data dictionary) and where you refuse (choosing the metric, deciding which table is authoritative, anything reaching an executive unverified). Having a clear line and being able to justify where you drew it is the signal.

Turning unstructured text into analysable data

Support tickets, survey free text, call transcripts, reviews, inspection comments and clinical notes were previously too expensive to analyse at scale and are now cheap to classify. This is genuinely new analyst territory, and often the fastest way to produce something nobody at your employer has: the reason codes behind a churn number rather than just the number.

Show it: Show that you measured the classifier instead of trusting it: hand-label a sample as ground truth, report agreement, keep a holdout, state the cost per thousand rows, and say what you did about the categories it confused. Then connect it to a decision, the top three ticket drivers, and the one that got fixed.

Measuring whether a self-serve AI deployment is actually correct

Someone at your employer has either bought a conversational analytics feature or will. The question that stops the room is "how do we know it is giving correct answers," and most organisations have no answer. An analyst who can build the evaluation becomes central to that programme immediately.

Show it: Describe a golden question set: real questions people ask, each with a verified answer and the SQL that produces it, re-run after every semantic-model change, with results tracked over time. Say what accuracy you measured and which question types failed, typically anything involving a period comparison, a non-obvious join, or a word the business uses differently from the schema.

Permissions, PII and what leaves the warehouse

A natural-language query tool is a new path to data, and permissions do not follow it automatically. If row-level security is enforced in the BI layer but the assistant queries the warehouse through a service account, people can answer questions they previously could not. In healthcare, financial services, education and government that is a compliance incident, not an inconvenience.

Show it: Say how access was enforced in a system you worked on (row-level security in the semantic model, warehouse roles, masking policies on PII columns, separate views for restricted fields) and name one dataset you deliberately kept out of a self-serve tool because the permission model could not be honoured. The exclusion is what makes it sound real.

Knowing where AI did not help, and saying so

Interviewers are tired of enthusiasm. A candidate who can name the limits sounds like someone who has used these tools on real data rather than on tutorials, and the limits are substantial: anything depending on knowing why a table is untrustworthy, any question whose correct answer lives in a business rule nobody wrote down, experiment design, root-causing a data quality problem, and persuading a sceptical stakeholder.

Show it: Give one specific example of asking an assistant for something and getting a wrong answer only business knowledge could catch, a metric defined differently in finance than in product, a backfill that made a historical trend look like a behaviour change, a table that stopped updating in March while still returning rows.

What a screen is looking for

These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.

Mistakes that cost people this job

Listing dashboards and tools as achievements: "built 14 Tableau dashboards," "proficient in SQL, Python, R, Tableau, Power BI, Excel, SPSS, SAS."

One line per piece of work naming the consequence with units: reports retired, hours saved, people using the output, the decision taken, the error found. Keep the tool list as one block near the bottom for the keyword screen, and only list tools you can be questioned about.

Writing a query in the live screen that runs, returns a believable number, and is wrong, usually because a join fanned out or NULLs were dropped, and not noticing.

Narrate a check after every join: row count against the base table, a total reconciled to something already trusted, three records traced end to end. Catching your own bug in front of the interviewer is evidence; a silent wrong answer is how the loop ends.

Applying only to tech companies and to remote postings, then concluding the market is closed.

Target the employers nobody is excited about in your own metro (health systems, insurers, utilities, county and state government, distributors, universities, banks, retail chains) plus contract-to-hire through staffing firms. Same work, a fraction of the applicants, production data on your resume in six months.

Building a portfolio of tutorial datasets: Titanic, Iris, a pre-cleaned Kaggle CSV, a Netflix EDA notebook.

One project on genuinely messy real data where you document the judgment calls (what you excluded, how you handled conflicting fields, which source you trusted and why) ending with a recommendation and a limitation. Better still, do real work for a real stakeholder: a non-profit, a local business, or a team at your employer living in spreadsheets.

Answering the case interview by naming a cause in the first fifteen seconds and then defending it.

Separate measurement from reality first (release, tracking change, pipeline failure, late data), then check baseline and seasonality, then segment to see whether the change is broad or concentrated. State the first query you would run and the result that would prove you wrong.

Leading the stakeholder interview with methodology and reaching the conclusion six minutes in.

Answer, recommendation, the one caveat that matters, then evidence on request. Rehearse it as a 90-second script against three pieces of your own work until it is automatic.

Claiming Python, R or machine learning on the resume to clear a keyword filter, with nothing behind it.

List only what survives a follow-up question. For most reporting jobs, deep SQL plus one BI tool plus genuinely strong Excel beats a thin spread across six tools. Count it yourself across twenty postings in your metro, Excel appears at least as often as Python, and frequently as a hard requirement.

Collecting certificates as a substitute for interview-ready skill: a Google certificate, a bootcamp, three Coursera specialisations, and no ability to write a window function unaided.

Treat certificates as structure for learning and as a checkbox for non-technical screens, not as evidence. Spend the same hours executing queries against a real database and finishing one project. If you want a certificate that occasionally moves a decision, pick the vendor one matching the employer's stack: PL-300 for Power BI, Tableau Certified Data Analyst, SnowPro Core.

Saying yes to every request in the job, and showing no evidence in the interview of ever having questioned one.

Have a story about asking what decision a request would inform and finding the real question was different, or cheaper to answer. Hiring managers probe for this, because an analyst who only takes orders becomes a bottleneck.

Misstating basic statistics in a product analyst loop, defining a p-value as the probability the hypothesis is false, calling a flat test "inconclusive" with no mention of power, or reading significance into one of ten slices.

Rehearse plain-English versions of p-value, confidence interval, power, peeking and multiple comparisons. Know that a flat result is often a real answer, and that slicing a test ten ways produces a false positive by construction.

Treating AI tools as either a threat to deny or a buzzword to sprinkle on the resume.

Take a specific position: where you let an assistant draft, where you refuse, one wrong answer you caught because you knew the business, and what you would document so the next self-serve answer is right. That is the version of the question employers are actually asking.

Applying to public-sector postings with a normal one-page private-sector resume.

Mirror the posting's wording literally against each listed minimum qualification, including dates and hours per week, because the first screen is a scored check for those phrases rather than a reading of your experience. Federal resumes are expected to be long. Then prepare for an identical-questions structured panel and a two-to-four-month clock.

Questions people ask

What does a data analyst actually do?

A data analyst turns a business question into a measurable one, gets the data, checks that it is right, and tells someone what to do about it. In practice that means writing SQL against a data warehouse, building and maintaining reports or dashboards in a tool like Power BI, Tableau or Looker, answering ad-hoc questions from a function such as marketing, finance or operations, defining and defending metrics so different teams report the same numbers, and investigating when a number looks wrong. The mix varies by job type: a reporting analyst spends most of the week on recurring reporting and stakeholder requests, a product analyst on experiments and product metrics, an analytics-engineering-leaning analyst on transformation models and documentation.

Do I need a degree or a certification to become a data analyst?

No licence and no certification is required: data analyst is an unlicensed occupation in the US with no board exam and no registration. Many postings list a bachelor's degree as preferred or required, and non-tech and public-sector employers apply that more literally than tech companies, which is why a degree in anything quantitative shortens the screen even when the subject is unrelated. Certifications are tie-breakers and HR checkboxes: Microsoft PL-300 where the employer uses Power BI, Tableau Certified Data Analyst where it is Tableau, SnowPro Core for Snowflake, the Google Data Analytics Professional Certificate as a foundation that no longer differentiates because so many people hold it. None of them substitutes for passing a live SQL exercise, which is the real gate.

How much SQL do I need to get hired as a data analyst?

Enough to write a correct multi-table query while a stranger watches and interrupts. Concretely: joins including anti-joins, aggregation with conditional logic, CTEs, window functions (ROW_NUMBER for deduplication, LAG for period comparisons, running totals, ranking within groups), date truncation and time zone handling, and cohort or funnel patterns. More important than syntax breadth is grain discipline, asking how many rows per key a table has before joining, and checking afterwards whether the join multiplied rows. Candidates who fail a SQL screen usually fail by returning a plausible wrong number and not noticing, not by forgetting syntax.

Has AI made data analyst jobs obsolete?

No, but it has changed which parts of the job carry value and made the easiest entry-level version harder to get. Text-to-SQL and assistant features inside Snowflake, Databricks, Power BI, Tableau and Looker can produce a competent draft query or a chart from a plain-English question, which genuinely thins out the work of being a human query interface for a stakeholder's requests. What has not moved is choosing what to measure, knowing which table is authoritative, catching an answer that is plausible and wrong, designing an experiment, and being the person a director trusts with a number. The practical response is to move toward the work the tools depend on (data modelling, metric definitions, documentation and a semantic layer) and to be able to validate what the tools produce.

What is the fastest realistic route into a first data analyst job?

An internal transfer at an employer you already work for, because you already hold the domain knowledge that takes an outsider a year to acquire and the hiring manager is comparing you with strangers. People move into analyst seats from customer support, claims processing, warehouse supervision, bank branch work, pharmacy tech, school administration and recruiting coordination routinely. Start by solving one real reporting problem for your own team, then ask the analytics manager what they wish someone would fix. If that is not available, the next best routes are contract-to-hire through a staffing firm that places into large enterprises, and an unglamorous local employer (an insurer, a health system, a county agency, a distributor) rather than a remote posting at a company everyone has heard of.

What is a data analyst paid?

It depends so heavily on industry, metro and which analyst job it actually is that a single average is misleading, and there is no clean federal occupation code for the title, which is the source of most bad figures online. Check the BLS Occupational Employment and Wage Statistics percentile tables for the code matching the real job, 15-2051 data scientists (where business intelligence analysts sit), 13-1161 market research analysts, 13-2051 financial and investment analysts, 15-2031 operations research analysts, 15-2041 statisticians, for your state and metro, and read the 25th to 75th percentile spread rather than the mean. Then read twenty live postings with ranges disclosed under pay-transparency laws in Colorado, California, Washington, New York, Illinois, Minnesota, Maryland, Hawaii, Vermont, New Jersey, Massachusetts or the District of Columbia for your exact title and level; levels.fyi for tech total compensation; and the OPM General Schedule and locality tables for government. Expect an external hire to land in the lower half of a posted range unless they match the stack exactly.

Do I need a portfolio, and what should be in it?

A portfolio does not get you hired; it gets you past the no-experience objection, and one project described well does that better than five. What makes a project count is real messy data you had to clean, a question somebody could act on, explicit documentation of the judgment calls (what you excluded, how you handled missing values, which of two conflicting fields you trusted and why) and a conclusion with a recommendation and a limitation. Tutorial datasets like Titanic and Iris signal nothing, because an assistant produces that notebook in a minute. Use your city or state open data portal, data.gov, CMS provider data, FRED, SEC EDGAR or transit GTFS feeds, and pick a domain matching the jobs you want. Real work for a real stakeholder (a non-profit, a local business, a spreadsheet-bound team at your employer) beats any self-directed project.

What is the difference between a data analyst, a data scientist and an analytics engineer?

An analytics engineer owns the transformation layer: dbt models, tests, documentation and the metric definitions everyone else queries, and is measured on whether the data is modelled and trustworthy. A data scientist at a tech company usually owns statistical modelling and experimentation and is expected to code at a higher level; at many non-tech employers the same title is used for work an analyst does elsewhere, so read the posting rather than the title. A data analyst sits between them, closer to the business question and the stakeholder, accountable for whether the answer is right and understood. Analytics engineering typically pays above a reporting analyst seat, and the gap is mostly SQL and modelling depth, which makes it the natural next step.

How long does the data analyst interview process take, and how many stages?

At a product or tech-adjacent company: four to six stages, recruiter screen, a live SQL screen of 30 to 60 minutes, a case or short take-home, a hiring manager conversation, one or two stakeholder interviews, sometimes a presentation, over two to five weeks. At a large non-tech employer: fewer stages but a panel of three to six people including the stakeholders you would serve, often in one block, over three to eight weeks. In the public sector: a scored application in which you must mirror the posting's wording for each minimum qualification, a structured panel where every candidate is asked identical questions, then references and a background check, taking two to four months. If a process stalls at an enterprise or an agency it is usually requisition approval or a panel calendar rather than your candidacy, and the recruiter will normally tell you which if asked directly.

What do analyst interviewers actually grade in a take-home?

Almost never model accuracy or chart design. They grade whether you stated the question you answered, whether you noticed what is wrong with the data, whether you said what you excluded and why, whether you reached a recommendation, and whether you named what would change your mind. A two-page write-up with a clear answer and documented assumptions beats a long notebook with no conclusion. Expect to defend your choices live for 20 to 30 minutes afterwards: since unmonitored take-homes stopped being informative, the defence is the real test, and an analysis you did not do yourself falls apart in the first three questions.

Put this on a resume in about a minute

Paste your history once and point it at the Data Analyst posting you are looking at. No account, no card.

Build my resume free More roles