Data & Analytics

How to get hired as a data scientist in 2026-27

The short answer

To get hired as a data scientist in 2026-27, decide which of four jobs under the title you are applying for, because the generalist version of the role has split apart: product and experimentation, applied machine learning, analytics and metrics, or research. No licence or certification gates the work, and outside research scientist seats a PhD is no longer the entry condition it once was; the real gate is a live SQL screen and a case. A typical loop runs four to six stages over four to eight weeks: recruiter screen, a SQL and Python screen, a take-home or live case, an experiment design and statistics round, sometimes an ML depth round, then stakeholders, and rejections cluster in the SQL screen and the experiment round rather than in modelling. AI made fitting a model cheap, which thinned the junior modelling seat, while leaving problem framing, causal inference and the measurement of AI systems as the skills employers are actually paying for.

Licence or credential requiredNone. Data scientist is an unlicensed, uncredentialed occupation in the US and most other markets: no board exam, no registration, no mandatory certification, no protected title. Anyone can use the title. The practical gate is a live technical screen plus a case, not a piece of paper. A PhD remains close to mandatory for research scientist and some applied scientist seats at large labs, and is genuinely irrelevant to most product data science roles.
How long it takes to become hireableA working estimate, not a rule. From a quantitative background (statistics, economics, physics, engineering, operations research, a numerate master's), roughly 6 to 12 months of deliberate work to be interview-ready. From zero, 18 to 30 months is more honest, and the realistic route runs through a data analyst or analytics engineer seat first. The slow parts are not the libraries. They are SQL you can write while someone watches, experiment design you can reason about out loud, and one piece of work where a real decision changed.
Typical hiring loopTech and tech-adjacent: application through Greenhouse, Ashby, Lever or Workday; recruiter screen of 20 to 30 minutes; a 45 to 60 minute technical screen that is usually SQL plus light Python; a take-home or a live case; an onsite of three to five rounds covering experiment design and statistics, a modelling or analytics case, sometimes coding, and a stakeholder round; then the hiring manager. Four to eight weeks end to end. Outside tech (hospital systems, insurers, banks, utilities, pharma, government, universities) the loop is usually shorter: a screen, one panel, sometimes a presentation, and often an internal candidate already in the running.
Who screens youFirst an ATS keyword match or a recruiter filtering on tool names and years. Then a data science manager, a principal or staff data scientist, or in smaller companies the head of analytics or an engineering lead. In regulated industries (banking, insurance, pharma, healthcare) a model risk or validation function may also review you, and they ask different questions: documentation, assumptions, monitoring, who signs off.
Pay: where to check instead of a quoted bandDo not trust an averaged figure from a content site. Start with the US BLS Occupational Employment and Wage Statistics entry for SOC 15-2051 Data Scientists, which gives national, state and metropolitan percentiles and is the only free source with a published methodology. Then read live postings in jurisdictions that require a pay range in the ad (Colorado, California, Washington, New York and Illinois among them; check which apply to the posting's location, because the rules change) because those are commitments for a specific level in a specific place. For public tech companies, levels.fyi maps titles to internal levels and reports base, bonus and equity, self-reported. For filed rather than self-reported data, the US Department of Labor Office of Foreign Labor Certification publishes LCA disclosure records with employer, job title, worksite and offered wage.
The actual technical barSQL through window functions and join grain. Python with pandas or Polars, scikit-learn, statsmodels, and a gradient boosting library. Statistics you can apply rather than recite: sampling, confidence intervals, power and minimum detectable effect, multiple comparisons, regression and its assumptions. Causal inference: randomised experiments first, then difference-in-differences, propensity scores, instrumental variables, regression discontinuity, and knowing when none of them will rescue a bad design. Enough engineering to put work somewhere it runs again: git, a notebook you can turn into a script, dbt or Airflow at least by sight.
What AI changed, honestlyLess at the core of the job than the hype suggests, and a lot around it. Model fitting got cheap and the junior modelling seat thinned. Problem framing, metric choice, causal reasoning and standing behind a number in front of an executive did not move at all. The genuinely new work is measuring AI systems themselves: evaluation sets, labelling and agreement, retrieval quality, and experiments on features whose output differs on every call. That work is statistics, which is why it landed on data scientists rather than on engineers.
Related titles that are the same job, or notProduct Data Scientist, Decision Scientist, Quantitative Analyst (in product rather than in finance) and Experimentation Scientist are usually the same job as product data science. Machine Learning Engineer and MLOps Engineer are a different job: production systems, not analysis. Applied Scientist and Research Scientist are research roles with a publication and PhD expectation. Analytics Engineer is a modelling and pipeline role, not a statistics role. Read the responsibilities and the tools, never the title.

Four jobs share this title. Pick one, then evidence it

The single most expensive mistake in a data science job search is applying to the title rather than to the job. Between roughly 2015 and 2021, "data scientist" meant a generalist who did everything from pulling the data to deploying the model. That version has mostly split apart. Data engineers took the pipelines, analytics engineers took the transformations and the semantic layer, machine learning engineers took production serving, and AI engineers took the LLM application stack. What is left under the data scientist title is four distinct jobs with four different interview loops.

You cannot prepare for all four at once, and a resume that tries to cover all four reads as a resume with no centre. The posting's verbs are the tell. "Design and analyse experiments" is a different job from "build and deploy models", which is a different job again from "build dashboards and define metrics". Pick the one you can defend for forty minutes, not the one with the most postings. Claiming the wrong one is how people end up failing in round three of a loop they should not have entered.

Claiming a specialisation is not a line on a resume. It is a set of artefacts an interviewer can inspect in ninety seconds, because the first depth question will test the seam between the claim and the evidence.

If you are claiming product and experimentation, the evidence is an experiment you designed rather than one you read about. That means a stated hypothesis, a primary metric chosen before the test and defended, a power calculation with an assumed baseline and a minimum detectable effect, a randomisation unit, a guardrail metric, a stopping rule, and a decision that followed, including the ones where the result was flat and the feature shipped anyway or did not. If you have never worked somewhere with an experimentation platform, you can still build this: run a test on something you control, or re-analyse a public dataset as a quasi-experiment and be explicit about what the design cannot support.

If you are claiming applied machine learning, the evidence is a model with an honest evaluation and a deployment story, however small. What baseline did you beat, measured how, on what split, and what was the business metric rather than the model metric. Say what it cost to be wrong in each direction. Say whether it shipped and, if it did not, why not. "Did not ship, because the operations team could only act on 40 cases a week and the lift was concentrated below that cut" reads better than a high AUC with no context.

If you are claiming analytics and metrics, the evidence is a metric definition that settled an actual disagreement, a dashboard somebody still opens, and a documented semantic layer. Name the tool, the dispute and the decision. If you are claiming research, the evidence is publications, preprints and code, and the dominant signals are your paper record and your advisor's network. Most readers of this guide are switching rather than starting fresh, and the switch that works is lateral and incremental: analyst to product data scientist by owning experiments where you already work; software engineer to ML-leaning data scientist by taking the model-adjacent work on your own team; academic to industry by reframing the research as a measurement and decision problem and cutting the methodology to one line. The switch that rarely works is a cold application into a new specialisation with a portfolio project as the only evidence.

What is actually true about the market in 2026-27

Honest framing first, because this role attracts more market mythology than any other. Data science did not die, and it did not keep growing at the rate of the late 2010s. The title stopped being a catch-all, demand redistributed toward specific competencies, and the easy entry door closed.

The entry-level squeeze is real and worth naming plainly. Postings that once said "0 to 2 years" now frequently say "3 to 5", and the number of bootcamp and master's graduates chasing the remaining junior seats is very high. Employers responded by hiring analysts and promoting internally rather than hiring juniors straight into data science. That is not a reason to give up. It is a reason to target the analyst or analytics engineer door and to stop applying blind to seats that were never going to open.

The second shift is where the jobs are. The concentration of data science work in consumer tech is lower than people assume. Hospital systems and payers, insurers, banks and credit unions, pharma and medical devices, utilities and energy, telecoms, logistics and freight, agriculture, defence contractors, state and federal agencies, large retailers and universities all run data science teams, hire year round, see a fraction of the applicant volume, and are frequently more willing to take someone with domain knowledge and decent technique than someone with perfect technique and no domain. The tradeoff is pay, pace and tooling maturity, in that order.

The third shift is the one everyone talks about and most people describe wrongly. AI assistants and AutoML made the modelling craft cheaper, which removed much of the work that used to justify a junior data scientist's first year. They did not remove the judgment: deciding what to measure, knowing which of four revenue tables is the audited one, telling the difference between a result and an artefact, and being the person an executive believes. At the same time a new category of work appeared, the measurement of AI systems, which is squarely statistical and squarely in this job. On balance the role moved rather than shrank: fewer seats for someone whose skill is fitting a model, more for someone whose skill is knowing whether a number is true.

A fourth thing to plan around: contract-to-hire and fixed-term roles are a larger share of this market than they were, particularly in healthcare, government and consultancies. Many people get their first data science title this way. It is a legitimate route, but take the conversion question into the interview. Ask what share of contractors converted to permanent staff last year, and treat a vague answer as an answer.

How hiring works, stage by stage

The loop below is the tech and tech-adjacent default. Non-tech employers compress it heavily, which matters because candidates over-prepare for a five-round onsite and then get a single panel with a hiring manager and a department head who mainly want to know whether you can talk to clinicians.

Stage zero is the application pile, and it is brutal at the top of the funnel. Postings at well-known employers routinely show hundreds of applicants on LinkedIn's own counter within days of going live. Your realistic paths past it are a referral, a very recent posting, a niche or non-tech employer, an internal transfer, or a recruiter who already has your profile. Apply early: many recruiters review in order of receipt and stop once they have a shortlist, so a posting that is three weeks old is often already in onsites. Sending two hundred applications with the same resume is the dominant failure strategy in this role.

The recruiter screen filters three things: are you real, are you in the band, are you in the right specialisation. Have a 90 second answer to "tell me about yourself" that names your specialisation, one domain and one result. Have a pay answer that gives a range anchored on the sources below rather than a number you invented. Ask what the loop is, in detail, and what the first project would be. Both answers are diagnostic.

The technical screen is usually SQL, sometimes with light Python, and it is where the most candidates die. It gets its own section below. If you are told to expect "a coding round", ask specifically whether it is SQL, pandas or algorithmic. Those are three different preparations and recruiters will tell you.

The case or take-home is the part companies vary most. A good take-home is three to five hours with a messy dataset and an open question. A bad one is twenty hours of unpaid work, and it is reasonable to ask the expected time and to decline when the answer is absurd. Whatever the format, the grading is almost never on the model. It is on whether you framed the question, checked the data, stated assumptions, chose a metric for a reason, noticed something the brief did not mention, and wrote a conclusion a non-technical person could act on. Put the recommendation on the first slide or in the first paragraph, every time.

The onsite is three to five rounds. Expect an experiment design and statistics round, a case or analytics round, a behavioural or stakeholder round, and depending on specialisation a modelling deep dive or a coding round. Senior candidates get a round on scoping and influence instead of a coding round, and increasingly a round on how they would evaluate an AI feature.

The stakeholder round is not a formality, and it is where senior candidates get rejected. The test is whether you can explain a technical result to someone who will act on it, admit uncertainty without being useless, and hold a position when a product manager pushes back. Prepare two stories: one where your analysis changed a decision, and one where someone overruled you and what you did next.

The SQL and coding screen: what is actually being tested

Most candidates who fail a data science loop fail in the first technical screen, and almost nobody fails on exotic syntax. They fail because they write a query that runs, returns a believable number, and is wrong, and they do not notice. The interviewer is grading the noticing.

What a competent interviewer watches: do you ask about the grain of each table before joining, do you check whether a join multiplied rows, do you handle NULLs deliberately, do you reach for a window function rather than a correlated subquery, do you narrate your reasoning, and do you sanity-check before announcing a result. A bug you catch in front of the interviewer is evidence of how you work. A wrong number delivered confidently is the most common way a loop ends.

Ask which dialect you are in before you start, then stay in it. Postgres, T-SQL, Snowflake and BigQuery differ on date functions, string handling and quoting. Using the wrong one unknowingly reads as inexperience. Saying "I usually write Postgres, is this BigQuery?" reads as the opposite.

The Python half, where there is one, is usually pandas or Polars data manipulation plus a short statistics question, not algorithmic puzzles. Know groupby with multiple aggregations, merges and what they do to row counts, reshaping between long and wide, time series resampling, and how to write a function with a test rather than a notebook cell that works once. If the posting is ML-leaning, expect to write a train and evaluate loop from memory, including the split, and to justify the split.

Practise against a real database, out loud. A local Postgres or DuckDB with a public dataset is enough. Typing SQL into a text editor without executing it is the main reason people feel ready and are not. DataLemur, StrataScratch and the HackerRank SQL track use roughly the right question shapes. The goal of practice is the point where you catch your own wrong answers before the interviewer does.

The experiment and causal inference round: where the job is actually decided

If you are interviewing for a product data science role, this round is the job. It is also where candidates with strong modelling backgrounds are most often rejected, because they can fit anything and cannot say what would have happened otherwise.

The round usually opens with an open design prompt: "we want to test a new onboarding flow, how would you set it up?" The strong answer has a shape. State the decision the test is meant to inform. Name a primary metric and defend it, including why the obvious one is wrong if it is. Name the guardrails you would not trade away. State the randomisation unit and why (user, session, account, geography, time slice) and what interference that unit does and does not protect you from. Do the power calculation out loud. Then state the stopping rule before you start, and say how you would check the test is healthy: sample ratio mismatch, pre-period A/A comparability, instrumentation sanity.

Do the arithmetic rather than gesturing at it, because this is the most commonly fumbled question in the loop. For a two-arm test on a conversion rate at 5% significance and 80% power, the standard approximation is about 16 x p x (1 - p) divided by the square of the absolute minimum detectable effect, per arm. With a 10% baseline and an MDE of 1 percentage point: 16 x 0.1 x 0.9 / 0.0001, which is 14,400 users per arm. Divide by daily eligible traffic, then round up to whole weeks so the test covers complete weekly cycles. Justify the MDE by what change would actually be worth shipping and maintaining, not by what is reachable in a week.

The follow-up questions are the real exam, and they are predictable. What do you do when the result is flat? (Answer with the confidence interval and whether it excludes the effect you powered for, not with "no significant difference".) What if the team wants to peek daily? (Either accept the inflated error rate or use a method designed for it, such as a sequential or always-valid approach, and know the names.) What if there are thirty metrics on the dashboard? (Multiple comparisons, a pre-registered primary, a different standard for exploratory slices.) What if users talk to each other, or supply is shared between treatment and control? (Interference. Cluster or switchback randomisation, geo tests, marketplace effects.) What if the effect decays? (Novelty and primacy, and why a two-day read is not a result.)

Then the harder half: when you cannot randomise. Pricing changes, a regulatory rollout, a feature already shipped to everyone, a campaign in one region. Know difference-in-differences and the parallel trends assumption you have to argue rather than prove. Know synthetic control and when it fits. Know propensity score matching and its central weakness, that it balances observables only. Know instrumental variables and why a plausible instrument is rare. Know regression discontinuity and the kind of cutoff that creates one. Know that the correct answer to some questions is "this design cannot support a causal claim, here is what it can support".

The failure modes are consistent. Reciting p < 0.05 without being able to size a test. Choosing a primary metric that is easy to move rather than the one that matters. Forgetting that a significant 0.2% lift on a metric nobody acts on is not a result. Treating a flat test as a failed test. And the big one: being unable to say what decision changes depending on the outcome, which means the experiment should not have been run.

One practical note. If you have never worked somewhere with an experimentation platform, learn the vocabulary anyway, because it is the shared language now: Statsig, Eppo, Optimizely, GrowthBook, LaunchDarkly and a long tail of in-house systems. Know what CUPED does and why variance reduction is worth real money, which is that it shortens tests. Know that a sample ratio mismatch is a symptom of broken assignment or logging, and that the usual response is to debug and discard rather than to analyse anyway. Knowing the terms signals you have been near a real experimentation practice even if your own tests were small.

The modelling round, and the metric question that ends interviews

For ML-leaning roles there is a deep dive on something you built, and it almost never goes where candidates prepare for. Interviewers rarely ask you to derive backpropagation. They ask how you knew it worked.

Open with the decision the model served, not the architecture. "We needed to decide which accounts the retention team called first, and they could call 300 a week" gives the interviewer the problem shape, the cost asymmetry and the constraint in one sentence. Then the baseline: what was happening before, and what a trivial rule achieved. A model that beats no stated baseline is not evidence of anything.

Then evaluation, where most of the grading happens. Why that metric. If you say AUC, expect "what does that tell the retention team", and have an answer about precision in the top slice of the ranked list, or expected value per call, or calibration if the score feeds a downstream threshold. Say how you split the data, and if there is any time dimension, say you split by time and why a random split would have leaked. Say how you checked for leakage explicitly, because leakage is the most common reason an impressive offline number means nothing: a feature computed after the outcome, a target encoded in an ID, a join that pulled in future state.

Then the part candidates skip: what happened afterwards. Did it ship. What did the online metric do against the offline estimate, and if there was a gap, what caused it. How was it monitored, what drifted, what retraining cadence, who got paged. If it never shipped, say so and say why. "We stopped because the operational cost of acting on it exceeded the lift" is a better answer than any AUC.

Technique questions that recur: class imbalance, and why resampling often helps less than choosing a threshold against a cost matrix; calibration, and when a well-ranked model is still unusable because the probabilities are wrong; regularisation and why; feature importance and its limits, including why SHAP explains the model rather than the world; cross-validation designs for grouped or time-ordered data; and the tradeoff between a gradient boosting model that works and a deep model that impresses. At most companies the honest answer is that gradient boosting on tabular data is still the right default, and saying so plainly reads as experience rather than as a lack of ambition.

Forecasting deserves its own mention, because it is one of the steadiest sources of data science work outside tech and is badly under-prepared for. Know the difference between a backtest with an expanding window and a random split. Know why you compare against a seasonal naive baseline. Know the common trap of forecasting a metric that is itself a decision output, so the forecast changes the thing it forecasts. Know what a prediction interval means to someone who has to hold inventory.

Pay: where to look, and what actually moves it

Quoted salary averages for this title are close to useless, because they blend four different jobs across industries whose pay differs enormously for identical work. Use sources with a stated method instead.

The US BLS Occupational Employment and Wage Statistics series publishes wage percentiles for SOC 15-2051 Data Scientists, nationally, by state and by metropolitan area. It surveys employers rather than collecting self-reports, it lags by roughly a year, and it excludes equity and most non-production bonuses, which means it understates large tech and is honest about the rest of the economy. Read the 10th, 25th, 50th, 75th and 90th percentiles for your metro rather than the mean. The BLS Occupational Outlook Handbook entry for the same code carries the current employment projection, which is a better thing to quote in an interview than a number from a listicle.

Live postings in pay-transparency jurisdictions are the most accurate free signal you can get, because they are what an employer committed to in writing for a specific level in a specific location. Several US states and cities require a range in the posting, Colorado, California, Washington, New York and Illinois among them, and the rules change, so check what applies to the role's location rather than assuming. Collect twenty real ranges for your target level and city and you have a better picture than any aggregator will give you.

For public technology companies, levels.fyi maps titles to internal levels and reports base, bonus and equity. It is self-reported and skews senior and well-paid, so read the level definitions and compare like with like rather than reading the headline total. One source almost nobody uses: the US Department of Labor Office of Foreign Labor Certification publishes LCA disclosure data listing employer, job title, worksite and offered wage for visa-sponsored roles, filed under penalty and searchable through several free front ends. For large employers it is the closest thing to audited pay data that exists publicly.

What actually moves the number, in rough order of effect. Industry first: quantitative finance, large technology and some health tech pay far above hospital systems, government, non-profits and universities for the same work. Level second, and level is negotiable in a way that salary within a level often is not, so push on the level before you push on the number. Equity third, and treat private-company equity as a lottery ticket with a strike price rather than as compensation; ask for the strike price, the preferred price, the total shares outstanding and the date of the last round before you value it at anything. Geography fourth, and remote roles are increasingly banded by your location rather than the company's. Specialisation fifth: experimentation and ML-leaning roles tend to out-earn analytics-leaning ones at the same level.

On negotiating: the most effective lever in this role is a competing process, and the second is evidence that your level is wrong. Both require you to have asked the loop and scope questions early. Saying "based on BLS percentiles for this metro and five posted ranges at this level, I am expecting X to Y" is a stronger move than naming a number with no provenance, and it has the advantage of being true.

The resume, the portfolio, and getting the first one

A data science resume is read in two passes. The recruiter pass is fifteen seconds for title, years, tools and industry. The hiring manager pass is ninety seconds on the top third of page one, looking for one thing: does this person produce results somebody acted on.

Write bullets as decision, action, measured outcome, in that order, with real units. "Designed and ran 14 checkout experiments; the two that shipped raised completion on a seven-figure monthly session base, worth roughly a quarter of the team's annual target" is a bullet. "Leveraged advanced analytics to drive business insights" is not a bullet, it is a smell. If confidentiality stops you sharing a number, give the shape: "cut manual review volume by roughly half" is fine, and far better than nothing. One page up to around eight years of experience, two beyond. A skills line, not a skills grid: an interviewer who sees forty tools concludes you have used none of them properly. List what you would be happy to be examined on and nothing else.

What gets ignored, consistently: course certificates, bootcamp badges, MOOC lists, generic "proficient in machine learning", an objective statement, a photograph in US applications, and projects on Titanic, Iris, MNIST, the Boston housing dataset, or any competition described only by leaderboard position. These do not merely fail to help. They signal that you have never worked on a problem somebody chose for a reason.

A portfolio, if you need one, is one to three pieces of work, each written as a short document rather than a notebook dump, each leading with the question and the recommendation. Acquire the data yourself where you can, because acquisition shows judgment and every clean public dataset has already been modelled a thousand times. Show the data checking. Show one thing you got wrong and fixed. Put the README first and make it readable by a hiring manager who will never open your code. A repo whose README opens with "I asked whether late pickups cause cancellations, and here is the answer" gets read. One that opens with "this project uses XGBoost and SHAP" does not. If you are coming from academia, cut the methodology: translate the paper into decision language, name the measurement problem, say what the result would change, and keep the publication list to one line unless you are applying for a research seat.

On ATS: the system is a keyword matcher, not an intelligence. Use the posting's exact nouns where they are true of you ("A/B testing", "causal inference", "dbt", "PySpark", "Snowflake"), use a single-column layout, use standard section headings, and submit a PDF unless told otherwise. Do not hide keywords in white text. It is detected and it ends the application.

If you do not already have a data science title, the realistic route in 2026-27 is almost never a cold application to a data scientist posting. It is a side door, and the side doors are well known. Internal transfer is the highest-probability path by a wide margin: if your current employer has a data team, find the analysis nobody is doing, do it well, and be known for it before any req opens. The analyst door is second, and a data analyst or analytics engineer seat two years from now puts you in a far stronger position than two more years of applying with no title. Contract-to-hire through staffing firms is third and larger than people expect, especially in health systems, insurers, government agencies and large enterprises. Structured programmes are fourth: graduate and rotational schemes, government fellowships, university research staff posts, and the data science arms of consultancies, all of which hire on a calendar, so the deadline matters more than your readiness.

On where to look, company career pages beat aggregators, which are stale and duplicated. Beyond the usual boards: hospital and health system portals, state and federal job sites, university HR pages, insurer and bank portals, utility and energy companies, national lab and defence contractor listings, and the careers pages of the specific vendors in your target domain. Each sees a fraction of the applicant volume of a tech posting for the same work. Finally, the thing that actually produces interviews is people who have seen your work. Write up one analysis properly and put it somewhere public. Answer a question well where practitioners are. Present at a local meetup. Message a data scientist at a target company with a specific question about something they built rather than a request for a referral. The conversion rate on cold applications here is low enough that two hours a week spent on being known beats twenty more applications.

Working with AI in this role

What a data scientist has to know about AI in 2026-27

The honest version first, because this role is surrounded by claims in both directions. AI did not eliminate data science and it did not leave it untouched. It removed a specific layer of the work and created a different one, and knowing which is which is itself an interview signal.

What got cheap: writing a model. An assistant plus a mature library produces a competent gradient boosting pipeline, a reasonable feature set and a plausible evaluation in minutes. AutoML had been doing a version of this for years and the assistants finished the job. The consequence is not that modelling stopped mattering. It is that modelling stopped being scarce, so it stopped being what you are paid for. The scarce things sit upstream and downstream: deciding what to measure, knowing which table is the audited one, telling a result from an artefact, and being the person an executive believes when the number is unwelcome. Those did not change at all, which is why a candidate who claims AI transformed the whole job reads as someone who has not done it.

What got created, and this is the part worth building toward: AI systems have to be measured, and measurement is statistics. Every company shipping an LLM feature hits the same wall, which is that nobody can say whether it is working. The answer is an evaluation set, a labelling process, a metric with defined ground truth, inter-rater agreement, a holdout, and an experiment design that survives an output that differs on every call. That is a data scientist's skill set applied to a new object, and it is one of the few areas of this field where demand clearly grew rather than redistributed. If you want one high-leverage thing to learn this year, it is this.

A caution about where the hype outruns the reality. Agentic systems that connect to a warehouse and run analyses end to end exist and are improving, and they are nowhere near trustworthy with an unsupervised answer on a real enterprise schema. Benchmarks built on realistic warehouses rather than tutorial databases, such as Spider 2.0 and BIRD for text-to-SQL, score far below what product demos imply, and the gap is mostly context: undocumented columns, four overlapping order tables, business rules that live in one person's head. Scores move, so read the current leaderboard before you form a view. Then, when an interviewer asks what you think about AI replacing analysis, you have a specific checkable answer rather than a posture.

Two professional cautions that come up in interviews and cost people offers. Do not claim fluency you cannot demonstrate: "I use Copilot" is not an answer, and interviewers now follow up with "what did it get wrong for you". And know your employer's rules about what data leaves the building. Pasting customer records, PHI, or anything under a confidentiality term into a consumer chatbot is a fireable mistake in a regulated industry, however convenient. In banking and insurance, models that inform decisions fall under model risk management expectations, with the US interagency guidance commonly cited as SR 11-7 as the reference point. In the EU, the AI Act places obligations on high-risk systems covering training and validation data, documentation and human oversight. The timing of those obligations has been amended since the Act was passed, so name the obligation rather than a date, and say the current text should be checked. An interviewer will respect that far more than a confident wrong deadline.

Evaluating LLM and AI systems as a measurement problem

This is the clearest new demand in the role. Teams ship a generative feature and then discover they have no way to say whether it is good, whether a prompt change improved it, or whether it regressed last Tuesday. Answering that is sampling, ground truth, agreement, confidence intervals and experiment design, which is why it landed on data scientists rather than on engineers. Postings name it directly: look for 'evals', 'LLM evaluation', 'model quality', 'groundedness', 'human-in-the-loop labelling'.

Show it: Describe a concrete evaluation you built or could build: a golden set of real inputs with verified correct outputs, a rubric a second person can apply to the same output and agree with, measured inter-rater agreement before you trust any judge, an LLM-as-judge validated against human labels rather than assumed, a holdout that never gets tuned on, and a regression suite re-run on every prompt or model change. Say what you measured and what failed. Name tooling if you have used it (LangSmith, Braintrust, Arize Phoenix, Weights & Biases Weave, Ragas, DeepEval) but lead with the method, because the method is the transferable part.

Running experiments on non-deterministic features

Classical A/B testing assumes a fixed treatment. A generative feature produces a different output on every call, the quality metric is often a human judgment rather than a click, and the effect you care about may be on trust or task completion rather than on conversion. Teams that apply a standard experiment template to these features get noisy results and then argue about them for a quarter. A data scientist who can design the right test here is useful on day one.

Show it: Talk through a design: pairwise preference or side-by-side comparison with randomised position, graded rubric scoring with multiple raters, interleaving for ranking changes, and the choice between measuring output quality offline and measuring downstream behaviour online. Say how you would handle variance from the model itself (fixing seeds or temperature where possible, increasing sample, blocking on input difficulty), and name the guardrails you would not trade for engagement: latency, cost per request, refusal rate, and an accuracy or safety floor.

Retrieval and grounding metrics

Most enterprise AI features are retrieval-augmented, and most of their failures are retrieval failures rather than generation failures. If the right document never came back, no prompt will fix the answer. Teams routinely spend weeks tuning prompts for a problem that lives one layer down, and the person who can separate the two saves the project.

Show it: Say how you would measure retrieval separately from generation: recall at k and nDCG against a labelled set of query-document pairs, then a groundedness or attribution check on whether the answer is actually supported by what was retrieved, plus a citation-correctness rate. Then say which you would fix first and why. Having a view on chunking, and on why the obvious chunk size is usually wrong for the document type, reads as hands-on.

Using LLMs as a labelling and feature-extraction tool, with measured agreement

Support tickets, survey free text, call transcripts, clinical notes, inspection comments and reviews were previously too expensive to analyse at scale and are now cheap to classify. This is genuinely new analytical territory, and often the fastest way to produce something nobody at your employer has: the reasons behind a churn number rather than the number. It is also the fastest way to produce a confident, wrong, unverified dataset, which is why the measurement discipline is the whole skill.

Show it: Show you measured rather than trusted: hand-label a sample as ground truth, report agreement between the model's labels and yours, keep a holdout, state the cost per thousand rows, and say what you did about the categories it systematically confused. Then connect it to a decision: the top three ticket drivers, and which one got fixed.

Knowing when not to build a model, and saying so

A large share of senior data science value is talking a team out of building something. The request arrives as 'can we build a model to predict X' and the right answer is often a rule, a report, a better process, or the observation that nobody can act on the prediction anyway. With model building now cheap, the volume of unnecessary models has gone up rather than down, and the judgment to refuse has become more valuable rather than less.

Show it: Have one story where you replaced a proposed model with something simpler and it was the right call, and one where you declined a project because no decision depended on the output. Be precise about the reasoning: the action threshold, the cost of being wrong, the operational capacity to act, the data that would have been needed and was not available. This is the clearest senior signal available in a behavioural round.

Documentation, governance and model risk in regulated industries

In banking, insurance, healthcare, pharma and government, a model that informs a decision carries documentation and oversight requirements, and AI features have pulled a lot of new work inside that perimeter. Validation functions ask about assumptions, data lineage, performance monitoring, bias testing and who signs off. Candidates from consumer tech routinely fail these panels, not on technique but on never having had to defend a model to someone whose job is to doubt it.

Show it: Name the artefacts: a model documentation pack with intended use, data lineage, assumptions and limitations; a monitoring plan with thresholds and owners; fairness or disparate impact testing where it applies; a sign-off record. Reference the frameworks by name rather than by date: the US interagency model risk management guidance (SR 11-7) in banking, the NIST AI Risk Management Framework as a voluntary structure, and EU AI Act obligations for high-risk systems covering data governance, documentation and human oversight. Say that the timing of the EU obligations should be checked against the current text, because it has moved.

Working with an assistant on your own analysis, with a stated line

Data scientists who get real speed from these tools work differently from those who do not. The pattern that works is specification first, small steps, verification at each step. The pattern that fails is asking for the whole analysis in one prompt and accepting the output. Interviewers can tell the difference within two questions, and an unverified generated query that produced a wrong number is a story many hiring managers have lived through.

Show it: Describe the workflow concretely: where you let it draft (boilerplate transformations, plotting code, regex, docstrings, a first-pass data dictionary, scaffolding a backtest) and where you refuse (choosing the metric, deciding which table is authoritative, interpreting a causal result, anything reaching an executive unverified). Then give one specific example of a generated answer that was plausible and wrong, how you caught it, and what the real number turned out to be. The catch is the credential, not the usage.

What a screen is looking for

These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.

Mistakes that cost people this job

Applying to the title rather than to one of the four jobs under it, with a resume that tries to cover product experimentation, ML, analytics and research at once.

Pick the specialisation you can defend for forty minutes, state it in the first line of the resume, and target postings whose verbs match. Three deep claims beat eleven shallow ones.

Leading a project description with the model: "I built an XGBoost classifier with SHAP explanations and achieved a high AUC."

Lead with the decision and the constraint, then the baseline, then the metric, then what happened in production. "The retention team could call 300 accounts a week and was picking them by tenure; ranking by predicted save value beat that baseline and the team kept using it."

Writing a query in the technical screen that runs, returns a believable number, and is wrong, then announcing it confidently.

Narrate the sanity check before you run it, check whether the join multiplied rows, and reconcile the total against something already trusted. A bug you catch in front of the interviewer is evidence. A wrong number delivered confidently ends the loop.

Treating a flat experiment result as a non-result, or reporting it as "no significant difference".

Report the point estimate with its confidence interval and say whether the interval excludes the effect you powered for. "We can rule out a lift above 0.4%, which is below the threshold that would justify the maintenance cost" is a decision. "Not significant" is not.

Describing a test design fluently and then being unable to size it when asked.

Practise the arithmetic out loud: per arm, roughly 16 x p x (1 - p) divided by the squared absolute MDE, at 5% significance and 80% power. State the baseline, justify the MDE by what would be worth shipping, then convert to whole weeks of traffic. This is the most commonly fumbled question in the loop.

A portfolio built on Titanic, Iris, MNIST, the Boston housing dataset, or a competition described by leaderboard position.

One or two projects on data you acquired yourself, framed as a question somebody would act on, written as a short document that leads with the recommendation. The acquisition and the framing are the signal. The modelling is not.

Claiming AI fluency as a line item: "experienced with LLMs and Copilot".

Give a specific story of a generated output that was plausible and wrong, how you caught it, and what the correct answer turned out to be. Then say where you let an assistant draft and where you refuse. Interviewers now ask exactly this follow-up.

Preparing only for a five-round tech onsite, then interviewing at a hospital system or an insurer and treating the domain questions as small talk.

For non-tech employers, prepare the domain: the data you would work with, the vocabulary, the regulatory constraint, and who acts on your output. Domain fluency outranks technique in those panels, and the loop is usually one panel, so there is no second chance.

Taking a take-home as a modelling exercise and submitting a notebook with twelve models and no conclusion.

Spend the time on framing, data checking, assumptions and the recommendation, and put the recommendation first. Use the simplest model that answers the question and say why you did not use a more complex one.

Letting a recruiter set the level before you have asked what the loop is and what the first project would be.

Ask the loop and scope questions in the screen, then anchor the level conversation on scope rather than on a number. Level is where the money is. Within-level negotiation moves far less.

Using a quasi-experimental method without stating the assumption it rests on, or refusing to use one at all.

Name the assumption every time: parallel trends for difference-in-differences, balance on observables only for propensity matching, relevance and exclusion for an instrument. And be willing to say a design cannot support a causal claim, which is stronger than a confident wrong answer.

Treating a PhD as the entry ticket, or treating its absence as a disqualification.

Check which job you are applying for. Research scientist seats effectively require one. Most product data science roles do not, and a PhD with no evidence of shipped decisions is a weaker candidate there than a three-year analyst with experiments to show.

Questions people ask

Do I need a PhD to get a data scientist job in 2026?

No, except for research scientist and some applied scientist seats at large labs, where a PhD and a publication record are effectively hard requirements. For product data science, applied machine learning and analytics-leaning data science roles, a PhD is neither required nor a strong advantage on its own. A master's in statistics, economics, operations research, computer science or another quantitative field is common but not mandatory. What is consistently required is demonstrated work: SQL you can write while someone watches, experiment design you can reason about out loud, and at least one piece of analysis that changed a decision.

Is data science still a good career in 2026 and 2027, or is it dying?

Data science is not dying, but it is no longer the uncapped growth field it was between 2015 and 2021. The generalist version of the role split apart: data engineers took the pipelines, analytics engineers took the transformations, machine learning engineers took production serving, and AI engineers took the LLM application stack. What remains under the data scientist title is more specialised and much more competitive at entry level, with postings that once said 0 to 2 years now frequently saying 3 to 5. Demand is strongest in experimentation and causal inference, in AI system evaluation, and across non-tech industries such as healthcare, insurance, banking, utilities, logistics and government, which hire year round and see a fraction of the applicant volume. For the official employment projection, read the BLS Occupational Outlook Handbook entry for SOC 15-2051 rather than a number quoted in a blog post.

What is the difference between a data scientist and a data analyst in 2026?

A data analyst answers what happened and what is happening, usually with SQL, a BI tool and a close stakeholder relationship. A data scientist is expected to answer what would happen if, which means causal inference and experiment design, or to build a model that makes a repeated decision at scale. The boundary is blurry and some employers use the titles interchangeably, which is why you read the responsibilities rather than the title. Pay typically differs by a meaningful margin at the same level, and the analyst-to-data-scientist ladder is the most common route into the role for people without a quantitative graduate degree.

What does the data scientist interview loop actually test in 2026-27?

The data scientist loop tests five things, in roughly this order of weight. First, SQL correctness under observation: join grain, NULLs, window functions, and whether you sanity-check your own output. Second, experiment design and causal reasoning: power calculations, metric choice, randomisation unit, threats such as sample ratio mismatch and interference, and what to do with a flat result. Third, problem framing in a case: can you turn a vague business question into a measurable one and state a recommendation. Fourth, for ML-leaning roles, a modelling deep dive focused on evaluation, leakage and what happened in production rather than on algorithms. Fifth, stakeholder communication: explaining a result to someone who will act on it and holding a position under pushback. Rejections cluster in the first two.

How has AI changed the data scientist job, honestly?

It made model building cheap and it made measuring AI systems a new source of work, while leaving the core of the job unchanged. Assistants and AutoML produce a competent modelling pipeline quickly, which removed much of the work that justified a junior data scientist's first year and raised the bar at entry level. It did not touch the judgment: choosing what to measure, knowing which table is authoritative, telling a result from an artefact, and defending a number to an executive. The genuinely new demand is evaluation of LLM and AI features, which is a statistics problem (ground truth sets, inter-rater agreement, holdouts, experiment design on non-deterministic outputs) and is the highest-leverage thing a data scientist can learn right now. Claims that AI replaced data analysis describe the version of the job where the data scientist was a human query interface, which was always the most replaceable version.

How much do data scientists earn, and where can I check a real number?

Do not use an averaged figure from a content site, because it blends four different jobs across industries whose pay differs enormously. Use the US BLS Occupational Employment and Wage Statistics entry for SOC 15-2051 Data Scientists for national, state and metro percentiles from an employer survey, noting that it excludes equity and most non-production bonuses. Then read live postings in jurisdictions that require a pay range in the ad, such as Colorado, California, Washington, New York and Illinois, checking which rules apply to the posting's location, because those ranges are commitments for specific levels. For public tech companies, levels.fyi gives level-banded totals including equity, self-reported. For filed rather than self-reported data, the US Department of Labor Office of Foreign Labor Certification publishes LCA disclosure records with employer, title, worksite and offered wage.

Which data science specialisation should I claim?

Claim the one you can evidence with artefacts an interviewer can inspect, not the one with the most postings. Product or decision data science (experiments, metrics, product decisions) is the largest category and is gated on causal inference. Applied machine learning (ranking, fraud, churn, pricing, forecasting) is gated on evaluation and deployment. Analytics or metrics data science is often a senior analyst job with a better title and a lower ceiling. Research scientist requires a PhD and publications. Pick one, write the resume for it, and keep one openable artefact per claim: an experiment writeup, a model evaluation document, a metric definition that settled a dispute, or a paper.

Do I need a certification to become a data scientist?

No. Data scientist is an unlicensed, uncredentialed occupation with no board exam, no registration and no protected title, and hiring managers consistently skip past course certificates and bootcamp badges on a resume. Certifications have marginal value in two situations only: a cloud platform certificate when the posting explicitly names that platform, and an employer whose HR screen uses them as a filter, which is more common in government and large enterprises. Time spent on a certificate is almost always better spent producing one piece of analysis that someone acted on.

How do I get a data science job with no experience?

Through a side door, because cold applications to data scientist postings convert poorly for people without the title. The highest-probability routes, in order: internal transfer at an employer that already has a data team, where the hiring manager has seen your work; a data analyst or analytics engineer seat first, then the internal ladder; contract-to-hire through a staffing firm, which is a large share of the market in healthcare, insurance and government; a structured graduate, rotational or fellowship programme, which hires on a calendar so the deadline matters more than your readiness; and only then open applications, targeted at non-tech employers and sent while the posting is still new. In parallel, make one piece of work public and specific enough that somebody can judge it.

What should a data scientist portfolio contain, and is Kaggle worth doing?

A data scientist portfolio should be one to three pieces of work, each written as a short document rather than a notebook dump, each leading with the question and the recommendation rather than the method. Acquire the data yourself where you can, because acquisition demonstrates judgment and every clean public dataset has already been modelled a thousand times. Show the data checking, show one thing you got wrong and fixed, and state what decision the result would change. Make the README readable by a hiring manager who will never open your code. Kaggle is useful as practice and weak as a credential, because it removes the two hardest parts of the real job, deciding what question to ask and deciding whether the data can answer it, and the part it does keep is the part that has become cheapest. Remove anything built on Titanic, Iris, MNIST or the Boston housing dataset, and remove competition results described only by leaderboard position.

Put this on a resume in about a minute

Paste your history once and point it at the Data Scientist posting you are looking at. No account, no card.

Build my resume free More roles