| Licence required | None for industry bioinformatics in the US or the EU. The UK is different: "Clinical Scientist" is a protected title and NHS clinical bioinformatics posts require HCPC registration. Elsewhere, regulation sits on the diagnostic laboratory and its director, not on you personally. |
|---|---|
| Typical degree | MS or PhD. A PhD is the standard gate for the "Scientist" job family at pharma and large biotech. BS or MS is normal for Bioinformatics Engineer, Bioinformatics Analyst, and pipeline or platform roles. |
| Time to qualify | MS route: about 2 years of coursework plus 1 to 3 years of applied experience. PhD route: 4 to 6 years, with a postdoc optional rather than expected for industry. UK NHS clinical route: a 3 year Scientist Training Programme, or the equivalence route if you already hold the competences. |
| Credentialed clinical routes | US diagnostic lab directorship: ABMGG board certification in Laboratory Genetics and Genomics, via an ACGME-accredited fellowship after a doctoral degree, roughly 2 years. UK: NHS Scientist Training Programme in Clinical Bioinformatics, then HCPC registration. Confirm current eligibility rules with ABMGG, the National School of Healthcare Science and the HCPC directly, because they are revised. |
| Pay data source | US: BLS OES 19-1029 (Biological Scientists, All Other) and O*NET 19-1029.01 (Bioinformatics Scientists), cross-checked against OES 15-2051 (Data Scientists) and 19-1021 (Biochemists and Biophysicists), because the work is split across those codes. For live numbers, read posted ranges in pay-transparency jurisdictions: California, Colorado, Washington, New York and Illinois. UK NHS: the published Agenda for Change bands. |
| Hiring stages and timeline | Typically 4 to 6 stages: recruiter screen, hiring manager deep dive on one project, a take-home or live technical exercise, a seminar or chalk talk for PhD-level Scientist titles, a cross-functional panel including a wet-lab collaborator, then references. One to three weeks at a small startup, four to eight at a mid-size biotech, eight to sixteen at large pharma. |
| The one artefact that moves the needle | A public repository with a containerised workflow, a CI job that runs a subsampled test dataset on every push, a README stating the biological question first, and a rendered report. Hiring managers open it, and most candidates do not have one. |
| Where the jobs are | Biotech and pharma R&D, clinical diagnostics and reference labs, CROs and CDMOs, agricultural and industrial biotech, academic and hospital core facilities, instrument and reagent vendors (Illumina, Oxford Nanopore Technologies, 10x Genomics, PacBio), and bioinformatics software and cloud platform companies. |
What the job actually is, and the four job families hiding behind one title
"Bioinformatics Scientist" is not one job. It is at least four, and the biggest single cause of wasted applications is pursuing the wrong one. Read the requisition's responsibilities section before the requirements section. The responsibilities tell you which family it belongs to, and the family tells you whether your background clears the screen.
The first family is the embedded analyst. You sit inside a research programme, a disease area, or a platform team, and you are the person who turns a sequencing run into a decision. The work is experimental design conversations, QC, differential analysis, figures that go into a go/no-go deck, and arguing with a bench scientist about whether a batch effect is real. Titles: Bioinformatics Scientist, Computational Biologist, Scientist I or II. This family is the most likely to require a PhD, because the value is scientific judgement rather than throughput.
The second family is the pipeline and platform builder. You own the workflows, the compute, the reference data, the containers, and the thing that lets twenty other people run an analysis without asking you. Titles: Bioinformatics Engineer, Computational Biology Software Engineer, Scientific Software Engineer, Genomics Platform Engineer. This family is much less credential-gated. Real engineering practice and a repository that proves it will usually beat a PhD with unrunnable code, and senior pay is competitive with the Scientist track.
The third family is the methods developer. You build new models, new statistics, new algorithms, and you publish. Titles: Senior Scientist, Principal Scientist, Machine Learning Scientist, Research Scientist. This is PhD-gated in practice, usually with a publication record in the specific method area, and it is the smallest of the four by headcount.
The fourth family is clinical bioinformatics, inside a CLIA-certified and CAP-accredited laboratory in the US, or an accredited genomics laboratory hub in the UK. The work is variant calling and interpretation support, pipeline validation, and the documentation that proves the pipeline does what you claim it does. It is more procedural than research bioinformatics, it is more stable through funding cycles, and it is the only part of the field with a genuine credentialed ladder. Candidates who assume industry research is the only destination never look at it, which is part of why it is easier to enter.
Do you need a PhD, and where credentials actually gate you
For a large share of the jobs, no. For the job titled "Scientist" at a mid-size or large biotech or pharma company, usually yes, and not because the work needs it. Job architecture at those companies defines levels partly by degree. A Scientist I requisition will often read "PhD in bioinformatics, computational biology, genetics or related field, or MS plus 4+ years of experience". That second clause is real and is routinely honoured. People talk themselves out of applying because they stop reading at the first half of the sentence.
Where the PhD genuinely matters is where you will be the scientific owner of a question nobody has answered: target discovery, novel method development, anything whose deliverable is a hypothesis rather than a number. The PhD is a proxy for having survived a multi-year project that did not work for the first three years. Hiring managers are testing for that, and they will accept other evidence of it, but you have to supply the evidence rather than expecting them to infer it from a job title.
If you have a wet-lab PhD and have been teaching yourself computation, you are strong for family one and weak for family two. Lead with the biology and prove the code runs. If you have a CS or physics degree and no biology, you are strong for family two and will struggle in family one until you can speak fluently about library prep, UMIs, strandedness, duplicate reads, and why a sample failed. Neither gap is fatal. Pretending the gap does not exist is.
A postdoc is no longer the default industry entry ticket in computational biology. Companies hire directly out of PhD programmes into Scientist I, and a postdoc taken purely as a credential, with no new skill or publication attached, can read as drift. Take one because a specific lab will teach you something you cannot get elsewhere.
The credential picture changes completely in clinical work, and this is the part most career advice gets wrong. In the US, nobody licenses a bioinformatician, but directing a clinical genomics laboratory generally requires board certification from the American Board of Medical Genetics and Genomics, obtained through an ACGME-accredited Laboratory Genetics and Genomics fellowship that you can only enter with a doctoral degree. A bioinformatician inside such a lab works under that director and needs no personal certification.
In the UK, "Clinical Scientist" is a title protected by law and registered by the Health and Care Professions Council. The main entry route is the NHS Scientist Training Programme, a three year salaried training post in Clinical Bioinformatics (genomics, physical sciences or health informatics specialisms) that leads to HCPC registration. There is also an equivalence route through the Academy for Healthcare Science for people who already hold the competences from elsewhere. Numbers of STP posts are small and recruitment runs once a year on a fixed national cycle, so if that is your route, find the opening date before you plan anything else. Check the current requirements with the National School of Healthcare Science and the HCPC rather than relying on an article, including this one.
How hiring actually works, stage by stage
Who screens you first depends on company size. At a startup under about fifty people, the hiring manager reads every application personally, often within days, and a good repository link can get you a call in under a week. At a large pharma company a recruiter or an automated screen runs first against the requisition text, and the hiring manager sees a shortlist. At an academic or hospital core facility, HR screens against a rigid published qualification list and a missing stated requirement is an automatic rejection however good you are, so apply only where you can tick every literal line.
Stage one, recruiter screen, 20 to 30 minutes. Logistics, work authorisation, salary expectations, a surface check that you have handled the data types in the requisition. Have a number ready and anchor it on the posted band rather than on your current pay. Several US states and cities restrict employers from asking salary history at all, so you are usually within your rights to decline that specific question and give a range instead.
Stage two, hiring manager call, 45 to 60 minutes. This is the real first filter. They will pick one project from your resume and go deep: what was the question, what did you do, what did you decide, what would you do differently. If you cannot defend one project in that kind of detail, nothing later in the loop rescues it.
Stage three, technical exercise. Two common shapes. A take-home: here is a public or subsampled dataset, do an analysis, send the code and a short write-up, usually with a stated budget of four to eight hours. Or a live session: screen-shared coding on a data-wrangling problem plus reasoning questions about a messy experimental design. Take-homes are more common here than in general software hiring because the job is literally to produce a defensible analysis document. Ask what the time budget is, and treat an open-ended exercise with no budget and no feedback as a signal about how the team works.
Stage four, the seminar or chalk talk. Standard for PhD-level Scientist roles, common in biotech, near-universal in academia. Thirty to forty-five minutes on your own work to an audience that includes people who do not compute. It is scored on whether a bench scientist followed you, not on whether the methods slide was impressive.
Stage five, the panel. Four to six one-on-ones including at least one wet-lab collaborator and often a program lead or a data engineer. The wet-lab interview is the one candidates underprepare and the one that most often sinks them. Those people are deciding whether working with you will be bearable, and whether you will tell them their experiment is underpowered before they spend the money rather than after.
Stage six, references and offer. Reference checking in this field is frequently informal and lateral: the hiring manager knows someone who knows you. Assume anything you have said publicly about a previous employer is discoverable.
Two variants worth knowing. Clinical diagnostic labs add a stage on regulatory and quality awareness, because an unvalidated change to a production pipeline is a reportable event rather than a hotfix. CROs and service providers add a stage on client communication, because you will be on calls with the customer's scientists and the deliverable is their confidence as much as their data.
- Startup: days to three weeks. The hiring manager reads everything, and a repository link is the strongest single asset.
- Mid-size biotech: four to eight weeks, structured loop, seminar for Scientist titles.
- Large pharma: eight to sixteen weeks, recruiter and automated screen first, keyword alignment with the requisition matters more here than anywhere else.
- Academic or hospital core facility: slow, HR screens against a literal qualification list, seminar required, pay is public and lower.
- Clinical diagnostic lab: adds a quality and regulatory conversation, and values documentation discipline over novelty.
The portfolio: what to build, in order
This is the part of the field where a public portfolio genuinely substitutes for pedigree, and most candidates build the wrong thing. A folder of notebooks with no README, no environment specification, no random seed and no stated question is worse than nothing, because it demonstrates exactly the habit the hiring manager is screening out. Build fewer things, finished.
Project one: a workflow that runs. Pick one real analysis (bulk RNA-seq differential expression, a small variant-calling workflow, an ATAC-seq peak pipeline) and implement it in Nextflow, Snakemake or WDL. Containerise every step with Docker or Apptainer. Pin versions. Include a test profile that runs on a subsampled dataset in under ten minutes on a laptop, and wire it to CI so a passing badge on the README proves it still runs today. That badge does more work than any bullet point on your resume. For a shortcut to good practice, read how nf-core structures its pipelines and follow it. Contributing one reviewed nf-core module is a checkable public credential, and the review itself teaches you the conventions.
Project two: an honest reanalysis. Take a public dataset from GEO, SRA, ENA, recount3, the GDC, GTEx, ENCODE or the Human Cell Atlas, and answer one specific question. The README opens with the question, the data (how many samples, what library chemistry, what depth, what genome build), and the answer in two sentences. Then the QC decisions and why you made them, then the analysis, then the limitations. If the result is null, say so. A clean null result with a correct power argument is a stronger hiring signal than a p-value fished out of a confounded design, and experienced reviewers see the difference immediately.
Project three: a benchmark. Compare two tools or two parameter choices on data where the truth is known: Genome in a Bottle benchmark sets for small variants, simulated reads where you control the ground truth, or a spike-in. Report precision, recall and F1 stratified by region difficulty, report the compute cost, and state a conclusion with its caveats. This proves you can evaluate rather than only execute, which is the single hardest thing to assess from a resume and the thing that separates a scientist from a tool operator.
Project four, optional and high value: a contribution to something people use. A merged pull request to samtools, bcftools, scanpy, Seurat, a Bioconductor package, an nf-core module. A Bioconductor package of your own that passed review is a particularly strong externally validated signal, because the review is rigorous and public and you did not control it.
What not to build: a dashboard nobody asked for, a general-purpose "bioinformatics toolkit" with three functions, a reproduced Kaggle tutorial, or a deep learning model trained on a dataset far too small to support it. These read as portfolio-filling, and reviewers have seen all of them many times.
- Every repository: README with the question first, environment pinned, one command to reproduce, a licence, and a short statement of what you would do with more data or more time.
- Use real file formats and get them right: FASTQ, BAM, CRAM, VCF, GTF, BED. Coordinate-system errors in a public repository get noticed.
- Name the genome build everywhere (GRCh38, T2T-CHM13, GRCm39) and the annotation release. Omitting it is a tell.
- If the data is controlled-access (dbGaP, UK Biobank, All of Us), do not put it in a public repository. Put the code there and say where the data lives and under what agreement.
The resume: what is read and what is skipped
A bioinformatics resume gets about forty seconds from a computational hiring manager on the first pass. They are looking for four things: the data types you have handled, the scale, what you decided, and whether you can ship. Everything else is noise to them, though some of it still has to be present for the keyword screen at larger companies.
Skipped almost entirely: a twenty-item skills block; "proficient in Python and R"; coursework; a generic objective statement; a publication list with no indication of what you did; tool logos; GPA after your first job.
Read closely: bullets shaped as "analysed X samples of Y data to answer Z, which led to decision W". Concrete scale (96 RNA-seq libraries, 40,000 cells, 1,200 exomes, 30x WGS on 300 samples). A named compute environment (Slurm, AWS Batch, AWS HealthOmics, GCP, Terra, DNAnexus, Seqera Platform). A named workflow language. And a GitHub link that is not empty.
For publications, state your contribution in a parenthetical. "Co-first author; built the single-cell pipeline and performed all differential abundance analysis" tells a hiring manager far more than the journal name. If you are sixth author on a consortium paper, say what you did or compress it into a list.
For the applicant tracking screen at larger companies, mirror the requisition's own nouns. If the posting says "single-cell RNA-seq", do not write only "scRNA-seq". Write both, once each. This is mechanical, it costs nothing, and it is the difference between a human seeing your application and not.
If you are transitioning from academia, cut the CV. Four pages with teaching, conferences and posters is correct for a faculty application and wrong for industry. Two pages maximum, projects reframed as deliverables, methods reframed as decisions, and the grant you were funded by removed unless the funder is the hiring company's partner.
What the interview really tests
Almost nobody is testing whether you can implement a sorting algorithm. The technical bar in bioinformatics interviews is reasoning under messy data, not algorithmic puzzles, and the failures cluster in four places.
Experimental design and confounding. "All the treated samples were run in batch two." What can you conclude, what can you not conclude, what would you ask for. This is the most common question and the most common failure. The right answer usually starts with "nothing cleanly, and here is why", not with a tool name.
Statistics as applied to this data. Multiple testing, and why FDR rather than Bonferroni for a 20,000-gene screen. What a p-value histogram should look like, and what a spike near 1 or a hump in the middle is telling you. Pseudoreplication in single-cell data: cells from one animal are not independent replicates, and the usual correct move is pseudobulk aggregation to the sample or donor level before differential testing. Normalisation choices and their assumptions. Power, calculated before the experiment rather than after.
File and reference literacy. 0-based half-open BED versus 1-based inclusive GTF, VCF and SAM. Strandedness and what happens when you set it wrong. Which transcript set you used and why MANE Select exists. Build mismatches and the hazards of liftover. Alt contigs and multi-mapping. This is not trivia. Wrong answers here corrupt results without raising an error, which is exactly why they are asked.
Debugging and ownership. "The pipeline ran green and the results look wrong. What do you check, in order?" They want a systematic answer: inputs and sample sheet first, then QC metrics, then a known-positive control, then one sample traced end to end by hand. A candidate whose first move is "I would rerun it" is being screened out.
Then the non-technical half, which decides more offers than the technical half. Can you explain a result to a bench scientist without condescension. Will you push back on a bad design before the money is spent. Can you be told your analysis is wrong without defending it. Hiring managers in this field describe a large part of the job as translation between the bench and the analysis, and they interview for it deliberately.
Prepare one project you can defend in complete depth, including what went wrong, what you got wrong, and what you changed. Candidates who present only successes read as either junior or evasive.
- Expect live coding at the level of: group and summarise a table, join two files on identifiers that do not quite match, parse a VCF field, write a function with a test. Pandas, dplyr, awk, occasionally SQL.
- Expect a take-home that is really a writing test. The write-up is scored as heavily as the code.
- Expect at least one question about reproducibility: could you rerun this analysis a year from now, and how.
- Expect one question about a time you disagreed with a wet-lab collaborator. Have a real one ready, with the outcome.
Where the jobs are in 2026 and 2027, and how to find them
The market is not what it was at the 2021 funding peak, and pretending otherwise helps nobody. Early-stage biotech hiring tightened after that peak and has not returned to it, and layoffs have been a regular feature of the sector since. What has held up better: clinical diagnostics and reference laboratories, where sequencing volume keeps growing; CROs and service providers, who absorbed work companies stopped doing in house; large pharma computational groups, slower to hire and also slower to cut; and the platform, instrument and tooling companies that sell into the sector regardless of who is doing the science.
That makes company selection part of the job search rather than an afterthought. Before you apply to a private company, check when it last raised and what it said the money was for, and check whether its lead programme has read out. Before you apply to a public one, the cash position and runway are in the filings. A role at a company with eight months of runway is a different decision from the same role at one with three years, and nobody will tell you which you are looking at unless you ask.
Geography still concentrates: Boston and Cambridge, the San Francisco Bay Area, San Diego, Seattle, Research Triangle, and the Maryland and DC corridor in the US; Cambridge, Oxford and London in the UK; Basel, Munich, Heidelberg, Copenhagen and Barcelona in Europe. Fully remote computational roles exist and are a real fraction of postings, but they skew senior, because the job depends on knowing what the bench is actually doing. Early in your career, physical proximity to a wet lab accelerates you, and hiring managers know it.
Where to look, concretely: company career pages directly (the aggregators lag), the Biostars and Bioconductor job boards, the nf-core and scverse community channels, Nature Careers and the society job boards (ISCB, ASHG), and the NHS Jobs site plus the national STP recruitment cycle if you are in the UK. Set alerts on company pages for the ten companies whose science you can actually discuss, and check them weekly. That beats volume applying by a wide margin.
Two underrated entry points. Core facilities at universities and hospitals hire continually, pay less, and give you breadth across assay types faster than any industry role will; two years there is a strong launchpad into industry. And instrument and reagent vendors hire field bioinformatics and application scientists, which puts you in front of dozens of labs a year and builds a network no internal role can match.
Work authorisation is a real variable in this field, because a large share of the candidate pool trained on student visas. Ask about sponsorship in the recruiter screen rather than at offer stage. Some employers sponsor routinely, some never do, and some will say so in the posting. Immigration rules and costs move, so check the current position with the employer and with official guidance rather than with an article or a forum thread. In the UK, note that NHS Scientist Training Programme posts and university posts have their own sponsorship rules that differ from private sector ones.
Areas worth orienting a portfolio towards, because the analysis is unsettled and the candidate pool is thinner than for bulk RNA-seq: single-cell and spatial transcriptomics (Visium, Xenium, MERFISH, CosMx); long-read sequencing (Oxford Nanopore, PacBio HiFi) for structural variants, methylation and hard regions; proteomics, especially DIA mass spectrometry, which postings ask for far more often than candidates list it; multi-omic integration; and clinical variant interpretation at scale. Immune repertoire analysis and cell therapy manufacturing analytics are both worth a look and rarely appear in career advice.
Getting the first job with no industry experience
Coming out of a PhD, a master's programme, or a career change, the order of operations that works is: finish one portfolio project properly, rewrite the resume around deliverables, then apply narrowly rather than broadly. Two hundred generic applications produce fewer interviews than twenty where you have read the company's papers and can name the data type they work with and the question it answers.
Referrals matter more here than in most fields because the community is small. Practical routes in: answer questions on Biostars or the Bioconductor support site under your real name; contribute to an open pipeline project and interact with the maintainers in public; present a poster at a meeting where industry recruits (ASHG, AGBT, ISMB, RECOMB, and the specialist meeting for your assay type); attend an nf-core hackathon, which are online and free and put you in a room with people who have open requisitions. None of this is networking theatre. It is doing visible work next to people who hire.
Internships and co-ops remain the highest-conversion route for students, and many large pharma and biotech companies run structured computational internships that convert to offers. Apply roughly nine to twelve months ahead. These close earlier than people expect.
Switching in from software engineering: take one serious genomics or molecular biology course and build the benchmark project. Your engineering practice is already ahead of most of the field. What you need to prove is that you will not produce a fast, clean, confidently wrong answer because you did not know what a duplicate read, a strandedness setting or a batch effect does to the result.
Switching in from the wet lab: your advantage is that you know what the data actually is and most of your competition does not. The gap is engineering credibility, and it closes with one public repository containing a workflow that runs reproducibly with tests. Target embedded analyst roles where biological judgement is the product, not platform roles where you will be compared against software engineers.
One more route people miss: the clinical laboratory. Genomics laboratories hire variant analysts and clinical bioinformatics staff continuously, train on the job, and value documentation discipline over novelty. It is one of the few places where a careful person with an MS and no industry history gets hired on their first application.
What a bioinformatics scientist needs to know about AI in 2026 and 2027
Bioinformatics is one of the few fields where the AI question is not hype. Machine learning has been load-bearing here for years. Two things genuinely changed the job recently, a third changed the daily work, and the one that gets the most press has changed it least. Saying that accurately in an interview is a stronger signal than enthusiasm, because the people interviewing you have read the same overclaiming papers you have.
What genuinely changed, first: structure prediction became infrastructure. AlphaFold2 and its successors, along with open models such as ESMFold, OpenFold, Boltz and Chai, moved predicted structures from a research result to a routine input. Employers no longer ask whether you can run them. They ask whether you know when a prediction is trustworthy: what pLDDT and PAE actually tell you, why a confident monomer prediction says nothing reliable about a complex interface, why disordered regions look the way they do, and when a predicted structure is not an acceptable substitute for an experimental one in a decision that costs money. Licensing is also a real workplace question here, since some of these models are released on terms that restrict commercial use, which is part of why companies run open alternatives. Check the current licence before assuming you can use a model at work.
What genuinely changed, second: sequence models became a standard tool for variant effect. Protein language models in the ESM family and genome-scale models such as Evo, Nucleotide Transformer, and Enformer and Borzoi for regulatory prediction now produce zero-shot fitness and pathogenicity estimates that get used in real pipelines. In clinical variant interpretation, computational predictors including AlphaMissense, SpliceAI and CADD feed the computational evidence criteria of the ACMG/AMP framework, and ClinGen's sequence variant interpretation working group has published calibrated thresholds and evidence strengths for using them. If you are interviewing anywhere near clinical genomics, know that calibration is the whole point and that applying a predictor at an uncalibrated cutoff is a recognised error rather than a style choice. Check the current ClinGen recommendation text instead of quoting a threshold from memory, because these are revised.
What changed less than claimed: single-cell foundation models. scGPT, Geneformer, scFoundation, UCE and their relatives were presented as general-purpose replacements for task-specific analysis. Independent benchmarking has repeatedly found that zero-shot performance on core tasks, cell type annotation and batch integration in particular, does not reliably beat well-tuned simple baselines, and in several published comparisons loses to logistic regression on highly variable genes. The honest position, and the one that reads as competence, is that they are promising in some transfer settings, they are not a default, and the correct response to a claim in either direction is to benchmark on your own data against a simple baseline. A candidate who says that and can describe how they would run the comparison beats a candidate who recites model names.
What changed in the daily work: coding assistants. Pipeline boilerplate, translating R to Python, drafting tests, generating plotting code, parsing configuration. All substantially faster with an assistant, and the companies hiring you assume you use one. The interview question is not whether you use them but how you review the output. The failure mode specific to this field is worth memorising, because interviewers probe it: assistants produce confidently wrong genomic coordinate conversions, mix 0-based and 1-based conventions, invent function arguments for fast-moving scanpy and Bioconductor APIs, and silently pick the wrong normalisation default. None of these raise an exception. All of them change the answer. The reviewable habits are unit tests that pin known coordinates, a known-positive control sample run through every pipeline change, and never merging generated code you cannot explain line by line.
What is being automated around the role: routine secondary analysis. A standard bulk RNA-seq run from FASTQ to count matrix to a differential expression table is close to a commodity, available as a managed service and increasingly assembled by an assistant in an afternoon. If your whole value proposition is executing a standard pipeline, that value is compressing. What is not automating: deciding what an experiment can and cannot answer, spotting the confound nobody designed out, choosing which of six defensible analyses addresses the biological question, and standing in front of a program team and owning the interpretation. Value is shifting from execution toward design and judgement, and hiring has followed.
Governance, which now appears in interviews and did not a few years ago: do not paste controlled-access data into a third-party model. dbGaP, UK Biobank, All of Us and most patient data carry data use agreements your convenience does not override. Clinical laboratories must validate any software component that affects a reported result, so an LLM in that path is either a validated component or a finding at your next inspection. Know your employer's policy and be able to say exactly what you would and would not send to an external API. Regulatory treatment of software in diagnostics has been actively moving in both the US and the EU, so check the current rule rather than repeating a date from memory or from an article.
Judging when a predicted protein structure is usable
Predicted structures are routine inputs to target assessment and design decisions. Teams get burned by treating a low-confidence region or a predicted complex interface as experimental truth.
Show it: In a project write-up, report pLDDT and PAE for the regions you relied on, state which conclusions the prediction does and does not support, and name the one you would have confirmed experimentally before spending money.
Zero-shot variant effect prediction and its calibration
Protein and genome language models feed real pipelines, and in clinical settings they feed formal evidence criteria in the ACMG/AMP framework. Using a predictor at an arbitrary cutoff is a recognised error.
Show it: Say which predictor you used, which calibrated threshold and evidence strength you applied, and where the calibration came from. Add that you would check the current ClinGen recommendation rather than reciting a number.
Benchmarking a foundation model against a simple baseline
The most useful thing you can do with single-cell foundation models right now is evaluate them honestly. Employers are tired of candidates who can name models and cannot evaluate one.
Show it: Publish a small benchmark: one task, one public dataset, one foundation model, one baseline (logistic regression on highly variable genes, Harmony, or scVI), metrics and compute cost reported, and a clear conclusion including the case where the baseline won.
Reviewing AI-generated bioinformatics code
Generated code fails silently in this field: coordinate conventions, strandedness, normalisation defaults, hallucinated API arguments. None raise an exception, all change the answer.
Show it: Keep tests in your repository that pin known coordinates and a known-positive control, and be ready to describe a specific time an assistant gave you plausible wrong output and how you caught it.
Data governance for model use
Controlled-access and patient data carry agreements that prohibit sending data to third-party services. This comes up in interviews now, especially at clinical labs and pharma.
Show it: State your rule plainly: what you would send to an external API, what you would run locally or in a compliant environment, and that you would read the data use agreement before deciding.
Validating software that touches a reported clinical result
In a CLIA or CAP environment the pipeline is a regulated artefact. Adding an unvalidated component, AI or otherwise, is a quality event rather than an optimisation.
Show it: If you have clinical experience, describe a validation you ran: the reference materials, the acceptance criteria, the documentation. If you do not, say you understand the pipeline is a validated instrument and would not change it outside change control.
Separating what AI changed from what it did not
Hiring managers are specifically screening for calibration on this question, and overclaiming costs offers.
Show it: Answer the inevitable "what do you think about foundation models in biology" with one thing that genuinely works, one thing that is contested and why, and the experiment you would run to settle it on your own data.
What a screen is looking for
These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.
- bioinformatics
- computational biology
- clinical bioinformatics
- NGS
- next-generation sequencing
- RNA-seq
- single-cell RNA-seq
- scRNA-seq
- whole genome sequencing
- WGS
- whole exome sequencing
- WES
- variant calling
- variant interpretation
- GATK
- samtools
- bcftools
- BWA
- STAR
- salmon
- DESeq2
- edgeR
- limma
- Seurat
- scanpy
- Bioconductor
- pseudobulk
- Nextflow
- nf-core
- Snakemake
- WDL
- Cromwell
- Galaxy
- Docker
- Singularity
- Apptainer
- Python
- R
- Bash
- SQL
- Git
- GitHub Actions
- CI/CD
- Slurm
- HPC
- AWS
- AWS Batch
- AWS HealthOmics
- Google Cloud
- Terra
- DNAnexus
- Seqera Platform
- ATAC-seq
- ChIP-seq
- methylation
- long-read sequencing
- Oxford Nanopore
- PacBio HiFi
- spatial transcriptomics
- Visium
- Xenium
- proteomics
- DIA mass spectrometry
- multi-omics
- GRCh38
- T2T-CHM13
- MANE Select
- VCF
- BAM
- CRAM
- GTF
- BED
- FASTQ
- statistical genetics
- GWAS
- differential expression
- batch effect
- FDR
- multiple testing correction
- benchmarking
- reproducible research
- pipeline validation
- machine learning
- AlphaFold
- protein language model
- ESM
- AlphaMissense
- SpliceAI
- CADD
- ACMG/AMP
- ClinGen
- CLIA
- CAP
- ISO 15189
- HCPC
- Scientist Training Programme
- GEO
- SRA
- TCGA
- GTEx
- dbGaP
- UK Biobank
- PhD computational biology
- MS bioinformatics
Mistakes that cost people this job
Treating the resume as a tool inventory: a twenty-item skills block listing every aligner you have ever installed.
Write bullets containing a question, a scale and a decision. "Analysed 96 RNA-seq libraries across 4 timepoints to test whether compound X engaged the pathway; showed the effect was a batch artefact, which stopped a planned 300-sample follow-up." That one bullet does more than the whole skills block.
Publishing a portfolio of notebooks nobody else can run: no README, no environment file, no seed, absolute paths to your laptop.
One repository, finished. README opening with the question and the answer, pinned environment, a container, a test dataset, one command to reproduce, and CI proving it still works. Reviewers open the README and the CI badge first and often nothing else.
Applying only to jobs titled "Scientist" because you have a PhD, or avoiding them because you do not.
Apply to the job family that matches what you are good at. If your strength is engineering, the Bioinformatics Engineer requisition pays comparably at senior levels and is open now. If a Scientist requisition says "or MS plus 4+ years" in the second half of the sentence, you are eligible. Apply.
Presenting only successes: every project in the portfolio and every story in the interview worked first time.
Bring one project where you were wrong, found out, and corrected it. Report a null result honestly in a public repository. Hiring managers weight this heavily because it predicts whether you will raise a problem or bury it.
Treating cells as independent replicates in single-cell differential expression, then presenting the resulting tiny p-values.
Aggregate to pseudobulk at the sample or donor level before testing across conditions, and be able to say why. This exact error is used as a trap question in interviews, and reviewers spot it instantly in portfolio projects.
Ignoring the genome build, the annotation release and the transcript set, or lifting over coordinates without checking what failed to map.
State GRCh38 or T2T-CHM13, the GENCODE or Ensembl release, and MANE Select where relevant, in every project and every methods conversation. Silent build mismatches are a leading cause of wrong answers and interviewers know it.
Preparing hard for the coding round and treating the wet-lab collaborator interview as a formality.
Prepare a two-minute jargon-free explanation of your main project aimed at a bench scientist, and one real story about disagreeing with a collaborator over an experimental design. That interview sinks more candidates than the technical round.
Leading with foundation models and model names to signal that you are current.
Lead with the question and the design. Mention models where you benchmarked them against a baseline and can report what happened. Overclaiming on single-cell foundation models in particular reads as not having read the benchmarking literature.
Submitting an academic CV: four pages, teaching, every poster, publications formatted for a tenure committee.
Two pages, industry format, projects as deliverables, and your specific contribution named in a parenthetical after each publication that matters.
Searching only for "Bioinformatics Scientist" on job boards, and only on aggregators.
Search Computational Biologist, Bioinformatics Engineer, Bioinformatics Analyst, Genomic Data Scientist, Scientific Software Engineer, Clinical Genomics Scientist and NGS Data Analyst as well, and set alerts on the career pages of ten companies whose science you can discuss. Aggregators lag the company pages.
Questions people ask
Do I need a PhD to get a bioinformatics job?
No, not for a large share of the field. A PhD is the standard gate for the "Scientist" job family at pharma and large biotech, because those companies define job levels partly by degree, and for methods development roles where you will publish. It is not required for Bioinformatics Engineer, Bioinformatics Analyst, pipeline and platform roles, clinical bioinformatics analyst roles, or most core facility positions, which commonly take a BS or MS plus demonstrated experience. Many Scientist requisitions also carry an explicit "or MS plus 4+ years" clause that is genuinely honoured.
Is a bioinformatics licence or certification required?
In the United States, no. Bioinformatics is not a licensed profession there and no certificate is required to be hired, though directing a clinical genomics laboratory generally requires board certification from the American Board of Medical Genetics and Genomics via an ACGME-accredited Laboratory Genetics and Genomics fellowship taken after a doctoral degree. The United Kingdom is different: "Clinical Scientist" is a title protected by law, and NHS clinical bioinformatics posts require registration with the Health and Care Professions Council, usually reached through the three year Scientist Training Programme or an equivalence route. Confirm current requirements with ABMGG, the National School of Healthcare Science and the HCPC directly, because they are revised.
What should a bioinformatics portfolio actually contain?
A Bioinformatics Scientist portfolio needs three things, finished rather than numerous. One containerised workflow in Nextflow, Snakemake or WDL with a test profile that runs on a laptop in under ten minutes and a CI job proving it still works. One honest reanalysis of a public dataset from GEO, SRA, GTEx, the GDC or ENCODE, where the README states the biological question first and reports limitations and null results. One benchmark comparing two tools or parameter choices against a known truth set such as Genome in a Bottle, reporting precision, recall and compute cost with a stated conclusion. A merged contribution to nf-core, Bioconductor, scanpy or samtools is a strong fourth.
What is the salary for a bioinformatics scientist?
Do not trust a single quoted band, because the spread between a core facility analyst and a principal scientist at a large pharma company is several-fold. In the United States, use Bureau of Labor Statistics OES data for SOC 19-1029 (Biological Scientists, All Other) and O*NET 19-1029.01 (Bioinformatics Scientists), cross-checked against OES 15-2051 (Data Scientists) and 19-1021 (Biochemists and Biophysicists), since the work is split across those codes. For current numbers, read posted ranges in pay-transparency jurisdictions: California, Colorado, Washington, New York and Illinois all require ranges in postings, and those postings are the most accurate live signal available. In the UK, NHS roles are paid on the published Agenda for Change bands.
What do bioinformatics interviews test?
Bioinformatics Scientist interviews test four technical areas and one non-technical one. Experimental design and confounding ("all the treated samples were in batch two, what can you conclude"). Applied statistics (FDR versus Bonferroni, reading a p-value histogram, why cells are not independent replicates in single-cell differential expression). File and reference literacy (0-based BED versus 1-based GTF and VCF, strandedness, genome builds, liftover hazards, MANE Select). Systematic debugging of a pipeline that ran green and produced wrong results. And, decisively, whether you can explain a result to a bench scientist and push back on a bad design before money is spent.
How long does the hiring process take?
One to three weeks at a small startup where the hiring manager reads applications directly, four to eight weeks at a mid-size biotech, and eight to sixteen weeks at large pharma. Academic and hospital core facilities are slower still, because HR screens against a fixed published qualification list before anyone technical sees your application. A typical loop is a recruiter screen, a hiring manager deep dive on one of your projects, a take-home or live technical exercise, a seminar if the title is Scientist, and a cross-functional panel including at least one wet-lab collaborator.
Can I move into bioinformatics from a wet-lab PhD?
Yes, and it is one of the most common routes. Your advantage is that you already know what the data is and how the library was made, which most of your competition does not. The gap is engineering credibility, and it closes with one public repository containing a workflow that runs reproducibly with tests and a README that opens with the biological question. Target embedded analyst roles such as Bioinformatics Scientist and Computational Biologist, where biological judgement is the product, rather than platform roles where you will be compared against software engineers.
Can I move into bioinformatics from software engineering?
Yes. Software engineers move into Bioinformatics Scientist work regularly, and the platform and pipeline roles are the right target. Your engineering practice will be ahead of most of the field. The gap is domain knowledge, and the failure mode is producing a clean, fast, confidently wrong answer because you did not know what a duplicate read, a strandedness setting or a batch effect does to the result. Take one serious genomics or molecular biology course, build a benchmarking project against a known truth set, and be able to describe a sequencing library preparation end to end.
Are AI foundation models replacing bioinformatics jobs?
Not the core of a Bioinformatics Scientist's job. What has compressed is routine secondary analysis: a standard bulk RNA-seq run from FASTQ to a differential expression table is close to commodity work now, available as a managed service. What has not compressed is deciding what an experiment can answer, spotting a confound, choosing among several defensible analyses, and owning the interpretation in front of a program team. Structure prediction and sequence-based variant effect prediction are genuine additions to the toolkit. Single-cell foundation models have under-delivered against their claims, and independent benchmarks often find well-tuned simple baselines competitive or better.
Which job titles should I search besides "Bioinformatics Scientist"?
Computational Biologist, Bioinformatics Engineer, Bioinformatics Analyst, Genomic Data Scientist, Scientific Software Engineer, Research Software Engineer, Clinical Genomics Scientist, Variant Scientist, NGS Data Analyst, Computational Genomics Scientist, and Machine Learning Scientist when the posting sits in drug discovery. The same work is advertised under all of these, and candidates who search one title see a small fraction of the market. Search company career pages directly as well, because aggregators lag them by days to weeks.
Put this on a resume in about a minute
Paste your history once and point it at the Bioinformatics Scientist posting you are looking at. No account, no card.
Build my resume free More roles