Cybersecurity

How to get hired as an AI security engineer in 2026-27

The short answer

An AI security engineer is hired on demonstrated judgement, not on a credential: no licence or certification gates the role in 2026-27, and the loop turns on two rounds. The first is a threat model of an agent you have never seen, where you mark the boundary at which untrusted text enters the model's context, bound what the agent's tools are allowed to do rather than promising to stop the injection, and accept the residual risk out loud with a detection attached. The second is hands-on: either a code review of an LLM application with planted flaws (retrieval that ignores tenancy, model output passed to a shell, a guardrail that is really a regular expression) or a live attempt to make a sandboxed agent leak a secret, written up with reproduction steps and a severity argument. Most candidates arrive with half the job, so the fastest route in is to fix your weaker half (mechanics if you came from security, authorization if you came from machine learning) and then move internally into the AI programme at the employer you already have, because that is how most of these roles are filled.

What the role is in 2026-27Making it safe for an organisation to ship and operate systems that take instructions from text it does not control. The work splits across threat review of AI features before they ship, review of the code and configuration that wires models to tools and data, the permission and identity model for agents, tenancy in retrieval pipelines, the supply chain for model weights and third-party connectors, adversarial testing with a measured attack suite, and detection on the tool-call layer. The deliverable is a bounded blast radius, not a list of jailbreak prompts.
Read the posting for this before anything elseWhich way the preposition runs. Security for AI means protecting models, agents and pipelines, and that is this job. AI for security means using models to accelerate detection, triage or code review, which is a security operations or detection engineering job wearing a fashionable title. Plenty of postings headed AI Security Engineer are the second thing and the title will not tell you which. The responsibilities section tells you within four lines, and applying to the wrong one wastes a screen.
Posting titles that mean this jobAI Security Engineer, AI/ML Security Engineer, GenAI Security Engineer, Machine Learning Security Engineer, Model Security Engineer, AI Red Team Engineer, Offensive Security Engineer (AI), AI Safeguards or Safety Systems Engineer, Security Engineer on an AI platform or AI infrastructure team, Product Security Engineer where the product is an AI product, MLSecOps Engineer, Security Research Engineer (AI). Searching only the exact phrase hides most of the market, and at startups the job is usually posted as Security Engineer with AI features buried in the responsibilities.
Credential gateThere is no licence to practise and no mandatory certification for an AI security engineer anywhere this job is hired. Two real gates exist. Government clearance (a US clearance, UK SC or DV, or the local equivalent) is hard-gated for defence, government and national security work and sets the timeline in months rather than weeks. And a frontier lab or a regulated employer may run enhanced background screening. Several issuers including ISACA, ISC2, SANS and GIAC, and CompTIA have launched AI-security credentials recently: most postings do not ask for any of them by name, the market moves faster than any syllabus, and you should check the current name, status and price on the issuer's own page before paying. Spend the same money on a GPU bill and publish what you break.
The two rounds that decide itAn agentic threat model delivered out loud (a system described in three minutes, then roughly 40 minutes to mark trust boundaries, rank threats by reachability and loss, propose controls that are properties of the design, and accept a residual risk with a compensating detection), and a hands-on round that is either a secure code review of an LLM application or a live red-team exercise against a sandboxed agent with a written finding at the end. Classic security depth and basic model mechanics are both examined somewhere in the loop, deliberately.
What gets you screened in from outsideOne public artefact that shows judgement rather than enthusiasm. In rough order of weight: a disclosed and fixed vulnerability in an AI product or an AI framework, with the reproduction and the patch; a written threat model of a real documented AI system, published; an attack suite with measured results (attack success rate before and after a mitigation, plus the false refusal cost); a tool other practitioners use; a verified bounty record on an AI vendor's programme; a placement in a public jailbreak or AI red team event. A certificate list and a prompt-engineering portfolio do not screen you in.
Level, and realistic time to hired by starting pointThis is a mid-level and senior title in practice: most requisitions ask for several years of security, software or platform engineering before the AI qualifier, and genuinely junior openings are rare. From an application security, cloud security or product security job: three to nine months, mostly spent on model and agent mechanics, and often resolved by an internal move. From machine learning or platform engineering: six to twelve months, spent on authorization, identity, web and cloud fundamentals, because that is where the loop rejects you. From security operations or GRC: twelve to twenty-four months, and the code fluency gap is the part that cannot be shortcut. From no technical background: this is not an entry-level role, and the honest route is a software or security engineering job first.
Pay, and where to check it rather than trust a bandThere is no occupation code for this title. The nearest authoritative US baselines are BLS OES 15-1212 Information Security Analysts, and 15-1221 Computer and Information Research Scientists for research-flavoured lab roles; both understate what a named AI security role pays. For a live number, read pay-transparency postings in Colorado, California, New York, Washington and Illinois, where the employer must publish a range, and compare the same employer's ranges for a senior security engineer and a senior machine learning engineer: AI security sits at or above the higher of the two, and at frontier labs and large technology companies a large share of the package is equity. Anchor on a source you can open, not on a band from a forum.

What an AI security engineer actually does, and which of the five jobs the posting means

The title is only a few years old, so the job behind it is not standardised and two postings with the same header can share almost no content. Before you prepare for a loop, work out which of five jobs you are applying to, because each weights the interview differently and each rewards a different artefact. Read the responsibilities, never the title.

The common core across all five is one sentence worth memorising, because it is also the answer to the opening interview question. An AI security engineer secures systems that take instructions from text nobody controls, act on them with real credentials, and cannot be made to behave deterministically. Every interesting problem in the role falls out of that: the model cannot reliably distinguish data from instruction, so authority has to live outside the model; the system is non-deterministic, so a test that passes once proves nothing and you need a measured attack suite instead; the model's output is attacker-influenced, so whatever renders or executes it is part of the attack surface.

The weekly shape, in the product-side version of the job, is less exotic than the title suggests. You read design documents for features that are two weeks from launch and write back the three things that must change. You review the pull request that adds a tool to an agent and find that the tool takes the account id from the model's text rather than from the authenticated session. You argue with a product manager about whether a refund agent needs a human click above fifty dollars. You maintain the attack suite that runs in continuous integration and explain to a team why their change moved attack success rate in the wrong direction, and by how much. You write the risk acceptance for the residual injection exposure, because it will not go to zero and someone senior has to sign that sentence. Perhaps one day a fortnight looks anything like research.

Be clear about what the job is not, because the mismatch burns screens. It is not trust and safety, which owns content policy, abuse and user harm with a largely non-engineering toolkit. It is not AI governance or AI assurance, which produces model cards, inventories and audit evidence and is hired against a GRC rubric. It is not machine learning research, and a candidate who talks mostly about adversarial examples on image classifiers reads as a decade out of date on what employers are paying to fix. And it is not prompt engineering: writing a clever jailbreak is the cheapest skill in this market and nearly every applicant has it.

The four routes in, and the half of the job each one has to fix

Almost everyone arrives with half of this job and loses the loop on the other half. Naming your half honestly, in week one, is worth more than any reading list, because the fix is different in kind and in duration.

Route A, by far the most common, is a security engineer moving across: application security, cloud security, product security, penetration testing, detection engineering. You already have the instincts the role is built on, which is the single most important thing: trust boundaries, authorization thinking, blast radius, the habit of asking who can reach this. What you are missing is mechanics, and the gap shows up in specific, catchable ways. You cannot say what a context window is in tokens, how retrieval actually selects documents, what fine-tuning changes that prompting does not, why temperature breaks your regression test, or what a guardrail model is and what it costs in latency and false refusals. Interviewers probe this not to test trivia but because without it your threat model will be confidently wrong: you will propose input sanitisation for prompt injection, which is the clearest single signal that someone has read the headlines and not the systems.

Route B is a machine learning or platform engineer moving into security. You know the mechanics cold, and you will lose the loop on fundamentals that feel beneath the title. You will be asked how you would design authorization for a tool that queries customer orders, and the answer has to include where the tenant identifier comes from and what verifies it. You will be asked what is wrong with a token, a redirect or a session, and a vague answer ends that round. The hardest part is cultural rather than technical: security engineering grades you on what an adversary can do with your design, not on whether your design works, and candidates from machine learning tend to narrate the happy path. Three to six months of deliberate web and cloud security work fixes the knowledge. The adversarial reflex takes longer and is best built by actually breaking things and writing them up.

Route C is a penetration tester or red teamer specialising into AI. The hands-on round will be your strongest and the constructive half your weakest. The prompt that catches route C is simple: fine, you got the agent to exfiltrate the key, now design the fix, and do not say sanitise the input. If your answer is a filter rather than a change in where authority lives, you will be read as a tester rather than an engineer, and the offer comes in at tester level.

Route D is GRC, audit or privacy moving in. Be honest with yourself: this is the longest road, and the shortest version of it is not a certificate. Target variant four (enterprise AI adoption security) rather than variant one, where your existing fluency in data classification, vendor review and regulatory mapping is genuinely scarce, and fix code fluency in parallel because the loop will still open a file and ask you what is wrong with it. Twelve to twenty-four months is realistic, and an internal move inside the employer you already have is much more likely than an external hire.

One structural fact beats all of this advice. Most of these roles are filled internally or from one step away, because the hiring manager is trying to find someone who can be trusted to say no to an engineering director in month two. If your current employer is standing up an AI programme, and most employers of size now are, you are already closer to the job than an external applicant with a better portfolio. Volunteer to be the person who reviews the first agent. Do it well, write it down, and you have both the experience and the reference.

How hiring works, stage by stage, and how it differs by employer

The loop is a security loop with two AI-specific rounds inserted, not a new kind of process. Knowing what each stage screens for tells you what to rehearse, and the two rounds people fail are predictable.

Two structural features are worth going in knowing. First, many employers have moved the exercise from a take-home to a live shared screen, for the obvious reason that a take-home about AI security is trivially completed by an AI. Expect a smaller exercise, watched, with harder questioning. Second, you will be asked how you use model assistance in your own work, and this is not a trap: the answer that lands is a working practice with a verification step attached, named tools, and one thing you no longer let it do. A flat refusal reads as unserious in this role specifically, and unqualified enthusiasm reads as someone who has not been burned yet.

The loop tests both halves of the role on purpose, and that is the design, not an accident. There will be a round that assumes you know what a model is and a round that assumes you know what authorization is. Candidates prepare one. The hiring bar is approximately: can this person sit between an AI team and a security team and be believed by both.

Employer type changes the loop more than seniority does. A frontier lab or a well-funded AI company runs the most rigorous version, often with a written exercise, a strong emphasis on measurement, and a values or mission round that is real rather than ceremonial. A large technology company runs five to seven rounds over a month, including a general security engineering round and often a classic coding round. An AI startup makes an offer in a week off two conversations and a short exercise, and the job will be AI security plus cloud security plus the compliance questionnaire plus whatever breaks. A bank, insurer or healthcare system adds background screening and a panel that cares about regulatory mapping, model inventories and third-party risk. A consultancy weights the writing sample heaviest, because the deliverable is a document a client pays for. Defence and government work is gated on clearance before any of this matters.

The agentic threat model round, worked end to end

This is the round that decides the job, so rehearse it until it is boring. The prompt is some version of: we are shipping an agent that reads inbound customer support tickets, can look up orders, search our knowledge base, and issue a refund, and it drafts replies that a human support agent sends. Threat model it. You have roughly 40 minutes and the grader is watching your process, not collecting a list.

Spend the first five minutes asking questions, and ask about money, data and reversibility rather than about technology. Who can create a ticket, and does that include anyone on the internet. What identity do the tools run as: the agent's own service account, or the support agent who is in the loop. Is the refund synchronous and irreversible, and is there a cap. Does order lookup take an order id or a customer identifier, and where does that value come from. Is the knowledge base single-tenant. Does anything render the model's draft as HTML or markdown, and can it contain a link or an image. Where do prompts and completions get logged, who can read that store, and for how long. What is the worst single loss you can imagine from this feature. Those questions alone put you ahead of most candidates, because most start drawing immediately.

Then draw the flow and mark exactly one thing first: the boundary where untrusted text enters the model's context. In this system it is the ticket body, and that is the whole game. A customer writes the ticket, so an attacker writes part of the prompt. Say out loud that from the model's point of view there is no reliable distinction between the instruction you wrote and the instruction the attacker pasted, and that you are therefore going to treat the model as a confused deputy and design the controls around it rather than inside it. The framing that has stuck in the field, usually credited to Simon Willison, is the lethal trifecta: access to private data, exposure to untrusted content, and a way to communicate outward. This system has all three legs, and your job is to remove or bound at least one of them on every path that matters.

Now walk the threats and rank them by who can reach them and what is lost, not by cleverness. The ranked top of the list should be: an instruction in a ticket body causes a refund to an attacker-chosen destination (reachable by anyone, direct financial loss); the order lookup tool accepts an identifier supplied by the model, so an injected ticket reads another customer's order history (reachable by anyone, cross-tenant data loss, likely reportable); exfiltration outward, where injected text makes the draft contain a link or image pointing at an attacker host with another customer's data in the query string, and the support console fetches it or the human clicks it; the prompt and completion log becomes an unclassified copy of customer personal data with wider access and longer retention than the ticket system it came from; cost and availability abuse through long inputs and loops; and the knowledge base as a poisoning target if anyone outside a small group can edit it. Say which of these you cannot assess from what you were told.

Controls are where mid separates from senior, and the rule is that a control must be a property of the system rather than something a future engineer has to remember. Take the authority out of the model: the refund tool receives the order id and the destination from the authenticated ticket record, server-side, and ignores anything the model says about either, so the model's only influence is the decision to call it. Bound the loss: a hard server-side cap per refund and per day, and a human click required above a threshold, enforced by the API and not by the prompt. Scope every tool call to the tenant from the ticket record at query time, not by filtering results afterwards. Close the outward leg: deny-by-default egress from the tool runtime, no fetching of model-supplied URLs, and render drafts as plain text with links stripped or explicitly reviewed. Classify the prompt log at the same level as the ticket store, with the same retention and the same access controls, and scrub it. Cap tokens, steps and spend per session. Then give the detection: alert on a refund issued without a matching human confirmation event, on tool-call sequences the design does not expect, and on retrieval that returned a document outside the ticket's tenant. Finish with the eval: a fixed suite of injection attempts for this feature running in continuous integration with attack success rate tracked per release, plus a rotating held-out set so the suite does not become the thing you optimise against.

Close by accepting a risk out loud, because that is the sentence that reads as senior. Something like: prompt injection on the ticket path cannot be eliminated, so I accept residual injection risk on the read-only low-value tools, bound it to a bad draft that a human reads, and I do not accept it on the refund path, which is why the cap and the confirmation are non-negotiable and why I would hold the launch over them. Then name what you want logged so that an incident can be reconstructed. The three most common failures in this round are a diagram with no boundary marked, an unranked list of threats, and a mitigation plan that amounts to better prompting and a filter.

The hands-on round: reviewing an LLM application, or breaking one live

The hands-on round comes in two shapes and some employers run both. The code review version hands you a few hundred lines of an agent or retrieval application, usually Python, usually with a framework you have seen, and asks what is wrong with it. The red-team version gives you a sandboxed target, a goal such as make the agent reveal the contents of a file it should not read or get it to call the payment tool, and between 45 and 90 minutes, ending in a written finding. Both are graded on process far more than on the count of issues found.

In the code review version, the planted flaws are drawn from a stable set. Retrieval that queries a shared index with no tenant predicate, or that filters results after retrieval instead of scoping the query. A tool whose parameters come straight out of the model's output and into a subprocess, a file path, an SQL string or an HTTP request. Generated code passed to eval or exec outside a sandbox. A system prompt holding a secret, or holding the authorization rule as a sentence of English. A tool description loaded from a third-party or remote source, so the thing that instructs your model is controlled by someone else. The response rendered as HTML or markdown with attacker-influenced links and images intact. An agent loop with no step, token or cost ceiling. A model artefact loaded from a user-supplied path in a format that deserialises arbitrary objects. No logging of tool calls, so an incident cannot be reconstructed. And a guardrail that is a regular expression or a keyword list, often with a comment above it claiming it blocks jailbreaks.

There will also be deliberate decoys, and calling them critical costs you. A low temperature setting is not a security control, but it is not a vulnerability either. A system prompt that is visible is usually not a finding worth a release hold, and saying so, with the reason that a system prompt is not a secret boundary, scores. Using a hosted model rather than a local one is a procurement question. Candidates who rate everything critical get read as someone engineering teams will learn to ignore, which is the most expensive reputation in this job.

Grading, as hiring managers describe it, follows a pattern. Did you ask about context before reviewing: who the users are, whether the data is multi-tenant, what the tools can do. Did you find the authority bug, meaning the one where the model decides something it should not. Did your severity reasoning mention reachability and loss rather than vulnerability class. Were your fixes controls rather than patches. Did you notice the missing observability, which most candidates do not. And would the way you phrased the feedback make the author defensive. Finding four issues with clean reasoning and systemic fixes beats finding eight and ranking them all critical.

In the red-team version, the mistake is to go straight to prompt tricks. Start by mapping what the agent can do, because the attack is almost always through a tool rather than through clever phrasing: list the tools, look for one that fetches, reads, writes or spends, and find out what reaches the model's context from outside your own messages. Try indirect injection before direct: plant the instruction in a document, a ticket, a filename, a web page, an image caption, wherever the system ingests content, because direct injection in your own chat turn is both the easiest to filter and the least interesting finding. Then chain it to something that leaves the system, which is where a real finding lives. And instrument yourself: keep the exact transcript, note how many attempts a technique needed, and when you have it working, run it three times, because a one-shot success against a non-deterministic system is an anecdote and interviewers know it.

The write-up is scored and most candidates rush it. A good one is short and has six parts: what the system is and which boundary failed; the reproduction, with the exact inputs and the success rate across attempts; what an attacker gets, in terms of data, money or access; the severity argument, naming reachability, prerequisites and blast radius; the fix at the level of a control, with the instance-level patch mentioned second; and what you would detect if the control fails. If you only have 10 minutes left, write that rather than find one more bug.

To practise this before anyone tests you, build the target yourself and break it, which teaches more than any range. Public practice does exist, and prompt-injection puzzle games and the HackAPrompt competition data are worth an evening to calibrate what phrasing does, but they train the cheap half of the skill. The expensive half is chaining an injection to a tool with real side effects, and you can only get that on a system you own or one you have written permission to test.

The technical ground the interview actually covers

This section is the study list, ordered by how often it comes up rather than by how interesting it is. Nothing here is exotic, which is itself the point: most of the role is ordinary security applied to an unfamiliar substrate, plus a small number of genuinely new problems.

Retrieval tenancy is the broken access control of AI security, and it is where real incidents cluster. Documents get indexed into a vector store stripped of the permissions they had in their source system, so an assistant answers a question using a file the asker could never have opened. Ask about it in every design review and look for it in every code review. The correct control is that retrieval is scoped by the caller's identity at query time, through a predicate the query cannot omit, and that the index records the source permissions rather than a copy of the text alone. Post-filtering is a bug, because the model has already seen the content by the time you filter.

Agent identity and authority is the other load-bearing area. Know the difference between an agent acting as its own service principal and an agent acting on behalf of a user, and be able to argue for which one a given feature needs and what each costs. Know why long-lived credentials in an agent runtime are worse than they look (the agent is reachable by attacker-influenced text, so its credentials are effectively internet-adjacent), and be ready with short-lived scoped tokens, per-tool scopes rather than one broad role, deny-by-default egress, a sandbox for anything that executes, and a human confirmation on irreversible actions. On the last one, have a view on where the threshold goes and why, because the product manager will push.

The supply chain has three parts and candidates usually know one. Model weights: formats that deserialise arbitrary objects on load are a remote code execution primitive, which is why the safer tensor-only formats exist and why loading weights from an untrusted source is a code execution decision, not a download. Model provenance and signing, and whether your registry can tell you which artefact is running in production and who approved it. Third-party tools and connectors, which in practice now means Model Context Protocol servers: the tool descriptions are instructions to your model, so a server you do not control is an injection channel, an update to a server you trusted yesterday is unreviewed code today, and the tokens those servers hold are a concentrated credential store. Add the dependency problem where assistants confidently recommend package names that were never published and attackers register them, a pattern the field has taken to calling slopsquatting.

Evaluation is the part that separates people who can hold the job from people who can pass the interview. If you cannot measure a control you cannot claim it works, and in a non-deterministic system a demonstration is not a measurement. Be able to describe an attack suite: a fixed set per feature so you can compare releases, a held-out rotating set so you do not optimise against your own tests, attack success rate as the primary metric, false refusal rate as the cost you pay for lowering it, per-release tracking in continuous integration, and a human review of a sample because automated graders drift. Know the open tooling by name and ideally by use: garak, PyRIT, promptfoo, the Adversarial Robustness Toolbox, model scanners for weight files, and whichever gateway your target employer runs. Naming one tool you have actually run beats listing six.

Frameworks and regulation come up more at regulated employers than at labs, and the trick is to be useful rather than encyclopaedic. Know MITRE ATLAS as the tactic and technique vocabulary for AI attacks, the OWASP Top 10 for LLM Applications and the wider OWASP generative AI work including its agentic threat material, the NIST AI Risk Management Framework with its generative AI profile, the NIST taxonomy of adversarial machine learning, ISO/IEC 42001 for an AI management system, and the joint CISA and NCSC guidelines for secure AI system development. On the EU AI Act, know the shape of the obligations rather than the calendar: duties on providers of general-purpose models and heavier duties on high-risk uses, covering governance of training and validation data, technical documentation, record keeping, human oversight, robustness and cybersecurity. Do not quote an application date from memory. That timetable has been amended since the Act entered into force, several US state AI laws have had their effective dates moved after passage, and a candidate who states a deadline confidently and wrongly in an interview has damaged the one impression that matters. Say the obligation, then say you would check the current date before relying on it. That reads as rigour, not ignorance.

Finally, keep the classic adversarial machine learning material in proportion. Model extraction, membership inference, training data extraction, data poisoning and backdoored weights are real, and you should be able to define each and say who actually bears that risk. For most employers, who are consuming a hosted model rather than training a frontier one, these sit below injection, tenancy and agent authority. If you lead with them you will sound like you prepared for an older version of this field. At a lab, a model provider or a company training on sensitive data they move up the list, and poisoning of a fine-tuning set or of a retrieval corpus is a live concern worth real preparation.

The resume, the portfolio, and what gets ignored

Resumes for this role are read by a security hiring manager, often in under a minute, looking for one thing: evidence that you have made a judgement call about an AI system and been right. Everything else on the page is context for that.

Write bullets that name the system, the decision and the consequence. Not responsible for AI security reviews, but reviewed the AI feature designs shipped last year and blocked two launches, one for cross-tenant retrieval and one for an agent with unbounded refund authority. Not improved guardrails, but built the injection test suite that runs on every release of the assistant, and give your own attack count, your own before-and-after attack success rate and the false refusal cost you accepted. Use your real numbers or no numbers: if you do not have the measurement, describe the artefact and the decision instead, which reads as honest and is checkable in conversation. An invented metric dies in the first follow-up question and takes the rest of the resume with it.

The portfolio question is unavoidable from outside, because a title this new has no accepted credential and the manager has to substitute something. One artefact is enough if it shows judgement rather than enthusiasm, and the ranking is fairly consistent: a disclosed and fixed vulnerability in an AI product or framework, with the reproduction and the patch, sits at the top. Then a published threat model of a real, documented AI system, which is the cheapest high-signal artefact available because it costs nothing but thought and is exactly what the hardest round tests. Then an attack suite with measured before-and-after results, which demonstrates the skill most teams are short of. Then a tool other practitioners use, a verified bounty record on an AI vendor's programme, or a placement in a public AI red team or jailbreak event. Several AI companies run both a conventional vulnerability programme and a separate model-safety or jailbreak programme with different rules, so read the current scope on the vendor's own page before you touch anything, and never test a production AI system you have no permission to test.

What gets ignored, and in some cases counts against you: a long certificate list with no artefact; prompt engineering as a skill line; a screenshot of a jailbreak in a chat window with no impact argument; passionate about AI safety and alignment as a summary line, which reads as a different job; a list of frameworks you have read rather than applied; and the word leveraged. Hiring managers for this role see a very high volume of applicants whose entire evidence is that they have used a chatbot a lot.

On keywords, match the posting's own vocabulary wherever it is honestly true of you, because the first pass is often automated and this field has three names for everything. If the posting says LLM security, write LLM security rather than only generative AI security. If it says red teaming, use that phrase. Keep both spellings of threat modelling where you can do it naturally, and include the specific tool and framework names you have actually used, because those are what a recruiter searches on.

Where the jobs are, what to search, and a 12-week plan

Demand for this role is concentrated rather than broad, which matters for how you search. The clusters are: frontier labs and large AI companies, which hire the deepest version of the role and have the highest bar; the large cloud providers, which are securing the AI platforms everyone else builds on and hire in volume; AI-native product companies past their first funding, which hire one or two generalists who do everything; security vendors building AI security products; large regulated enterprises standing up an AI programme, which is the fastest-growing source of openings and the least glamorous work; consultancies and boutique AI red team shops, which hire for writing and client presence as much as for depth; and defence, government and national security, gated on clearance. Postings also hide inside existing teams: a platform security team that picked up the AI platform, a product security team that picked up the copilot.

Search on the work, not the title, and keep an alert running on each phrase because these postings fill fast. Useful strings: AI security engineer, AI/ML security, GenAI security, LLM security, model security, AI red team, adversarial testing, AI safeguards, agent security, MLSecOps, and security engineer combined with agent or inference or model registry. Then read the responsibilities for which of the five variants it is and which half of the loop will be heavy.

On pay, resist the urge to anchor on a number from a forum. There is no occupation code for this title, so the honest sources are the ones you can open: BLS OES 15-1212 for information security analysts as a floor and sanity check, 15-1221 for research-flavoured lab roles, and the published ranges on postings in jurisdictions with pay transparency rules, including Colorado, California, New York, Washington and Illinois. The useful comparison is internal: for the same employer, this role generally pays at or above its senior security engineer band, and at labs and large technology companies a large share of total compensation is equity whose value you should discount rather than assume. When a recruiter asks for your expectation, name the band from their own published posting if there is one.

Here is a 12-week plan that produces interviews rather than a reading list. It assumes six to ten hours a week and it is built around producing two artefacts, because artefacts are what get you screened in.

Working with AI in this role

What an AI security engineer must know about AI in 2026-27

For this role the usual version of this section would be circular, since AI is the subject matter. The useful version is sharper: which parts of the AI story are real enough to build a career on, which are marketing, and what has changed in how this job is actually done. Getting that calibration right is itself an interview question, because many hiring managers ask some version of what do you think is overhyped here, and both the cynic and the enthusiast fail it.

What is real, and will still be real across 2027. Prompt injection has no fix, only architecture. That is not a temporary state of the art problem: a model that follows instructions in text cannot reliably distinguish your instruction from an attacker's, so every serious defence is about where authority lives, what tools can do, and what leaves the system. Agents made this material rather than theoretical, because an agent has credentials and side effects. Retrieval over internal documents turned permission inheritance into the most common real-world AI data leak, and it is boring plumbing rather than research. The tool and connector ecosystem, which in practice now means Model Context Protocol servers, created a genuine new supply chain in which third-party text instructs your model and third-party servers hold your tokens. Prompt and completion logs became a large new personal data store at nearly every company that shipped an assistant, usually without the classification the source data had. And measurement became the only credible way to claim a control works, because nothing here is deterministic. Those six things are the job.

What is oversold, and saying so carefully is a signal of competence rather than scepticism. Guardrail products sold as an AI firewall are useful as defence in depth and are not a boundary: they are classifiers with a false negative rate, and designing as though a filter were an authorization control is the most common expensive mistake in production systems right now. Fully autonomous AI attackers remain a thinner story than the vendor decks suggest, although the honest intermediate position matters: model-driven analysis and fuzzing have genuinely found previously unknown vulnerabilities in widely used open source software, Google's Big Sleep work on a previously unknown SQLite bug and machine-generated fuzz targets in OSS-Fuzz being documented examples worth reading in the original, and phishing, reconnaissance and low-grade malware development have clearly been made cheaper. What has not happened is an autonomous adversary that decides what matters in your specific environment. Meanwhile the volume of fluent, confident, wrong vulnerability reports arriving in bug bounty and open source security inboxes has grown into a real operational problem, something the curl maintainers in particular have documented publicly, and triaging that flood quickly and politely is now part of several versions of this job. Classical adversarial machine learning is also oversold relative to employer risk: real, well studied, and below injection and tenancy for anyone consuming a hosted model rather than training their own.

And here is the part candidates miss: AI changed how this job is done, not only what it is about. Design review volume went up sharply, because every product team is now shipping an AI feature, so the role lives or dies on reusable patterns rather than bespoke reviews, which is why the strongest answer to almost any interview prompt is a control, a library or a default rather than a finding. Red teaming became partly automated, so building and maintaining the harness is more of the work than inventing attacks by hand. Your own assistant use is examined: expect to be asked, and answer with a practice and a verification step, such as using it to generate attack variants and to summarise a log, never to decide a severity or to approve a design. Review of generated code matters here too, since a growing share of the code arriving on your desk was drafted by a model, and its characteristic failures are clean, idiomatic and missing an authorization check nobody prompted for. Finally, the field's vocabulary is still settling, so define your terms once at the start of an interview answer rather than assuming the panel uses them the way your last employer did.

Designing around prompt injection instead of promising to stop it

This is the single discriminating question in the role. An input filter, a sanitiser or a phrase allowlist as the headline mitigation marks a candidate as someone who read coverage rather than systems, and interviewers use it as a fast negative signal. The correct model is that the language model is a confused deputy: it holds credentials, it takes instructions from strangers, and it cannot tell the two apart. Everything else follows from accepting that.

Show it: In the threat model round, say early that injection cannot be eliminated and that you will therefore bound authority and close exfiltration paths. Apply the lethal trifecta test out loud (private data, untrusted content, outward communication) and name which leg you are removing on each path. Give concrete controls: tool parameters sourced from authenticated context rather than from model output, server-side caps on irreversible actions, human confirmation above a threshold enforced by the API, deny-by-default egress, plain-text rendering. Then accept the residual risk in one sentence with a detection attached. On the resume, a published threat model that does this is worth more than any certificate.

Agent authority, identity and blast radius

Agents turned AI security from a content problem into an access control problem. An agent with one broad role, a long-lived credential and unrestricted egress is an internet-reachable confused deputy holding your production permissions, and that single design error is behind most of the serious agentic incidents that get written up. Employers hiring this title in 2026-27 are overwhelmingly doing so because they are shipping or adopting agents and someone has to own that boundary.

Show it: Be able to argue service principal against on-behalf-of for a specific feature and say what each costs. Specify per-tool scopes rather than one role, short-lived scoped tokens, a sandbox for anything that executes code, deny-by-default egress from the tool runtime, step and spend ceilings, and a confirmation threshold for irreversible actions with a number you can defend. In the hands-on round, trace every tool parameter back to its origin and flag anything the model decided that changes who or what an action affects. In a portfolio, design the permission model for an agent, then show the detection that catches it when the model is wrong.

Retrieval tenancy and permission inheritance

This is the most common real AI data leak and the least glamorous. Documents get indexed without the access controls they had in the source system, so the assistant answers from a file the asker could never open, and nobody notices until someone asks about salaries or an acquisition. It is also the planted bug in a large share of code review exercises, because it is the AI-flavoured version of broken access control and it is invisible to scanners.

Show it: Ask in the first minute of any review whether the corpus is multi-tenant and whether source permissions came along with the text. Insist on scoping at query time through a predicate the query cannot omit, and explain why post-filtering is a bug (the model has already read the content). Mention re-indexing when permissions change and what happens to a document after it is deleted at source. If you have built one, say how you tested it: a per-tenant canary document and a query that should never return it, running in continuous integration.

Tool, connector and model supply chain review

A third-party tool description is an instruction to your model from someone you do not employ, and the server behind it holds tokens for systems you care about. An update to a connector you reviewed last month is unreviewed code today. On the model side, weight formats that deserialise arbitrary objects make loading a checkpoint a code execution decision rather than a download. This is the area where security engineers have the most transferable advantage and where AI teams most often have no process at all.

Show it: Describe a concrete connector review: who publishes it, what scopes it requests against what it needs, what credentials the server stores, whether tool descriptions are pinned or fetched at runtime, what happens on update, and whether its responses reach your model's context unlabelled. For models, require provenance for weights, pin and verify artefacts, prefer tensor-only formats, scan files before load, and be able to say which artefact is in production and who approved it. Mention hallucinated dependency names, the slopsquatting pattern, and the lockfile discipline that closes it.

Measuring a control instead of demonstrating one

Nothing in this stack is deterministic, so a successful demonstration proves almost nothing and a single blocked attack proves less. The teams actually reducing risk have a suite, a baseline and a trend, and the teams that are not have a slide. Interviewers screen for whether a candidate thinks in distributions, because the alternative is a mitigation that works in the demo and regresses silently two releases later.

Show it: Talk in attack success rate against false refusal rate, and treat the trade-off as explicit rather than embarrassing. Describe a fixed per-feature suite for release comparison plus a rotating held-out set so you are not optimising against your own tests, per-release tracking in continuous integration, and a sampled human review because automated graders drift. Name tools you have actually run, such as garak, PyRIT or promptfoo, and give one number with its context. If you have an attack suite with before-and-after results published, lead the resume with it.

Being useful about governance without quoting dates you cannot verify

At banks, insurers, healthcare systems, public sector bodies and any employer selling into the EU, part of this job is translating obligations into controls, and panels include people who read the texts for a living. The failure mode that costs offers is not ignorance, it is confident wrongness about a timetable: the EU AI Act's application dates have been amended since the Act entered into force, and several US state AI laws have had their effective dates moved after passage. Stating a deadline from memory in the room is a risk with no upside.

Show it: Speak in obligations and controls rather than calendars: governance of training, validation and testing data, technical documentation, record keeping and logging, human oversight, robustness and cybersecurity, reporting of serious incidents. Then add the sentence that scores: I would check the current application date before relying on it, because that timetable has moved. Map an obligation to something concrete you would build, such as a retained tool-call log that can reconstruct an agent action, or a model inventory tied to the registry rather than to a spreadsheet. Know ISO/IEC 42001, the NIST AI Risk Management Framework and its generative AI profile, MITRE ATLAS and the OWASP generative AI material well enough to pick the right one for a given audience.

What a screen is looking for

These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.

Mistakes that cost people this job

Answering the prompt injection question with input sanitisation, a phrase allowlist or a guardrail product.

Say that injection cannot be eliminated and that the fix is architectural: authority outside the model, tool parameters from authenticated context, server-side caps and human confirmation on irreversible actions, deny-by-default egress, output treated as untrusted. Mention a classifier only as defence in depth with a false negative rate. This one answer separates the two halves of the applicant pool faster than anything else in the loop.

Arriving with only half the job and assuming the AI half is the hard one.

Name your weaker half in week one and fix that. If you came from security, build a real multi-tenant retrieval application with tools so the mechanics are concrete. If you came from machine learning, get to where you can find broken access control in unfamiliar code and explain what goes wrong in an OAuth flow. The loop is deliberately built to find whichever half you skipped.

Applying to postings without checking whether the job is security for AI or AI for security.

Read the first three responsibilities. Threat model, review, guardrail and agent mean this job. Detection content, triage automation and copilot for the security operations centre mean a detection or security operations job with a fashionable title. Both are fine careers and they prepare differently, so filter before you spend a screen.

Leading with jailbreak demonstrations and clever prompts as the portfolio.

Lead with impact and architecture. A jailbreak with no impact argument is the cheapest artefact in this market and every applicant has one. Show indirect injection through ingested content chained to something that exits the system, with a reproduction, a success rate across at least three runs, a severity argument and a control-level fix.

Rating everything critical in the code review round, including the deliberate decoys.

Rate by who can reach it and what is lost, and say out loud which items are not findings and why. A visible system prompt is usually not a release blocker, and explaining that a system prompt is not a secret boundary scores better than flagging it. Then name the one thing you would hold the launch for.

Treating the model's own text as a source of parameters for a tool.

Make it a rule you say in every design review: the model may decide to act, but the identifiers that determine who or what is affected come from the authenticated context, server-side. Refund the order in the ticket record, not the order number the model produced.

Leading with adversarial examples, model extraction and membership inference at an employer that consumes a hosted model.

Match the risk to the employer. Define those attacks accurately, then say they rank below injection, retrieval tenancy and agent authority for a company that does not train its own models, and move them up the list only where fine-tuning on sensitive data or frontier training is actually happening. Proportion reads as judgement.

Quoting a regulatory deadline from memory, especially an EU AI Act application date.

State the obligation without the calendar: governance of training and validation data, documentation, logging, human oversight, robustness and cybersecurity. Then add that you would verify the current application date, because that timetable has been amended since the Act entered into force and several US state AI laws have had effective dates moved after passage. Confident and wrong in the room is the expensive outcome.

Claiming a mitigation works because you watched it block an attack.

Give a measurement. Attack success rate across a fixed suite before and after, the false refusal cost you accepted, the held-out set you kept aside, and where it runs in continuous integration. In a non-deterministic system a single blocked attempt is an anecdote, and interviewers screen for whether you know that.

Buying a new AI security certificate and treating it as the credential that unlocks the role.

No certificate gates this job and most postings do not name one. Spend the money and the weeks on one artefact instead: a published threat model, an attack suite with measured results, or a disclosed and fixed finding. If a specific employer's posting names a specific certificate, get that one, and check its current status on the issuer's own page because this corner of the market changes fast.

Looking only outside while your own employer is standing up an AI programme.

Volunteer to review the first agent, the first assistant deployment or the first connector at the employer you already have. These roles are mostly filled internally or one step away, because the manager is hiring someone who can be trusted to block a launch. Doing one review well gives you the experience, the reference and the story at once.

Questions people ask

What does an AI security engineer actually do?

An AI security engineer secures systems that take instructions from text nobody controls, act on them with real credentials, and cannot be made to behave deterministically. In practice that means reviewing AI feature designs before they ship, reviewing the code that wires models to tools and data, owning the permission model for agents and their tools, making sure retrieval respects the permissions of the documents it indexes, reviewing the supply chain for model weights and third-party connectors, maintaining a measured attack suite that runs on every release, and building detection at the tool-call layer. The deliverable an AI security engineer is judged on is a bounded blast radius, not a collection of jailbreak prompts.

Do I need a machine learning background to become an AI security engineer?

An AI security engineer does not need to be able to train a model, but does need enough mechanics to avoid being confidently wrong: context windows, tokenisation, how retrieval selects documents, what fine-tuning changes that prompting does not, what a guardrail model costs in latency and false refusals, and why non-determinism breaks ordinary testing. Most people hired into this title come from security engineering rather than from machine learning, because the role rests on authorization thinking and blast radius, which are harder to acquire than the mechanics. Budget three to nine months to close the AI gap from a security job, and six to twelve months to close the security gap from a machine learning job, because the second direction is the one the interview loop rejects more often.

What certifications or qualifications do I need for an AI security engineer job?

No licence and no mandatory certification gate an AI security engineer role in any market where it is hired. Several issuers including ISACA, ISC2, SANS and GIAC, and CompTIA have launched AI-security credentials recently, most postings do not ask for any of them by name, and you should check the current name, status and price on the issuer's own page before paying and treat it as a resume filter rather than a qualification. The two real gates are government clearance for defence and national security work, which sets the timeline in months, and enhanced background screening at some labs and regulated employers. An AI security engineer is screened in by one artefact that shows judgement, such as a published threat model, an attack suite with measured results, or a disclosed and fixed vulnerability in an AI product.

What does the AI threat modelling interview test?

The threat model round tests whether an AI security engineer has a repeatable process that works out loud on a system they have never seen. You get a system described in two or three minutes, usually an agent with tools and some customer-facing input, and about 40 minutes to ask five questions about money, data, reversibility, tenancy and rendering, mark the boundary where untrusted text enters the model's context, rank threats by reachability and loss rather than by cleverness, propose controls that are properties of the design rather than rules a future engineer must remember, accept one residual risk with a compensating detection, and name what you would log. The three failures that end this round for an AI security engineer are a diagram with no trust boundary marked, an unranked list of threats, and a mitigation plan that comes down to better prompting and a filter.

How should I answer 'how would you stop prompt injection?'

An AI security engineer should open by saying that prompt injection cannot be eliminated, because a model that follows instructions in text cannot reliably separate your instruction from an attacker's, so the work is containment rather than prevention. Then give the architecture: take authority out of the model so tool parameters come from authenticated context rather than from model output, bound irreversible actions with server-side caps and a human confirmation enforced by the API rather than the prompt, deny egress by default from anything the model can reach, treat model output as untrusted input to shells, databases and renderers, and remove one leg of the lethal trifecta (private data, untrusted content, outward communication) on every path. Mention classifiers and filters only as defence in depth with a false negative rate, and finish by accepting the residual risk in writing with a detection attached, because that is the answer that reads as senior.

What belongs on an AI security engineer resume, and what gets ignored?

An AI security engineer's resume should lead with decisions and artefacts: designs reviewed and launches blocked with the reason given, a cross-tenant retrieval bug found and the control that closed the class, an injection test suite built with the attack success rate before and after plus the false refusal cost, an agent permission model designed, a connector review process stood up. Keep the classic security spine visible as well (authorization, identity, cloud IAM, code review, incident response) because the loop examines it, and name tools you have genuinely run such as garak, PyRIT or promptfoo rather than listing everything you have read about. What gets ignored or counts against an AI security engineer: prompt engineering as a skill line, a certificate list with no artefact, a jailbreak screenshot with no impact argument, passionate about AI safety as a summary, and any number you cannot defend in the follow-up question.

Who is hiring AI security engineers in 2026-27?

AI security engineers are hired by a fairly concentrated set of employers: frontier labs and large AI companies, which run the deepest version of the loop; the major cloud providers securing the AI platforms everyone else builds on; AI-native product companies past their first funding, which hire one generalist to cover AI, cloud and compliance at once; security vendors building AI security products, where the loop weights software engineering; large regulated enterprises standing up an AI programme, which is the fastest-growing and least glamorous source of openings; consultancies and boutique red team shops, which hire for writing and client presence; and defence and government, gated on clearance. Many openings are also hidden inside existing teams, so an AI security engineer should search on the work (agent, inference, model registry, LLM, red team) rather than on the title.

Is AI security engineer a junior-friendly role?

AI security engineer is structurally a mid-level and senior title, and genuinely junior openings are rare, because the job involves telling an engineering director that a launch needs to change and being believed. Most requisitions ask for several years of security, software or platform engineering before the AI qualifier, and the usual entry is sideways rather than upwards: an application security, cloud security, product security, penetration testing or platform engineering job first, then an internal move into the AI programme at the same employer, which is how a large share of these roles are actually filled. For someone early in their career, the realistic plan is to get a software or security engineering job, make yourself the person who reviews the first AI feature there, and publish one artefact in the meantime so that the title change is a formality rather than a leap.

What does an AI security engineer get paid, and where can I check?

There is no occupation code for AI security engineer, so any single band you see quoted is someone's estimate rather than a measurement. The checkable sources are BLS OES 15-1212 for information security analysts as a floor and sanity check, 15-1221 for research-flavoured roles at labs, and the published salary ranges on actual postings in jurisdictions with pay transparency rules including Colorado, California, New York, Washington and Illinois. The useful comparison is internal rather than across the market: for a given employer, an AI security engineer generally sits at or above the senior security engineer band, and at frontier labs and large technology companies a large share of the package is equity whose value should be discounted rather than assumed.

How is an AI security engineer different from an AI red teamer, an application security engineer or an AI governance analyst?

An AI security engineer owns the security of AI systems end to end, which means both breaking them and designing the controls and defaults that keep the next one safe. An AI red teamer is the adversarial half of that, running time-boxed attack campaigns and writing findings, and usually hands the fix to someone else. An application security engineer owns the security of the product's code and design generally, and in many companies picks up AI features as one more area rather than as a specialism. An AI governance or AI assurance analyst owns inventories, model documentation, policy and audit evidence against a GRC rubric and rarely opens the code. The distinction matters when applying, because an AI security engineer posting will test both a threat model and a hands-on exercise, while the other three weight one of those much more heavily than the rest.

Put this on a resume in about a minute

Paste your history once and point it at the AI Security Engineer posting you are looking at. No account, no card.

Build my resume free More roles