| What the job is in 2026-27 | Turning pixels into a decision a machine or a person acts on, inside a fixed latency, power and cost budget, with a measured error rate at a chosen operating point. Model training is a minority of the week. Data curation, evaluation design, calibration, integration and deployment are the majority. |
|---|---|
| One title, four different jobs | Perception engineer (geometry, calibration, tracking, sensor fusion, ROS 2, real time), vision modelling or applied research (datasets, training, ablations, sometimes publications), vision deployment or edge (TensorRT, quantisation, GStreamer, throughput on a board), and machine vision engineer on a factory line (optics, lighting, PLC handshakes, uptime). Read the posting for which one it is before you tailor anything. |
| Where the hiring sits | Robotics and warehouse automation, defence and dual-use plus geospatial, manufacturing and semiconductor inspection, medical imaging and digital pathology, autonomous trucking and driver assistance, and drone-based inspection in agriculture, construction and energy. Consumer photo apps and checkout-free retail hire fewer computer vision engineers than they did. |
| Credential gate | None. No licence, no mandatory certification, no required degree anywhere this job is hired. A master's degree is common and a PhD is effectively expected only for titles containing research or scientist. The real screen is a portfolio with numbers in it. |
| The eligibility wall that decides more outcomes than skill | US export control. Defence, aerospace and much geospatial computer vision work is ITAR or EAR controlled, which in practice restricts it to US persons (citizens and lawful permanent residents, among a few other categories), and many roles additionally require a security clearance that takes months to adjudicate. If you cannot meet that, target commercial robotics, inspection, medical imaging or media instead, and do not spend weeks applying into a wall. |
| The loop | Typically four to six stages over three to eight weeks: recruiter screen, hiring manager project deep dive, a coding or take-home exercise, then an onsite covering project depth, classical geometry, deep learning and data strategy, deployment and systems, and behaviour. Defence adds a clearance wait measured in months. Medical device employers add a regulatory and validation conversation. |
| What gets you screened in from outside | One watchable artefact with defensible numbers. In rough order of weight: a deployed system you can show running on video you collected, with a metric table and a latency or power figure on named hardware; a calibration or geometry writeup with reprojection error in pixels; a distillation project that trades a large teacher model for a small fast student and reports the tradeoff; then papers, which matter mainly for research titles. |
| Pay, and where to check it rather than trust a band | There is no occupation code for this exact title. The nearest authoritative US baselines are BLS OES 15-1252 Software Developers, 15-2051 Data Scientists, and 15-1221 Computer and Information Research Scientists for research-titled roles, each with metro tables that matter more than the national figure. For a live number, read pay ranges in postings from states that require them (California, Colorado, Washington, New York and Illinois among others; check which applies to the role's location), and treat self-reported aggregators as a sanity check rather than a source. Clearance and hard real-time embedded experience both move the number. |
What a computer vision engineer actually does, and which of the four jobs the posting means
The deliverable is not a model. It is a decision: this pallet is at this pose, this weld is defective, this frame contains a vehicle, this region of the slide warrants a pathologist's attention, this crop row needs spraying. A computer vision engineer owns the path from a sensor to that decision, and is judged on the error rate at an operating point somebody chose for a business reason, at a latency the rest of the machine can tolerate, on hardware that was costed before you arrived.
That shape explains how the time actually goes. In most production computer vision roles the week is dominated by data and labels (deciding what to collect, finding the label noise, building a held-out set that is not contaminated), by evaluation (does this number correspond to the thing the product cares about), and by integration (the camera driver, the timestamps, the preprocessing that has to match exactly between training and the robot, the board that thermally throttles once the enclosure heats up). Training runs are the part that gets talked about and the smaller part of the job. Candidates who prepare only for the training part interview badly, because the interviewer has spent the last year on the other parts.
Four distinct jobs hide under this title, and conflating them is a common reason a good candidate gets a polite no. Perception engineering in robotics and autonomy is geometry-first: calibration, coordinate frames, 6-DoF pose, tracking, multi-sensor fusion, time synchronisation, real-time constraints, and usually C++ and ROS 2. Vision modelling or applied research is dataset-first: architectures, training recipes, ablations, benchmark discipline, and sometimes publications. Vision deployment, often called edge AI, is throughput-first: ONNX, TensorRT or OpenVINO, quantisation, pipeline plumbing in GStreamer or DeepStream, and a frames-per-second number on a specific chip. Machine vision on a factory line is a different trade again, closer to controls engineering: lenses, lighting, exposure, global versus rolling shutter, a PLC handshake, a reject gate, and uptime measured in months.
You can tell them apart from the posting's nouns, and you should, because the resume you send and the project you lead with change completely. Geometry nouns mean perception. Benchmark nouns mean research. Chip and runtime nouns mean deployment. Plant-floor nouns mean machine vision. If a posting mixes all four, it is a small team and you are expected to do all of it, which is the most interesting version of the job and the one where a generalist portfolio wins.
One more honest note on scope. A large share of the open work is not greenfield modelling at all. It is taking a system that already works at the two sites it was built for and making it work at the fourteenth site, where the lighting is different, the camera was mounted five degrees off, and the operator has started wiping the lens with a sleeve. That work is unglamorous, pays well, and is what senior candidates get hired for. If you can tell a story about generalising a working system across sites or devices, lead with it.
- Perception signals in a posting: extrinsics, intrinsics, hand-eye calibration, time sync or PTP, SLAM, VIO, point cloud, 6-DoF pose, ROS 2, C++, real-time.
- Research signals: ablation, benchmark, state of the art, publications at CVPR, ICCV, ECCV, NeurIPS or ICRA, self-supervised pretraining, diffusion, dataset scaling.
- Deployment signals: TensorRT, ONNX, OpenVINO, TFLite or LiteRT, INT8, quantisation-aware training, Jetson Orin, Hailo, Ambarella, GStreamer, DeepStream, fps, watts.
- Machine vision signals: GigE Vision, GenICam, telecentric lens, backlight, strobe, line scan, global shutter, PLC, parts per minute, false reject rate, Cognex, Keyence, HALCON.
- Medical imaging signals: DICOM, NIfTI, PACS, whole slide imaging, nnU-Net or MONAI, reader study, sensitivity at a fixed specificity, 510(k), ISO 13485, IEC 62304.
- Defence and geospatial signals: EO and IR imagery, SAR, STANAG 4609 or KLV metadata, counter-UAS, SWaP, ITAR, clearance, GEOINT.
- Mixed-signal postings at startups mean one person owns pixels to decision. That is the job to apply to if your portfolio spans data, training and a board.
- If the posting never names a camera, a chip, a latency or a metric, ask in the screen. Teams that cannot answer those questions usually do not have a shipped system yet, which changes what the job is.
Where computer vision hiring has concentrated in 2026-27
The pattern across the last few years states in one sentence: hiring moved to places where a camera removes one repetitive human judgement and somebody can calculate the saving, and to places with an industrial or defence budget behind them. It moved away from consumer photo and video features, where general-purpose models absorbed the work, and away from checkout-free retail, which pulled back from large grocery formats to smaller venues. If you are choosing where to point a job search, point it at a sector with a measurable unit of waste.
Robotics and warehouse automation carries a large share of commercial perception openings. The demand sits in piece picking and depalletising, autonomous mobile robots and forklifts, inspection and cleaning robots, agricultural and construction machines, and surgical and lab robotics. The problems are relentlessly geometric: calibrate a camera to an arm, estimate the pose of an object you were not trained on, find a grasp point, track something through occlusion, keep the extrinsics honest after a month of vibration, and fail safely when the scene is ambiguous. Humanoid companies hire into this pool too and advertise loudly, but the production volume and the steady headcount are in the boring machines that move boxes. Interviews here test transforms and failure handling harder than they test architectures.
Defence, dual-use and geospatial hires heavily and is the most gated pool of the six. The work is detection and tracking in electro-optical and infrared video, counter-drone perception, maritime and ground target recognition, change detection in satellite and synthetic aperture radar imagery, and running all of it on an airframe inside a size, weight and power budget that would horrify a cloud engineer. Employers include newer prime contenders, the established primes, and the geospatial analytics firms. The gate is export control and clearance, not skill. If you are a US person, this sector will usually interview you faster than any other and will pay a premium for a clearance you already hold. If you are not, most of its roles are closed to you and you should spend the time elsewhere.
Manufacturing and semiconductor inspection is the steadiest pool and the least visible, because the jobs are at automotive suppliers, electronics assemblers, food and pharmaceutical packers, battery plants and fab tool vendors rather than at companies whose names you see in technology news. The defining constraints are that defects are rare, that the line will not slow down for you, and that a missed defect (an escape) and a false reject have very different costs which someone can quantify in currency per unit. The modern version of this work is anomaly detection trained mostly on good parts, plus a rigorous argument about thresholds. It also requires something most machine learning candidates have never done: changing the lighting, the lens or the shutter so the defect is visible at all, which is frequently a bigger win than any model change.
Medical imaging and digital pathology hires computer vision engineers continuously, with a longer and more documented development cycle. Radiology triage and quantification, whole slide imaging in pathology, ophthalmology screening, ultrasound guidance, surgical video and dermatology all have active commercial programmes. Two things make this sector different. First, the data is heavily governed: de-identification, institutional review, site agreements, and a gigapixel file format you have to learn. Second, validation is a regulated artefact rather than a slide, so you will work with clinical collaborators on reader studies, operating points chosen by clinical consequence, and submissions under a quality management system, including how a model update is pre-authorised rather than shipped quietly. If the product is sold in the EU, the AI Act's high-risk obligations touch data governance and post-market monitoring; its phasing has already been amended more than once, so cite the obligation and say the timing needs checking rather than quoting a date. Expect to be asked how you would split a dataset by patient and by site, and expect that answer to matter more than your architecture choice.
Autonomous driving still hires, but the shape changed. Driverless ride-hailing consolidated to a few operators, and much of the remaining perception headcount sits in autonomous trucking, in driver assistance programmes at vehicle makers and tier one suppliers, in the silicon vendors that supply them, and in the simulation and data engine teams that feed all of the above. The work is long-tail triage more than architecture invention: find the scenario class that fails, mine for it, label it, retrain, prove the regression suite did not move the wrong way, and do it inside a functional safety process where ISO 26262 and the safety of the intended functionality standard, ISO 21448, shape what you are allowed to claim. Candidates who treat it as a benchmark competition interview poorly here.
Then the smaller pools, which are real and often less competitive. Drone and fixed-camera inspection of infrastructure, solar, rail, pipelines and crops. Mining and heavy equipment. Sports and broadcast analytics. Live-event and retail analytics on edge boxes. Document and receipt understanding, which quietly became a vision-language model problem. Media generation companies, which hire computer vision people for data pipelines, dataset curation, safety filtering and evaluation rather than for model research. AR and spatial computing, which hires a small number of very strong geometry and tracking people. None of these will have fifty openings, but a tailored application into one of them competes against a much shorter queue than a generic robotics application does.
- Pick a sector before you build a portfolio project. A defect-detection demo and a 6-DoF pose demo signal to different employers and are not interchangeable.
- Robotics asks: can you keep a perception stack honest after a month on a real machine. Prepare a drift, calibration or occlusion story.
- Defence asks: are you eligible, and can you make it work in a SWaP budget. Eligibility is the first screen and there is no way around it.
- Inspection asks: can you reason about escapes against false rejects in money, and will you change the optics instead of the model when that is the cheaper fix.
- Medical asks: can you design an evaluation a regulator and a clinician will both accept, with splits by patient and site and an operating point chosen by clinical cost.
- Autonomous driving asks: can you run a data engine loop on the long tail and prove you did not regress something that matters.
- Smaller pools (inspection drones, sports, documents, generative media data) have shorter queues and often a faster loop. Candidates underrate them.
- Public sector and research institutes also hire, with slower processes and more documentation. Worth applying to if you want a geometry-heavy role without a product deadline.
What gates the role: no licence, four real eligibility walls, and the degree question
Nothing licenses a computer vision engineer. There is no board, no registration, no mandatory certification, and no employer in this field that will reject you for lacking a specific certificate. Vendor credentials exist (NVIDIA's deep learning and Jetson courses, cloud machine learning certifications, machine vision certifications from industry associations and from Cognex or MVTec training), and they are worth exactly what they signal: that you sat through a curriculum. They do not substitute for a measured result on a board. If you are choosing between a certificate and four weekends building a deployed demo, build the demo.
The degree question has a precise answer. A bachelor's degree in a numerate subject plus a strong portfolio clears most product roles. A master's is the modal qualification and helps mainly because it buys supervised time on a real dataset. A PhD is effectively required only where the title contains research or scientist, where the job is publishing or inventing, or in a small number of labs that use it as a filter. Self-taught entry works in this field, but it is harder than in web development for a concrete reason: the prerequisites include linear algebra you can actually use, projective geometry, probability, and comfort with hardware, and no amount of framework familiarity papers over missing geometry in an interview.
The first eligibility wall is export control, and it decides more outcomes in computer vision than any skill gap. A large share of defence, aerospace, counter-drone and geospatial work falls under ITAR or EAR, which in practice restricts it to US persons, a category that includes lawful permanent residents and not only citizens, and many of those roles additionally require a security clearance. Clearance adjudication runs in months rather than weeks, and the higher tiers can include a polygraph and a long interim period. Two practical consequences: if you are eligible, an existing clearance is the single most valuable line on your resume and belongs near the top, and if you are not eligible, filter those postings out of your search rather than discovering the wall at the offer stage.
The second wall is a regulated-domain track record. Medical imaging employers shipping a cleared product want someone who has worked inside a quality management system and can talk about design controls, software lifecycle documentation, traceability from requirement to test, change control for model updates, and the difference between a research result and a validated one. You do not need a licence to do this work, and nobody will test your clinical knowledge, but a candidate who has never seen a design history file competes badly against one who has. The entry route is the research or prototyping side of a medical imaging company, then moving toward the submission side.
The third wall is practical access to data and hardware. This field punishes candidates with no device. You cannot credibly claim deployment experience without having put a model on something: a Jetson Orin Nano developer kit (the original Jetson Nano is end of life, so buy a current board), a single-board computer with an accelerator module, a used industrial camera off the secondary market, a cheap depth camera, a drone. The spend is in the low hundreds of dollars, less secondhand, and it changes what you can say in an interview from an opinion into a measurement. Cloud GPUs by the hour cover training; nothing covers embedded except an embedded board.
The fourth wall is work authorisation in the ordinary sense, and it interacts with the first. In the United States, early-career candidates on student visas face a smaller commercial pool because the export-controlled share of computer vision is large. The sectors that remain fully open are robotics outside defence, inspection, medical imaging, retail and logistics analytics, consumer devices and media. Outside the United States, the equivalent concentrations are industrial inspection and automotive perception in Germany and Japan, robotics and automotive in China and Korea, defence and autonomy in the United Kingdom, Israel and increasingly across the EU, each with their own national security gates.
- No licence, no required certification, no required degree. Portfolio and measured results are the gate.
- PhD matters for research and scientist titles and for a handful of labs. It is not a requirement for product perception, deployment or machine vision roles.
- An active clearance is the highest-value single line on a defence computer vision resume. State the level and status plainly and factually.
- If you are not export-control eligible, prune those postings from the search. The wall is legal, not negotiable, and recruiters cannot waive it.
- Buy one embedded board. Being able to say a measured frame rate and wattage for your own model, on a board you own, beats any certificate on the resume.
- Medical device employers screen for quality system exposure. If you lack it, enter through the research or prototype side of an imaging company.
- Vendor courses are fine as structure, weak as signal. Finish them only if they end in something you can demo.
- Mathematics prerequisites are real: projective geometry, linear algebra, probability and optimisation. Interviewers ask you to do the algebra, not to name it.
How hiring works, stage by stage, and how it differs by employer
Who reads your application depends on employer size. At a robotics or perception startup under roughly 200 people, the perception lead reads resumes directly, clicks the portfolio link, and decides in under two minutes. That is unusual in engineering hiring and it is the most exploitable fact in this field: a watchable clip in the first line of your application is seen by the person who makes the decision. At large technology companies and at primes, a recruiter screens on keywords first, which is why the exact strings (TensorRT, ROS 2, CUDA, DICOM, GigE Vision) have to appear in your resume text. At integrators and plant-side manufacturers, an engineering manager who also owns the line reads it, and willingness to travel to customer sites is a real screen.
The common loop is four to six stages across three to eight weeks. A recruiter screen that checks eligibility, location, salary expectations and the shape of your experience. A hiring manager call that is mostly a project deep dive and is where most rejections happen. A technical exercise, either a live coding session or a take-home dataset task. Then an onsite or virtual panel of three to five rounds: project depth again at greater length, classical geometry and image formation, deep learning and data strategy, deployment and systems, and a behavioural or team-fit conversation. Senior loops add a design round where you architect a perception system out loud, and a cross-functional round with robotics, hardware, product or clinical staff.
The variants matter. Startups compress the loop and weight the take-home or a paired working session heavily; they will often ask you to spend two to four hours on their actual problem, and they decide fast. Large technology companies keep a general algorithmic coding round that has little to do with vision, then test machine learning depth and system design; prepare for both or you will fail the round that is not about your specialty. Primes and defence contractors move slowly, front-load paperwork, and sometimes make an offer contingent on clearance processing, which means a start date months out. Medical device employers add a round with clinical or regulatory staff and ask how you would validate, not just how you would train. Factory-side employers may take you to a line and ask what you would change about the lighting, which is a genuine test and not small talk.
The take-home, when there is one, is almost always the same exercise in different clothes: here is a small, messy, partly mislabelled dataset, build something that detects or segments or classifies the thing, and report your results. Candidates lose this by optimising the headline metric. It is graded on whether you split the data in a way that does not leak (by scene, by video, by site, by patient, never randomly across frames of the same clip), whether you established a trivial baseline before the neural network, whether you looked at the images and said something true about them, whether you chose an operating point and justified it, whether you reported runtime on stated hardware, and whether your README lets someone reproduce it in one command. Write down what you cut for time. Interviewers read that line and respect it.
Live coding in this field has its own flavour. Expect to implement something small in plain NumPy with no library help: intersection over union, non-maximum suppression, a bilinear resize, a 2D convolution, a RANSAC line or plane fit, connected components, a simple Kalman update, or converting between a rotation matrix and a quaternion. These are not trick questions, they are a check that you understand what the library does. Practise writing non-maximum suppression from memory in fifteen minutes; it appears constantly. Larger employers will also ask an ordinary data structures question, so keep that muscle alive even though it has little to do with the job.
On timelines and process hygiene: ask in the recruiter screen how many stages there are, who is on the panel, and whether the exercise is paid or timeboxed. Decline unbounded take-homes politely and offer a two hour version or a walkthrough of existing work instead; good teams accept. If a clearance is involved, ask explicitly whether the offer is contingent on it and what happens to the start date, because the answer determines whether you can afford to wait.
- Startups: the perception lead reads your link personally. Put a clip first, name the hardware, state one number. Loop is fast and take-home heavy.
- Big technology companies: recruiter keyword screen, a general coding round, then machine learning depth and design. Prepare the non-vision round too.
- Defence and primes: eligibility first, slow paperwork, clearance timelines in months. Ask whether the offer is contingent and when the start date lands.
- Medical device: expect a validation and regulatory conversation, and questions about patient-level and site-level splits.
- Industrial and integrators: travel and uptime come up. A site visit or a live line question is normal and is a real test.
- Take-home grading is about the split, the baseline, the error analysis, the operating point and the runtime. Not the headline metric.
- Always state the hardware and the time taken in a take-home report. A table with latency on a named device separates you from most submissions.
- Live coding: non-maximum suppression, intersection over union, bilinear resize, convolution, RANSAC fit, Kalman update, rotation conversions. Write them from scratch.
- Ask what the current system's error rate and latency are. The answer tells you whether the team ships, and interviewers mark the question as senior.
- Expect three to eight weeks commercially, and longer with a clearance. Run applications in batches so offers arrive close together.
The portfolio: what a demo-able project actually has to show
Computer vision is one of the few engineering fields where a portfolio is read rather than skimmed, because the output is visual and a hiring manager can evaluate it in half a minute without reading your code. That cuts both ways. A project showing a working system on real footage is the strongest application artefact available to you, and a notebook with a confusion matrix and no video is nearly worthless. Build for the thirty second test: can the person deciding whether to interview you see the thing working, on data that looks like the real world, with the failure cases shown honestly.
A project that gets replies has six parts. A clip at the top of the README, under a minute, no music, showing the system running on video rather than hand-picked stills. A one-paragraph statement of the decision it supports and who would act on it. A metric table with the split named and justified, and the operating point chosen rather than the best number found. A latency and hardware line, in frames per second and milliseconds, with the device, the precision and ideally the power draw. A failure section with three examples and a sentence each about cause. And a reproduction command that works on a clean machine. Nothing else is required, and adding more usually dilutes it.
Getting data is the step people skip, so be concrete about it. Your phone shoots usable video; a used GigE or USB3 industrial camera and a fixed lens cost less than a monitor on the secondary market; a cheap depth camera gives you point clouds; and the thing you are detecting can be something you own and can damage on purpose. Label it in CVAT or Label Studio, keep the dataset in FiftyOne so you can actually look at errors in bulk, and use public datasets for pretraining or comparison rather than as the project itself.
Choose the project to match the sector you are aiming at. For inspection: pick a defect you can create at home (bad 3D printer layers, cracked solder on a scrap board, mis-seated caps on bottles, damaged produce), collect a few thousand images with controlled lighting, deliberately keep defects rare so the problem is realistic, train an anomaly detector on good samples plus a small supervised comparison, and report escapes against false rejects at two thresholds with the cost argument for each. Then deploy it to a board at line rate and publish the throughput. That project demonstrates exactly what an inspection employer needs and almost no applicants submit it.
For robotics perception: calibrate a cheap camera properly with a ChArUco board and publish the intrinsics plus a reprojection error in pixels, then do 6-DoF pose estimation of a known object and output a grasp or approach point, wrap it in a ROS 2 node, and run it for an hour while logging drift. If you have a second camera or a depth sensor, show the extrinsic calibration and what happens to the pose error when the baseline changes. Add one honest limitation: at what distance does your depth error exceed the tolerance a gripper needs, and why stereo depth error grows with the square of range. That single paragraph signals geometry competence more clearly than any benchmark.
For the modelling and vision-language side: build a distillation pipeline. Use a strong vision-language model or an open-vocabulary detector to auto-label a corpus, add a human verification pass on a sample so you can report label quality, train a small student model, and publish the tradeoff table: teacher accuracy and cost per thousand frames against student accuracy, latency and cost on a board. This is the dominant production pattern now, and demonstrating it end to end, including what the auto-labels got wrong and how you caught it, is directly relevant to most teams' current work.
What kills a portfolio is specific and repeated. Training a stock detector on a public dataset with one command and reporting the number the tutorial produced. A hosted demo link that is dead when clicked. Claims with no units. Only still images. No failure cases, which reads as either dishonest or inexperienced. And one commercial trap: several popular detector packages ship under copyleft licences such as AGPL-3.0, which matters the moment you describe your work as a product or a company uses your repo as a template, so check the current licence terms of anything you build on and state which licence you used. Interviewers at product companies notice when a candidate knows that.
On presentation: host on GitHub, keep the README as the entry point, put the clip in the first screen of text, and write one page about the decisions rather than a research paper. If you can afford one more artefact, write a short technical post about the thing that surprised you and link it from the resume. Hiring managers in this field read those, and a post about why a model lost recall after INT8 quantisation and what you did about it is more persuasive than any summary of your skills.
- Clip first, under a minute, real video, failures included. The thirty second test decides whether anything else gets read.
- Name the split and defend it. Random frame splits across the same video are the most common way candidates fool themselves.
- Report at an operating point you chose, not the best point you found. Say what the threshold costs in the other direction.
- Always publish hardware, precision and throughput together: resolution, precision, frames per second, end-to-end milliseconds, device, and watts if you can measure them.
- Three failure cases with a cause each. Interviewers trust a candidate who shows the breaking point more than one who hides it.
- One command to reproduce, tested on a clean machine. A broken setup script is a quiet rejection.
- Check the licence of anything you build a product story on, including popular detector repositories, and state it.
- Collect your own data for at least one project. Public-dataset-only portfolios all look the same and none of them show you can handle a camera.
- Label in CVAT or Label Studio, browse errors in FiftyOne, and show one screenshot of the label noise you found.
- A short writeup about one hard bug (preprocessing mismatch, quantisation loss, timestamp skew) is worth more than three more demos.
What the interview really tests, round by round
The project deep dive is where most candidates are lost, and it is also the most predictable round. Pick one project and prepare it at three depths: sixty seconds, five minutes, and thirty minutes with a whiteboard. The interviewer will push on five things in roughly this order: what decision the system supported and what the error cost, how you built the evaluation set and whether it leaked, what the dominant failure mode was and how you found it, what the simpler approach was and why you did not take it, and what you would do with ten times the data or ten times the compute. Have a number for every claim and a limitation you volunteer before you are asked. If your honest answer is that you inherited the evaluation set, say that and then say what was wrong with it.
The geometry round is the filter that separates computer vision engineers from general machine learning engineers, and it is pure fundamentals. Expect to write the pinhole projection, build the intrinsic matrix, and say what happens to the focal length and principal point when you resize, crop or letterbox an image, because that is a real production bug and interviewers know it. Expect distortion models, planar calibration with a checkerboard or ChArUco target, what a good reprojection error looks like in pixels and what a suspiciously low one implies, the relationship between homography, essential and fundamental matrices and when each applies, PnP with RANSAC, stereo baseline against depth error, triangulation, and a transform chain where you have to compose and invert poses between named frames without losing the convention. Practise saying the pose of the camera in the world frame rather than camera to world, because the sloppy phrasing is where errors hide.
The deep learning round is less about architectures than candidates expect and more about diagnosis. Common ground: loss selection and class imbalance when the positive class is one in ten thousand, anchor-based against query-based detection and what non-maximum suppression does to crowded scenes (and why query-based detectors drop it), augmentations that silently break the task (horizontal flips on an asymmetric defect, or on a medical finding where left and right mean different things, or colour jitter when the decision is a colour), batch normalisation at small batch sizes, what a train and validation gap tells you against what it does not, label noise and how you would quantify it, confidence calibration, and the question that comes up constantly: mean average precision went up and the product got worse, what happened. Be able to answer that one crisply, because it is the whole discipline in one question.
The data strategy round is a scenario. You have three hundred defective parts and two million good ones. You have fifty labelled surgical videos and a thousand unlabelled ones. You have one site's data and need to ship to fourteen. A strong answer moves through: what decision and what error asymmetry, what the evaluation set must look like before you collect anything else (held out by site, lot, patient or scene), a baseline that does not require labels, pretrained features or a foundation model as a starting point, where the annotation budget goes first, active learning or hard negative mining to spend it well, synthetic data and where it stops helping, and the point at which the right answer is to change the sensor or the lighting instead of the model. Candidates who jump straight to an architecture fail this round even when the architecture is correct.
The deployment round wants a budget that adds up. You will be given a constraint (30 frames per second on a mid-range embedded module, under 10 watts, no internet, six cameras) and asked to make the numbers work. Walk the full path: exposure and sensor readout, transfer over the interface, colour conversion and resize, inference, post-processing including non-maximum suppression and tracking, then the action. Know the standard levers and their costs: FP16 and INT8 quantisation and what each can cost in recall, quantisation-aware training when post-training quantisation loses too much, resolution against small-object recall, frame skipping with tracking to fill the gaps, batching and why it helps throughput and hurts latency, pruning, running the detector on a region of interest from a cheap first stage, and moving preprocessing onto the hardware encoder or the image signal processor. Say which number you would measure first and on what device.
There is usually a debugging round, and it is often the one that earns the offer because almost nobody prepares for it. The prompt: it works in my notebook and fails on the robot. Enumerate causes in order of likelihood, out loud, as a checklist: preprocessing mismatch (channel order, resize interpolation, normalisation constants, letterbox padding), a different camera or lens or exposure than the training data, automatic white balance or gain changing between sessions, timestamp misalignment between sensors, resolution and aspect ratio handling, quantisation loss, thermal throttling, a threshold tuned on the wrong distribution, and only then genuine distribution shift. Then say how you would isolate it: run the deployed binary on a training frame and compare tensors numerically. That answer, delivered calmly, reads as someone who has shipped.
Senior candidates get a design round and a judgement round. The design round is a system: design perception for a shelf-scanning robot, a weld inspection cell, a drone that finds corrosion, a triage tool for chest imaging. Structure it as requirements and error costs, sensor and optics choice, the data plan, the model and the fallback, the evaluation and its regression suite, the deployment budget, the monitoring that tells you it has drifted, and the human in the loop. The judgement round is one question in different forms: what did you kill, or ship late, or refuse to claim. Have a true story where you declined to state a number you could not defend, because in this field that is the most valuable professional habit you can demonstrate.
- Prepare one project at sixty seconds, five minutes and thirty minutes. Volunteer the limitation before the interviewer finds it.
- Be able to write the intrinsic matrix and adjust it for a crop, a resize and a letterbox. This is asked because it is a real bug.
- Know why stereo depth error grows with the square of range, and what that means for a gripper tolerance.
- Compose and invert transforms between named frames without slipping conventions. Say the full phrase every time.
- Have a crisp answer for: average precision improved and the product got worse. Operating point, class mix, evaluation leakage, small objects, or the metric was never the product metric.
- Name augmentations that break a specific task. Flips on asymmetric defects and on laterality in medical images are the canonical examples.
- In data strategy, design the evaluation set before the training set, and split by site, lot, patient or scene.
- Keep a deployment budget in your head: readout, transfer, preprocess, inference, post-process, act. Put milliseconds on each.
- Memorise the deployment debugging checklist, preprocessing mismatch first. It wins rounds.
- For INT8 and FP16, speak in measured tradeoffs on a device you used rather than in general claims.
- In design rounds, start with error cost and end with monitoring and the human fallback. Skipping either reads as junior.
- Have a story about refusing to claim a number. It is the most senior thing you can say in this field.
- Ask the interviewer what their current error rate and latency are, and what the failure mode they cannot fix is. Their answer tells you whether to take the job.
The resume, and what gets ignored
One page up to roughly ten years of experience, two if you have publications and patents worth listing. The top of the page carries three things: the sector you work in, the stack in exact strings a recruiter will search for, and if applicable a clearance line stated factually. Then experience, in bullets that each name a decision the system made, a measured result at a stated operating point, and the hardware or constraint it ran under. Then one portfolio link that is live. Education and publications last. No skills grid with forty logos, no profile paragraph of adjectives, no photograph.
The highest-leverage change most candidates can make is replacing capability bullets with measured ones. The pairs below are the shape to copy, not numbers to borrow: use your own, and only ones you can reconstruct. Before: built a deep learning model for defect detection using PyTorch and OpenCV. After: cut escapes on a high-speed capping line by an order of magnitude at a 2 percent false reject rate, with an anomaly model trained on good samples, running INT8 on a Jetson Orin NX inside the line's cycle time, with a PLC reject handshake. Before: worked on object detection for autonomous robots. After: raised pallet pocket detection recall by nine points on 21,000 held-out frames from six sites absent from training, by fixing a lens distortion mismatch and mining hard negatives, holding end-to-end latency under 20 ms on an Orin Nano.
The units that make a computer vision resume legible: frames per second and milliseconds for throughput and latency; watts for a power budget; recall at a fixed false positive rate per image or per hour, which is more useful than mean average precision alone; escapes and false rejects for inspection; millimetres and degrees for pose error; pixels for reprojection error; sensitivity at a fixed specificity for clinical work; absolute trajectory error for localisation; dataset sizes and how the held-out set was separated; number of sites, devices or patients generalised across; and cost per thousand frames if you ran anything through a hosted model. Every number should be one somebody could ask you to defend, and you should want them to.
What gets ignored, consistently. Course lists and specialisation certificates. Projects on MNIST, CIFAR, Fashion-MNIST or any tutorial dataset. Phrases like proficient in OpenCV, which tells the reader nothing. A wall of library names with no result attached. GPA after a few years of work. Papers you read rather than wrote. Summaries written in the third person. Long descriptions of team process. And the generic claim of experience with computer vision and machine learning, which appears on nearly every resume in the pile and therefore distinguishes none of them.
Applicant tracking reality in this field is simpler than in most. Recruiters and keyword filters search for exact tool and format strings, so the ones that apply to you must appear in prose rather than only in a graphic: PyTorch, CUDA, TensorRT, ONNX, OpenCV, C++, Python, ROS 2, Jetson, GStreamer, DICOM, GigE Vision, point cloud, calibration, segmentation, object detection, tracking, quantisation. Write them inside bullets where they are doing work rather than in a keyword block, which a human reader discounts.
Publications and patents: one line each, with venue and year, newest first, and no abstracts. If you have a highly cited paper, say the citation count once; otherwise leave it out. For research titles, a short list of first-author work beats a long list of middle-author work. For product titles, one line saying you publish is enough, and the deployment bullets matter more.
A cover letter is optional almost everywhere and useful in exactly one form: 150 words that name the product, name the specific perception problem you believe it has, say which of your projects maps to it, and link the clip. That letter gets read at startups because it proves you looked at what they build. A letter that could be sent to any company is worse than no letter.
- Lead each bullet with the decision and the number, not with the framework.
- Prefer recall at a fixed false positive rate over a bare mean average precision figure. It shows you think in operating points.
- Name the device, the precision and the throughput for anything you deployed. That trio is the senior signal in this field.
- State the held-out set and what separated it: sites, lots, patients, scenes, devices, time periods.
- Put a clearance line near the top when relevant, stated plainly with level and status, and nothing more.
- One live portfolio link. Check it the day you apply. A dead link reads worse than none.
- Delete MNIST, CIFAR and tutorial projects entirely. They actively cost you credibility.
- Write tool strings into bullets so a keyword search finds them in context rather than in a logo grid.
- Publications and patents: one line each, no abstracts, newest first.
- A 150 word letter naming the company's actual perception problem works at startups. A generic letter does not.
Where to look, how to reach the person who decides, and a 12-week plan
A large share of computer vision roles never appear on the big aggregators, because robotics, defence and industrial employers post to their own career pages and to trade channels. Search company pages directly once you have a sector list. Beyond that: the ROS Discourse jobs area and robotics community boards for perception roles; startup job boards attached to accelerators for early teams; the recruiting sessions and job boards attached to CVPR, ICRA, IROS, NeurIPS and MICCAI, which are genuinely effective because hiring managers staff them; machine vision trade events and association member directories for inspection work; radiology and informatics conferences for medical imaging; and the partner and integrator lists published by camera, sensor and edge silicon vendors, which are effectively a map of every company doing this work in your region.
Search strings that surface the real roles: perception engineer, computer vision engineer, machine vision engineer, image processing engineer, deep learning engineer vision, edge AI engineer, embedded vision, AI engineer robotics, visual inspection engineer, imaging scientist, imaging algorithms engineer, GEOINT or imagery developer, medical imaging AI engineer, and sensor fusion engineer. Add a hardware or format noun to filter by sector: Jetson, TensorRT, ROS 2, GigE Vision, DICOM, SAR, point cloud. Also search the sector's product words rather than the job words, because a bin-picking company sometimes posts a title you would never guess.
The highest-yield move in this field is direct contact, and it works because the output is visual. Find the perception lead or the engineering manager, send five sentences: what you built, the number it achieved, the hardware it ran on, a link to the clip, and one specific sentence about their problem. No attachment, no paragraph about passion. A clip of a working system is self-verifying and almost nobody sends one. Referrals matter too, and the easiest source is people you overlap with at conferences and in open source issue threads rather than cold network requests.
Do not neglect the internal route, which fills a large fraction of these jobs. A controls or manufacturing engineer moving into machine vision, a software engineer moving into the perception team, a data scientist moving into imaging, an embedded engineer moving into edge inference, a radiology technologist or laboratory scientist moving into an imaging annotation and validation role: all of these happen regularly and all of them are easier than an external hire, because the employer already knows you and you already know the domain, which is the half external candidates lack. If you are at a company with a camera anywhere in its product or its factory, start there.
A twelve week plan that assumes a job and limited evenings. Weeks one and two: pick one sector, read thirty real postings in it, extract the nouns into a list, write your resume against that list, and buy one embedded board. Weeks three to six: build project one end to end, collecting your own data, with an evaluation set split honestly, a chosen operating point, and a clip. Weeks seven and eight: deploy it, quantise it, measure frames per second and watts on the board, write the one-page report including what quantisation cost you, and publish both. Weeks nine and ten: drill interviews. Write non-maximum suppression and intersection over union from scratch until it is boring, calibrate a camera twice and report reprojection error, derive the projection equation on paper, work ten deployment debugging scenarios out loud, and rehearse the project pitch at three lengths. Weeks eleven and twelve: apply in one batch of twenty-five tailored applications plus eight direct emails to named perception leads, and keep a second batch ready for the rejections.
Realistic timelines by starting point, stated without optimism. From an adjacent engineering job (software, embedded, controls, data) with a strong project: typically two to five months of searching. From a machine learning job without vision specifics: similar, but spend the first month on geometry and one month on a deployment project, because those are the two gaps the loop finds. From a fresh master's with a real thesis project and one deployed demo: three to six months, faster if you target a sector with a shorter queue. From self-taught with no degree in a numerate field: expect longer, aim first at inspection, integrators and smaller product companies where a working demo on a board is close to the whole evaluation, and plan on two projects rather than one. With a clearance already in hand and US eligibility: faster than any of the above, often weeks to first interviews.
On pay, check rather than guess. The nearest authoritative United States baselines are the BLS Occupational Employment and Wage Statistics tables for 15-1252 Software Developers, 15-2051 Data Scientists, and 15-1221 Computer and Information Research Scientists, each of which publishes metro-level figures that matter far more than the national number for a field this geographically concentrated. For live market data, read the posted ranges in jurisdictions that require them, and compare three employers in the same metro and sector rather than across the whole field, because defence, medical device, industrial and frontier-lab compensation structures differ enough that a single average misleads. Equity at robotics startups and bonus structures at primes change the comparison more than base salary does, so ask for the full structure before comparing two offers.
- Search company career pages directly for robotics, defence and industrial employers. Many never post anywhere else.
- Conference recruiting (CVPR, ICRA, IROS, NeurIPS, MICCAI) and machine vision trade events are staffed by hiring managers, not recruiters.
- Vendor partner and integrator directories from camera and edge silicon suppliers are a map of every local company doing this work.
- Send five sentences and a clip to the perception lead. Build, number, hardware, link, one specific sentence about their problem.
- Internal moves fill many of these roles. Controls, embedded, software, data and imaging technologist backgrounds all transfer.
- Twelve weeks: two weeks of posting research and resume, four weeks building, two weeks deploying and measuring, two weeks drilling, two weeks applying in batch.
- Apply in batches of twenty-five so offers arrive within a window you can negotiate across.
- For pay, use BLS OES metro tables plus posted ranges in states that require them, and compare within sector and metro rather than across the field.
What a computer vision engineer must know about AI in 2026-27
This role's version of the AI question is not whether to learn AI. It is a calibration question, and hiring managers ask it directly: what do foundation models actually solve in vision now, and what do they not. Get it wrong in either direction and you fail the round. The candidate who says nothing has changed is unemployable, and the candidate who says a vision-language model can run the inspection line has never measured a latency budget.
What genuinely changed, and will still be true through 2027. Generic perception became cheap. Promptable segmentation (the Segment Anything family, including its video version) and open-vocabulary detectors mean a usable mask or box for a common object no longer needs a bespoke dataset, with one catch worth saying out loud in an interview: a promptable segmenter returns regions, not labels, so you still need a detector, a classifier or a prompt that names the thing. Self-supervised backbones of the DINO lineage produce features strong enough that a linear probe or a small head on frozen features is a serious baseline, and often the right production answer for a niche class with few labels. Monocular depth models turned relative depth from one cheap camera into something usable, while metric scale remains the hard part, so claiming a monocular metric depth to millimetres is still a red flag. Vision-language models absorbed most document understanding, OCR, captioning and visual question answering, so building a bespoke receipt parser is usually the wrong decision now, though dense tables and small print still need a validation step and sometimes a specialist model. Feed-forward reconstruction models and 3D Gaussian splatting changed how quickly you get geometry out of captures and made synthetic data from real scenes practical. And the biggest structural change: the dominant production pattern is teacher and student, where a large model or an ensemble labels data and a small, fast, quantised model ships.
What did not change, which is where you get paid. Calibration, coordinate frames, synchronisation and timestamps are exactly as they were, and foundation models have nothing to say about them. Optics and lighting still determine whether a defect is visible, and still beat model changes as the cheapest fix. Evaluation design, splitting and leakage still decide whether your number means anything, and auto-labelling made label quality a sharper problem rather than a softer one, because a systematic teacher error propagates silently through a whole dataset. Latency, power, memory and cost budgets on embedded hardware are unmoved, and a model that needs a data centre is not an answer when the decision has to be made in 20 milliseconds on 8 watts with no network. Rare-defect and long-tail problems remain unsolved by scale, because the thing you care about is by definition thin in the pretraining distribution. And accountability did not change: when the robot drops the box, a person has to explain why, and no model output explains itself.
Say the honest thing about where the hype outran the field. Vision-language-action models for robotics are a real and fast-moving research direction, and a small number of teams run them in production for constrained tasks, but most shipping robots today still run classical perception with learned components inside it, and claiming otherwise to a perception lead will cost you. Likewise, running a general multimodal model on every frame is almost never the production answer at industrial frame rates; it is a labelling tool, a cold-start tool, a fallback for the rare ambiguous case, and sometimes a reasoning layer sitting on top of fast perception. The correct phrasing in an interview is that foundation models moved the floor and the baseline, so the engineering value migrated to data, evaluation and deployment. That sentence, with a measurement behind it, is the answer to the calibration question.
One practical consequence for your career: the skills that commoditised are the ones that used to be the resume. Fine-tuning a detector on a labelled dataset is no longer a differentiator, because the workflow is a command. The skills that appreciated are the ones a model cannot do for you: deciding what to measure, designing the held-out set, finding the label noise, choosing the operating point against a cost asymmetry, closing the gap between a notebook and a board, and knowing when the right answer is a different lens rather than a different network. Build the portfolio around those and you are building toward where the budget is.
Using foundation models as a data engine, then distilling into something that ships
This is the dominant production pattern in computer vision now, and most teams are either doing it or want to. A large open-vocabulary detector, a promptable segmenter or a vision-language model labels a corpus at a cost per thousand frames; a human verifies a sample; a small student model trains on the result and runs quantised on the device at a latency the product can afford. Teams need someone who can run this loop without poisoning the dataset, because a teacher's systematic bias becomes a student's permanent blind spot, and a candidate who has only ever trained on human labels will not see it coming.
Show it: Build one distillation project end to end and publish the table: teacher accuracy and cost per thousand frames, human agreement on the verification sample, student accuracy, student latency and power on a named board. Say explicitly what the teacher got wrong (small objects, unusual lighting, a class it conflated) and how you detected it, which should involve a human-labelled gold set the teacher never touched. In an interview, be able to explain why you verify a sample rather than trusting the teacher, and how you would split the annotation budget between teacher output and human labels.
Knowing where promptable and open-vocabulary models replace a dataset, and where they do not
The practical decision that saves a team weeks is whether a problem needs a bespoke labelled dataset at all. Common objects, a rough mask, a cold start, a demo for a customer next Tuesday: promptable segmentation and open-vocabulary detection get you there in hours. A rare defect class specific to one plant, a part your customer manufactures, a clinical finding, anything where the decision boundary is subtle and the vocabulary is not in the pretraining data: those still need collected data, careful labelling and a trained model. Getting this call right is an interview question in disguise, because it reveals whether you have shipped or only read.
Show it: Have a concrete pair of examples from your own work: a task you solved zero-shot that would previously have taken a month, and a task where zero-shot failed and you can say why with a number attached. Say that zero-shot output is a labelling accelerant and a baseline rather than a product, and name the check you ran to decide. In a take-home, use a promptable or open-vocabulary model to build a baseline fast, say so in the report, and spend the saved time on error analysis, which is exactly the judgement the grader is looking for.
Frozen pretrained features as the first serious baseline
With strong self-supervised backbones available, a linear probe or a small head on frozen features often reaches most of the achievable accuracy in an afternoon with a few hundred labels, and sometimes it is the right thing to ship because it is stable, cheap to retrain and easy to reason about. Candidates who skip straight to full fine-tuning on a thousand images waste a week and often overfit. Interviewers in low-label domains (medical imaging, industrial inspection, defence) probe for this instinct specifically, because labels are the scarce resource in all three and a candidate who burns the annotation budget before trying a frozen-feature baseline is expensive to employ.
Show it: In any project with few labels, report the frozen-feature baseline alongside the fine-tuned model: accuracy, number of labels used, training time, and how the gap closed as you added labels. That curve is the artefact, because it shows where the next hundred labels are worth more than the next architecture change. In an interview, be able to say when you would not freeze (a large domain gap from the pretraining data, such as thermal, radar, microscopy or some medical modalities) and what you would try instead, which is usually partial unfreezing or continued pretraining on unlabelled in-domain images.
Making it fit: quantisation, runtime and the budget that has to add up
The constraint foundation models did not relax is the device. Most hiring in robotics, inspection, defence and consumer hardware ends with a model running on an embedded module inside a latency, power, thermal and bill-of-materials budget. Teams interview for this specifically because it is where projects die: a model that was accurate in the notebook loses recall on small objects after INT8 conversion, the pipeline drops frames because preprocessing is on the CPU, the board throttles in a sealed enclosure. This is the skill that most clearly separates a candidate who has deployed from one who has only trained.
Show it: Deploy something yourself and publish the numbers: model, input resolution, precision, frames per second, end-to-end milliseconds including pre and post-processing, watts, device. Then publish what quantisation cost you, by class if possible, and what you did about it (quantisation-aware training, keeping a sensitive layer in higher precision, raising input resolution, a two-stage region-of-interest crop). In the interview, build the latency budget out loud from readout to action, and name which number you would measure first on real hardware rather than estimate.
Evaluation and label quality in a world of generated labels
Everything above makes evaluation the load-bearing skill. When labels can come from a model, the held-out set is the only thing standing between a team and a confident wrong number, and the classic failure (splitting frames randomly across the same video, or across the same patient, lot or site) is now compounded by teacher-generated labels that correlate their own mistakes. Employers in regulated and high-cost-of-error settings screen for this hard, because the whole regulatory and safety argument rests on the evaluation being honest.
Show it: In any project or take-home, state the split and the unit of independence first (scene, video, session, site, patient, production lot, device) and say what would leak under a naive split. Keep a small human-labelled gold set no model ever touched and report on it separately. Quantify label noise rather than asserting it: relabel a sample, report agreement, and show how much of your error bar is annotation disagreement. Then choose an operating point by cost rather than by F1 and say what it costs in the other direction.
Honest calibration about robotics foundation models and per-frame multimodal inference
Perception leads use this as a judgement test. Vision-language-action models and general multimodal models are genuinely important and genuinely not yet the answer for most real-time, power-constrained, high-consequence perception. A candidate who overclaims reads as someone who follows announcements; a candidate who dismisses the direction reads as someone who stopped learning. The hireable position is specific: here is what I would use a large model for today, here is what I would not, here is the measurement that would change my mind.
Show it: Be able to name where a large multimodal model earns its latency in a system (labelling, cold start, the rare ambiguous case escalated from a fast path, a supervisory or reasoning layer, operator-facing summarisation) and where it does not (every frame at line rate). Bring measurements you took yourself: what a hosted multimodal call cost you per thousand frames and in milliseconds, against your student model on a board. Then say what evidence would make you move the boundary, which is the sentence that makes the whole answer credible.
What a screen is looking for
These are the terms that a resume screen, human or automated, is matching against for this role. Use the ones that are true of you, in the words the posting uses.
- Computer vision engineer
- Perception engineer
- Machine vision engineer
- Image processing engineer
- Embedded vision
- Edge AI engineer
- Imaging algorithms engineer
- Sensor fusion engineer
- Python
- C++
- NumPy
- OpenCV
- PyTorch
- CUDA
- TensorRT
- ONNX Runtime
- OpenVINO
- LiteRT
- DeepStream
- GStreamer
- Open3D
- Object detection
- Instance segmentation
- Semantic segmentation
- Multi-object tracking
- Optical flow
- Anomaly detection
- Open-vocabulary detection
- Promptable segmentation
- Segment Anything (SAM, SAM 2)
- Grounding DINO
- DINOv2
- CLIP
- Vision transformer (ViT)
- YOLO
- RT-DETR
- Mask R-CNN
- U-Net
- nnU-Net
- MONAI
- Vision-language model (VLM)
- Vision-language-action (VLA)
- Knowledge distillation
- Pseudo-labelling
- Self-supervised learning
- Linear probing
- Active learning
- Hard negative mining
- Synthetic data generation
- Domain adaptation
- Label noise
- Dataset curation
- FiftyOne
- CVAT
- Label Studio
- Camera calibration
- Intrinsic parameters
- Extrinsic calibration
- Hand-eye calibration
- Lens distortion
- ChArUco
- Reprojection error
- Pinhole camera model
- Epipolar geometry
- Homography
- Essential matrix
- PnP
- RANSAC
- Triangulation
- Bundle adjustment
- Stereo vision
- Monocular depth
- LiDAR
- Point cloud registration
- 6-DoF pose estimation
- Structure from motion
- SLAM
- Visual inertial odometry
- 3D Gaussian splatting
- Kalman filter
- Sensor fusion
- Time synchronisation
- Coordinate frame transforms
- Quaternions
- ROS 2
- rosbag
- Isaac ROS
- Isaac Sim
- NVIDIA Jetson
- Jetson Orin
- Hailo
- INT8 quantisation
- FP16
- Quantisation-aware training
- Model pruning
- Real-time inference
- Frames per second (fps)
- Power budget
- SWaP
- GigE Vision
- GenICam
- Machine vision lighting
- Telecentric lens
- Line scan camera
- Global shutter
- Image signal processing (ISP)
- Cognex
- Keyence
- MVTec HALCON
- PLC integration
- False reject rate
- Escape rate
- Defect detection
- Visual inspection
- OCR
- Document understanding
- DICOM
- NIfTI
- PACS
- Whole slide imaging
- Digital pathology
- Medical image segmentation
- Reader study
- Sensitivity and specificity
- Operating point selection
- Precision-recall curve
- mAP
- IoU
- Dice coefficient
- Non-maximum suppression
- Model calibration
- Error analysis
- Data leakage
- Cross-site validation
- Regression testing
- Drift detection
- Weights & Biases
- Docker
- Linux
- CMake
- Geospatial imagery
- Synthetic aperture radar (SAR)
- Electro-optical and infrared (EO/IR)
- Change detection
- GEOINT
- ITAR
- Security clearance
- ISO 26262
- ISO 21448
- ISO 13485
- IEC 62304
- Design controls
- Verification and validation
Mistakes that cost people this job
Submitting a portfolio of public-dataset notebooks with no video, no hardware and no failure cases.
Build one project where you collected the data yourself, show it running on real video in a clip under a minute, publish a metric table with the split named, and state frames per second and watts on a specific board. One project like that outperforms six notebooks, because a hiring manager can verify it in half a minute without reading your code.
Splitting frames randomly across the same videos, patients, lots or sites, then reporting the number.
Split by the unit of independence before you train anything: scene, clip, session, site, production lot, patient, device. Say in your report what would leak under a naive split and in which direction the number would move. If you inherited a contaminated evaluation set at work, say so in the interview and say what you did about it. Interviewers trust that answer more than a clean-sounding number.
Reporting mean average precision alone and treating it as the product metric.
Choose an operating point for a stated cost asymmetry and report recall at a fixed false positive rate, or escapes against false rejects, or sensitivity at a fixed specificity. Then say what the threshold costs in the other direction and who pays that cost. Average precision is a model comparison tool; the business runs at one threshold.
Chasing a newer architecture when the problem is the data, the labels or the lighting.
Look at the images first, in bulk, including every error the current system makes. The highest-value fixes in production computer vision are usually a relabelled subset, a changed lens or light, a corrected preprocessing mismatch, or a few thousand mined hard negatives. Being the candidate who says I changed the backlight and the problem went away is a strong signal, not a weak one.
Interviewing for a perception job without being able to do the geometry out loud.
Practise until it is automatic: write the projection equation and the intrinsic matrix, adjust them for a crop, a resize and a letterbox, explain distortion and planar calibration, say what a credible reprojection error is in pixels, compose and invert transforms between named frames, and explain why stereo depth error grows with the square of range. This round is the filter between a computer vision engineer and a general machine learning engineer, and no amount of model knowledge substitutes for it.
Claiming a general vision-language model as the production answer at industrial frame rates.
Position large multimodal models where they actually pay: auto-labelling, cold start, the rare ambiguous case escalated from a fast path, operator-facing summarisation, and evaluation. Then give the per-frame latency and the cost per thousand frames you measured, against a small quantised student on a board. Perception leads use this exact question to sort candidates who have shipped from candidates who have read.
Saying nothing has really changed because the fundamentals still matter.
Name what changed and what it did to your workflow: promptable segmentation and open-vocabulary detection removed the bespoke dataset from many generic tasks, strong self-supervised features made a frozen-feature baseline serious, monocular depth models made one cheap camera useful for relative depth, vision-language models absorbed most document understanding, and teacher-student distillation became the default pipeline. Then say what stayed: calibration, optics, evaluation, budgets, the long tail. Both halves are required.
Ignoring deployment entirely and treating the notebook as the deliverable.
Put something on a device. Measure end-to-end latency with pre and post-processing included, measure power, measure what INT8 cost you by class, and learn the debugging checklist that starts with preprocessing mismatch. Deployment experience is the clearest differentiator in this market and a current developer board costs a fraction of a week's pay.
Applying for months into export-controlled defence and aerospace postings without being eligible.
Read the eligibility line before you apply. If a posting requires US person status or an active clearance and you do not have it, that is a legal wall a recruiter cannot waive, so redirect the effort into commercial robotics, inspection, medical imaging, logistics or media, which are fully open. If you are eligible, do the opposite and lean into it, because that sector interviews fast and pays for eligibility.
Sending the same resume to a perception role, a research role and a machine vision role.
Write three versions. Perception leads with calibration, frames, tracking and real-time integration. Research leads with datasets, ablations, benchmarks and publications. Machine vision leads with optics, lighting, line rate, uptime and the PLC interface. The underlying work may be the same; the first six lines have to speak the sector's language or the reader concludes you are applying everywhere.
Overclaiming a number you cannot reconstruct.
For every figure on your resume, be able to state the dataset, the split, the threshold, the hardware and the date. If you cannot, remove the number or weaken the claim to what you can defend. Being caught unable to explain your own metric ends a loop immediately, and volunteering the limitation before you are asked is the most reliable way to read as senior in this field.
Questions people ask
What does a computer vision engineer actually do?
A computer vision engineer builds the path from a camera or other imaging sensor to a decision a machine or a person acts on, and is accountable for the error rate of that decision at a chosen operating point, inside a fixed latency, power and cost budget. In practice a computer vision engineer spends most of the week on data and labels, evaluation design, calibration and integration, and a minority of it on training models. Typical outputs are a detector or segmenter running on an embedded board at a measured frame rate, a calibration procedure with a reprojection error somebody can check, an evaluation suite that catches regressions before a release, and a written argument about why the chosen threshold is right given what a miss and a false alarm each cost.
Do you need a PhD to become a computer vision engineer?
No. A computer vision engineer working on products does not need a PhD, and most do not have one: a bachelor's degree in a numerate field plus a portfolio with measured results clears the majority of postings, and a master's degree is the modal qualification. A PhD is effectively expected only where the title says research or scientist, where the job is publishing or inventing new methods, and at a small number of labs that use it as a filter. What the role does require, degree or not, is working knowledge of projective geometry, linear algebra and probability, because the geometry interview round cannot be passed with framework familiarity alone.
Which industries are hiring computer vision engineers in 2026 and 2027?
Computer vision engineer hiring has concentrated in six places: robotics and warehouse automation (piece picking, mobile robots, agricultural and construction machines, surgical and lab robotics), defence, dual-use and geospatial work, manufacturing and semiconductor inspection, medical imaging and digital pathology, autonomous trucking and driver assistance rather than driverless taxis, and drone-based inspection of infrastructure, energy and crops. Smaller but real pools include sports and broadcast analytics, document understanding, generative media companies that need data and evaluation people, and spatial computing. Consumer photo and video features and checkout-free retail hire a computer vision engineer less often than they did, because general-purpose models absorbed the first and the economics never worked at scale for the second.
What should a computer vision portfolio contain?
A computer vision engineer's portfolio should contain one project carried end to end rather than several partial ones, and it has six parts: a clip under a minute showing the system running on real video including failures, a paragraph stating what decision it supports, a metric table with the split named and justified, a latency and hardware line (frames per second, milliseconds end to end, precision, device, ideally watts), three failure cases with a cause each, and a one-command reproduction that works on a clean machine. Collect your own data for at least one project, because public-dataset-only portfolios all look alike and none of them demonstrate that a computer vision engineer can handle a camera, a lens and uncontrolled lighting.
Have vision-language models made computer vision engineers obsolete?
No, but they moved what a computer vision engineer is paid for. Promptable segmentation, open-vocabulary detection, strong self-supervised features, usable relative monocular depth and vision-language models mean a generic perception task now has a decent baseline in an afternoon with no bespoke dataset, so fine-tuning a detector on labelled data is no longer a differentiator. The work that appreciated is the work those models do not do: deciding what to measure, designing an evaluation set that does not leak, finding and quantifying label noise (including noise a teacher model generated), choosing an operating point against a real cost asymmetry, and getting a model inside a latency, power and cost budget on hardware. A computer vision engineer who can run a teacher-student distillation loop and publish the tradeoff table is in more demand now than one who could only train from human labels.
What do computer vision interviews actually test?
A computer vision engineer's interview loop usually has six distinguishable tests: a project deep dive that pushes on the metric, the split, the dominant failure mode and the simpler alternative you rejected; a classical geometry round covering projection, intrinsics under crop, resize and letterbox, distortion, calibration, homography against essential matrix, PnP with RANSAC, and transform composition between named frames; a deep learning round that is mostly diagnosis rather than architecture, including the standard question of why average precision improved while the product got worse; a data strategy scenario such as three hundred defective parts against two million good ones; a deployment round where a latency and power budget has to add up on a named device; and a debugging round where it works in the notebook and fails on the robot. Live coding tends to mean writing non-maximum suppression, intersection over union, a bilinear resize or a RANSAC fit in plain NumPy.
How much does a computer vision engineer get paid?
There is no single authoritative band for a computer vision engineer, because no occupation code covers the title, and anyone quoting one number across robotics startups, defence primes, medical device firms and frontier labs is averaging incompatible things. Use the US Bureau of Labor Statistics Occupational Employment and Wage Statistics tables for 15-1252 Software Developers, 15-2051 Data Scientists and 15-1221 Computer and Information Research Scientists, reading the metro figures rather than the national ones because this field is geographically concentrated. Then read posted ranges in the states that require them, compare three employers in the same metro and sector, and ask for the full structure, because equity at a robotics startup and bonus and clearance premiums at a prime move a computer vision engineer's total compensation more than base salary does.
Do you need a security clearance to work in computer vision?
Only for part of the field, but that part is large. A computer vision engineer working in defence, aerospace, counter-drone or much geospatial imagery work usually needs to be a US person, because the work falls under ITAR or EAR export control, and many of those roles additionally require an active security clearance that takes months to adjudicate and can involve a polygraph at higher tiers. No clearance is needed for commercial robotics, industrial inspection, medical imaging, logistics, retail analytics, consumer devices or media. The advice runs both ways: a computer vision engineer holding a clearance should put it near the top of the resume because it is the highest-value single line there, and one who is not eligible should filter those postings out rather than discovering the wall at offer stage.
Do computer vision engineers need to know C++, or is Python enough?
A computer vision engineer needs Python always and C++ for one part of the field, and the posting's nouns tell you which. Python alone is sufficient for most modelling, research and data-centric roles, and for a good share of medical imaging and inspection work where the deployment target runs a Python pipeline or a vendor runtime. C++ is a hard requirement for real-time perception in robotics, autonomy and embedded products, where the production path is C++ with CUDA kernels, a ROS 2 node graph and a TensorRT or ONNX Runtime engine, and Python exists only for training and tooling. A computer vision engineer who wants access to the highest-paying embedded and autonomy roles should be able to read and write modern C++, profile it, and reason about memory copies between host and device, because that is where the latency goes.
How do you switch into computer vision from software engineering or data science, and how long does it take?
The reliable route into a computer vision engineer role from an adjacent job is one deployed project plus one closed knowledge gap, and the gap differs by starting point: from software or embedded engineering it is modelling and evaluation discipline, so build a project with an honest split and a chosen operating point; from data science or machine learning it is geometry and hardware, so calibrate a camera, publish the reprojection error, and put a quantised model on a board; from controls, manufacturing or test engineering it is deep learning mechanics, and machine vision on a line is the shortest path because your plant-floor knowledge is the half external candidates lack. Expect two to five months of searching from an adjacent engineering job with one strong deployed project, three to six months from a fresh master's with a thesis project and a demo on hardware, and longer if you are self-taught without a numerate degree, in which case plan on two projects and aim at inspection, integrators and smaller product companies first. An internal move is easier than an external one in every case, so if your current employer has a camera anywhere in its product or its factory, a computer vision engineer role there is the fastest version of this switch.
Put this on a resume in about a minute
Paste your history once and point it at the Computer Vision Engineer posting you are looking at. No account, no card.
Build my resume free More roles