Medical Imaging AI
Imaging models built around the two things that decide whether they ever reach a patient: a clearly defined intended use, and clinical validation that holds up on scanners and sites the model has never seen.
Medical imaging is the area where the distance between a strong research result and a deployed product is largest. Published models routinely achieve impressive figures on curated datasets and then fail on a different scanner in a different hospital, which is exactly the situation deployment creates.
Medical imaging AI is the development of models that detect, segment, measure or triage findings in radiological, pathological or other clinical images, built to a defined intended use, validated on data from sites and devices not used in training, and designed to support a clinician's decision within a documented workflow.
Intended use determines everything
Before any data is touched, the intended use statement is written and agreed with your clinical and regulatory leads. It is the sentence the entire project is judged against:
- What the device does. Detect, segment, measure, triage or prioritise, stated precisely.
- For which population. Age range, indication, imaging protocol and any exclusions.
- In what workflow position. Before, alongside or after the clinician, and whether output is advisory or acted on directly.
- With what claim. What performance is asserted, in which population, and against which reference standard.
A triage tool that reprioritises a worklist and a diagnostic aid that influences a report are different regulatory propositions with different evidence requirements. Teams that leave this vague until the end usually discover their evidence does not support the claim they wanted to make.
Ground truth is a decision, not a fact
Who defines the reference standard, and how disagreement between readers is resolved, determines what the model learns and what its accuracy figure means. Multiple independent readers with adjudication is the credible approach, and it costs real clinical time that has to be planned and funded from the start.
How we build imaging models
Generalisation is the whole problem
Models learn scanner artefacts, acquisition protocols and site-specific quirks with enthusiasm. We hold out entire sites and device types from training and validate there, because a model validated on a random split of one hospital's data tells you almost nothing about how it behaves in the second hospital.
Guard against shortcut learning
Imaging models find shortcuts: laterality markers, annotation burn-ins, scanner identifiers, even the presence of a chest drain implying the condition that required it. We test explicitly for these, because a model achieving excellent numbers by reading a marker is a well-documented pattern in the literature and a catastrophic one in deployment.
Report subgroup performance
Accuracy by age, sex, device, site, and by clinical subgroup where the data supports it. Aggregate performance that hides a failure in one population is the kind of finding that ends a deployment after it has begun, and it is far cheaper to surface during development.
Design for the clinician, not around them
Output that fits the reading workflow: overlays on the image, measurement with the region shown, priority flags in the worklist. Clinicians need to see why, verify quickly, and disagree easily, and the disagreements are among the most valuable data the system produces.
Build the evidence as you go
Development documentation, data provenance, version control of datasets and models, and the validation protocol agreed before results are generated. Regulatory submissions are constructed from this material, and reconstructing it afterwards is slow and considerably less credible.
Plan for post-deployment monitoring
Scanner upgrades, protocol changes and population shift all degrade performance quietly. Monitoring inputs and outputs, with a defined recalibration and reporting path, is expected by regulators and is the difference between a device that stays safe and one that is assumed to be.
The regulatory reality
We are engineers, not your regulatory consultants, and we work alongside whoever holds that role. What we do is build so the path stays open:
| Aspect | What we produce | Who owns it |
|---|---|---|
| Intended use and claims | Drafted with your clinical lead | Your regulatory function |
| Data provenance and ethics | Documented lineage and approvals referenced | Your research governance |
| Validation protocol | Pre-registered design, external sites | Joint |
| Technical documentation | Development record, versioning, testing | Us, handed over |
| Clinical evaluation | Performance evidence in the intended population | Your clinical team |
| Post-market monitoring | Monitoring design and thresholds | Joint, operated by you |
Where a product is intended for market rather than internal research use, engage your regulatory adviser before development starts, not after the model performs well.
How the engagement runs
Intended use and validation design are fixed before results exist, so the evidence answers the right question.
Intended use and data governance
Intended use statement agreed, data provenance and approvals confirmed, reference standard and adjudication process designed.
Dataset construction
Multi-site data assembled, reading and adjudication completed, holdout sites and devices reserved and never touched.
Model development
Training with shortcut testing, subgroup analysis, and internal validation on data separate from the external holdout.
External validation
Performance measured on held-out sites and devices against the pre-registered protocol, with subgroup reporting.
Workflow integration and handover
Integration into the reading workflow, monitoring design, and the technical documentation package.
What you receive
A model with evidence that survives an external site, and the documentation to support what comes next.
Intended use statement
Agreed with your clinical and regulatory leads, with the performance claim it supports.
Curated multi-site dataset
With provenance, reading protocol, adjudication record and reserved external holdout.
Trained model
With versioning, reproducible training and shortcut-learning test results.
External validation report
Performance on unseen sites and devices, with subgroup breakdowns.
Workflow integration
Overlays, measurements or worklist priority delivered into the reading environment.
Technical documentation
Development record, testing, risk considerations and post-deployment monitoring design.
Is this the right engagement?
Worth being direct. Medical Imaging AI is the wrong spend in some situations, and those are listed rather than buried.
Good fit if
- Imaging data is available from more than one site or device, with governance approvals in place.
- Clinicians can commit time to reading and adjudicating a reference standard.
- The intended use can be stated precisely and a regulatory owner exists.
- Deployment is advisory with a clinician in the loop.
- There is appetite for external validation before any clinical claim.
Choose something else if
- Data comes from a single scanner at a single site with no prospect of more.
- No clinical time is available to establish a reference standard.
- The intention is autonomous diagnosis without clinician review.
- Ethics or data governance approvals are not in place and not being sought.
Frequently asked questions
Marked up with FAQPage schema so these answers can surface directly in search results and inside AI assistant responses.
Why do published imaging models fail in deployment?
Because they were validated on data resembling their training data. Models learn scanner artefacts, acquisition protocols and site-specific quirks readily, so performance falls when the scanner or the hospital changes. We hold out entire sites and devices from training and validate there, which produces lower and far more honest numbers.
How do you establish ground truth?
With multiple independent readers and a documented adjudication process for disagreements. Who defines the reference standard and how disagreement is resolved determines what the accuracy figure means, and it costs real clinical time that has to be planned and funded rather than assumed.
Do we need regulatory approval?
It depends on the intended use and market. Internal research and workflow tools sit differently from products making clinical claims. We are engineers rather than regulatory consultants and work alongside whoever holds that role, building the documentation and validation evidence so the path stays open.
What is shortcut learning and how do you prevent it?
It is when a model achieves good results by reading something incidental: a laterality marker, a burned-in annotation, a scanner identifier, or a treatment device implying the condition that required it. We test for it explicitly, because it is well documented in the literature and produces excellent metrics with no clinical validity.
How long does a medical imaging AI project take?
Sixteen to twenty-four weeks for development and external validation, with data governance and reading often the longest poles. Any timeline that omits multi-site validation is describing a research prototype rather than something intended to reach a patient.
Often paired with this
Most clients combine two or three engagements from the Computer Vision pillar. These are the ones that most often run immediately before or after.
Intelligent Document Processing and OCR
Document extraction at volume with field-level confidence, validation and a review queue that only sees exceptions.
Read more →Object Detection and Image Classification
Detection and classification models trained on your own images, with honest labelling and cost-weighted thresholds.
Read more →Pose Estimation and Activity Recognition
Movement and activity understood over time, using skeletal data that needs no identifiable footage.
Read more →Is this the right engagement?
Tell us what you are trying to build. If a different service fits better, or if you do not need us at all, we will say so.