Life sciences · Imaging
Detection from diagnostic imaging.
A radio-diagnostic tool reporting 96.4% accuracy against radiologist-adjudicated ground truth.
What that figure is, and is not
Accuracy is a weighted average of two very different mistakes, and the weighting comes from prevalence rather than from the model. In a population where the condition is rare, a system that called every study negative would score well and find nothing.
We do not think an accuracy number belongs on a website on its own, and we are not going to pretend otherwise. The operating characteristics behind it, and the prevalence they were measured at, go to you under agreement. If a competitor shows you accuracy alone and nothing else, ask them the same question you should ask us.
Generalisation is mostly a hardware question
Imaging models fall over across scanner vendors, acquisition protocols and reconstruction settings more than they fall over across populations. A validation set from one site with one scanner fleet will flatter any model, ours included. Multi-site and multi-vendor validation is not optional on an imaging claim.
The radiologist still writes the report
The system supports a reporting decision, it does not make one. The interface is deliberately built so the output appears as an input to the read rather than as a conclusion sitting at the top of the worklist.
Cancer detection from diagnostic imaging, at 96.4% accuracy.
- Measured against
- Radiologist-adjudicated ground truth.
- Output
- Presented inside the reporting workflow, as an input to the read.
- Deployment
- On-premise or in-country, integrated with the existing imaging estate.
- Monitoring
- Live performance compared against reported outcomes, with a defined trigger for re-validation.
- Ask us for
- Sensitivity, specificity, AUC and the prevalence the accuracy was measured at.
- Detail
- Method, validation design and performance data are shared under agreement.
After go-live
Performance measured before deployment is a statement about the past.
Scanner fleets get replaced. Protocols get revised. Case mix shifts with referral patterns and with whatever the local screening programme is doing that year. A model that was validated properly two years ago is not automatically the same model in effect today.
So monitoring is part of the deployment rather than an add-on: live output compared against reported outcomes, with a threshold agreed in advance at which the model is re-validated or withdrawn. We would rather write that threshold into the contract than discover it during an incident review.
Radiology department and vendor enquiries.
Integration detail and validation data are shared under agreement.