AI in Medicine: Why MRCP PACES Still Matters
Every MRCP PACES candidate has had the thought — usually somewhere around their third attempt to elicit a murmur in a study group at 11 pm: machines can already read ECGs, mammograms and retinal photographs better than many specialists, so why am I travelling to a UK hospital to be graded on auscultation and fundoscopy?
It is a fair question, and the answer is more interesting than "tradition". The Royal Colleges are not ignoring artificial intelligence — they are implicitly betting on it. As AI absorbs the routine pattern-recognition that once defined junior doctoring, the value of a physician concentrates in exactly the skills PACES was built to test: synthesis under uncertainty, safe prioritisation, and the ability to sit with a frightened human being and build a plan together.
This guide walks through what AI genuinely does well in 2025, where it fails in ways every PACES candidate should understand, and what all of this means for how you prepare.
The AI Revolution in Medicine Is Real — Respect It
You do not need a computer science degree to appreciate how far clinical AI has come. The headline results every physician should know:
Radiology: In Sweden's MASAI randomised trial, AI-supported mammography screening detected roughly 20–29% more cancers while cutting radiologist screen-reading workload by over 40%. Chest X-ray triage tools now flag pneumothorax, consolidation and nodules within seconds.
Cardiology: A single-lead AI-enhanced ECG can estimate left ventricular ejection fraction (AUC >0.9 in validation cohorts, Mayo Clinic data) and unmask silent paroxysmal atrial fibrillation. The Apple Heart Study enrolled over 400,000 people and confirmed AF in about a third of those who received a smartwatch rhythm alert.
Ophthalmology: Autonomous diabetic retinopathy screening (pioneered by IDx-DR) has been FDA-approved since 2018 — an algorithm making an independent referral decision without a physician over-read.
Stroke: AI-based CT interpretation platforms deployed across the NHS (e.g. e-Stroke/Viz.ai-style tools) have substantially shortened onset-to-treatment and transfer times for thrombectomy candidates by automating flagging and referral.
The ward round: Ambient AI scribes that draft consultation notes from ambient audio are scaling through NHS pilots and are already changing outpatient clinic workflows in the UK and US.
Biomedical science: AlphaFold predicted the structures of essentially all known proteins — work recognised with the 2024 Nobel Prize in Chemistry.
The pattern is clear: wherever the task is image or signal classification from good-quality data, machines are now genuinely excellent. So why does the exam still exist?
Where AI Stumbles — The Lessons Every PACES Candidate Should Internalise
The failures of medical AI are not footnotes; they are the entire justification for your exam fee.
1. Detection is not diagnosis
The most instructive cautionary tale remains the Epic sepsis prediction model. Embedded in hundreds of US hospitals, its external validation (Wong et al., JAMA Internal Medicine 2021) found an AUC of ~0.63, sensitivity around a third at the deployed threshold, and alert fatigue on an enormous scale. Why? The model learned statistical associations in historical data — it did not understand the patient. Diagnosis is a hypothesis-driven, context-laden process; PACES examiners are scoring precisely that process.
2. Dataset shift eats algorithms alive
A model trained in one health system, demographic mix or coding culture underperforms when transplanted. Medicine's most dangerous phrase — "the patient in front of me" — is exactly the situation where historical averages mislead. Your elderly Station 5 patient with frailty, polypharmacy and atypical presentations is the population AI handles worst.
3. Bias is not theoretical
Obermeyer and colleagues (Science, 2019) showed a widely used US risk-prediction algorithm using healthcare cost as a proxy for need systematically underestimated the sickness of Black patients — correcting it would nearly triple their enrolment in extra care programmes.
Pulse oximetry (Sjoding et al., NEJM 2020) demonstrated occult hypoxaemia was almost three times more common in Black patients due to the device's calibration — a reminder that even "simple" measurement pipelines encode bias, long before you reach deep learning.
4. AI cannot access the world outside its data
The pillbox the patient stopped opening. The job loss driving the non-adherence. The daughter quietly signalling from behind the bed curtain. The subtle way a patient says "fine, doctor" while their face says otherwise. None of this exists in the electronic record that feeds the model — and all of it changes management.
5. Accountability cannot be delegated
An output requires a clinician to own the decision. The GMC's updated Good Medical Practice (2024) makes clear that doctors must be competent in the technologies they use — including AI — and remain responsible for decisions supported by them. "The algorithm said so" has never been a defence, and never will be.
The Skills PACES Tests Are Precisely the Skills AI Cannot Replace
Look at what the exam actually rewards and map it against machine capability:
| What PACES rewards | Why AI struggles with it |
|---|---|
| Targeted data gathering — building a hypothesis-driven history in 14 minutes | Machines can only query data that exists; the clinical interview creates the data |
| Differential diagnosis under uncertainty | AI interpolates from training distributions; clinical judgement extrapolates from an individual |
| Physical examination — interpreting signs in context (and knowing when findings are artefact) | Sensor-limited, context-blind; the subtleneurological examination remains far beyond robotics |
| Managing patients' concerns — ICE, chunking, checking, shared decisions | Language models produce fluent words; concern-management requires accountability and relationship |
| Safe prioritisation in acute scenarios (Station 5) | Triage is a values-laden judgement involving capacity, frailty, prognosis and family — not a classification task |
| Treating the patient as a person | The defining gap. There is no gradient descent pathway to compassion |
This is why the exam is not winding down. AI raises the floor of pattern recognition; PACES certifies the ceiling of human judgement.
Why the Colleges Are Doubling Down, Not Winding Down
Think about what the consultant of 2035 actually does in an AI-saturated NHS:
Receives AI-flagged ECGs, imaging and deterioration alerts — and must triage the true positives from the noise using pre-test probability and bedside assessment
Supervises algorithm outputs for patients the model has never "seen" — the frail, the atypical, the multi-morbid
Explains AI-informed decisions to patients: "The computer flagged this, here is what it means, here is what I recommend" — a consultation skill, not a computational one
Catches the errors — which requires independent competence, not deference to the tool
The doctor who cannot examine, cannot synthesise, and cannot communicate has nothing left to supervise the machine with. PACES is, in effect, certifying your ability to remain the safety net around AI.
What This Means for Your Revision — Five Practical Takeaways
1. Stop collecting findings; practise synthesis
The examiner who asks "what else would you like to examine?" and "how would you manage this patient?" is testing the integration that no dashboard provides. Every practice case should end with a two-sentence problem representation ("a 68-year-old with asymmetrical resting tremor and reduced arm swing...") and a ranked differential with a next investigation — not a recital of findings.
2. Train the human differentiators disproportionately
Chunk-and-check, purposeful silence, acknowledging emotion before information, negotiating a plan — these feel "soft" and are precisely what the marking scheme rewards and what machines cannot do. Ironically, the more medicine automates, the higher the proportion of your score lives here.
3. Use AI to prepare for the humans — not to replace practising on them
AI revision tools are superb for drilling guideline recall, generating differential lists to critique (never to copy uncritically — hallucinated citations are a real hazard), and rehearsing consultation flow in the blank days between study groups. But the exam tests you against real, unpredictable humans. Close every loop with live practice.
4. Build basic AI literacy for the viva-style questions
A modern Station 5 or discussion prompt can quite reasonably include: "This patient's ECG was automatically flagged by an AI tool as showing AF with low ejection fraction risk — how do you respond?" A safe answer structure:
Treat the AI output as one data point, not a diagnosis
Corroborate clinically — symptoms, pulse check, repeat ECG, focused examination
Apply pre-test probability and guideline-based confirmation (e.g. ECHO for EF)
Communicate transparently with the patient about what was found and what happens next
Retain clinical accountability for every resulting decision
5. Frame failures of AI as lessons in your own safety-checking
When you read that a sepsis model missed two-thirds of septic patients, translate it: what would have caught those cases? A worried nurse. A soft examination finding. A gut sense that the patient looks unwell. That translation exercise — from algorithmic failure to bedside vigilance — is exactly the mindset PACES rewards.
A Sample Viva Exchange Worth Rehearsing
Examiner: "With AI reading scans better than radiologists, why examine patients at all?"
Strong answer: "AI excels at classification tasks on good-quality data, and I would absolutely use it as a decision-support tool. But detection is not diagnosis — external validations like the Epic sepsis model show how poorly algorithms can perform on real populations, and biased training data has caused documented harm. The clinical examination generates information the record doesn't contain, contextualises everything the machines produce, and lets me take responsibility for the patient in front of me — which, professionally, I cannot delegate. So I'd use AI to augment my assessment, and my bedside skills to verify and interpret it."
Notice the structure: acknowledge the technology, demonstrate knowledge of its limits, then re-centre the patient and your accountability. That is a consultant-level answer in any station.
The Bottom Line
The question was never will AI change medicine? — it already has. The question is what remains valuable in a doctor when pattern recognition is cheap? The answer is the PACES curriculum in miniature: gathering information no sensor can reach, reasoning when the data conflict, prioritising safely under pressure, and communicating decisions you personally stand behind.
So when you next struggle to elicit that sign, remember: you are not training for a world before AI. You are training for the world after it — as the one clinician in the room whom the machine cannot replace, and the one it must ultimately answer to.
Key Evidence and Sources for Further Reading
MASAI trial — AI-supported vs standard mammography screening (Lancet Digital Health)
Attia ZI et al. — AI-ECG for low ejection fraction (Nature Medicine, 2019)
Perez MV et al. — Apple Heart Study, smartwatch AF detection (NEJM, 2019)
Wong A et al. — External validation of a sepsis prediction model (JAMA Internal Medicine, 2021)
Obermeyer Y et al. — Dissecting racial bias in a clinical risk algorithm (Science, 2019)
Sjoding MW et al. — Racial bias in pulse oximetry (NEJM, 2020)
GMC Good Medical Practice (2024) — professional obligations around technology and accountability
AlphaFold — protein structure prediction (Nature, 2021; Nobel Prize in Chemistry, 2024)
Join the Discussion
Share your thoughts and insights with the medical community
Comments
Delete Comment
Are you sure you want to delete this comment? This action cannot be undone.