Cardiology holds two things clinical AI rarely has together: a randomized trial in which an AI alert changed diagnoses in routine care, and the largest consumer screening study ever run, with more than 400,000 participants. Between them sits a decade of work turning the humblest test in the specialty — the ECG — into a detector of conditions it was never designed to reveal. This guide reads what is authorized and what the strongest trials measured, with every figure tied to a primary source. As of July 2026.
What the FDA record shows
On the FDA's AI-Enabled Medical Device List 1, cardiovascular is the second-largest specialty — and a distant second. A peer-reviewed audit of the machine-learning devices authorized in 2024 found cardiovascular took 6.5% (11 of 168), against radiology's 74.4% 6. Cardiology is a meaningful but minority share of the cleared landscape, and the tools cluster into a few families that a taxonomy of 1,016 authorizations makes visible: signal analysis (chiefly the ECG), image analysis and measurement (echocardiography, cardiac CT and MRI), and acquisition guidance 7.
That last family produced a notable first: the FDA has authorized real-time AI guidance for cardiac-ultrasound image acquisition, cleared via the De Novo pathway in 2020, which coaches a less-experienced operator to capture diagnostic-quality echocardiography views 1. It is a good emblem of where cleared cardiology AI adds value — lowering the skill floor for a capture or a measurement, with a clinician still interpreting the result.
The signal in the ECG
The most important idea in cardiology AI is that a standard 12-lead ECG carries signals a human reader cannot see. The founding demonstration trained a network on 44,959 patients to detect low ejection fraction (35% or below) — weak pumping function usually measured by echocardiography — from the ECG alone. Tested on 52,870 patients, it reached an AUC of 0.93, with 86.3% sensitivity and 85.7% specificity 2. Strikingly, patients who screened positive but had normal function at the time went on to develop low ejection fraction at about four times the rate of those who screened negative — the model was seeing something real and early 2.
A development result is a promise; a randomized trial is a test of it. The EAGLE trial delivered one. It cluster-randomized primary-care teams and ran the low-EF algorithm on the ECGs of 22,641 adults, alerting clinicians in the intervention arm. New diagnoses of low ejection fraction within 90 days rose to 2.1% in the AI arm versus 1.6% in usual care (odds ratio 1.32, 95% CI 1.01-1.61, P=0.007) 3. The absolute increase is modest, and that is the honest reading: the alert surfaced cases that would otherwise have been missed or delayed, without transforming the base rate.
| ECG-AI evidence | Design | Scale | Headline result | Source |
|---|---|---|---|---|
| Low-EF detection | Development + test cohorts | 44,959 train / 52,870 test | AUC 0.93; sensitivity 86.3%; specificity 85.7% | 2 |
| EAGLE deployment | Pragmatic cluster RCT | 22,641 adults | New low-EF dx 2.1% vs 1.6%; OR 1.32; P=0.007 | 3 |
| Ambulatory rhythm | Deep neural network | 91,232 ECGs / 53,549 patients | 12 classes at AUC 0.97; F1 0.837 vs cardiologist 0.780 | 4 |
The third row points to the other mature ECG use: rhythm. A single-lead deep neural network trained on 91,232 recordings from 53,549 patients classified 12 rhythm types at an AUC of 0.97, with an F1 of 0.837 that exceeded the average cardiologist's 0.780 4. Automated rhythm reading — the task behind ambulatory monitors and wearables — is one of the better-validated capabilities in the specialty.
Rhythm and the wrist
That capability is what put screening on consumers' wrists, and it is where the evidence needs the most careful reading. The Apple Heart Study enrolled 419,297 participants and used a smartwatch to flag an irregular pulse suggestive of atrial fibrillation. The results are as instructive for their limits as their reach: only 0.52% of participants received a notification; among those who were notified and then wore a simultaneous ECG patch, 34% had atrial fibrillation on the patch, and 84% of notifications were concordant with atrial fibrillation during simultaneous recording 5.
Hold those numbers side by side. An 84% concordance when the watch and a reference ECG are recording together says the signal detection is good. A 34% confirmation on delayed patch monitoring says something different but equally important: atrial fibrillation is intermittent, so a true earlier alert often finds no arrhythmia days later — the lower number reflects the rhythm's nature as much as the watch's accuracy. Neither figure makes a notification a diagnosis. It is a prompt to seek testing, and it behaves like a clinical decision support signal, not a verdict.
What a notification does and doesn't prove
The wrist-screening evidence carries three cautions that generalize to most consumer cardiology AI. First, the tested population shapes the number: the Apple cohort skewed young — most participants were well under the age where atrial fibrillation is common — so the notification rate and the confirmation rate would both differ in an older, higher-risk group. Second, a screening tool needs a denominator: performance stated only among people who were flagged omits the false negatives among the unflagged, which a wrist-worn sensor cannot easily count. Third, a positive screen starts a workup rather than ending one — its value depends on the confirmatory pathway it triggers, not on the alert alone. The operating point matters here as much as any headline figure, and the full discipline for weighing these studies is in our guide on how to read an AI validation study.
As with the rest of the field, cleared cardiology AI is assistive and signal- or image-based; none of the FDA-authorized devices yet use large language models 7. The Good Machine Learning Practice principles from the FDA, Health Canada, and the MHRA (2021) frame the development expectations these devices are held to 8.
The acquisition frontier
Detection and rhythm classification are the mature uses, but the fastest-moving family is acquisition — AI that helps a human capture and measure a study in the first place. The FDA's authorization of real-time guidance for cardiac-ultrasound acquisition 1 is the emblem: software that coaches an operator's probe toward a diagnostic-quality echocardiography view, lowering the skill floor for a scan that once required a trained sonographer. The value here is access — a competent image captured at the point of care, in a clinic or an emergency bay, by an operator without formal sonography training. That access can matter most where cardiology expertise is scarce — rural clinics, emergency departments, lower- resource systems — which is also where an over-trusted, unconfirmed image carries the most risk. The taxonomy of authorized devices shows the same shape across cardiac imaging: measurement and quantification, image analysis, and acquisition support, with a clinician interpreting the result 7. None of it is autonomous, and the Good Machine Learning Practice principles frame what such tools are expected to demonstrate before and after clearance 8.
Questions before you rely on it
Whether a cardiology AI tool earns trust turns on questions the demo rarely covers, and each maps to the evidence above.
- Is the output a marker or a diagnosis? An AI-ECG flags low ejection fraction or a rhythm 24; it points toward an echocardiogram or a monitor rather than replacing one. Confirm what the tool claims and what still needs confirmation.
- What is the effect size in a population like yours? EAGLE's absolute increase in new low-EF diagnoses was about half a percentage point 3, and the Apple cohort skewed young 5. Ask for the number that matters in your age and risk mix, rather than the headline from a different one.
- Where is the denominator? Consumer screening figures often describe only the people who were flagged. A tool's false negatives — the cases it stayed silent on — belong in the evaluation, and a wrist sensor cannot easily count them.
- At what operating point? A rhythm or dysfunction flag runs at a threshold with a specific sensitivity and specificity. Ask for both at the deployed setting, and for the alert volume your clinicians will field.
- What confirmatory pathway does a positive trigger? A screen's value lives in the workup it starts. A flag without a clear route to an ECG, an echocardiogram, or a clinic visit generates anxiety and cost without benefit — it behaves as clinical decision support, and needs a decision behind it.
The through-line matches the other specialties: the trials tell you a capability exists, and a local evaluation tells you what it does in your clinics, at your operating point, for your patients. A cheap test that surfaces a treatable condition earlier is worth having — once you have measured what it costs to chase every alert it raises. That downstream cost — the confirmatory tests, the clinic visits, the reassurance of the worried-well — is the figure a local pilot measures and a product demo omits.
How to read this guide
Four cautions travel with everything above. First, a marker is not the disease: an AI-ECG flags low ejection fraction or a rhythm, which then needs confirmation and management — the model surfaces a signal, it does not treat a patient. Second, effect sizes are honest-sized: EAGLE's half-a-percentage-point increase is a real gain from a cheap test rather than a revolution, and it should be read as such. Third, consumer numbers need a denominator and a population: a notification rate and a confirmation rate mean little without the age mix and the false-negative picture behind them. Fourth, status is perishable — the device list grows, new trials report, and every figure here is dated to when it was pulled.
Because clearance status, reimbursement, and liability turn on these distinctions, confirm the current regulatory status and evidence for any specific tool — and its cleared indication — with your compliance or regulatory counsel before acting; a clearance is not the same as validation quality in your own population. We revisit this guide on a 180-day cycle and whenever the FDA authorizes a new cardiovascular device or a trial reports. For the field-wide effect numbers, see our clinical AI trial results tracker, and for the neighboring specialties, our guides on AI in radiology and AI in pathology.
Sources and method
This guide draws on the FDA's AI-Enabled Medical Device List 1; the founding AI-ECG detection study 2 and its EAGLE randomized deployment 3; a single-lead arrhythmia deep neural network 4; the 419,297-participant Apple Heart Study 5; a 2024 authorization audit 6; a taxonomy of 1,016 authorizations 7; and the tri-regulator Good Machine Learning Practice principles 8. Every figure is tied to the primary source cited beside it, and each perishable number carries its date. Statuses and counts are current as of July 2026.