Detailed results
End-to-end sensor results, EC57 benchmark tables, sensitivity by episode length, false-episode analysis and the comparison with the previous model.
Sensor strips, end to end, all 755 held-out strips
| Configuration | Inconclusive | Classified | Se | Sp | PPV | NPV | FP / FN |
|---|---|---|---|---|---|---|---|
| 32-beat windows, on 0.6 / off 0.3 | 44.5 % | 419 | 98.8 | 98.0 | 97.0 | 99.2 | 5 / 2 |
| 16-beat windows, on 0.6 / off 0.3 | 32.2 % | 512 | 98.9 | 94.1 | 89.6 | 99.4 | 20 / 2 |
| 16-beat windows, on 0.8 / off 0.5 | 31.8 % | 515 | 98.7 | 98.6 | 96.9 | 99.4 | 5 / 2 |
| hybrid 32-beat else 16-beat, on 0.7 / off 0.3 | 31.5 % | 517 | 98.2 | 98.0 | 95.9 | 99.1 | 7 / 3 |
The 32-beat window refuses strips below about 66 beats per minute (fewer than 33 beats in 30 s); the 16-beat window removes that. Remaining refusals are strips whose engine beats all exceed 200 per minute (artefacts) and strips in the probability band.
Clean strips only (expert-flagged artefacts removed, 424 strips), 32-beat model, per strip: Se 94.5 %, Sp 99.2 %, PPV 98.7 %. On the same 424 strips the previous model’s delivered output gives Se 99.4 %, Sp 96.1 %, PPV 94.3 %, but all 47 patients were in its training set.
Where the inconclusive strips come from. Of the 240 refusals, 199 have no plausible window: 139 are expert-flagged artefacts, 24 strips are only about 10 s long (this set mixes 10-s and 30-s recordings; a 16-beat window cannot exist in 10 s below ~100 bpm), and the rest are unflagged artefacts (saturated signal, spike trains at 240 “beats” per minute). Restricted to recordings of at least 25 s that the expert did not flag, the product rule gives:
| Subset | N | Inconclusive | Classified | Se | Sp | FP / FN |
|---|---|---|---|---|---|---|
| all strips | 755 | 31.8 % | 515 | 98.7 | 98.6 | 5 / 2 |
| strips >= 25 s | 725 | 29.9 % | 508 | 98.7 | 98.6 | 5 / 2 |
| strips >= 25 s, expert-clean | 525 | 14.1 % | 451 | 98.7 | 99.0 | 3 / 2 |
On readable 30-s strips the inconclusive rate is 14 %, inside the proposed 15 % limit.
Confidence intervals, held-out sensor strips (product configuration)
Model 9, 16-beat windows, AF if the highest window probability is at least 0.8, not AF if at most 0.5. All 755 held-out strips of 51 patients; 12 patients contribute AF strips. Wilson and Clopper-Pearson intervals treat strips as independent; the cluster bootstrap resamples patients (2,000 draws) and is the honest interval when strips of one patient are correlated.
| Metric | k / n | Point | Wilson 95 % | Clopper-Pearson 95 % | Patient cluster bootstrap 95 % |
|---|---|---|---|---|---|
| Sensitivity | 157 / 159 | 98.7 % | 95.5 - 99.7 | 95.5 - 99.8 | 92.3 - 100 |
| Specificity | 351 / 356 | 98.6 % | 96.8 - 99.4 | 96.8 - 99.5 | 97.4 - 99.7 |
| PPV | 157 / 162 | 96.9 % | 93.0 - 98.7 | 92.9 - 99.0 | 83.3 - 99.6 |
| NPV | 351 / 353 | 99.4 % | 98.0 - 99.8 | 98.0 - 99.9 | 98.2 - 100 |
| Inconclusive | 240 / 755 | 31.8 % | 28.6 - 35.2 | 28.5 - 35.2 | 19.7 - 46.2 |
The strip-level lower bounds meet the proposed acceptance limits (sensitivity >= 90 %, specificity >= 95 %). The patient-level bounds are wide because only 12 held-out patients have AF; the pivotal study needs on the order of 100 subjects with AF and 100 without.
MIT-BIH AF Database, EC57, RR features only
| Episodes | Test episodes | Episode Se / +P | Duration Se / +P | Se < 10 s | 10-30 s | 30-60 s | >= 60 s |
|---|---|---|---|---|---|---|---|
| 32-beat model | 964 | 93.9 / 46.8 | 96.5 / 92.0 | 50 % | 89 % | 100 % | 100 % |
| 16-beat model | 1543 | 96.2 / 45.4 | 96.0 / 92.5 | 61 % | 100 % | 100 % | 100 % |
| cascade: 16 proposes, 32 confirms | 1220 | 93.5 / 55.7 | 95.9 / 94.0 | 46 % | 89 % | 100 % | 100 % |
Reference: 293 episodes in 23 records. False episodes have a median length of 33 s; 63 of 410 exceed one minute, two exceed five minutes. Positive predictivity restricted to detected episodes of at least 60 s is 83.9 %. The false windows are statistically indistinguishable from AF on RR intervals (irregularity index 0.65 vs 0.68, RR CV 0.205 vs 0.208) and sit in eight records with frequent ectopy; the pending engine-native test adds P-wave, QRS and quality features to this benchmark.
Cross-check of the EC57 scorer against the reference WFDB tools
Our episode and duration statistics are computed by a Python implementation of the EC57 definitions. To verify it, the model-8 AFDB episodes were written as WFDB rhythm annotations and scored with epicmp, the reference program of the WFDB software package (built from source on the analysis machine), gross statistics by sumstats.
| Scorer | Episode Se | Episode +P | Duration Se | Duration +P | Reference / test episodes |
|---|---|---|---|---|---|
| Our scorer, from record start | 93.9 | 46.8 | 96.5 | 92.0 | 293 / 964 |
epicmp -f 0, from record start |
95 | 45 | 97 | 92 | 291 / 918 |
epicmp, EC57 standard (first 5 minutes excluded) |
95 | 45 | 97 | 92 | 291 / 912 |
Agreement within about one point on every statistic; the episode-count differences come from how episodes touching each other or the record boundaries are merged. Submission numbers will be taken from epicmp.
Sensitivity by heart-rate band (window level)
| Set | < 60 bpm | 60-100 | 100-130 | > 130 |
|---|---|---|---|---|
| Sensor held-out, Se | n/a | 98.1 % | 98.4 % | 97.4 % |
| AFDB, Se | 99.3 % | 95.4 % | 98.1 % | 95.7 % |
| AFDB, Sp | 99.9 % | 97.6 % | 92.6 % | (65 windows) |
Beat detection, unchanged engine, MIT-BIH Arrhythmia Database
bxb rule, 150 ms window, all reference beats counted: Se 98.1 %, +P 98.7 % (109 176 reference beats).
Previous model, same scoring
Per-beat Random Forest as delivered, on MIT-BIH which was in its training set: EC57 episode 91.6 / 58.6, duration 92.9 / 86.5.
External comparison: ArNet2 (Technion AIM lab, pretrained, CC BY-NC, comparison only)
ArNet2 is a deep network on 60-beat RR windows trained on about 51,000 hours of private Holter data. Run with its own runner on the AFDB beat annotations and scored with the same EC57 procedure (consecutive positive windows merged):
| Model on AFDB | Window-level Se / Sp | Episode Se / +P | Duration Se / +P | Se by episode length <10 s / 10-30 / 30-60 / >= 60 |
|---|---|---|---|---|
| ArNet2, threshold 0.5 | 97.5 / 98.6 | 74.4 / 96.6 | 93.9 / 98.5 | 14 / 14 / 63 / 97 % |
| ArNet2, threshold 0.3 | 99.3 / 97.7 | 87.0 / 79.7 | 96.9 / 96.5 | 36 / 53 / 95 / 99 % |
| ours, 32-beat model | 95.6 / 96.7 | 93.9 / 46.8 | 96.5 / 92.0 | 50 / 89 / 100 / 100 % |
| ours, 16+32 cascade | 93.5 / 55.7 | 95.9 / 94.0 | 46 / 89 / 100 / 100 % |
At matched duration sensitivity ArNet2 holds 3-4 points more duration precision and produces far fewer, longer episodes, at the cost of missing most episodes under 60 seconds. Its window specificity on AFDB, 98.6 %, is the level the morphology features are expected to bring ours to. Whether AFDB was part of its development data is not stated. The raw-ECG model of the same toolkit could not be run: its exported graph fails on a public record under the repository’s own runner.
Work in progress, night of 15 September (preliminary, will be replaced)
PhysioNet/CinC 2017, 300-record validation sample, frozen rule (max window probability, 0.8 / 0.5). Single-lead 9-61 s strips from a consumer device, labels normal / AF / other / noisy. Sample used for development; the 8,528-record set is being processed as the test set.
| Class | n | AF | not AF | inconclusive | Se / Sp |
|---|---|---|---|---|---|
| AF | 47 | 34 | 5 | 8 | Se 87.2 % |
| normal | 148 | 2 | 133 | 13 | Sp 98.5 % |
| other rhythm | 65 | 2 | 56 | 7 | Sp 96.6 % |
| noisy | 40 | 3 | 22 | 15 | reported separately |
Overall on the three interpretable classes: Se 87.2 % (95 % CI 73-94), Sp 97.9 % (95-99), inconclusive 15.6 %. Sensitivity is below the 95 % target on this sample, so every error was reviewed on the waveform:
- Two of the five misses are questionable labels: one is a bigeminal rhythm (every second beat early, identical pattern throughout) labelled AF; one is a regular rhythm at 70 bpm without clear P waves. One miss is a poor-quality strip with 37 % of the signal marked out-of-signal by the engine; that strip should be inconclusive, not “not AF” (a quality gate is being added). Two misses are genuine: irregular rhythms where the model’s probability stayed at 0.5.
- Of the four false alarms, two are “regularly irregular” rhythms (bigeminy-like alternation of short and long RR intervals), one is sinus arrhythmia with a very noisy baseline, and one is a mostly regular strip whose last two windows were irregular.
Figure. A CinC 2017 strip labelled AF; the RR intervals alternate 970 / 600 ms with two beat morphologies, a bigeminal pattern. The model correctly did not call AF.
Figure. A genuine miss: irregular RR intervals without P waves; the window probabilities reached only 0.48.
Aggregating windows by median or mean instead of maximum removed all false alarms but cut sensitivity to about 70 %, so the maximum rule stays. Actions taken tonight: (1) two new RR features that detect regularly-irregular patterns (median RR difference at lag 2 and lag 3 relative to lag 1: about 0 for bigeminy or trigeminy, about 1 for AF; on this sample the AF median is 0.92 against 1.50 for normal and 1.35 for other rhythms); (2) a quality gate that routes strips with out-of-signal windows to inconclusive; (3) retraining with these features (experiment 11) and an experiment that adds the previous model’s per-beat AF output as a feature (experiment 10), both scored on the held-out sensor strips, the CinC sample, the full CinC set and AFDB with engine beats.
Previous model’s own rhythm output on AFDB (engine run on 10,000 of 14,041 minutes so far). EC57 episode Se 93.7 %, +P 74.3 %; duration Se 43.8 %, +P 84.3 %. Inside long AF the engine’s state machine relabels most of the time as “sinus arrhythmia” or “irregular”; in the two all-AF records only 11-13 % of the AF time is labelled AF. Final numbers follow when the run completes.