Data and audit
LTAFDB, MIT-BIH, a clinical AF set and sensor strips for training; AFDB and held-out sensor patients for testing; a label audit found and fixed a loader bug and two classes of bad windows.
Sources and roles
| Set | Role | Size | Beat source | Labels |
|---|---|---|---|---|
| Long-Term AF Database (PhysioNet) | training 60 records, tuning 24 records | 84 records, ~24 h each | reference annotations | interval, cardiologist-reviewed |
| MIT-BIH Arrhythmia Database | training | 48 records | engine | interval, PhysioNet |
| Clinical AF set (single-lead, 52 patients) | training | 218 segments | engine | whole segment, physician |
| Sensor strips, 97 patients | training | 6 156 windows | engine | whole strip, original developer |
| MIT-BIH AF Database | test, never used in training | 23 records, 10 h | PhysioNet detections | interval, cardiologist |
| Sensor strips, 47 patients | test, patients disjoint from training | 755 strips | engine | whole strip, original developer |
The sensor split is a deterministic hash of the patient identifier; zero patient overlap was verified. Windows are labelled AF when at least 80 % of their span lies inside an annotated AF interval, not AF at most 20 %, discarded in between.
Label audit
Out-of-fold predictions by patient over 111 629 training windows, plus physiological checks, found:
- A loader bug: about two hours at the start of one LTAFDB record precede its first rhythm annotation and were silently treated as non-AF although the rhythm is irregular; 11 601 windows removed, loader fixed.
- Unflagged artefacts: 754 training windows with engine beat rates above 200 per minute, mostly in one sensor set; dropped by rule.
- Impossible labels: 205 windows labelled AF with practically regular RR intervals (whole-segment labels covering non-AF stretches); dropped by rule.
- Two LTAFDB records with atrial bigeminy and supraventricular tachycardia that look like AF on RR intervals: genuine hard negatives, kept.
Model-based disagreements were not used for cleaning, to avoid circularity. Retraining on the corrected data raised cross-validated specificity from 93.5 % to 95.5 %.
Known limitations of the labels
The sensor strip labels were produced by the original developer from an episode list, whole strip, with his engine output at hand. Cardiologist re-annotation at interval level, two blinded readers plus adjudication, is the planned replacement; the export for it exists.