Project / 002

ECG Arrhythmia Classification & TinyML Deployment

A thesis project connecting rigorous ECG model evaluation with constrained on-device inference.

A CNN-based ECG arrhythmia classifier evaluated across patients and deployed for on-device inference on ESP32.

Category
AI & ML

PROJECT / 002

01MIT-BIH ECG (MLII)
02Pan-Tompkins beats
0313-feature LDA / CNNs
04Rigorous inter-patient test
05ESP32 inference

PUBLIC RECORD · REV / 01

ECG classifiers can look convincing without proving that they generalize to patients the model has not seen. The second constraint was practical: carrying the trained model from a research environment onto an ESP32-class device.

How the pieces connect

The work uses a CNN with an inter-patient evaluation framework, then converts the trained network with TensorFlow Lite for TinyML deployment. This keeps evaluation discipline and deployment constraints in the same research pipeline.

Implemented capabilities

  • BSc thesis, team of four
  • Standard vs rigorous inter-patient evaluation
  • 13-feature LDA in 272 bytes
  • Two CNN baselines on TensorFlow Lite Micro
  • ESP32 deployment checked on 30 real beats
  • On-device patient adaptation in 10.3 ms

The thesis links model development, inter-patient testing, TensorFlow Lite conversion, and ESP32 inference instead of stopping at an offline accuracy result.

Overview

A BSc thesis in Computer Science and Engineering at Bangladesh University of Professionals, written by a team of four — Ayman Tazwar, Morium Chowdhury, Mohammed Ashik, and Md. Tafsir Un Nahian — and supervised by Rashed Mazumder, Ph.D., Associate Professor at the Institute of Information Technology, Jahangirnagar University. Accepted in July 2026.

The task is inter-patient ECG beat classification: train on one set of patients, then classify beats from patients the model has never seen, into the three AAMI EC57 classes — normal (N), supraventricular ectopic (S), and ventricular ectopic (V). It is the clinically relevant setting and the one published results most often inflate. A 2025 systematic review cited by the thesis found that only 4.1% of 122 papers combine inter-patient evaluation, the AAMI standard, and embedded feasibility. This work targets that intersection, and ends on real hardware rather than an offline score.

What was built

  • A 13-feature Linear Discriminant Analysis classifier (LDA-13): RR-interval timing, beat morphology, higher-order statistics, and QRS geometry, with Ledoit-Wolf shrinkage, equal class priors, and SMOTE oversampling of the S class.
  • Two evaluation protocols run side by side: the standard one used across the literature, and a rigorous one that seals the test set before any model-selection decision.
  • Two CNN baselines: a 12,595-parameter RR-dominant hybrid CNN, and a reproduction of a published 50,011-parameter attention CNN.
  • All three classifiers deployed to an ESP32 and run on the same 30 real test beats; the LDA and the attention CNN match their CPU reference on all 30.
  • On-device patient adaptation: the ESP32 re-fits the LDA from 25 labelled beats of a new patient.

Architecture

The LDA pipeline runs identically in Python and in C on the ESP32. The 13-value feature vector is the only interface between signal processing and classification.

The LDA-13 pipeline in seven panels: raw MLII signal, bandpass and normalisation, Pan-Tompkins R-peak detection, the 250-sample beat window, the 13 standardised features, SMOTE on the training fold, and the three discriminant scores with the argmax prediction.
Fig. — The LDA-13 pipeline, from raw MLII to an N/S/V label (thesis Fig. 3.1)

The decision that shaped the thesis

Under the standard protocol, the reproduced attention CNN reaches a supraventricular (S) F1 of 0.831 — within 0.013 of the best published inter-patient result. Choose its checkpoint on held-out patients instead, and the same architecture averages 0.374 ± 0.256 across five seeds. The LDA drops too, from 0.576 to 0.090 ± 0.07 under five-fold GroupKFold, but it has no checkpoint to select, so its estimate is stable.

Two bar charts of SVEB F1: the attention CNN falls from 0.831 under the standard protocol to 0.374 plus or minus 0.256 under the rigorous one, with five seed results spread from about 0.05 to 0.75; the LDA falls from 0.576 to 0.090.
Fig. — Protocol choice, not model capacity, dominates the reported S-F1 (thesis Fig. 4.4)

That gap is the thesis's central finding: on this benchmark, how a model is selected moves the reported score more than which model it is. It is also why the deployed baseline is a 272-byte linear model — it gives a deterministic, reproducible estimate at a tiny fraction of the cost, not a higher headline number.

The thesis traces the S-precision bottleneck to the LDA acting as a detector of absolute heart rate rather than of relative prematurity. Record 209 contributes 40.6% of all training S-beats, and capping its contribution makes S-precision worse, not better — the limit is in the single-lead feature space, not in class imbalance.

Validation on hardware

All figures below are measured on the ESP32-WROOM-DA at 240 MHz, as medians over 30 real DS2 test beats (10 N, 10 S, 10 V).

Metric                            LDA-13        RR-dominant CNN  Attention CNN (int8)
--------------------------------  ------------  ---------------  --------------------
Inference latency                 131 µs        37.64 ms         251.5 ms
Model weights                     272 B         57 KB            114 KB
Real-time headroom per beat       ~5,300×       ~18×             ~2.8×
Runtime needed                    none          TFLite Micro     TFLite Micro
Agreement with the CPU reference  30/30         —                30/30
S-F1, standard protocol           0.576         0.446            0.832
S-F1, rigorous protocol           0.090 ± 0.07  0.134            0.374 ± 0.256

The LDA's 272 bytes are 68 floats — a 3×13 weight matrix, three biases, and the scaler — exported to a C header and checked against the Python predictions before the header is written. It needs 264 bytes of stack and no heap. The full LDA binary is 288 KB, 22% of flash.

The 131 µs is compute on the device. In the demonstration setup each beat window reaches the ESP32 as ASCII over a 115,200-baud serial link, and that transfer — about 223 ms per beat for the LDA — dominates the end-to-end time. A device that samples the ECG itself would not pay it.

A ten-second strip of MIT-BIH record 200 with each detected beat marked and labelled N or V by the deployed LDA code; two beats whose prediction disagrees with the reference annotation are ringed.
Fig. — LDA-13 inference on DS2 record 200, rendered from the deployed feature-extraction and scoring code; ringed beats disagree with the annotation (thesis Fig. 7.2)
Three log-scale bar charts comparing LDA-13 and the RR-dominant CNN on the ESP32: latency 131 microseconds against 37,640, flash 272 bytes against 57 kilobytes, and rigorous S-F1 0.090 against 0.134.
Fig. — LDA-13 against the RR-dominant CNN, measured on the ESP32 (thesis Fig. 7.5)

The attention CNN needed two fixes before it ran correctly on TensorFlow Lite Micro: the converted model loaded and executed without error but called about 93% of beats normal. The build log below covers both failures and the weight-preserving reformulations that resolved them.

On-device patient adaptation

The adaptation sketch re-fits the LDA's class means on the ESP32 from 25 labelled beats chosen from a new patient's learning period. The fit takes 10.3 ms, the adapted model is 1,104 bytes, and inference afterwards stays at 12 µs per beat.

Grouped bars of ventricular (V) F1 before and after on-device adaptation: MIT-BIH record 200 from 0.948 to 0.953, INCART I01 from 0.991 to 0.993, and INCART I02 from 0.689 to 0.949.
Fig. — On-device adaptation, measured on the ESP32 (thesis Fig. 7.6)

On INCART record I02, the worst-transferring record, V-F1 recovers from 0.689 to 0.949. On records that already transfer well the gain is marginal, and S-class adaptation is limited by how few S-beats a learning period contains.

Limitations

  • Not a clinical device. This is an engineering research prototype evaluated on public databases.
  • S-precision remains the open problem. Under the rigorous protocol no single-lead model here reaches clinically adequate S-precision.
  • Cross-site false alarms. Zero-shot on INCART, S-recall transfers (0.866) but S-precision is 0.126 — roughly seven false S alarms per true one. Calibrating to INCART's class prior does not fix it.
  • Single lead, no P-wave features. P- and T-wave features were tested and degraded inter-patient performance.
  • The rigorous estimate is itself fragile. Record 209's concentration of S-beats makes cross-validation unstable on this dataset.
  • Adversarial training not reproduced. The attention CNN's architecture was reproduced, not its subject-invariance training, which is what the original authors credit for their higher score.
  • CNN timings cover inference only. For the CNNs, preprocessing is assumed to run elsewhere; the LDA's figures include on-device feature extraction.
  • 01Python
  • 02TensorFlow Lite
  • 03CNN
  • 04TinyML
  • 05ESP32