Project / 002
ECG Arrhythmia Classification & TinyML Deployment
A thesis project connecting rigorous ECG model evaluation with constrained on-device inference.
A CNN-based ECG arrhythmia classifier evaluated across patients and deployed for on-device inference on ESP32.
PROJECT / 002
PUBLIC RECORD · REV / 01
01The challenge
ECG classifiers can look convincing without proving that they generalize to patients the model has not seen. The second constraint was practical: carrying the trained model from a research environment onto an ESP32-class device.
02The approach
How the pieces connect
The work uses a CNN with an inter-patient evaluation framework, then converts the trained network with TensorFlow Lite for TinyML deployment. This keeps evaluation discipline and deployment constraints in the same research pipeline.
03Build scope
Implemented capabilities
- BSc thesis, team of four
- Standard vs rigorous inter-patient evaluation
- 13-feature LDA in 272 bytes
- Two CNN baselines on TensorFlow Lite Micro
- ESP32 deployment checked on 30 real beats
- On-device patient adaptation in 10.3 ms
04The result
The thesis links model development, inter-patient testing, TensorFlow Lite conversion, and ESP32 inference instead of stopping at an offline accuracy result.
05Technical record
Overview
A BSc thesis in Computer Science and Engineering at Bangladesh University of Professionals, written by a team of four — Ayman Tazwar, Morium Chowdhury, Mohammed Ashik, and Md. Tafsir Un Nahian — and supervised by Rashed Mazumder, Ph.D., Associate Professor at the Institute of Information Technology, Jahangirnagar University. Accepted in July 2026.
The task is inter-patient ECG beat classification: train on one set of patients, then classify beats from patients the model has never seen, into the three AAMI EC57 classes — normal (N), supraventricular ectopic (S), and ventricular ectopic (V). It is the clinically relevant setting and the one published results most often inflate. A 2025 systematic review cited by the thesis found that only 4.1% of 122 papers combine inter-patient evaluation, the AAMI standard, and embedded feasibility. This work targets that intersection, and ends on real hardware rather than an offline score.
What was built
- A 13-feature Linear Discriminant Analysis classifier (LDA-13): RR-interval timing, beat morphology, higher-order statistics, and QRS geometry, with Ledoit-Wolf shrinkage, equal class priors, and SMOTE oversampling of the S class.
- Two evaluation protocols run side by side: the standard one used across the literature, and a rigorous one that seals the test set before any model-selection decision.
- Two CNN baselines: a 12,595-parameter RR-dominant hybrid CNN, and a reproduction of a published 50,011-parameter attention CNN.
- All three classifiers deployed to an ESP32 and run on the same 30 real test beats; the LDA and the attention CNN match their CPU reference on all 30.
- On-device patient adaptation: the ESP32 re-fits the LDA from 25 labelled beats of a new patient.
Architecture
The LDA pipeline runs identically in Python and in C on the ESP32. The 13-value feature vector is the only interface between signal processing and classification.

The decision that shaped the thesis
Under the standard protocol, the reproduced attention CNN reaches a supraventricular (S) F1 of 0.831 — within 0.013 of the best published inter-patient result. Choose its checkpoint on held-out patients instead, and the same architecture averages 0.374 ± 0.256 across five seeds. The LDA drops too, from 0.576 to 0.090 ± 0.07 under five-fold GroupKFold, but it has no checkpoint to select, so its estimate is stable.

That gap is the thesis's central finding: on this benchmark, how a model is selected moves the reported score more than which model it is. It is also why the deployed baseline is a 272-byte linear model — it gives a deterministic, reproducible estimate at a tiny fraction of the cost, not a higher headline number.
The thesis traces the S-precision bottleneck to the LDA acting as a detector of absolute heart rate rather than of relative prematurity. Record 209 contributes 40.6% of all training S-beats, and capping its contribution makes S-precision worse, not better — the limit is in the single-lead feature space, not in class imbalance.
Validation on hardware
All figures below are measured on the ESP32-WROOM-DA at 240 MHz, as medians over 30 real DS2 test beats (10 N, 10 S, 10 V).
Metric LDA-13 RR-dominant CNN Attention CNN (int8)
-------------------------------- ------------ --------------- --------------------
Inference latency 131 µs 37.64 ms 251.5 ms
Model weights 272 B 57 KB 114 KB
Real-time headroom per beat ~5,300× ~18× ~2.8×
Runtime needed none TFLite Micro TFLite Micro
Agreement with the CPU reference 30/30 — 30/30
S-F1, standard protocol 0.576 0.446 0.832
S-F1, rigorous protocol 0.090 ± 0.07 0.134 0.374 ± 0.256The LDA's 272 bytes are 68 floats — a 3×13 weight matrix, three biases, and the scaler — exported to a C header and checked against the Python predictions before the header is written. It needs 264 bytes of stack and no heap. The full LDA binary is 288 KB, 22% of flash.
The 131 µs is compute on the device. In the demonstration setup each beat window reaches the ESP32 as ASCII over a 115,200-baud serial link, and that transfer — about 223 ms per beat for the LDA — dominates the end-to-end time. A device that samples the ECG itself would not pay it.


The attention CNN needed two fixes before it ran correctly on TensorFlow Lite Micro: the converted model loaded and executed without error but called about 93% of beats normal. The build log below covers both failures and the weight-preserving reformulations that resolved them.
On-device patient adaptation
The adaptation sketch re-fits the LDA's class means on the ESP32 from 25 labelled beats chosen from a new patient's learning period. The fit takes 10.3 ms, the adapted model is 1,104 bytes, and inference afterwards stays at 12 µs per beat.

On INCART record I02, the worst-transferring record, V-F1 recovers from 0.689 to 0.949. On records that already transfer well the gain is marginal, and S-class adaptation is limited by how few S-beats a learning period contains.
Limitations
- Not a clinical device. This is an engineering research prototype evaluated on public databases.
- S-precision remains the open problem. Under the rigorous protocol no single-lead model here reaches clinically adequate S-precision.
- Cross-site false alarms. Zero-shot on INCART, S-recall transfers (0.866) but S-precision is 0.126 — roughly seven false S alarms per true one. Calibrating to INCART's class prior does not fix it.
- Single lead, no P-wave features. P- and T-wave features were tested and degraded inter-patient performance.
- The rigorous estimate is itself fragile. Record 209's concentration of S-beats makes cross-validation unstable on this dataset.
- Adversarial training not reproduced. The attention CNN's architecture was reproduced, not its subject-invariance training, which is what the original authors credit for their higher score.
- CNN timings cover inference only. For the CNNs, preprocessing is assumed to run elsewhere; the LDA's figures include on-device feature extraction.
06Technology
- 01Python
- 02TensorFlow Lite
- 03CNN
- 04TinyML
- 05ESP32
07Related work