Writing / Build Log

It loaded, it ran, and it called 93% of heartbeats normal

An attention CNN quantised, converted, loaded, and ran on an ESP32 without a single error — and called about 93% of heartbeats normal. Two TensorFlow Lite Micro operators were failing silently. Here is how the graph gave them away, and the weight-preserving fixes.

This comes from our BSc thesis on inter-patient ECG classification — a team of four at Bangladesh University of Professionals. Most of the thesis is about evaluation. This is the part where a model that was right on the CPU was wrong on the microcontroller, and nothing said so.

Everything succeeded

The model was worth deploying. Under the standard evaluation protocol it reached a supraventricular F1 of 0.831 — within 0.013 of the best published inter-patient result — and the int8 version scored 0.83 on the CPU.

Post-training quantisation succeeded. Conversion succeeded. The model loaded on the ESP32, allocated a 25 KB tensor arena, and ran at 141 ms per beat, comfortably inside the thesis's 694 ms real-time budget per beat.

Then it predicted class N — normal — for about 93% of beats.

No converter warning. No runtime error. No crash. A classifier that answers "normal" almost every time will even look plausible on an ECG, where most beats are normal. The only signal was that the device and the CPU disagreed about the same beats.

Read the graph, not the logs

With no error to chase, the next step was to parse the converted model's FlatBuffer and list the operators TensorFlow Lite had actually emitted. Two things in that list did not belong in a model with static input shapes.

Failure one: attention that computes its shape at runtime. The CBAM spatial-attention branch averages over channels with keepdims=True. With a dynamic batch dimension, the converter lowers that reshape into a SHAPE → STRIDED_SLICE → PACK chain that works out the output shape at runtime. On TensorFlow Lite Micro that chain returned incorrect values, so the attention gate collapsed to a near-constant map and suppressed the features the classifier depends on.

Failure two: a broken dilated convolution. The residual blocks use convolutions with dilation 3. The converter implements dilation as SPACE_TO_BATCH_ND → CONV_2D → BATCH_TO_SPACE_ND. In this architecture the first valid convolution collapses one axis to height 1, and for those degenerate tensors the int8 SPACE_TO_BATCH_ND kernel produced incorrect output.

Both are silent: the operators exist, are registered, and run. They just return the wrong numbers.

Two fixes that keep every trained weight

Retraining was not the answer; the weights were fine. Both problem operators could be re-expressed with statically shaped primitives, and the reformulated layers have no parameters of their own, so the trained weights transfer exactly.

The attention fix swaps each keepdims reduction for a plain reduction followed by an explicit expand_dims:

# before: the converter emits SHAPE / STRIDED_SLICE / PACK
avgp = reduce_mean(x, axis=-1, keepdims=True)

# after: a single EXPAND_DIMS
avgp = expand_dims(reduce_mean(x, axis=-1), axis=-1)

The channel-attention Reshape((1, 1, C)) became two nested expand_dims the same way. After that, no SHAPE, PACK, or STRIDED_SLICE remained anywhere in the graph.

The dilation fix is arithmetic. A convolution with kernel width 3 and dilation 3 is identical to a dense convolution with kernel width 1 + (3 − 1) × 3 = 7 whose weights are zero except at taps 0, 3, and 6. Copy each trained 1×3 kernel into those three positions of a 1×7 kernel, zero the other four, and the network runs through the plain CONV_2D kernel with no SPACE_TO_BATCH_ND at all. Over eight random inputs, the largest output difference between the original and the reformulated float32 models was 3.6 × 10⁻⁷ — single-precision round-off.

The price of being correct

Three ways to express the same dilated convolution, measured on the same ESP32:

Dilation implementation    Predictions           Latency per beat  Test-set S-F1 (int8)
-------------------------  --------------------  ----------------  --------------------
SPACE_TO_BATCH_ND, 3 taps  broken, almost all N  141 ms            —
Native dilated CONV_2D     correct               806 ms            0.827
Sparse 1×7 CONV_2D         correct               251 ms            0.832

The broken path was the fastest, which is its own warning: a speed number means nothing until the output is checked. The sparse kernel spends seven multiply-accumulates per output where three would do, because TensorFlow Lite Micro cannot skip the zero weights — so it cannot match the broken path's 141 ms. It is still 3.2× faster than the native dilated kernel with identical outputs, and it is what we deployed.

What we would tell ourselves next time

  • A clean conversion is not a correct model. Check device outputs against the CPU reference on real inputs before looking at a latency number.
  • When nothing errors, read the operator list. Ops you did not expect — here SHAPE, PACK, STRIDED_SLICE, SPACE_TO_BATCH_ND — are where to look first.
  • Prefer statically shaped, boring operators on a microcontroller. Both fixes replaced something clever with something plain that the runtime handles well.
  • Fix the graph, keep the weights. Both reformulations are exact, so no accuracy was traded for correctness.

There is a larger lesson in the thesis that this model sits inside: the same attention CNN that scores 0.831 under the standard protocol averages 0.374 when its checkpoint is chosen on held-out patients. Getting it to run correctly on the ESP32 made it deployable. It did not make it the most reliable model — the project record covers that part.