Diego Garcia

● building · rung 08, detection

Karyo

A Vision Transformer built from scratch that looks for dividing cells.

engine
hand-written, from a scalar autodiff engine up
parameters
≤ 25M
hardware
one MacBook
KARYOH&Eabstractnot tissue(x, y, score)384 px

fig. 01

What it is

Finding dividing cells, on a tumour type it has never seen.

Karyo takes H&E histology images and marks mitotic figures, cells caught mid-division, as points. Pathologists count these to grade tumours.

The core problem is domain shift. Evaluation is leave-one-tumour-type-out: train on six tumour types, test on a seventh the model has never seen, across species, labs and scanners.

It exists to find out how far a small, hand-built model gets on a domain it has never seen. It is a research model on a public benchmark, MIDOG++. It is not a diagnostic tool.

The name: karyon, Greek for nucleus, and karyokinesis, the nucleus dividing.

7.5 µm · matching radius
64 px · window overlap
8 px · model stride

fig. 02

Architecture

One forward pass, with the shape at every boundary. Flip the toggles. Hover or tap a block.

mode
backbone

S = 384 · B = 32 · grid 48×48

training window

(B, 3, 384, 384)
(B, 576, 192)
(B, 576, 192)
(B, 581, 192)

(B, 581, 192) → (B, 581, 192)

  • ·pre-norm
  • ·3 heads, d_head 64
  • ·MLP 192 → 768 → 192, exact GELU
  • ·12 × 444,864
transformer blocks
5,338,368
(B, 581, 192)
(B, 581, 192)
(B, 576, 192)
(B, 192, 48, 48)

stride-8 feature map · both backbones meet here

(B, c, S/8, S/8). The two heads are identical in structure, differing only in input width.

(B, 1, 48, 48)
(B, 2, 48, 48)

heatmap logits and offsets run in parallel, and rejoin in decoding

parameter budget · ViT6,188,995 / 25,000,000 cap
025M

backbone subtotal 5,597,952 · total counted at 384 px

fig. 03

Pipeline

The same weights, two kinds of window. Training on the left, inference on the right.

Split, by case

  • Leave one tumour type out.
  • Train and val: the other six. 20% of each tumour's cases go to val, fixed seed 0.
  • Test: the held-out tumour's test cases only.
  • A leak assertion by case ID runs at every launch.

training

Sampled windows

(32, 384, 384, 3) uint8

  • One third near a mitosis, one third near an imposter, one third in tissue.
  • Centre jitter ±S/4 per axis. Shifted to stay inside the image, no padding.
  • Imposters steer sampling only. They are never a target.
  • 156 steps × 32 = 4,992 windows per epoch. 60 epochs = 9,360 steps.

inference

Sliding grid

(16, 512, 512, 3) uint8

  • 512 px windows over the whole region.
  • 64 px overlap, stride 448.
  • Plus one window flush with each far edge.

Preprocessing

(B, 3, S, S) float in [0, 1]

  • Training only, on GPU, in this order: one of the 8 dihedral symmetries (points move with the pixels); HED stain jitter, scale ±0.05 and bias ±0.05 per stain channel; brightness ±0.1 and contrast ±0.1; Gaussian blur, σ ~ U(0, 1) px, 5×5 kernel; clamp to [0, 1].
  • Both: normalise with per-channel mean and std, computed once from 512 training windows and stored in the checkpoint.

Model

(B, 1, S/8, S/8) + (B, 2, S/8, S/8)

  • The same weights see 384 px random windows in training and 512 px grid windows at inference.
  • Out: heatmap logits and offsets on the stride-8 grid.

training

Targets and loss

  • Heatmap (B, 1, 48, 48): Gaussian, σ = 2 cells, max where they overlap, exactly 1 at the point's cell.
  • Offset (B, 2, 48, 48) and mask (B, 1, 48, 48), set at point cells only.
  • CenterNet focal loss (α = 2, β = 4) plus L1 offset loss at point cells, weighted 1 : 1.

inference

Decoding

  • Sigmoid, then peaks: a cell that is the max of its 3×3 neighbourhood and ≥ 0.01.
  • Point = ((col + ox) · 8, (row + oy) · 8) in window pixels.
  • Inner-edge rule: drop detections within 32 px of any window edge that is not the image border.

training

Checkpoint gate

  • 0.6 · F1 + 0.3 · FROC-AUC + 0.1 · AP, averaged over the six val tumour types with equal weight.
  • Evaluated every 5 epochs. The highest gate becomes best.pt: weights, normalisation stats and τ.

inference

Merge

  • Shift to image coordinates, concatenate per case.
  • Greedy suppression from the highest score down, dropping anything closer than 15 px.

inference

Hungarian matching

(x, y, score) per region

  • Matching within 7.5 µm, converted to pixels per case.
  • Threshold swept over 0.01 to 0.99 in steps of 0.01, matching redone at each.
  • τ is picked on training-domain val. Test is read once at that τ.
  • No calibration step: the raw heatmap probability is the score.

fig. 04

The ladder

Rungs 0 to 14. A rung is done when its tests pass, not when the loss goes down.

  1. 00protocoldone
  2. 01autodiffdone
  3. 02MLPdone
  4. 03vectorised enginedone
  5. 04attentiondone
  6. 05transformer blockdone
  7. 06ViTdone
  8. 07pretrainingclosed: negative
  9. 08detectioncurrent
  10. 09conv stem, 2D RoPE, multi-level necknext
  11. 10training recipenext
  12. 11masked-image pretrainingnext
  13. 12foundation-model distillationnext
  14. 13all seven folds, whole-slide testnext
  15. 14deploymentnext

fig. 05

Principles

Why it is built this way.

  1. 01

    Protocol fixed before any model

    Leave-one-domain-out, split by case. The held-out domain touches nothing: no pretraining, no normalisation stats, no early stopping, no threshold. The code asserts it by case ID.

  2. 02

    A frozen NumPy reference

    The NumPy engine from the first three rungs is never modified. Later code is checked against it. Gradient checks run in float64.

  3. 03

    Planted bugs

    The tests get tested. Break the code on purpose, then confirm a test fails.

  4. 04

    Gates written before experiments run

    Each experiment's keep-or-drop number is written down before it runs. Decisions use training-domain validation only. A failed gate is recorded as a result.

fig. 06

Stack

Six items.

Python
 
NumPy
the from-scratch engine
PyTorch
on Apple MPS
pytest
 
Ultralytics YOLO
baseline only
MacBook Pro M5
24 GB

fig. 07

Related writing

Results, experiments and failures live in Briefs and Logs.

Nothing tagged Karyo yet. The slide is still on the stage.