Diego Garcia

Back to briefs

> karyo

A Transferred Threshold Invents One Gap and Hides Another

Three mitosis detectors, one unseen tumour type. At the threshold chosen on validation, a 20M YOLO beats a 6M CNN by 0.105 F1 and the CNN cannot be told apart from a ViT. Read each model at its own best threshold and the first gap closes to 0.03 while the second opens to 0.16. Validation had the order right all along.

Oct 202610 min read

Leave-one-domain-out evaluation of a detector has a step that rarely gets a sentence in the methods section. The model outputs scores; F1 needs a threshold; the held-out domain must not be touched, so the threshold is chosen on validation data from the training domains and carried over. That rule is correct. It also means every held-out F1 is a joint measurement of two things: how well the model separates objects from background on the new domain, and whether its score scale landed in the same place there. Tables of held-out F1 are read as if they measured the first.

This brief measures how much of each they contain, on one fold of MIDOG++.

Setup

MIDOG++ has 503 cases across seven tumour types, each case a 2 mm² region with mitotic figures annotated as points. Canine lymphoma is held out: 278 training cases from the other six tumours, 71 validation cases from the same six (1,166 figures), and 12 lymphoma cases (840 figures) read once at the end. A detection counts if it lies within 7.5 µm of a figure; F1 is computed from pooled counts.

Three detectors, all trained from scratch on the six training tumours, three seeds each: a ResNet with a heatmap head (6.1M parameters), a ViT-Tiny/16 with the same head (6.2M), and YOLO11m (20.1M). For each run the threshold is the one that maximises pooled validation F1.

Because every candidate detection was saved, the held-out F1 can also be computed at every other threshold. The best of those is a diagnostic, not a result: it is selected on the 12 cases it is scored on and could never be reported as performance. It answers one question only: what did the detections contain before the threshold was applied.

Two Readings of the Same Detections

The held-out F1 curve peaks far to the left of the threshold that validation picks

Detection F1 against score threshold, canine lymphoma held out. Mean of 3 seeds, band = min to max.

F1score thresholdCNN (6.1M)0.21 F1 lost to the threshold00.20.40.60.800.51threshold 0.54val 0.7320.5470.760YOLO11m (20.1M)0.08 F1 lost to the threshold00.51threshold 0.37val 0.7500.6520.733Karyo-T ViT (6.2M)0.11 F1 lost to the threshold00.51threshold 0.51val 0.5470.4840.595
validation, the six training tumoursheld-out tumour, in the model's colour
Figure 1: Each panel is one model. The grey curve is validation, and the vertical line marks its maximum, which is the threshold the protocol uses. The coloured curve is the held-out tumour. Its maximum sits 0.14 to 0.29 lower on the score axis, and the protocol reads it on the falling slope.
Validation F1Held-out F1, validation thresholdHeld-out F1, own best thresholdLost to the threshold
CNN0.7320.547 ± 0.0300.760 ± 0.0030.213 [0.173, 0.263]
YOLO11m0.7500.652 ± 0.0180.733 ± 0.0060.081 [0.053, 0.134]
ViT0.5470.484 ± 0.0340.595 ± 0.0040.111 [0.062, 0.182]

Mean ± standard deviation over three seeds. Intervals are a paired bootstrap over the 12 held-out cases, 10,000 draws, with the best threshold chosen again inside each draw.

The loss is not a selection artefact of a spiky curve. The CNN stays within 0.01 of its best held-out F1 for every threshold from 0.19 to 0.30, YOLO from 0.07 to 0.21.

What matters more than the size of the loss is that it differs by model, which changes every comparison:

Difference in F1ValidationHeld-out, validation thresholdHeld-out, own best threshold
YOLO11m minus CNN+0.018 [+0.004, +0.035]+0.105 [+0.060, +0.148]−0.026 [−0.055, +0.001]
CNN minus ViT+0.185 [+0.157, +0.218]+0.063 [−0.036, +0.169]+0.161 [+0.088, +0.249]

Validation says YOLO and the CNN are close and both are far ahead of the ViT. The standard held-out reading says something else on both counts: a clear YOLO lead, and a CNN whose advantage over the ViT has an interval spanning zero. With the threshold removed from the comparison, the held-out tumour agrees with validation. The 0.105 was manufactured by the threshold and the 0.161 was buried by it.

The threshold-free metrics were saying this already. Held-out average precision is 0.816 for the CNN, 0.764 for YOLO and 0.615 for the ViT; FROC-AUC is 0.470, 0.446 and 0.253.

What Moves Is the Score of True Figures

On the unseen tumour the detectors go quiet: recall falls, precision rises

Recall and precision against score threshold. The vertical line is the threshold chosen on validation.

recallprecisionscore thresholdCNN (6.1M)00.510.710.3900.5100.510.760.94YOLO11m (20.1M)0.730.5200.510.770.88Karyo-T ViT (6.2M)0.500.3500.510.620.80
validation, the six training tumoursheld-out tumour, in the model's colour
Figure 2: At the validation threshold, recall on the held-out tumour is 0.15 to 0.32 lower than on validation, and precision is 0.11 to 0.19 higher. All three detectors under-call the unseen tumour.

The usual picture of domain shift is a model confused by unfamiliar tissue and firing on the wrong things. That is not what happens here. At the validation threshold the CNN produces 0.95 false positives per mm² on lymphoma against 1.8 on validation; YOLO 2.5 against 1.8; the ViT 3.0 against 2.6. False positives barely change. What falls is the score assigned to true figures: recall at that threshold drops from 0.71 to 0.39 for the CNN, 0.73 to 0.52 for YOLO, 0.50 to 0.35 for the ViT.

The detectors still rank lymphoma figures above lymphoma background. They are less sure of them, and a fixed threshold converts lower confidence into silence. The same signature appeared earlier in crop classifiers trained on the same split: mitosis recall fell from 0.76 to 0.28 on held-out lymphoma while precision rose from 0.77 to 0.83.

This also explains why YOLO loses least. Its validation threshold already sits low (0.33 to 0.41, against 0.51 to 0.57 for the CNN) and its held-out curve falls slowly from there. That is a property of how its scores are distributed, and it has deployment value: with one frozen threshold, YOLO is the safer model. It is not evidence that it finds more mitoses.

Density Alone Does Not Do This

Lymphoma is the densest tumour in the dataset: 35 figures per mm² against 8.2 across validation. A dense domain rewards a lower threshold whatever the model, because each false positive is diluted by more true positives. That would make the loss an effect of prevalence, with no shift in the scores needed.

Density does not explain it: only the unseen tumour pays for the pooled threshold

F1 a domain would gain from its own best threshold, against its mitotic-figure density. 7 tumours, 3 models.

0.000.050.100.150.200.2505101520253035mitotic figures per mm² of tissueF1 lost at the pooled val thresholdthe other five training tumoursat most 0.04canine mast cell tumourtrained on, 26 figures per mm², loses nothingcanine lymphomaheld out, 35 figures per mm²
CNN (6.1M)YOLO11m (20.1M)Karyo-T ViT (6.2M)
Figure 3: Each point is one tumour type for one model: the F1 it would gain from its own best threshold over the pooled validation threshold. Six of the tumours were trained on. Points are offset horizontally by model so none is hidden.

Validation contains its own control. Canine mast cell tumour has 26 figures per mm², three times the validation average, and all three models lose at most 0.006 there at the pooled threshold. The other five training tumours lose at most 0.04, and that figure is inflated, since a per-tumour optimum on 7 to 23 regions is itself optimistic. Canine lung is the training tumour with the lowest recall at the pooled threshold (0.26 to 0.52) and it loses 0.02.

So among tumours the model has seen, neither high density nor low recall moves the optimum. Lymphoma has both, and it is the one tumour the model has not seen. One fold cannot separate "unseen" from "lymphoma".

Seed Variance Is Threshold Variance

In all three models, the seed whose validation threshold came out lowest has the highest held-out F1. The CNN's three seeds chose 0.57, 0.51 and 0.53 and scored 0.530, 0.581 and 0.529. Across seeds the standard deviation of held-out F1 is 0.030 at the validation threshold and 0.003 at the best one; for the ViT, 0.034 and 0.004. Almost all of what looks like training noise on the held-out domain is the second decimal of a threshold.

What This Does and Does Not Show

A 2026 audit of RetinaNet on MIDOG++ already reported that F1 and calibration diverge under leave-one-domain-out. What these runs add is scale relative to the comparisons people draw: on this fold the threshold term is as large as the difference between architectures, it is not shared across models, and it can reverse an ordering that validation and every threshold-free metric agree on.

It does not show that the threshold rule should change. Choosing the threshold on held-out data would be leakage, and none of the best-threshold numbers above is a result.

What is still open

  • This is one fold with 12 held-out cases. Whether the other six tumours behave this way when held out is unknown.
  • The loss needs a dense tumour and deflated scores together. Lymphoma supplies both, so the data cannot say which one a new domain must have.
  • If scores on an unseen domain deflate while false positives hold still, a threshold could in principle be set from unlabelled quantities of the new domain, such as the density of detections. Any such rule would have to be chosen on a training domain held out from training, never on the test domain.
  • The MIDOG++ paper reports RetinaNet at 0.73 on lymphoma when trained on it and 0.57 when it is held out. The published numbers cannot say how much of that 0.16 is threshold.