> karyo
A Transferred Threshold Invents One Gap and Hides Another
Three mitosis detectors, one unseen tumour type. At the threshold chosen on validation, a 20M YOLO beats a 6M CNN by 0.105 F1 and the CNN cannot be told apart from a ViT. Read each model at its own best threshold and the first gap closes to 0.03 while the second opens to 0.16. Validation had the order right all along.
Leave-one-domain-out evaluation of a detector has a step that rarely gets a sentence in the methods section. The model outputs scores; F1 needs a threshold; the held-out domain must not be touched, so the threshold is chosen on validation data from the training domains and carried over. That rule is correct. It also means every held-out F1 is a joint measurement of two things: how well the model separates objects from background on the new domain, and whether its score scale landed in the same place there. Tables of held-out F1 are read as if they measured the first.
This brief measures how much of each they contain, on one fold of MIDOG++.
Setup
MIDOG++ has 503 cases across seven tumour types, each case a 2 mm² region with mitotic figures annotated as points. Canine lymphoma is held out: 278 training cases from the other six tumours, 71 validation cases from the same six (1,166 figures), and 12 lymphoma cases (840 figures) read once at the end. A detection counts if it lies within 7.5 µm of a figure; F1 is computed from pooled counts.
Three detectors, all trained from scratch on the six training tumours, three seeds each: a ResNet with a heatmap head (6.1M parameters), a ViT-Tiny/16 with the same head (6.2M), and YOLO11m (20.1M). For each run the threshold is the one that maximises pooled validation F1.
Because every candidate detection was saved, the held-out F1 can also be computed at every other threshold. The best of those is a diagnostic, not a result: it is selected on the 12 cases it is scored on and could never be reported as performance. It answers one question only: what did the detections contain before the threshold was applied.
Two Readings of the Same Detections
The held-out F1 curve peaks far to the left of the threshold that validation picks
Detection F1 against score threshold, canine lymphoma held out. Mean of 3 seeds, band = min to max.
| Validation F1 | Held-out F1, validation threshold | Held-out F1, own best threshold | Lost to the threshold | |
|---|---|---|---|---|
| CNN | 0.732 | 0.547 ± 0.030 | 0.760 ± 0.003 | 0.213 [0.173, 0.263] |
| YOLO11m | 0.750 | 0.652 ± 0.018 | 0.733 ± 0.006 | 0.081 [0.053, 0.134] |
| ViT | 0.547 | 0.484 ± 0.034 | 0.595 ± 0.004 | 0.111 [0.062, 0.182] |
Mean ± standard deviation over three seeds. Intervals are a paired bootstrap over the 12 held-out cases, 10,000 draws, with the best threshold chosen again inside each draw.
The loss is not a selection artefact of a spiky curve. The CNN stays within 0.01 of its best held-out F1 for every threshold from 0.19 to 0.30, YOLO from 0.07 to 0.21.
What matters more than the size of the loss is that it differs by model, which changes every comparison:
| Difference in F1 | Validation | Held-out, validation threshold | Held-out, own best threshold |
|---|---|---|---|
| YOLO11m minus CNN | +0.018 [+0.004, +0.035] | +0.105 [+0.060, +0.148] | −0.026 [−0.055, +0.001] |
| CNN minus ViT | +0.185 [+0.157, +0.218] | +0.063 [−0.036, +0.169] | +0.161 [+0.088, +0.249] |
Validation says YOLO and the CNN are close and both are far ahead of the ViT. The standard held-out reading says something else on both counts: a clear YOLO lead, and a CNN whose advantage over the ViT has an interval spanning zero. With the threshold removed from the comparison, the held-out tumour agrees with validation. The 0.105 was manufactured by the threshold and the 0.161 was buried by it.
The threshold-free metrics were saying this already. Held-out average precision is 0.816 for the CNN, 0.764 for YOLO and 0.615 for the ViT; FROC-AUC is 0.470, 0.446 and 0.253.
What Moves Is the Score of True Figures
On the unseen tumour the detectors go quiet: recall falls, precision rises
Recall and precision against score threshold. The vertical line is the threshold chosen on validation.
The usual picture of domain shift is a model confused by unfamiliar tissue and firing on the wrong things. That is not what happens here. At the validation threshold the CNN produces 0.95 false positives per mm² on lymphoma against 1.8 on validation; YOLO 2.5 against 1.8; the ViT 3.0 against 2.6. False positives barely change. What falls is the score assigned to true figures: recall at that threshold drops from 0.71 to 0.39 for the CNN, 0.73 to 0.52 for YOLO, 0.50 to 0.35 for the ViT.
The detectors still rank lymphoma figures above lymphoma background. They are less sure of them, and a fixed threshold converts lower confidence into silence. The same signature appeared earlier in crop classifiers trained on the same split: mitosis recall fell from 0.76 to 0.28 on held-out lymphoma while precision rose from 0.77 to 0.83.
This also explains why YOLO loses least. Its validation threshold already sits low (0.33 to 0.41, against 0.51 to 0.57 for the CNN) and its held-out curve falls slowly from there. That is a property of how its scores are distributed, and it has deployment value: with one frozen threshold, YOLO is the safer model. It is not evidence that it finds more mitoses.
Density Alone Does Not Do This
Lymphoma is the densest tumour in the dataset: 35 figures per mm² against 8.2 across validation. A dense domain rewards a lower threshold whatever the model, because each false positive is diluted by more true positives. That would make the loss an effect of prevalence, with no shift in the scores needed.
Density does not explain it: only the unseen tumour pays for the pooled threshold
F1 a domain would gain from its own best threshold, against its mitotic-figure density. 7 tumours, 3 models.
Validation contains its own control. Canine mast cell tumour has 26 figures per mm², three times the validation average, and all three models lose at most 0.006 there at the pooled threshold. The other five training tumours lose at most 0.04, and that figure is inflated, since a per-tumour optimum on 7 to 23 regions is itself optimistic. Canine lung is the training tumour with the lowest recall at the pooled threshold (0.26 to 0.52) and it loses 0.02.
So among tumours the model has seen, neither high density nor low recall moves the optimum. Lymphoma has both, and it is the one tumour the model has not seen. One fold cannot separate "unseen" from "lymphoma".
Seed Variance Is Threshold Variance
In all three models, the seed whose validation threshold came out lowest has the highest held-out F1. The CNN's three seeds chose 0.57, 0.51 and 0.53 and scored 0.530, 0.581 and 0.529. Across seeds the standard deviation of held-out F1 is 0.030 at the validation threshold and 0.003 at the best one; for the ViT, 0.034 and 0.004. Almost all of what looks like training noise on the held-out domain is the second decimal of a threshold.
What This Does and Does Not Show
A 2026 audit of RetinaNet on MIDOG++ already reported that F1 and calibration diverge under leave-one-domain-out. What these runs add is scale relative to the comparisons people draw: on this fold the threshold term is as large as the difference between architectures, it is not shared across models, and it can reverse an ordering that validation and every threshold-free metric agree on.
It does not show that the threshold rule should change. Choosing the threshold on held-out data would be leakage, and none of the best-threshold numbers above is a result.
What is still open
- This is one fold with 12 held-out cases. Whether the other six tumours behave this way when held out is unknown.
- The loss needs a dense tumour and deflated scores together. Lymphoma supplies both, so the data cannot say which one a new domain must have.
- If scores on an unseen domain deflate while false positives hold still, a threshold could in principle be set from unlabelled quantities of the new domain, such as the density of detections. Any such rule would have to be chosen on a training domain held out from training, never on the test domain.
- The MIDOG++ paper reports RetinaNet at 0.73 on lymphoma when trained on it and 0.57 when it is held out. The published numbers cannot say how much of that 0.16 is threshold.