Diego Garcia

Back to logs

> karyo

Colour Is Worth 0.24 AP on the Training Tumours and 0.10 on the Unseen One

Four mitosis classifiers trained in colour, evaluated on greyscale copies of the same crops. On the tumours they trained on, removing colour costs 0.24 average precision. On a held-out tumour it costs 0.10, and the domain gap in ranking changes sign.

Oct 20265 min read

The default explanation for a domain gap in H&E histology is stain: laboratories and scanners render the same tissue in different purples, and the remedy is stain augmentation or normalisation. A cheap way to ask how much of a model's evidence is colour, and whether that evidence survives a change of domain, is to take colour away at test time and see what is left.

Measurement

Three-class crop classifiers (mitotic figure, imposter, background) on MIDOG++, canine lymphoma held out, three seeds per arm. Four arms that differ in everything except the question: ViT-Tiny/8 on 64 px crops, the same without stain jitter in training, ViT-Tiny/16 on 128 px crops, and the latter initialised from a self-supervised encoder. Each model is scored on its validation crops (the six training tumours, unseen cases) and on held-out lymphoma, once as trained and once with luminance copied to all three channels.

Take colour away and the domain gap in ranking reverses

Mitotic-figure average precision of four colour-trained classifiers, on colour and on greyscale input. Mean of 3 seeds.

0.500.550.600.650.700.750.800.85colour inputgreyscale inputtraining tumours0.800.56training tumourscolour was worth 0.24held-out lymphoma0.750.65held-out lymphomacolour was worth 0.10
training tumours, one line per armheld-out lymphoma, one line per arm
Figure 1: One line per arm. In colour, the training tumours rank better than the held-out one in every arm (0.80 against 0.75). In greyscale the order is reversed in every arm (0.56 against 0.65). The lines cross because colour carries more than twice as much of the ranking on the tumours the model was trained on.
ArmColour AP lost, training tumoursColour AP lost, held-out lymphoma
ViT-Ti/8, 64 px0.2410.096
ViT-Ti/8, no stain jitter0.2360.134
ViT-Ti/16, 128 px0.2370.079
ViT-Ti/16, self-supervised init0.2340.097

The asymmetry is in ranking only. Balanced accuracy, which depends on which class wins the argmax, falls by 0.21 to 0.32 on both sides: greyscale is out of distribution for a colour-trained model and its decisions degrade everywhere. What differs is how much of the ordering of crops by mitosis score was built on colour.

Reading

The 0.05 gap in average precision between training tumours and lymphoma is not a deficit of shape evidence. On shape and texture alone these models rank lymphoma mitoses better than the mitoses of tumours they trained on. The gap is accounted for by colour: evidence worth 0.24 at home delivers 0.10 on the new tumour.

That is different from a stain shift, and the usual remedy does not touch it. Training the first arm without stain jitter does not hurt on lymphoma: the difference is +0.033 F1 [−0.004, +0.069] and +0.013 AP [−0.008, +0.032] in favour of no jitter. The errors point the same way. With lymphoma held out, mitosis recall falls from 0.76 to 0.28, and 68% of the true mitoses are labelled as imposters, at most 4% as background. The model sees a cell of interest and declines to call it a mitosis.

A reading consistent with all three observations: in the training tumours, a mitotic figure is partly recognised by being darker and bluer than its neighbours. In lymphoma, where most nuclei are dark and dense, that contrast carries less information, so a cue that is valid at home is uninformative on the new tumour without the stain having moved at all.

Limits: only one jitter strength (±5% in HED space) was tested, so this does not show that stronger stain augmentation fails. The greyscale average precisions are the colour values at each model's selected checkpoint plus the measured differences. The task is classification of crops centred on a known point, not detection.

What is still open

Whether a model trained on lymphoma also gets only 0.10 from colour there, which would make this a property of the tumour, or gets the full 0.24, which would make it a property of transfer. That model was not evaluated in greyscale.