Diego Garcia

Back to logs

> karyo

Ten Hours of LeJEPA Buy 0.04 Over Six Colour Numbers

A ViT-Tiny pretrained with LeJEPA on 174,110 histology tiles shows no collapse, an isotropic embedding and a loss that falls for all 100 epochs. Its class token predicts tumour type at 0.755. Six colour statistics of the same window score 0.717, and the untrained encoder 0.714.

Oct 20266 min read

LeJEPA replaces the usual anti-collapse machinery (teacher networks, stop-gradients, centring) with one regulariser, SIGReg, that pushes the embedding distribution towards an isotropic Gaussian. Two of its selling points are that training is stable without tuning and that the training loss tracks downstream quality, so a run can be monitored without labels.

The first held here. The second did not.

The Run

ViT-Tiny/16 with four registers (5.5M parameters), pretrained on H&E tiles from 278 MIDOG++ training cases, canine lymphoma excluded. Each tile yields two global views at 224 px and six local views at 96 px, all cut from one 320 px source region, with stain jitter of ±5% in HED space and ±10% brightness and contrast. SIGReg weight 0.05, 100 epochs of 25,000 tiles, 10.6 hours.

Every diagnostic the method offers looks right. The prediction loss falls from 0.150 to 0.016. The SIGReg statistic falls from 3.54 to 0.94, just under its Gaussian null of 1.06. The per-dimension standard deviation of the embedding rises from 0.85 to 0.997. Nothing collapsed.

The objective improves for 100 epochs; both probes peak by epoch 40

LeJEPA pretraining of a ViT-Tiny on 174,110 H&E tiles, one run. Linear probes on frozen features every 5 epochs. Shaded: after epoch 40.

Prediction lossthe invariance term0.000.050.100.150401000.016SIGReg statisticGaussian null 1.061230401000.939Class-token probebalanced accuracy, chance 0.330.340.380.420401000.419 at epoch 400.372Patch-token probemitosis in this patch, AP0.100.150.200401000.212 at epoch 350.150epoch
Figure 1: The two left panels are what the method optimises; the two right panels are linear probes on frozen features, scored on validation cases the encoder never saw. Both probes peak by epoch 40 and decline while the objective keeps improving. Rank correlation between loss and probe over the 20 probe points: +0.19 for the class token, −0.39 for the patch tokens. A loss that tracked quality would give values near −1.

Against a Floor

The probes in Figure 1 are low in absolute terms, which by itself says little: the class-token probe asks whether the cell at the centre of a 224 px tile is a mitosis, and one cell is a small part of a tile. The informative comparison is against two encoders that learned nothing: the same architecture at random initialisation, and six numbers, the mean and standard deviation of each RGB channel.

100 epochs of pretraining move each probe 0.02 to 0.05 past an untrained encoder

Linear-probe balanced accuracy on unseen validation cases. The floors are six numbers: RGB mean and standard deviation.

0.30.40.50.60.70.8balanced accuracyTumour type from the class token6-way, chance 0.1670.7170.755Crop class from the centre patch token3-way, chance 0.3330.5950.697Crop class from the class token3-way, chance 0.3330.3340.372
colour statisticsuntrained encoderpretrained, epoch 40pretrained, epoch 100
Figure 2: Three probes, each with a colour floor (bar), the untrained encoder (orange) and the pretrained encoder at epochs 40 and 100 (blue). Across all three, 100 epochs of pretraining add 0.02 to 0.05 to what random weights already give.

On tumour type, the untrained class token and the six colour numbers are level (0.714 and 0.717): a random ViT's class token is already a colour summary of its input. Pretraining adds 0.04. On the centre patch token, colour of the centre 32 px gives 0.595, random weights 0.645, pretraining 0.697. On crop class from the class token, the pretrained encoder ends at 0.372 against a chance level of 0.333.

Downstream agrees. Fine-tuned on mitosis classification with three seeds per arm, the pretrained encoder scores 0.549 F1 on held-out lymphoma against 0.597 for random initialisation: a difference of −0.048 [−0.083, −0.019] by paired case bootstrap, with average precision tied at −0.006 [−0.022, +0.009]. In-domain the difference is −0.012 [−0.023, −0.002]. The pretrained arm reaches its best validation score at epochs 12 to 21 against 48 to 75 for random initialisation: it converges faster, to no better a place.

Reading

An objective that asks for views of one tile to agree, with embeddings spread into a Gaussian, is satisfied by any statistic that is constant within a tile and varied across tiles. In H&E the cheapest such statistic is stain: every view inherits its tile's colour, and jitter of ±5% does not remove it. The numbers above are what that solution would look like. They do not prove it is the solution found: the run that would test it, the same pretraining with colour decorrelated across views, was not made.

The in-domain results in the LeJEPA paper that motivated this run were obtained with ConvNeXt-V2 Nano and ResNet-34. This is a ViT-Tiny, where nothing in the architecture prefers local structure over a global colour summary.

Limits: one pretraining run, so the probe curves carry no interval; probes trained on 3,000 windows and scored on 1,500; one augmentation strength.

What is still open

Whether the probes would keep rising with the loss if colour were made uninformative across views, or whether a class-token objective on tiles this homogeneous has little else to find.