> karyo
Ten Hours of LeJEPA Buy 0.04 Over Six Colour Numbers
A ViT-Tiny pretrained with LeJEPA on 174,110 histology tiles shows no collapse, an isotropic embedding and a loss that falls for all 100 epochs. Its class token predicts tumour type at 0.755. Six colour statistics of the same window score 0.717, and the untrained encoder 0.714.
LeJEPA replaces the usual anti-collapse machinery (teacher networks, stop-gradients, centring) with one regulariser, SIGReg, that pushes the embedding distribution towards an isotropic Gaussian. Two of its selling points are that training is stable without tuning and that the training loss tracks downstream quality, so a run can be monitored without labels.
The first held here. The second did not.
The Run
ViT-Tiny/16 with four registers (5.5M parameters), pretrained on H&E tiles from 278 MIDOG++ training cases, canine lymphoma excluded. Each tile yields two global views at 224 px and six local views at 96 px, all cut from one 320 px source region, with stain jitter of ±5% in HED space and ±10% brightness and contrast. SIGReg weight 0.05, 100 epochs of 25,000 tiles, 10.6 hours.
Every diagnostic the method offers looks right. The prediction loss falls from 0.150 to 0.016. The SIGReg statistic falls from 3.54 to 0.94, just under its Gaussian null of 1.06. The per-dimension standard deviation of the embedding rises from 0.85 to 0.997. Nothing collapsed.
The objective improves for 100 epochs; both probes peak by epoch 40
LeJEPA pretraining of a ViT-Tiny on 174,110 H&E tiles, one run. Linear probes on frozen features every 5 epochs. Shaded: after epoch 40.
Against a Floor
The probes in Figure 1 are low in absolute terms, which by itself says little: the class-token probe asks whether the cell at the centre of a 224 px tile is a mitosis, and one cell is a small part of a tile. The informative comparison is against two encoders that learned nothing: the same architecture at random initialisation, and six numbers, the mean and standard deviation of each RGB channel.
100 epochs of pretraining move each probe 0.02 to 0.05 past an untrained encoder
Linear-probe balanced accuracy on unseen validation cases. The floors are six numbers: RGB mean and standard deviation.
On tumour type, the untrained class token and the six colour numbers are level (0.714 and 0.717): a random ViT's class token is already a colour summary of its input. Pretraining adds 0.04. On the centre patch token, colour of the centre 32 px gives 0.595, random weights 0.645, pretraining 0.697. On crop class from the class token, the pretrained encoder ends at 0.372 against a chance level of 0.333.
Downstream agrees. Fine-tuned on mitosis classification with three seeds per arm, the pretrained encoder scores 0.549 F1 on held-out lymphoma against 0.597 for random initialisation: a difference of −0.048 [−0.083, −0.019] by paired case bootstrap, with average precision tied at −0.006 [−0.022, +0.009]. In-domain the difference is −0.012 [−0.023, −0.002]. The pretrained arm reaches its best validation score at epochs 12 to 21 against 48 to 75 for random initialisation: it converges faster, to no better a place.
Reading
An objective that asks for views of one tile to agree, with embeddings spread into a Gaussian, is satisfied by any statistic that is constant within a tile and varied across tiles. In H&E the cheapest such statistic is stain: every view inherits its tile's colour, and jitter of ±5% does not remove it. The numbers above are what that solution would look like. They do not prove it is the solution found: the run that would test it, the same pretraining with colour decorrelated across views, was not made.
The in-domain results in the LeJEPA paper that motivated this run were obtained with ConvNeXt-V2 Nano and ResNet-34. This is a ViT-Tiny, where nothing in the architecture prefers local structure over a global colour summary.
Limits: one pretraining run, so the probe curves carry no interval; probes trained on 3,000 windows and scored on 1,500; one augmentation strength.
What is still open
Whether the probes would keep rising with the loss if colour were made uninformative across views, or whether a class-token objective on tiles this homogeneous has little else to find.