Benchmarks
All numbers below are from our own evaluation harness unless stated
otherwise. Mean-6 is the unweighted mean of HellaSwag, PIQA, WinoGrande,
ARC-Easy, ARC-Challenge, and LAMBADA. HellaSwag, PIQA, ARC-Easy, and
ARC-Challenge are reported as acc_norm; WinoGrande and LAMBADA
as acc.
No benchmark data is trained on, and no checkpoint is selected on a benchmark. Where our own held-out estimates have turned out to be wrong, we say so — see littlerock-1M-arithmax.
Boris-1.3-125M, pass by pass
Seven continued-pretraining passes on the Boris-125M base checkpoint, each measured individually.
| Task | Boris-125M | +FineWeb-Edu | +DCLM | +FW-Edu ×3 | +FineWeb-Edu | Boris-1.3-125M |
|---|---|---|---|---|---|---|
| HellaSwag | 29.33 | 29.40 | 29.24 | 29.50 | 29.79 | 29.68 |
| PIQA | 59.74 | 60.72 | 60.72 | 61.43 | 60.61 | 61.32 |
| WinoGrande | 49.72 | 50.36 | 50.59 | 50.28 | 51.70 | 52.72 |
| ARC-Easy | 41.75 | 41.41 | 41.54 | 41.79 | 43.01 | 42.51 |
| ARC-Challenge | 23.89 | 24.74 | 23.72 | 24.40 | 25.09 | 24.23 |
| LAMBADA | 22.86 | 23.17 | 24.63 | 24.74 | 23.23 | 25.79 |
| Mean-6 | 37.88 | 38.30 | 38.41 | 38.69 | 38.91 | 39.38 |
The seven passes, in order: FineWeb-Edu 0.6B (~6.9h), DCLM-baseline 1.0B (~11.9h), FineWeb-Edu 0.1B (~1.2h), FineWeb-Edu 0.1B (~1.2h), FineWeb-Edu 0.1B (~1.2h), FineWeb-Edu-leaning 0.91B (~7.8h, required a restart), DCLM-baseline 0.6B (~7.1h). Wall-clock figures other than pass 6 are estimated.
Boris-1.3-75M
| Task | Boris-1.3-75M |
|---|---|
| HellaSwag | 27.57 |
| PIQA | 59.41 |
| WinoGrande | 51.54 |
| ARC-Easy | 40.57 |
| ARC-Challenge | 23.46 |
| LAMBADA | 18.16 |
| Mean-6 | 36.79 |
Training: 1.55B tokens of FineWeb-Edu as base, then three continued passes — DCLM-baseline 1.5B, FineWeb-Edu 0.3B, FineWeb-Edu 0.6B — for ~3.95B tokens total.
What the mixture does
The consistent finding across both sizes: DCLM-baseline improves the fluency and coherence tasks (LAMBADA, WinoGrande) but tends to cost ARC-Easy and ARC-Challenge. At 125M, ordering the passes so FineWeb-Edu comes first and DCLM arrives later avoided the ARC regression that the 75M run showed. That ordering effect is the single most useful thing these runs produced.
littlerock-1M-arithmax on ArithMark-3
| Set | Score | Measured by |
|---|---|---|
| Public | 40.4% | OpenCerebral |
| Private held-out (memorization) | 35.5% | AxiomicLabs |
| Private variety (rephrased) | 21.5% | AxiomicLabs |
| Base model, pre-finetune | 25.2% | OpenCerebral |
Chance is 25%. The full analysis of why the third row happened is on the littlerock page.
