Phase E · Continued pre-training and new corpus
In progress in the Development project: continued pre-training of the Saga 3B base is complete, and work continues on a 15.45-billion-parameter model built from that base.
The Development project (EHPC-DEV-2026D09-196) has been running since 18 September 2026: first multi-node training measurements, then two stages of continued pre-training (22 September, 21.3 billion tokens). Validation loss on an unchanged set fell from 2.5105 to 2.4840. The next stage, on the unused remainder of the existing corpus, was stopped on 23 September after about 9 billion tokens because the loss did not move. Further continued pre-training therefore needs new material.
New corpus (ready on 25 September): deduplication reduced 167.9 million documents from five pools to 151.1 million, which the 64k tokenizer turns into 105.4 billion tokens. A full scan against the evaluation texts found coverage below 1% in each of the exam's three parts, lower than in the Saga corpus. Continued pre-training ran in stages of 20 billion tokens with about 92% new material, against a required minimum of 60%. Stages 3.1–3.3 lowered validation loss from 2.4841 to 2.4704, and a learning-rate decay from the end of stage 3.3 brought it to 2.4667, the lowest value of the Saga 3B base in the project. Stages 3.4 and 3.6 in the FSDP setup did not lower the loss; the stage 3.5 mixture was used for growth trials.
Since 27 September the Development project has been used to prepare the growth. We tested the recipe first on small models of the same architecture. Doubling the number of layers at once proved unstable at the 5.8-billion scale. Two methods that preserve the model's function were stable: inserting a copy with zeroed output after every third layer, and exactly doubling the width. This is how the 15.45-billion-parameter model (48 layers × 5120) was built directly from the 3B base on 29 September. After 6.3 billion tokens of training, ending with a learning-rate decay, its validation loss is 2.4944 against 2.4667 for the base. A small-model test on 1 October explains this result: a model grown this way catches up with the control only after about 12–15% of the tokens the base was trained on, and then pulls ahead. At 58% its loss was 0.17 lower than the control's, and the lead was larger at a higher learning rate.