SKALDBETA

Project documentation

Data, training
and method

How Saga 3B was built — a Polish base model in the Skald family, trained from scratch on the JUPITER supercomputer.

Status from the report

Editorial snapshot: 19 September 2026 · 18:17:19 Europe/Warsaw. Reading the current published report requires JavaScript.

Job-history reading: 19 September 2026, 18:17:19 (Warsaw). Measurements from the FAKTY document: 19 September 2026, 18:17:09 (Warsaw).

The full Saga 3B training run is complete.

Recipe, settings and evaluation from the model card (README.md) and receptura-podsumowanie.json, read on 19 September 2026. This page describes the base model, which predicts how text continues. Links to source files require access to the private repository.

Project phases

Reproduction, tokenizer selection, corpus-composition comparison and full Saga 3B training are complete. Continued pre-training on the new corpus is complete; work on a 15.45-billion-parameter model built from Saga 3B is in progress in the Development project, and its main training in the Fast Lane allocation awaits dates from JSC. The weights stay private. Phase statuses refresh from the published report.

  1. Complete

    Phase A · Reproduction and scaling

    Reproduction and scaling measurements are complete.

    Validation loss in the JUPITER trial: 2.9847; local reference: 2.9867. Difference −0.0020. Trial duration: 13 min 11 s.

    Largest measured scale in the report: 64 GPUs; efficiency relative to the baseline configuration: 89.5%; throughput: 9.0 million tokens/s. These measurements describe our training trials.

  2. Complete

    Phase B · Tokenizer selection

    The 64k tokenizer was selected, with a mean of 1.4299 bits per character.

    Saga uses its own byte-level BPE tokenizer, trained on Polish text, with 65,536 tokens. Bits per character allow comparison between models with different tokenizers.

  3. Complete

    Phase C · Corpus composition

    The corpus-composition comparison is complete.

    The shares of five pools were selected after six smaller ablation runs, evaluated on a separate exam. The recipe below describes the completed training file.

  4. Complete

    Phase D · Training Saga 3B

    The full Saga 3B training run is complete.

    Training ended on 14 September 2026, after 186,688 updates in six jobs on 64 GH200 GPUs. It processed 195,756,556,288 tokens, or about 195.8 billion.

    Report values and the separate technical trial

    The report’s plan fields contain 186,688 updates, 195.8 billion tokens, and an estimate of 68 hours on 64 GPUs. The time estimate is not a measurement of the completed training run.

    Separate technical trial: 13,883 tokens/s/GPU, 50.2% MFU, and 76.1 GiB of memory per GPU. Averages for the full run appear in the training settings section.

  5. In progress

    Phase E · Continued pre-training and new corpus

    In progress in the Development project: continued pre-training of the Saga 3B base is complete, and work continues on a 15.45-billion-parameter model built from that base.

    The Development project (EHPC-DEV-2026D09-196) has been running since 18 September 2026: first multi-node training measurements, then two stages of continued pre-training (22 September, 21.3 billion tokens). Validation loss on an unchanged set fell from 2.5105 to 2.4840. The next stage, on the unused remainder of the existing corpus, was stopped on 23 September after about 9 billion tokens because the loss did not move. Further continued pre-training therefore needs new material.

    New corpus (ready on 25 September): deduplication reduced 167.9 million documents from five pools to 151.1 million, which the 64k tokenizer turns into 105.4 billion tokens. A full scan against the evaluation texts found coverage below 1% in each of the exam's three parts, lower than in the Saga corpus. Continued pre-training ran in stages of 20 billion tokens with about 92% new material, against a required minimum of 60%. Stages 3.1–3.3 lowered validation loss from 2.4841 to 2.4704, and a learning-rate decay from the end of stage 3.3 brought it to 2.4667, the lowest value of the Saga 3B base in the project. Stages 3.4 and 3.6 in the FSDP setup did not lower the loss; the stage 3.5 mixture was used for growth trials.

    Since 27 September the Development project has been used to prepare the growth. We tested the recipe first on small models of the same architecture. Doubling the number of layers at once proved unstable at the 5.8-billion scale. Two methods that preserve the model's function were stable: inserting a copy with zeroed output after every third layer, and exactly doubling the width. This is how the 15.45-billion-parameter model (48 layers × 5120) was built directly from the 3B base on 29 September. After 6.3 billion tokens of training, ending with a learning-rate decay, its validation loss is 2.4944 against 2.4667 for the base. A small-model test on 1 October explains this result: a model grown this way catches up with the control only after about 12–15% of the tokens the base was trained on, and then pulls ahead. At 58% its loss was 0.17 lower than the control's, and the lead was larger at a higher learning rate.

  6. Planned

    Phase F · Growing Saga

    Fast Lane allocation awarded for the main training of the 15.45-billion-parameter model built from Saga 3B. JSC will set the dates.

    Allocation EHPC-AIF-2026FL01-1001: 50,000 GPU-hours for 3 months, awarded on 25 September 2026. Plan from the proposal: first increase the depth from 36 to 72 layers (about 5.8 billion parameters), then the model width (about 15 billion). Each step is preceded by a pilot with a criterion recorded before it starts; if a step worsens the evaluation, the smaller working model stays. The 15.45-billion model was built differently from the proposal: with 48 layers and doubled width, because doubling the layers at once was unstable (see phase E). Its main training falls within the Fast Lane allocation.

  7. Planned

    Phase G · Evaluation and publication

    The Saga 3B weights stay in a private repository. What remains to be published is the report on the method and the recipe.

    The final checkpoint was evaluated. Model repository — restricted access ↗

Corpus recipe

The completed corpus combines five pools of Polish text after global deduplication. The shares refer to tokens in the training file.

216.8 billion

Tokens in the training file.
Separate validation file: 10 million tokens.

Training-file composition from the completed recipe
Pool and sourcesToken shareTokens in filePasses over the pool
Web: FineWeb-2 + HPLT 2.0 cleaned89.6%194.17 billion≈1.04
FinePDFs, additional quality filter8.9%19.29 billion1
Wikipedia1.1%2.34 billion4
Wikisource0.4%0.78 billion4
Wolne Lektury0.1%0.21 billion4

Shares are rounded separately to 0.1 percentage points, so they sum to 100.1%. Passes describe copies in the training file, not guaranteed exposure of every token during training. Exact values: receptura-podsumowanie.json.

Deduplication

328 million documents

328,007,101 documents were processed. 53,819,391 duplicates were removed: about 53.8 million documents and 33.1 billion tokens. A further 30 FineWeb-2 documents containing a literal end-of-text marker were rejected.

Data split

Validation before mixing

Whole unique documents were reserved for validation before the pools were mixed. Unused tails of those documents did not enter training. The validation file contains 10 million tokens.

How duplicates were removed

A 64-bit hash identified candidates, followed by verification on normalised decoded text. Sources listed earlier in the mixture definition took precedence. Deduplication also covered repetitions between sources.

Provenance and source terms

The corpus uses Polish subsets of FineWeb-2, HPLT 2.0 cleaned and FinePDFs; Wikipedia from 1 November 2023; Wikisource from 1 September 2026; and the Wolne Lektury corpus. Sources and their terms of use are documented in the model card.

Training settings

Settings for the completed run and its average measurements come from the base-model card.

Model

About 3 billion parameters

A decoder-only transformer in the GPT-2 layout: 36 layers, 20 heads, width 2,560. Context: 2,048 tokens. LayerNorm without bias, GELU, learned positions and tied input/output embeddings.

Hardware and run

64 GH200 GPUs

16 JUPITER Booster nodes, six jobs. Training finished on 14 September 2026 at update 186,688. Code based on nanoGPT, bf16 mixed precision, torch.compile and DDP.

Batch and optimisation

Each update covered 1,048,576 tokens in sequences of 2,048 tokens. AdamW optimiser: β₂ = 0.95, weight decay 0.1, and gradient clipping at 1.0.

Learning-rate schedule

Warm-up over 3,000 updates, then a constant rate of 3e-4 and linear decay to 3e-5 over the last 15% of training.

Measurements from the full run

Mean throughput: 12,329 tokens/s per GPU; mean MFU: 46.6%. Mean training loss fell from 5.32 at the start to 2.08 in the last 2,000-update window.

Source: model card — Training.

Evaluation and limits

The final checkpoint was evaluated after 186,688 updates. Lower perplexity and fewer bits per character indicate better text prediction on a given set.

Base-model evaluation from the model card
SetPerplexityBits per characterShared 13-token windows
Validation slice: 1,048,576 tokens12.06——
Encyclopedia after the data cut-off10.290.7260.80%
Web text after the cut-off11.210.6550.41%
Prose after the cut-off15.451.0140.05%

A base model

The base version of Saga 3B predicts how text continues. It is not instruction-tuned or aligned for conversation. The table describes this base model.

History and technical trials

The local Canto 33M model from August 2026 remains a reference for the training process. Its results do not describe Saga 3B. Scaling trials and full training have separate measurements; reported throughput describes our code and configurations, not the performance of JUPITER as a whole or its host.

Source: model card — Evaluation.

Publication

The Saga 3B model repository on Hugging Face remains private. The corpus recipe, the measurements and the work journal are public; the weights are not.

The Saga 3B weights stay in a private repository. What remains to be published is the report on the method and the recipe.

decisioninterface/saga-3b — restricted access ↗

The package contains 55 float32 safetensors weight files in Transformers GPT-2 format. The card documents the lossless conversion of the training export and the comparison of logits and text continuations.

Weight licence stated in the model card: Apache 2.0. The repository remains private. The weight licence does not replace the terms of the source data. Details: model card and condensed recipe.

Epos remains a future large mixture-of-experts model planned outside the current allocation.

Resources and acknowledgement

Allocation EHPC-AIF-2026PG01-1218, EuroHPC AI Factory — Playground, provides 5,000 GPU-hours. The full Saga 3B training run used this allocation on JUPITER Booster at Jülich Supercomputing Centre, with NVIDIA GH200 GPUs.

Allocation EHPC-DEV-2026D09-196 (EuroHPC Development Access), awarded on 10 September 2026, provides 14,000 GPU-hours, or 3,500 node-hours, for 12 months from the start date specified by JSC. It is accounted for separately and supports development of training across multiple nodes.

Allocation EHPC-AIF-2026FL01-1001 (EuroHPC AI Factory — Fast Lane), awarded on 25 September 2026, provides 50,000 GPU-hours on JUPITER Booster for 3 months. JSC will set the start and end dates; there are no extensions. It is accounted for separately and supports growing Saga 3B into a model of about 15 billion parameters.

We acknowledge EuroHPC JU for awarding the project ID EHPC-AIF-2026PG01-1218 access to JUPITER Booster hosted by Jülich Supercomputing Centre (JSC), Germany.

Na zasobach EuroHPC (JUPITER Booster) wykonano: faza A: przebiegi wielowezlowe i pomiary skalowania; faza B: przebiegi potwierdzajace wybor slownika. Poza przydzialem EuroHPC powstaly: trening slownikow i pomiar plodnosci, rozwoj filtrow korpusu, lokalna linia odniesienia oraz kod pakujacy i raportujacy.

Translation of this historical record: EuroHPC resources (JUPITER Booster) were used for phase A runs across multiple nodes and scaling measurements, and phase B runs confirming the tokenizer choice. Tokenizer training and fertility measurements, corpus-filter development, the local reference model, and packaging and reporting code were developed outside the EuroHPC allocation.