NEWThird EuroHPC allocation · Fast Lane · 50,000 GPU-hours

Building Saga.

Saga 3B — a Polish language model trained from scratch on the JUPITER supercomputer, with its own tokenizer and its own corpus. This page is its record: every run and every measurement, including the failed ones.

Infrastructure and allocations
EuroHPC JUJülich Supercomputing CentreJUPITER BoosterNVIDIA GH200Decision Interface
01 — Numbers

Scale you can verify

These figures describe our training and its configuration on JUPITER Booster. They are not a performance assessment of JUPITER or its host.
Phase A2 · scaling efficiency at 64 GPUs
89.5%

Aggregate throughput of 9.0 M tok./s on 64 GPUs, relative to ideal scaling from the first point of the series (4 GPUs). This is a measurement of our workload.

Scaling chart and all points — Polish / PL ↗
Target model
3 B

parameters · Saga 3B

Tested scale
64

GH200 GPUs in one training run

Training plan
216.8 B

tokens of exposure · Saga 3B

Planned time
68 h

on 64 GPUs · 206,800 steps

A1 · Reproducing the reference
2.9847

Validation loss on JUPITER compared with 2.9867 from the local 33M baseline.

Difference −0.002013 min 11 s
Probe · 3B model
13,883 tok./s per GPU

A technical probe of the Saga 3B configuration, measuring throughput and memory, separate from the full training run.

MFU · 50.2%76.1 GiB per GPU
Second EuroHPC allocation
14,000 GPU-h

Development Access EHPC-DEV-2026D09-196: 3,500 node-hours for 12 months. Awarded on 10 September 2026; first job on 18 September. Counted separately from the 5,000 GPU-hour Playground allocation and intended for multi-node training development.

Third EuroHPC allocation · Fast Lane
50,000 GPU-h

AI Factory Fast Lane EHPC-AIF-2026FL01-1001 on JUPITER Booster for 3 months. Intended to grow Saga 3B into a model of about 15 billion parameters in two steps; each step is preceded by a pilot with a criterion recorded before it starts. Counted separately from the Playground and Development allocations.

Awarded on 25 September 2026Start and end dates set by JSC · no extensions
02 — Project status

Where we are

Report timestamp

Reproduction and scaling: complete. The report selects the 64k tokenizer for Saga. Corpus-composition comparison: complete. The full Saga 3B training run is complete according to the report. Work on a 15.45-billion-parameter model built from Saga 3B is in progress in the Development project. Its main training in the Fast Lane allocation awaits dates from JSC.

Editorial snapshot: 7 September 2026 · 19:24:40 Europe/Warsaw. Reading the current published report requires JavaScript.

Task history readout: 7 September 2026, 19:24:40 (Warsaw). Measurements and the plan retain their FAKTY source date: 7 September 2026, 17:46:39 (Warsaw).

  1. AComplete

    Reproduce and scale

    Reproduction and scaling: complete.

  2. BComplete

    Choose the tokenizer

    The report selects the 64k tokenizer for Saga.

  3. CComplete

    Corpus composition

    Corpus-composition comparison: complete.

  4. DComplete

    Train Saga 3B

    The full Saga 3B training run is complete according to the report.

  5. EIn progress

    Continued pre-training and new corpus

    In progress in the Development project: continued pre-training of the Saga 3B base is complete, and work continues on a 15.45-billion-parameter model built from that base.

  6. FPlanned

    Growing Saga

    Fast Lane allocation awarded for the main training of the 15.45-billion-parameter model built from Saga 3B. JSC will set the dates.

  7. GPlanned

    Publication

    The Saga 3B weights stay in a private repository. What remains to be published is the report on the method and the recipe.

Completed, running and failed jobs remain in the same project record.Full job registry — Polish / PL
03 — Models and tokenizer

From the baseline to the target

The August baseline, the selected tokenizer and Saga 3B trained on JUPITER.
Phase B · tokenizer selectionComplete

Tokenizer selected.
64k for Saga.

Bits per character · lower = better
TokenizerEncycl.WebProseMean
32k1.28041.11171.99231.4615
48k1.27711.09931.91851.4316
64k · selected1.26281.09251.93431.4299

64k selected for Saga 3B, with a mean of 1.4299 bits per character. Bits per character allow comparison across different vocabularies. This is a tokenizer result, not an assessment of the finished model.

Domain evaluation · indicative results

An exam built from texts outside training

The earlier evaluation sets shared passages with the training data, so we built a new exam solely from texts written after the data cut-off date. The 33M results remain part of the historical record.

The results are indicative and have not been independently confirmed. We do not make claims about the target model's capabilities.

Historical baselineAugust 2026 · local hardware

Canto 33M

A reference for the training process.

The recorded learning curve, evaluation and generated samples belong to the small local model. They do not describe the target Saga 3B model.

View the original curve — Polish / PL ↗
2.9867minimum validation loss
11,000completed training steps
Target modelUnconfirmed · phase D

Saga 3B

The latest report does not confirm the status of the full Saga 3B training run.

The report contains no learning curve for the full Saga training run.

216.8 billiontokens to process · planned
68 hplanned time on 64 GPUs

206,800 planned steps. This is the training budget, not the number of unique corpus tokens.

04 — Data and recipe

Three sources, measured composition

The corpus combines Polish text from three sources, with deduplication across them. Corpus-composition ablations compare how data mixtures affect training.Corpus recipe ↗
FineWeb-2Polish web
HPLT v2web, second filter
FinePDFstext from PDF files
Step 1Deduplicationacross sources
Step 2 · phase CComposition ablationssix training runs
Planned exposure
216.8 billion tokens

Token exposure over training steps, not the size of the unique corpus: training can process the same material more than once. The corpus recipe will be public.

05 — Method

An open journal. Open recipe, open measurements, open mistakes.

Weights

Private

The Saga 3B weights stay in a private repository. What remains to be published is the report on the method and the recipe.

Corpus

Public recipe

Sources, filtering and corpus composition described so that they can be reproduced. Source data have their own licensing terms; the weight license does not replace them.

Family

Skald

Skald is the model family. Canto is the local 33M reference model. Epos is a future large mixture-of-experts model planned outside this allocation.

Chat with
Saga 3B.

Test version. Access with a code.

Open the chat