Canto 33M
The recorded learning curve, evaluation and generated samples belong to the small local model. They do not describe the target Saga 3B model.
View the original curve — Polish / PL ↗Saga 3B — a Polish language model trained from scratch on the JUPITER supercomputer, with its own tokenizer and its own corpus. This page is its record: every run and every measurement, including the failed ones.
Aggregate throughput of 9.0 M tok./s on 64 GPUs, relative to ideal scaling from the first point of the series (4 GPUs). This is a measurement of our workload.
Scaling chart and all points — Polish / PL ↗parameters · Saga 3B
GH200 GPUs in one training run
tokens of exposure · Saga 3B
on 64 GPUs · 206,800 steps
Validation loss on JUPITER compared with 2.9867 from the local 33M baseline.
A technical probe of the Saga 3B configuration, measuring throughput and memory, separate from the full training run.
Development Access EHPC-DEV-2026D09-196: 3,500 node-hours for 12 months. Awarded on 10 September 2026; first job on 18 September. Counted separately from the 5,000 GPU-hour Playground allocation and intended for multi-node training development.
AI Factory Fast Lane EHPC-AIF-2026FL01-1001 on JUPITER Booster for 3 months. Intended to grow Saga 3B into a model of about 15 billion parameters in two steps; each step is preceded by a pilot with a criterion recorded before it starts. Counted separately from the Playground and Development allocations.
Reproduction and scaling: complete. The report selects the 64k tokenizer for Saga. Corpus-composition comparison: complete. The full Saga 3B training run is complete according to the report. Work on a 15.45-billion-parameter model built from Saga 3B is in progress in the Development project. Its main training in the Fast Lane allocation awaits dates from JSC.
Reproduction and scaling: complete.
The report selects the 64k tokenizer for Saga.
Corpus-composition comparison: complete.
The full Saga 3B training run is complete according to the report.
In progress in the Development project: continued pre-training of the Saga 3B base is complete, and work continues on a 15.45-billion-parameter model built from that base.
Fast Lane allocation awarded for the main training of the 15.45-billion-parameter model built from Saga 3B. JSC will set the dates.
The Saga 3B weights stay in a private repository. What remains to be published is the report on the method and the recipe.
| Tokenizer | Encycl. | Web | Prose | Mean |
|---|---|---|---|---|
| 32k | 1.2804 | 1.1117 | 1.9923 | 1.4615 |
| 48k | 1.2771 | 1.0993 | 1.9185 | 1.4316 |
| 64k · selected | 1.2628 | 1.0925 | 1.9343 | 1.4299 |
64k selected for Saga 3B, with a mean of 1.4299 bits per character. Bits per character allow comparison across different vocabularies. This is a tokenizer result, not an assessment of the finished model.
The earlier evaluation sets shared passages with the training data, so we built a new exam solely from texts written after the data cut-off date. The 33M results remain part of the historical record.
The results are indicative and have not been independently confirmed. We do not make claims about the target model's capabilities.
The recorded learning curve, evaluation and generated samples belong to the small local model. They do not describe the target Saga 3B model.
View the original curve — Polish / PL ↗The report contains no learning curve for the full Saga training run.
206,800 planned steps. This is the training budget, not the number of unique corpus tokens.
| Set | Perplexity | Bits per character | Overlap |
|---|
| Model | Encyclopedia | Web | Prose |
|---|
Token exposure over training steps, not the size of the unique corpus: training can process the same material more than once. The corpus recipe will be public.
The Saga 3B weights stay in a private repository. What remains to be published is the report on the method and the recipe.
Sources, filtering and corpus composition described so that they can be reproduced. Source data have their own licensing terms; the weight license does not replace them.
Skald is the model family. Canto is the local 33M reference model. Epos is a future large mixture-of-experts model planned outside this allocation.
Test version. Access with a code.