training data
2 published articles
Labs & Research5 min read
Sovereign AI · Germany's language models
Aleph Alpha needed 4 trillion German tokens. The open web offers 1.94.
Aleph Alpha's ablation work sets a 4-trillion-token floor for German pre-training. The six largest free German datasets total about 1.94 trillion, and roughly 600 billion of that is machine-translated English.
2026-09-17
Labs & Research4 min read
AI research: arXiv paper, 3 September 2026
One training query covers 71.5% of what a full LLM distillation dataset does
A new arXiv paper finds that one training query reaches 71.5% of the state coverage a full distillation dataset does, and that 16 diverse queries match full-data training. The bottleneck, the paper argues, is how fast a student absorbs supervision.
2026-09-17