Skip to content
All publications

Research note

Published

From 100M to 600M German Tokens

What the public quantum-1.6-pilot record supports about continued pretraining—and why more reported training did not establish reliable generation.

Published

23 July 2026

Review status

Not peer-reviewed

DOI

Not available

This project research note is not an academic publication. It summarizes only the linked public record and does not fill gaps in training logs, raw evaluations or release provenance.

01

Research question

The public experiment asks whether continued pretraining on 500 million additional German tokens can improve language-pattern consistency while preserving the 49,295,872-parameter architecture, 512-token context and frozen quantum-1 tokenizer.

This note evaluates only what the linked public artifacts support. It does not reconstruct the private training run or infer missing measurements.

02

Documented method

The public configuration describes weights-only initialization from the earlier Quantum base stage, a fresh optimizer, scheduler and step counter, and continued training on German FineWeb2-HQ material.

The configured target is approximately 30,518 optimizer steps at an effective 16,384 tokens per step. No final public run manifest verifies the actual completed-step count, resolved environment or every configured detail.

  • Architecture and parameter count held constant
  • 512-token context held constant
  • Custom 16,384-token tokenizer held constant
  • 500M additional German tokens reported by the release card

03

Publicly reported observation

The quantum-1.6-pilot Hugging Face card reports validation loss 3.348852 and perplexity 28.4700 and publishes an F16 GGUF artifact with a manifest and SHA-256 checksum.

These next-token metrics do not measure factual accuracy, instruction following, downstream-task performance or production readiness.

04

Negative result

The public diagnosis and model cards continue to describe factual unreliability, repetition, incomplete text and incoherent continuations. The available evidence therefore does not show that the additional reported training established reliable generation.

This is a result about this small experimental setup. It is not a general conclusion that continued pretraining is ineffective or that compact models cannot be useful.

05

Evidence boundary

No versioned raw generation record, standardized downstream benchmark, complete training log, final data manifest or independently replayable metric report is public.

The strongest supported conclusion is that the documented workflow produced a public continued-pretraining artifact with reported held-out metrics while important generation limitations remained.