Skip to content

DATA & TRAINING

Configured methods are not completed runs.

This record distinguishes dataset provenance, configured targets, smoke-test evidence and absent production or training artifacts across the Quantum model line.

PRIMARY DATASET

epfml/FineWeb2-HQ

deu_Latn

A German-language subset derived from public web material distributed through FineWeb2-HQ.

The dataset card identifies ODC-By 1.0 and Common Crawl terms. Those terms do not automatically clear rights in every source document or license model weights, code, tokenizers or visual assets.

Inspect dataset card
01Partial evidence

quantum-1-pilot

Source revision
Configured as mutable `main`
Configured target
100M training, 1M validation and 1M test tokens at 512-token context.
Public observation
The model card reports approximately 100M training tokens; no final data manifest is public.
02Partial evidence

quantum-1.6-pilot

Source revision
Configured as mutable `main`
Configured target
500M new training, 2M validation and 2M test tokens at 512-token context.
Public observation
The release card reports 500M additional tokens; final document counts, split hashes and overlap statistics are not public.
03Partial evidence

quantum-1-echelon

Source revision
c0c06e94fd3a44ae9e802b2b0fc533817601eb5e
Configured target
8B training, 10M validation and 10M test tokens at 2,048-token context.
Public observation
A smoke run saw 5,001 documents, accepted 1,559 and produced 1,380,886 packed tokens. The production run had not started.

PIPELINE CONTROLS

Documented controls.

  • Language, script, length, quality, URL, digit and symbol filtering
  • Deterministic split assignment and configured seeds
  • Exact-document controls and reliance on upstream MinHash deduplication in the Echelon pipeline
  • Packed token storage with checkpoint and resume handling

KNOWN LIMITATIONS

Evidence still missing.

  • Exact fingerprints do not remove every near-duplicate or semantically overlapping document.
  • Public web data may retain personal, harmful, copyrighted or otherwise sensitive material.
  • The pilot configurations do not resolve their mutable dataset revision in a final manifest.
  • No public record links final accepted documents and split hashes to either pilot GGUF.
  • No production Echelon dataset or completed Echelon training run is public.

Questions about provenance, privacy, rights or removal should be sent to the public project email with the affected source and enough information to identify it. Contact: cikemgil@rappidai-research.com