Chunked JSONL built from public documentation by a deterministic script. Each pack carries its QA report, its manifest, and a notice of where the text came from and under what licence. Download one, check it against its source, and judge the format for yourself.
First line of the pack's chunks.jsonl, text cut for display. Omitted keys: url, html_url, ingested_at, merged_headings, chars, token_method, licence URL.
01 — The packs
Three public corpora, one script.
Each corpus was pinned to a single commit and chunked by the same pipeline (182d7ecd0cda63e5), then re-checked by a separate verification script. The numbers below come from each pack's qa_report.json.
chunks.jsonlThe records: one JSON object per line, with text, source path and URL, file hash, heading path, line range and licence.
schema.jsonThe JSON Schema every record is validated against (forge.rag_chunk.v1.1).
qa_report.jsonThe QA report: counts, coverage, source-line and source-URL checks, duplicates, size statistics, and the run's time, cost and model calls.
sample.mdThe first 20 records in a readable form.
README.mdChunking policy and limits.
manifest.jsonA sha256 for each of the five files above.
SOURCE-AND-LICENSE.mdSource repository, pinned commit and licence basis. The two MIT packs include the complete MIT notice.
The text inside every record is upstream documentation, reused under the licence stated for that pack. The pipeline adds structure and metadata only.
The three runs made 0 model calls, as recorded in each qa_report.json.
03 — Check it yourself
Verify the download, then the data.
Put the zips you downloaded and CHECKSUMS.sha256 in one folder, then run:
shasum -a 256 -c CHECKSUMS.sha256
The checksum file also lists the other packs and RELEASE-README.md. Lines for files you didn't download will show as unreadable; the ones you have should read OK. Then open a pack's sample.md, pick a record, and compare it with the source file at the pinned commit using the path and line range stored on the record.
Characters ÷ 4 on whitespace-normalised text. Indentation-heavy code is undercounted.
L2
Some records fall outside the size band.
Over 600 estimated tokens: 5 in Ollama, 1 in llama.cpp, 0 in PEPs, each labelled. Flagged as short: 5, 4 and 5.
L3
Near-duplicates are flagged, not removed.
In testing, exact copies, whitespace-only variants and 3%-edited copies were all caught. Copies with 10% edits were caught only 31–43% of the time.
L4
Public documentation only, in three formats.
.md, .mdx and .rst. Some MDX syntax (frontmatter, top-level import / export lines, top-level comments and single-line tags) is left out of records; code inside fences is unaltered.
L5
Checks are automated.
The QA report comes from a verification script. This page makes no claim that a person has reviewed these packs, and no claim about how they perform in any downstream use.