FORGED GOODS  /  Dataset Prep Free sample packs

Dataset Prep — free samples

Three free sample datasets.

Chunked JSONL built from public documentation by a deterministic script. Each pack carries its QA report, its manifest, and a notice of where the text came from and under what licence. Download one, check it against its source, and judge the format for yourself.

3 packs132 source files1,195 records0 model calls

A record — 1 of 532 · Python PEPs pack · abridged

{
 "id": "6cde6f4427a20512",
 "text": "PEP: 1\nTitle: PEP Purpose and
   Guidelines\nAuthor: Barry Warsaw, …",
 "source": {
  "path": "peps/pep-0001.rst",
  "sha256": "e3a9cba0eb60…"
 },
 "provenance": {
  "repo": "python/peps",
  "commit": "ff16962a22fd…",
  "heading_path": "(root)",
  "line_start": 1,
  "line_end": 30
 },
 "stats": { "est_tokens": 294 },
 "license": {
  "spdx": "CC0-1.0 OR public-domain
   (per-file Copyright section)"
 }
}
First line of the pack's chunks.jsonl, text cut for display. Omitted keys: url, html_url, ingested_at, merged_headings, chars, token_method, licence URL.

01 — The packs

Three public corpora, one script.

Each corpus was pinned to a single commit and chunked by the same pipeline (182d7ecd0cda63e5), then re-checked by a separate verification script. The numbers below come from each pack's qa_report.json.

Pack 1

Ollama documentation

Source
ollama/ollama · docs/ · commit 16b4376aeadbec58a18b9817d49c37b1b64e33d0
Licence
MIT — Copyright (c) Ollama. Full notice inside the pack.
Source lines found
6,496 of 6,496 present in a record
Sources resolved
61 of 61 pinned URLs
In size band
96.2% of records in 120–600 estimated tokens
Near-duplicates
1 pair, kept and flagged
Download Ollama pack (.zip)

forge-dsprep-sample1-ollama.zip · 117,610 bytes
sha256 5f58d37b0e65a5ce0ee5a8c0610fa787b0c6d970884e5aaa40d4186da52bb4ff

Pack 2

llama.cpp documentation

Source
ggml-org/llama.cpp · docs/ · commit c9064dded732d81f90b34e6b33d4fbd77cbfa058
Licence
MIT — Copyright (c) 2023-2026 The ggml authors. Full notice inside the pack.
Source lines found
7,707 of 7,707 present in a record
Sources resolved
51 of 51 pinned URLs
In size band
98.7% of records in 120–600 estimated tokens
Near-duplicates
10 pairs, kept and flagged (templated model pages)
Download llama.cpp pack (.zip)

forge-dsprep-sample2-llamacpp.zip · 185,208 bytes
sha256 29a72feb4413455e2f257e71e1069d9b5306835ab971b76d930fda1cb2782d09

Pack 3

Python Enhancement Proposals (20 selected)

Source
python/peps · 20 selected PEPs · commit ff16962a22fdc5e2095e0cbc5c243ea76e34fb52
Licence
Public domain / CC0-1.0, from each file's own Copyright section (14 public domain; 6 public domain or CC0-1.0-Universal). Not MIT.
Source lines found
14,835 of 14,835 present in a record
Sources resolved
20 of 20 pinned URLs
In size band
99.1% of records in 120–600 estimated tokens
Near-duplicates
0 pairs
Download PEPs pack (.zip)

forge-dsprep-sample3-peps.zip · 295,703 bytes
sha256 08ed238fc53ee142c51d2491550cf9f3e3e83b9b21bc727276b4fd8cf8f8d6c6

02 — Inside each pack

Seven files, one folder.

  • chunks.jsonlThe records: one JSON object per line, with text, source path and URL, file hash, heading path, line range and licence.
  • schema.jsonThe JSON Schema every record is validated against (forge.rag_chunk.v1.1).
  • qa_report.jsonThe QA report: counts, coverage, source-line and source-URL checks, duplicates, size statistics, and the run's time, cost and model calls.
  • sample.mdThe first 20 records in a readable form.
  • README.mdChunking policy and limits.
  • manifest.jsonA sha256 for each of the five files above.
  • SOURCE-AND-LICENSE.mdSource repository, pinned commit and licence basis. The two MIT packs include the complete MIT notice.

The text inside every record is upstream documentation, reused under the licence stated for that pack. The pipeline adds structure and metadata only.

The three runs made 0 model calls, as recorded in each qa_report.json.

03 — Check it yourself

Verify the download, then the data.

Put the zips you downloaded and CHECKSUMS.sha256 in one folder, then run:

shasum -a 256 -c CHECKSUMS.sha256

The checksum file also lists the other packs and RELEASE-README.md. Lines for files you didn't download will show as unreadable; the ones you have should read OK. Then open a pack's sample.md, pick a record, and compare it with the source file at the pinned commit using the path and line range stored on the record.

04 — Limits

What these packs do and don't cover.

  • L1
    Token counts are estimates.

    Characters ÷ 4 on whitespace-normalised text. Indentation-heavy code is undercounted.

  • L2
    Some records fall outside the size band.

    Over 600 estimated tokens: 5 in Ollama, 1 in llama.cpp, 0 in PEPs, each labelled. Flagged as short: 5, 4 and 5.

  • L3
    Near-duplicates are flagged, not removed.

    In testing, exact copies, whitespace-only variants and 3%-edited copies were all caught. Copies with 10% edits were caught only 31–43% of the time.

  • L4
    Public documentation only, in three formats.

    .md, .mdx and .rst. Some MDX syntax (frontmatter, top-level import / export lines, top-level comments and single-line tags) is left out of records; code inside fences is unaltered.

  • L5
    Checks are automated.

    The QA report comes from a verification script. This page makes no claim that a person has reviewed these packs, and no claim about how they perform in any downstream use.