# Dataset Prep — free sample packs

Three downloadable corpora made from public technical documentation. Each pack is chunked JSONL with source metadata, a QA report and a licence notice so you can inspect the format and trace a record back to its source.

## Packs

| Pack | Records | Source | Licence basis |
| --- | ---: | --- | --- |
| Ollama documentation | 264 | 61 files from `ollama/ollama` `docs/`, commit `16b4376aeadbec58a18b9817d49c37b1b64e33d0` | MIT — Copyright (c) Ollama |
| llama.cpp documentation | 399 | 51 files from `ggml-org/llama.cpp` `docs/`, commit `c9064dded732d81f90b34e6b33d4fbd77cbfa058` | MIT — Copyright (c) 2023–2026 The ggml authors |
| Python Enhancement Proposals | 532 | 20 selected PEPs from `python/peps`, commit `ff16962a22fdc5e2095e0cbc5c243ea76e34fb52` | Public domain / CC0-1.0 under each file's Copyright section |

## Inside every ZIP

- `chunks.jsonl` — one JSON record per line, with text, source path and URL, file hash, heading path, line range and licence.
- `schema.json` — the JSON Schema used to validate the records.
- `qa_report.json` — coverage, source-line and source-URL checks, duplicates and size statistics.
- `sample.md` — the first 20 records in a readable form.
- `README.md` — chunking policy and limits.
- `manifest.json` — SHA-256 values for the distribution files.
- `SOURCE-AND-LICENSE.md` — source repository, pinned commit and licence basis; the MIT packs include the complete notice.

## Verify a download

Put the ZIP files you downloaded together with `CHECKSUMS.sha256`, then run:

```
shasum -a 256 -c CHECKSUMS.sha256
```

Files you have should report `OK`. You can then open `sample.md`, choose a record, and compare it to the source file at the pinned commit using the path and line range stored on that record.

## Limits

- Token counts are estimates: characters divided by four after whitespace normalisation.
- Near-duplicates are flagged, not removed.
- The samples cover `.md`, `.mdx` and `.rst` public documentation.
- The QA checks are automated. The packs make no claim about downstream model performance.
