Open data
The Cultivation Corpus
Years of commercial mushroom growing, transcribed and released as a dataset. Free, no signup, cite it however you like.
- Videos
- 68
- Words
- 131,872
- Hours
- 15.4
- Coverage
- 2020–26
Most cultivation data is written after the fact, by people describing a process. This is the opposite. It is what was said during the work, on a farm that had to make payroll. It is messier than a textbook and more specific than one.
Files
- corpus.jsonl
One record per video with the full restored transcript and metadata.
- chunks.jsonl
Passages of roughly 1200 characters, each deep-linked to the second it was said. Built for retrieval.
- corpus.csv
The same records flattened, for spreadsheets and quick inspection.
- README.md
Schema, provenance, statistics, and the caveats that matter.
Honest caveats
- Most records began as automatic speech recognition. Punctuation and capitalization were restored programmatically, and every restored passage was checked to confirm no word was added, removed, or substituted. Wording is Michael’s.
- Species names and Latin binomials are where machine transcription fails most often. Treat those spellings as approximate.
- Timestamps are accurate to the caption cue, not to the individual word.
Prefer to read rather than download? The same material is in the Cultivation Library.
