Free US shipping on every order30-day refund · Ships from Phoenix

Open data

The Cultivation Corpus

Years of commercial mushroom growing, transcribed and released as a dataset. Free, no signup, cite it however you like.

Videos
68
Words
131,872
Hours
15.4
Coverage
2020–26

Most cultivation data is written after the fact, by people describing a process. This is the opposite. It is what was said during the work, on a farm that had to make payroll. It is messier than a textbook and more specific than one.

Files

  • corpus.jsonl

    One record per video with the full restored transcript and metadata.

  • chunks.jsonl

    Passages of roughly 1200 characters, each deep-linked to the second it was said. Built for retrieval.

  • corpus.csv

    The same records flattened, for spreadsheets and quick inspection.

  • README.md

    Schema, provenance, statistics, and the caveats that matter.

Honest caveats

  • Most records began as automatic speech recognition. Punctuation and capitalization were restored programmatically, and every restored passage was checked to confirm no word was added, removed, or substituted. Wording is Michael’s.
  • Species names and Latin binomials are where machine transcription fails most often. Treat those spellings as approximate.
  • Timestamps are accurate to the caption cue, not to the individual word.

Prefer to read rather than download? The same material is in the Cultivation Library.