Mora / Datasets / "speech recognition" / MLCommons/peoples_speech
People's SpeechFree and public
The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license.… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples speech.
- Published on
Hugging FaceMLCommons/peoples_speech
- Price
- Free
- License
- cc-by-2.0
- Allows
- Commercial, with attribution
- Size
- 1M<n<10M rows
- Made for
- automatic-speech-recognition
- Languages
- en
- Downloads
- 34,311
- Last updated
- 2024-11-20
What each record holds
| Field | Type |
|---|---|
| id | string |
| audio | audio |
| duration_ms | int32 |
| text | string |
Mora did not check this dataset. "Allows" reads the declared license only.