Mora / Datasets / "speech recognition" / MLCommons/unsupervised_peoples_speech

Unsupervised Peoples SpeechFree and public

The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers.

Published on
Hugging FaceMLCommons/unsupervised_peoples_speech
Price
Free
License
not stated
Allows
Unknown: read the license before you use it
Made for
automatic-speech-recognition, audio-classification
Languages
eng
Downloads
15,406
Last updated
2025-02-27

Mora did not check this dataset. "Allows" reads the declared license only.