Mora / Datasets / "speech recognition" / speechcolab/gigaspeech

GigaspeechFree and public

GigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training. The transcribed audio data is collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc.

Published on
Hugging Facespeechcolab/gigaspeech
Price
Free
License
apache-2.0
Allows
Commercial
Size
10M<n<100M rows
Made for
automatic-speech-recognition, text-to-speech, text-to-audio
Languages
en
Downloads
16,132
Last updated
2026-02-07
Access
The publisher asks you to accept its terms first

What each record holds

FieldType
segment_idstring
speakerstring
textstring
audioaudio
begin_timefloat32
end_timefloat32
audio_idstring
titlestring
urlstring
sourceclass_label
categoryclass_label
original_full_pathstring

Mora did not check this dataset. "Allows" reads the declared license only.