Mora / Datasets / speechcolab/gigaspeech
GigaspeechFree and public
GigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training. The transcribed audio data is collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc.
- Published on
Hugging Facespeechcolab/gigaspeech
- Price
- Free
- License
- apache-2.0
- Allows
- Commercial
- Size
- 10M<n<100M rows
- Made for
- automatic-speech-recognition, text-to-speech, text-to-audio
- Languages
- en
- Downloads
- 16,132
- Last updated
- 2026-02-07
- Access
- The publisher asks you to accept its terms first
What each record holds
| Field | Type |
|---|---|
| segment_id | string |
| speaker | string |
| text | string |
| audio | audio |
| begin_time | float32 |
| end_time | float32 |
| audio_id | string |
| title | string |
| url | string |
| source | class_label |
| category | class_label |
| original_full_path | string |
Mora did not check this dataset. "Allows" reads the declared license only.