Mora / Datasets / Orders / Lab 7B362A
Order from Lab 7B362A
Low resource languages
Native-speaker recordings with human transcripts in under-served languages (Wolof, Fula, Kinyarwanda, Quechua and others), each clip tied to a speaker record with dialect, native language and a written consent covering commercial AI training. Alongside it, human-translated parallel sentence pairs for the same languages, with who translated and by what method.
How answering works
Any number of people fill one order. The lab receives one dataset, in its own columns.
- Your agent reads the order
It looks at what you hold and says which columns you can fill, and what is missing.
- Mora checks what you send
Every record: no copies, no personal data, not already public, not on Mora from someone else.
- The lab pays per record it accepts
The lab named the price. You fill as much of the order as you can.
The columns the lab wants
You do not need every column. Mora says which ones each dataset fills.
| Column | Type | What it must hold |
|---|---|---|
clip_id | id | One recording |
audio | audio | FLAC or WAV, 16 kHz or better, mono, never re-encoded from MP3 |
duration_s | number | Length of the clip in seconds |
language | label | ISO 639-3 code of the language spoken |
dialect | label | Region and Glottocode, declared by the speaker |
speaker_id | id | One real person, the same id on every clip of theirs |
native_language | label | The speaker's first language, to tell native from second-language speakers |
gender | label | Speaker gender, self-declared |
age_band | label | Speaker age range |
consent_id | id | Link to the consent record: form version, date signed, form language, scope (AI training, commercial use, resale), paid or not, how to withdraw |
speech_type | label | Read or spontaneous |
prompt_id | id | For read speech, the sentence that was read |
transcript | text | Verbatim text in the language's own script, code-switching marked |
transcriber_id | id | Who transcribed |
reviewer_id | id | The second native speaker who checked the transcript |
device | label | Recording device and setting |
pair_id | id | One parallel sentence pair |
src_text | text | Source sentence |
src_lang | label | ISO 639-3 code of the source |
tgt_text | text | Translated sentence |
tgt_lang | label | ISO 639-3 code of the target |
original_side | label | Which side was written first |
translator_id | id | Who translated |
method | label | Human, post-edited, or machine |
domain | label | News, health, religious, conversation and so on |
How much, in the lab's words
2,000 transcribed hours of speech (200 hours per language across 10 languages, at least 300 native speakers per language) and 1 million human-translated sentence pairs (100,000 per language pair)
- recordings
- speakers
- consent
- transcripts
- parallel-text
- prompts