Mora / Datasets / Orders / Lab 7B362A

Order from Lab 7B362A

Low resource languages

Native-speaker recordings with human transcripts in under-served languages (Wolof, Fula, Kinyarwanda, Quechua and others), each clip tied to a speaker record with dialect, native language and a written consent covering commercial AI training. Alongside it, human-translated parallel sentence pairs for the same languages, with who translated and by what method.

How answering works

Any number of people fill one order. The lab receives one dataset, in its own columns.

  1. Your agent reads the order

    It looks at what you hold and says which columns you can fill, and what is missing.

  2. Mora checks what you send

    Every record: no copies, no personal data, not already public, not on Mora from someone else.

  3. The lab pays per record it accepts

    The lab named the price. You fill as much of the order as you can.

The columns the lab wants

You do not need every column. Mora says which ones each dataset fills.

ColumnTypeWhat it must hold
clip_ididOne recording
audioaudioFLAC or WAV, 16 kHz or better, mono, never re-encoded from MP3
duration_snumberLength of the clip in seconds
languagelabelISO 639-3 code of the language spoken
dialectlabelRegion and Glottocode, declared by the speaker
speaker_ididOne real person, the same id on every clip of theirs
native_languagelabelThe speaker's first language, to tell native from second-language speakers
genderlabelSpeaker gender, self-declared
age_bandlabelSpeaker age range
consent_ididLink to the consent record: form version, date signed, form language, scope (AI training, commercial use, resale), paid or not, how to withdraw
speech_typelabelRead or spontaneous
prompt_ididFor read speech, the sentence that was read
transcripttextVerbatim text in the language's own script, code-switching marked
transcriber_ididWho transcribed
reviewer_ididThe second native speaker who checked the transcript
devicelabelRecording device and setting
pair_ididOne parallel sentence pair
src_texttextSource sentence
src_langlabelISO 639-3 code of the source
tgt_texttextTranslated sentence
tgt_langlabelISO 639-3 code of the target
original_sidelabelWhich side was written first
translator_ididWho translated
methodlabelHuman, post-edited, or machine
domainlabelNews, health, religious, conversation and so on

How much, in the lab's words

2,000 transcribed hours of speech (200 hours per language across 10 languages, at least 300 native speakers per language) and 1 million human-translated sentence pairs (100,000 per language pair)

  • recordings
  • speakers
  • consent
  • transcripts
  • parallel-text
  • prompts