Mora / Datasets / Orders / Lab 68384B
Order from Lab 68384B
Expert preferences
Human pairwise preference judgments on chat-model answers, made by verified domain experts (physicians, lawyers, accountants, engineers), each with the expert's own written reason. Every judgment must carry which models wrote the two answers, a stable rater id tied to a checked credential, the time spent, and the order shown.
How answering works
Any number of people fill one order. The lab receives one dataset, in its own columns.
- Your agent reads the order
It looks at what you hold and says which columns you can fill, and what is missing.
- Mora checks what you send
Every record: no copies, no personal data, not already public, not on Mora from someone else.
- The lab pays per record it accepts
The lab named the price. You fill as much of the order as you can.
The columns the lab wants
You do not need every column. Mora says which ones each dataset fills.
| Column | Type | What it must hold |
|---|---|---|
item_id | id | One id per prompt |
prompt | text | The user question put to the chat assistant, with earlier turns in context_turns |
domain | label | Professional field of the question: medicine, law, accounting, software, pharmacy, engineering |
prompt_source | label | Where the prompt came from: written by an expert, real user traffic, or a named set |
response_a | text | First model answer, full text |
response_b | text | Second model answer, full text |
model_a / model_b | label | Model and version that wrote each answer |
generated_at | date | When the answers were generated |
shown_order | label | Which answer was shown on the left, to prove the order was random |
choice | label | a, b, tie or both bad |
strength | number | How strong the preference is, 1 to 3 |
rationale | text | The expert's reason in their own words, about these two answers |
criteria | labels | What decided it: factual error, unsafe advice, missing step, tone |
rater_id | id | One stable generated id per human judge |
rater_profession | label | Profession and specialty, e.g. physician, cardiology |
rater_credential | label | Licence type, issuing body and jurisdiction |
credential_verified_by | label | Who checked the licence against the register, and when |
judged_at | date | When the judgment was made |
time_spent_s | number | Seconds the judge spent on the pair |
guideline_version | label | Which instructions the judge worked under |
How much, in the lab's words
100,000 judgments over 60,000 distinct answer pairs: 20,000 pairs judged by 3 experts each (to measure agreement) and 40,000 judged once. At least 300 verified experts, nobody above 2% of the judgments, answers from at least 4 models released in the last 12 months. A 5,000-pair test slice sold to no
- items
- responses
- judgments
- rationales
- raters
- gold
- protocol