Mora / Datasets / Orders / Lab 68384B

Order from Lab 68384B

Expert preferences

Human pairwise preference judgments on chat-model answers, made by verified domain experts (physicians, lawyers, accountants, engineers), each with the expert's own written reason. Every judgment must carry which models wrote the two answers, a stable rater id tied to a checked credential, the time spent, and the order shown.

How answering works

Any number of people fill one order. The lab receives one dataset, in its own columns.

  1. Your agent reads the order

    It looks at what you hold and says which columns you can fill, and what is missing.

  2. Mora checks what you send

    Every record: no copies, no personal data, not already public, not on Mora from someone else.

  3. The lab pays per record it accepts

    The lab named the price. You fill as much of the order as you can.

The columns the lab wants

You do not need every column. Mora says which ones each dataset fills.

ColumnTypeWhat it must hold
item_ididOne id per prompt
prompttextThe user question put to the chat assistant, with earlier turns in context_turns
domainlabelProfessional field of the question: medicine, law, accounting, software, pharmacy, engineering
prompt_sourcelabelWhere the prompt came from: written by an expert, real user traffic, or a named set
response_atextFirst model answer, full text
response_btextSecond model answer, full text
model_a / model_blabelModel and version that wrote each answer
generated_atdateWhen the answers were generated
shown_orderlabelWhich answer was shown on the left, to prove the order was random
choicelabela, b, tie or both bad
strengthnumberHow strong the preference is, 1 to 3
rationaletextThe expert's reason in their own words, about these two answers
criterialabelsWhat decided it: factual error, unsafe advice, missing step, tone
rater_ididOne stable generated id per human judge
rater_professionlabelProfession and specialty, e.g. physician, cardiology
rater_credentiallabelLicence type, issuing body and jurisdiction
credential_verified_bylabelWho checked the licence against the register, and when
judged_atdateWhen the judgment was made
time_spent_snumberSeconds the judge spent on the pair
guideline_versionlabelWhich instructions the judge worked under

How much, in the lab's words

100,000 judgments over 60,000 distinct answer pairs: 20,000 pairs judged by 3 experts each (to measure agreement) and 40,000 judged once. At least 300 verified experts, nobody above 2% of the judgments, answers from at least 4 models released in the last 12 months. A 5,000-pair test slice sold to no

  • items
  • responses
  • judgments
  • rationales
  • raters
  • gold
  • protocol