Datasets

Sending a dataset

The four calls a provider's agent makes: create the listing, send records in batches, finish. Mora draws the sample.

A provider's agent builds the dataset. Nobody fills in a form, and nobody picks the sample.

The flow

  1. Find the records. read_source on a connected tool, or an export the person gives the agent.
  2. Create the listing. create_listing, with a standard and its parts, or your own fields.
  3. Send the records. submit_records, up to 200 at a time, with people taken out.
  4. Finish. finish_dataset. Mora draws ten records at random: that is the sample labs test.

Create the listing

json
{ "name": "create_listing", "arguments": {
  "title": "Tier 2 support escalations",
  "recordType": "customer-support-conversations",
  "system": "Zendesk", "years": "2021 to 2026", "volume": "About 88,000 threads",
  "contents": "The customer problem, every reply, and what fixed it.",
  "owns": "yes", "personal": "no", "contracts": "yes",
  "standard": "support-conversations",
  "parts": {
    "tickets":      { "present": true, "records": "88,000", "measured": { "turns": "6 per ticket" } },
    "resolution":   { "present": true, "measured": { "resolved": "91%" } },
    "satisfaction": { "present": false },
    "knowledge":    { "present": false }
  }
} }
json
{ "id": "7d048849-fd79-4ec3-be33-dc7e4e499626", "state": "Saved. Mora reviews it before anyone can see it." }

Send records

json
{ "name": "submit_records", "arguments": {
  "listing": "7d048849-fd79-4ec3-be33-dc7e4e499626",
  "part": "tickets",
  "records": [
    { "ticket_id": "T-88412", "customer_id": "p_7f3a91", "opened_at": "2026-03-02",
      "messages": ["no heat on first floor", "igniter replaced, heating again"], "status": "solved" }
  ]
} }
json
{ "received": 1, "total": 1, "next": "Send the next batch, or call finish_dataset when everything is in." }

For a listing that follows a standard, records go part by part and use that part's field names and no others. That is what makes datasets from different companies line up.

A batch that still holds personal data, or a field the part does not have, is refused whole. The answer says what to fix.

Finish

json
{ "name": "finish_dataset", "arguments": { "listing": "7d048849-fd79-4ec3-be33-dc7e4e499626" } }
json
{ "total": 88000, "sampled": 10, "schema": "kept",
  "state": "Mora drew 10 records at random out of 88000. A person at Mora reads the sample before the listing goes live." }

Calling finish_dataset again draws a new sample. If the listing had no schema, Mora reads one off the sample: one field per key, with its type.

What a lab then sees

The listing with its parts, its schema, and the sample. Never the provider's name. A lab makes an offer for the parts it wants, and once it is accepted reads the dataset: see Training on a dataset.