Mora / Datasets / Orders / Lab 6F3856

Order from Lab 6F3856

Tool use traces

Logs of agents that ran in production for real users: the request, the tool catalogue, every tool call with its raw arguments and full result, and the human correction when the agent got it wrong. Failures kept, nothing synthetic, nothing public, with rights to train commercial models.

How answering works

Any number of people fill one order. The lab receives one dataset, in its own columns.

  1. Your agent reads the order

    It looks at what you hold and says which columns you can fill, and what is missing.

  2. Mora checks what you send

    Every record: no copies, no personal data, not already public, not on Mora from someone else.

  3. The lab pays per record it accepts

    The lab named the price. You fill as much of the order as you can.

The columns the lab wants

You do not need every column. Mora says which ones each dataset fills.

ColumnTypeWhat it must hold
session_ididOne agent run, from the request to the last step
deployment_ididWhich production agent and tool catalogue the session ran on
started_attimestampWhen the session began
model_idtextThe model and version that wrote the calls
user_requesttextWhat the person asked, as typed
toolsjsonThe tool catalogue of the deployment: tool_name, description, input_schema (JSON Schema), version
call_ididPairs one call with its result
step_indexnumberOrder of the call inside the session
tool_nametextThe tool that was called
argumentsjsonThe arguments exactly as the model wrote them, even malformed
resultjsonWhat the API returned, whole, including errors
statuslabelok or error, with error_code
latency_msnumberHow long the call took
correction_kindlabelWhat the human did: edited arguments, rejected, redid by hand
corrected_argumentsjsonThe call as the human fixed it
corrected_byidGenerated id of the person who corrected
corrected_attimestampWhen the correction was made, after the call
outcomelabelResolved, undone within 7 days, or escalated to a person

How much, in the lab's words

200,000 sessions (about 1.5 million tool calls) from at least 20 deployments and 300 distinct tools, from the last 18 months, with at least 30,000 sessions holding a real human correction; plus 2,000 sessions kept exclusive for evaluation

  • sessions
  • tools
  • calls
  • corrections
  • outcome