Datasets
Training on a dataset
How a lab's agent reads a licensed dataset and turns it into training or evaluation examples.
Once the provider has accepted your offer and Mora has verified both sides, your agent reads the dataset with get_dataset and builds examples from it. The dataset arrives in the shape of its standard, so the same code works on data from any company.
Read it
{ "name": "get_dataset", "arguments": { "listing": "7d048849-fd79-4ec3-be33-dc7e4e499626" } }{
"manifest": {
"title": "Marketplace web app, with its commit history",
"standard": "software-project",
"parts": [
{ "id": "code", "fields": ["repo_id", "path", "language", "content", "commit_sha"], "records": 169 },
{ "id": "commits", "fields": ["commit_sha", "author_id", "authored_at", "message", "diff"], "records": 29 }
],
"missing": ["pull-requests", "issues", "ci", "discussion"],
"train": [
{ "name": "Write the change", "input": "The commit message and the files as they were before", "output": "The diff", "needs": ["code", "commits"] }
]
},
"records": [
{ "part": "commits", "commit_sha": "bedc9be", "author_id": "p_1", "message": "Agents: MCP server, JSON API, llms.txt and a skill file", "diff": "..." }
],
"next": null
}- Records come 200 at a time. Pass
nextback asafteruntil it isnull. partfilters to one part.trainlists only the recipes this dataset has the parts for. A repository with no issues cannot teach "resolve the issue", and the manifest does not pretend it can.
Lay it out
Save each part as its own JSON Lines file: one record per line.
dataset/
manifest.json
code.jsonl
commits.jsonlThis is the layout Hugging Face and most training tools read directly.
Build examples
A recipe says what goes in and what the model must produce. For "Write the change":
import json
commits = [json.loads(l) for l in open("dataset/commits.jsonl")]
with open("train.jsonl", "w") as out:
for c in commits:
out.write(json.dumps({"messages": [
{"role": "user", "content": f"Make this change to the repository:\n{c['message']}"},
{"role": "assistant", "content": c["diff"]},
]}) + "\n")train.jsonl is the chat format fine-tuning services accept. Hold some examples back to measure the model before and after.
Evaluate before you train
The sample is free to read with get_sample, before any offer. Run your current model on it first: if it already writes these diffs, you do not need this data.
The recipes
Each standard comes with its recipes. list_standards returns them under train.