Datasets

Training on a dataset

How a lab's agent reads a licensed dataset and turns it into training or evaluation examples.

Once the provider has accepted your offer and Mora has verified both sides, your agent reads the dataset with get_dataset and builds examples from it. The dataset arrives in the shape of its standard, so the same code works on data from any company.

Read it

json
{ "name": "get_dataset", "arguments": { "listing": "7d048849-fd79-4ec3-be33-dc7e4e499626" } }
json
{
  "manifest": {
    "title": "Marketplace web app, with its commit history",
    "standard": "software-project",
    "parts": [
      { "id": "code", "fields": ["repo_id", "path", "language", "content", "commit_sha"], "records": 169 },
      { "id": "commits", "fields": ["commit_sha", "author_id", "authored_at", "message", "diff"], "records": 29 }
    ],
    "missing": ["pull-requests", "issues", "ci", "discussion"],
    "train": [
      { "name": "Write the change", "input": "The commit message and the files as they were before", "output": "The diff", "needs": ["code", "commits"] }
    ]
  },
  "records": [
    { "part": "commits", "commit_sha": "bedc9be", "author_id": "p_1", "message": "Agents: MCP server, JSON API, llms.txt and a skill file", "diff": "..." }
  ],
  "next": null
}
  • Records come 200 at a time. Pass next back as after until it is null.
  • part filters to one part.
  • train lists only the recipes this dataset has the parts for. A repository with no issues cannot teach "resolve the issue", and the manifest does not pretend it can.

Lay it out

Save each part as its own JSON Lines file: one record per line.

text
dataset/
  manifest.json
  code.jsonl
  commits.jsonl

This is the layout Hugging Face and most training tools read directly.

Build examples

A recipe says what goes in and what the model must produce. For "Write the change":

python
import json

commits = [json.loads(l) for l in open("dataset/commits.jsonl")]
with open("train.jsonl", "w") as out:
    for c in commits:
        out.write(json.dumps({"messages": [
            {"role": "user", "content": f"Make this change to the repository:\n{c['message']}"},
            {"role": "assistant", "content": c["diff"]},
        ]}) + "\n")

train.jsonl is the chat format fine-tuning services accept. Hold some examples back to measure the model before and after.

Evaluate before you train

The sample is free to read with get_sample, before any offer. Run your current model on it first: if it already writes these diffs, you do not need this data.

The recipes

Each standard comes with its recipes. list_standards returns them under train.