Datasets
Standards
Datasets are made of parts. A listing states every part, there or not, and what was measured.
A standard says what a complete dataset of one kind holds. Few companies have every part: a solo developer has code and commits, no Jira and no Slack. A listing says which parts it holds, a lab pays for those, and the missing ones are stated instead of hidden.
In a listing
Send the standard's id and every part, there or not:
"standard": "software-project",
"parts": {
"code": { "present": true, "records": "412 files", "measured": { "tests": "31% of files, suite passes", "builds": "yes" } },
"commits": { "present": true, "records": "1,840 commits", "measured": { "span": "4 years" } },
"pull-requests": { "present": false },
"issues": { "present": false },
"ci": { "present": false },
"discussion": { "present": false }
}- A required part that is missing is refused.
measuredtakes the measures of that part, by id. Measure them: run the tests before saying they pass.- Unknown parts and measures are ignored.
Asking for data
A lab asks in the same terms:
{ "name": "ask_for_data", "arguments": {
"what": "Python web services with their test history",
"mustHave": "At least 2 years of commits, CI runs with failures",
"standard": "software-project", "parts": ["code", "commits", "ci"]
} }When no standard fits, send your own schema in fields instead.
The standards
Software projectsoftware-project
Training and testing coding agents on real work: a change, why it was made, and whether it held.
| Part | Needed | Fields | Measures |
|---|---|---|---|
codeThe repository at each commit | required | repo_id, path, language, content, commit_sha | languages Share of code per languagelines Counted without vendored and generated filestests Share of source files with a test, and whether the suite runsbuilds A clean checkout installs and builds |
commitsEvery change with its message and diff | required | commit_sha, author_id, authored_at, message, diff | count Number of commitsspan First to last commitauthors Distinct authors, as ids |
pull-requestsProposed changes, review comments, and whether they merged | optional | pr_id, title, description, author_id, review_comments, merged, merged_at | count Numberreviewed Share with at least one review comment |
issuesBug reports and tasks from Jira, Linear or GitHub, linked to the change that closed them | optional | issue_id, title, description, type, status, created_at, resolved_at, fixed_by_commit | count Numberlinked Share with the commit or pull request that closed them |
ciEach CI run with its result and logs | optional | run_id, commit_sha, status, failed_tests, log_excerpt, duration_s | count Numberfailures Share that failed: failures are what a model learns from |
discussionSlack or Teams threads about the work | optional | thread_id, author_id, posted_at, text, refers_to | count Numberlinked Share that refer to an issue, a pull request or a commit |
Training recipes. Write the change (needs code, commits). Fix the failing build (needs code, commits, ci). Resolve the issue (needs code, commits, issues). Review the change (needs pull-requests).
Support conversationssupport-conversations
Support agents and their evaluation: a customer problem and how it was really solved.
| Part | Needed | Fields | Measures |
|---|---|---|---|
ticketsEach conversation from first message to last | required | ticket_id, customer_id, opened_at, channel, messages, status | count Numberturns Medianlanguages Share per language |
resolutionWhat fixed it, and whether the customer came back | optional | ticket_id, resolution, resolved_at, reopened, escalated_to | resolved Share with a stated resolutionreopened Share reopened within 30 days |
satisfactionThe score the customer gave | optional | ticket_id, score, comment | rated Share of tickets with a score |
knowledgeThe articles agents answered from | optional | article_id, title, body, updated_at | count Number |
Training recipes. Answer the customer (needs tickets). Solve it (needs tickets, resolution). Grade an answer (needs tickets, satisfaction).
Field service recordsfield-service
Diagnosis and repair agents: a fault, what the technician found, and whether the fix held.
| Part | Needed | Fields | Measures |
|---|---|---|---|
work-ordersEach job: complaint, findings, what was done | required | order_id, site_id, equipment_model, complaint, diagnosis, work_done, visited_at | count Numberspan First to last visitdiagnosed Share with a stated finding |
outcomeWhether the customer called back for the same fault | optional | order_id, callback_within_30_days, repeat_of | known Share of jobs with a known result |
readingsEquipment readings before the visit | optional | site_id, read_at, sensor, value, unit | covered Share of jobs with readings in the 48 hours before |
partsWhat was replaced | optional | order_id, part, quantity | covered Share |
Training recipes. Diagnose (needs work-orders). Diagnose from readings (needs work-orders, readings). Will it hold (needs work-orders, outcome).
Speech audiospeech-audio
Speech recognition in hard conditions: noise, jargon, people talking over each other.
| Part | Needed | Fields | Measures |
|---|---|---|---|
audioThe recordings | required | clip_id, audio, duration_s, sample_rate, recorded_at, source_id | hours Total durationquality Median signal-to-noise ratio |
transcriptsWhat was said, by whom it was checked | optional | clip_id, text, transcribed_by, reviewed | covered Share of hoursreviewed Share of transcripts |
speakersWho speaks when | optional | clip_id, speaker_id, start_s, end_s | covered Share of hours |
Training recipes. Transcribe (needs audio, transcripts). Who speaks when (needs audio, speakers).