Datasets

Standards

Datasets are made of parts. A listing states every part, there or not, and what was measured.

A standard says what a complete dataset of one kind holds. Few companies have every part: a solo developer has code and commits, no Jira and no Slack. A listing says which parts it holds, a lab pays for those, and the missing ones are stated instead of hidden.

In a listing

Send the standard's id and every part, there or not:

json
"standard": "software-project",
"parts": {
  "code":          { "present": true, "records": "412 files", "measured": { "tests": "31% of files, suite passes", "builds": "yes" } },
  "commits":       { "present": true, "records": "1,840 commits", "measured": { "span": "4 years" } },
  "pull-requests": { "present": false },
  "issues":        { "present": false },
  "ci":            { "present": false },
  "discussion":    { "present": false }
}
  • A required part that is missing is refused.
  • measured takes the measures of that part, by id. Measure them: run the tests before saying they pass.
  • Unknown parts and measures are ignored.

Asking for data

A lab asks in the same terms:

json
{ "name": "ask_for_data", "arguments": {
  "what": "Python web services with their test history",
  "mustHave": "At least 2 years of commits, CI runs with failures",
  "standard": "software-project", "parts": ["code", "commits", "ci"]
} }

When no standard fits, send your own schema in fields instead.

The standards

Software projectsoftware-project

Training and testing coding agents on real work: a change, why it was made, and whether it held.

PartNeededFieldsMeasures
code
The repository at each commit
requiredrepo_id, path, language, content, commit_sha
languages Share of code per language
lines Counted without vendored and generated files
tests Share of source files with a test, and whether the suite runs
builds A clean checkout installs and builds
commits
Every change with its message and diff
requiredcommit_sha, author_id, authored_at, message, diff
count Number of commits
span First to last commit
authors Distinct authors, as ids
pull-requests
Proposed changes, review comments, and whether they merged
optionalpr_id, title, description, author_id, review_comments, merged, merged_at
count Number
reviewed Share with at least one review comment
issues
Bug reports and tasks from Jira, Linear or GitHub, linked to the change that closed them
optionalissue_id, title, description, type, status, created_at, resolved_at, fixed_by_commit
count Number
linked Share with the commit or pull request that closed them
ci
Each CI run with its result and logs
optionalrun_id, commit_sha, status, failed_tests, log_excerpt, duration_s
count Number
failures Share that failed: failures are what a model learns from
discussion
Slack or Teams threads about the work
optionalthread_id, author_id, posted_at, text, refers_to
count Number
linked Share that refer to an issue, a pull request or a commit

Training recipes. Write the change (needs code, commits). Fix the failing build (needs code, commits, ci). Resolve the issue (needs code, commits, issues). Review the change (needs pull-requests).

Support conversationssupport-conversations

Support agents and their evaluation: a customer problem and how it was really solved.

PartNeededFieldsMeasures
tickets
Each conversation from first message to last
requiredticket_id, customer_id, opened_at, channel, messages, status
count Number
turns Median
languages Share per language
resolution
What fixed it, and whether the customer came back
optionalticket_id, resolution, resolved_at, reopened, escalated_to
resolved Share with a stated resolution
reopened Share reopened within 30 days
satisfaction
The score the customer gave
optionalticket_id, score, comment
rated Share of tickets with a score
knowledge
The articles agents answered from
optionalarticle_id, title, body, updated_at
count Number

Training recipes. Answer the customer (needs tickets). Solve it (needs tickets, resolution). Grade an answer (needs tickets, satisfaction).

Field service recordsfield-service

Diagnosis and repair agents: a fault, what the technician found, and whether the fix held.

PartNeededFieldsMeasures
work-orders
Each job: complaint, findings, what was done
requiredorder_id, site_id, equipment_model, complaint, diagnosis, work_done, visited_at
count Number
span First to last visit
diagnosed Share with a stated finding
outcome
Whether the customer called back for the same fault
optionalorder_id, callback_within_30_days, repeat_of
known Share of jobs with a known result
readings
Equipment readings before the visit
optionalsite_id, read_at, sensor, value, unit
covered Share of jobs with readings in the 48 hours before
parts
What was replaced
optionalorder_id, part, quantity
covered Share

Training recipes. Diagnose (needs work-orders). Diagnose from readings (needs work-orders, readings). Will it hold (needs work-orders, outcome).

Speech audiospeech-audio

Speech recognition in hard conditions: noise, jargon, people talking over each other.

PartNeededFieldsMeasures
audio
The recordings
requiredclip_id, audio, duration_s, sample_rate, recorded_at, source_id
hours Total duration
quality Median signal-to-noise ratio
transcripts
What was said, by whom it was checked
optionalclip_id, text, transcribed_by, reviewed
covered Share of hours
reviewed Share of transcripts
speakers
Who speaks when
optionalclip_id, speaker_id, start_s, end_s
covered Share of hours

Training recipes. Transcribe (needs audio, transcripts). Who speaks when (needs audio, speakers).