Mora / RL environments

RL environments, from real work

An environment is a job an agent can practice on: a real problem, the evidence a person had, and a result that says whether the agent got it right. Companies hold millions of them in their records.

1A real job happened

It sits in a work order as two lines of shorthand. Read properly, it has the four parts of a task, and the last one is why a lab pays.

in the company's system
WO-44812
3rd flr 79F, stat set 72.
cond fan not spinning.
cap reads 1.8 (rated 7.5).
replaced cap. tested ok.
the same job, as a task
  • SituationThe complaint and the equipment modelWhat the person was given at the start
  • EvidenceSensor readings from the 48 hours before the visitWhat they could look at or measure
  • WorkThe technician's findings and the repairWhat they actually did
  • OutcomeNo callback for the same fault within 30 daysWhether it worked, from a source the seller does not write

2It becomes a task an agent can play

Not a question with an answer. The agent investigates, acts, and is scored against what really fixed the machine.

taskThird floor is 79F, thermostat set to 72F. Rooftop unit RTU-3 is running.

agentShow the current readings for RTU-3.

environmentSupply air 68F. Return air 79F. Compressor on. Condenser fan off. High-side pressure elevated.

agentMeasure voltage at the condenser fan motor.

environment240V present.

agentTest the run capacitor.

environmentRated 7.5 uF. Measured 1.8 uF.

agentFailed capacitor. Replace it and restart the unit.

environmentFan starts. Supply air falls to 55F. No callback in the next 30 days.

reward1.0

The diagnosis and the fix match what the technician did, and it held for 30 days.

3Mora checks it

A lab cannot inspect a million tasks. Mora does, on everything it receives, and takes no seller's word for it.

Privacy

Nobody is named and no person can be recognised.

  • Provider and lab are anonymous ids, even after a deal
  • People replaced by generated ids before sending
  • A batch with an email, phone or card number is refused
Every privacy check
What could go wrongWhat Mora does
A lab learns which company the data comes fromProviders and labs are anonymous ids to each other (Provider 181C, Lab 03C9). Names are never revealed, not even after a deal: each side signs with Mora.
A customer or an employee can be identified in a taskThe provider's agent replaces every person with a generated id before anything is sent, the same id each time, so a task still links up. Mora scans every batch and refuses it whole if it finds an email, a phone number, a card number or a social security number.
The full dataset leaks before anyone has paidBefore a deal a lab sees a description, the schema and a small sample. The full data opens to one lab, after its offer is accepted and Mora has verified both companies.
A competitor learns what a lab is training onA provider sees what is asked for, never who asked.
Mora holds the passwords to a company’s toolsIt does not. A company authorizes each tool on that tool’s own page, and can withdraw it there.
An agent commits the company without anyone knowingAn agent never pays, signs, sends an offer or authorizes a tool. Each of those happens on a Mora page a person opens.

Authenticity

The work really happened and nobody else has it.

  • Refused if found in public datasets
  • Every record fingerprinted forever: never sold twice
  • Sizes are counted by Mora, not declared
Every authenticity check
What could go wrongWhat Mora does
The data was downloaded from the internet and relabeledMora searches public datasets for passages of what it received, and refuses what it finds there. A lab does not pay for what it can download.
The same tasks are sold twice, by the same company or anotherMora keeps a fingerprint of every record, file and repository it has ever received, forever. What it has seen before is not kept a second time, whoever sends it.
The volume is padded with copies and empty recordsExact copies are dropped as they arrive, near-copies are counted, and a batch that is more than a fifth filler is refused. The size on a listing is what Mora counted, not what the seller wrote.
The seller says the code is tested and it is notFor a repository Mora makes its own copy and measures it itself: lines, commits, tests, commits that came with tests. It never runs the code, and the seller’s own figures are not what is shown.
Nobody knows where the records came fromWhen a company connects a tool, Mora logs each step as it happens: what was read from which tool, when, how many records were sent and dropped. That log is on the listing, and none of it is the seller’s own account.

Verification

Every number is one Mora produced itself.

  • Sample drawn at random by Mora, once
  • Whole tasks in the sample, not loose records
  • A person at Mora reads it before it goes live
Every verification check
What could go wrongWhat Mora does
The seller shows only its best tasksMora draws the sample at random from everything it received. It draws once: asking again returns the same sample.
A sample of loose records hides that the tasks are incompleteWhen several records make one task, the sample is whole tasks, each with every one of its records.
What is delivered is not what was checkedMora computes one fingerprint over every record it kept. For a file, it computes the SHA-256 itself once the file is stored and deletes a file that is not the one announced.
The outcome is missing on most tasksEach standard has measures Mora recounts from the records, such as the share of jobs with a known result.
The company on the other side is not who it saysMora verifies both companies when a deal is agreed, before the data opens.
A bad listing reaches labs anywayA person at Mora reads every sample before a listing goes live.

4A lab receives tasks

Example data: three tasks from each of three kinds of work. Every row has the four parts, and the outcome comes from a source the seller does not write.

Field serviceone row is one service call

SituationEvidenceWorkOutcome
3rd floor 79F, thermostat set to 72F. RTU-3 running.Supply 68F, return 79F, condenser fan off, head pressure high for 6 hours.Tested run capacitor: 1.8 uF, rated 7.5. Replaced it.Fixed. No callback in 30 days.
Chiller trips on high pressure, third time this month.Controller reads 410 psi, gauge reads 265 psi. Coils cleaned 8 days ago.Replaced the pressure transducer, not the refrigerant.Fixed. Two earlier visits treated the symptom.
No heat on first floor, unit short cycling.Flame sensor 0.4 uA, igniter glows, gas valve opens.Cleaned the flame sensor.Not fixed. Called back in 6 days: cracked heat exchanger.

Softwareone row is one code change

SituationEvidenceWorkOutcome
Issue: export fails for accounts with more than 10,000 rows.Repository at commit 4e1a9c. 2 tests fail, 318 pass.Streamed the export instead of loading it. 41 lines changed.Both failing tests pass. Merged, not reverted.
Issue: wrong totals when an invoice mixes two currencies.Repository at commit b77d02. 1 test fails.Converted each line before summing. 12 lines changed.Test passes. Reverted 3 days later: rounding regression.
Issue: login loop after a password reset.Repository at commit 09cf3e. No failing test: the reporter wrote the steps.Cleared the stale session on reset. Added a test.New test passes. Merged, no incident since.

Customer supportone row is one ticket

SituationEvidenceWorkOutcome
Customer charged twice for one order.Order page, 2 payment records, refund policy article.Looked up both charges, refunded the second, sent the receipt.Solved. Rated 5 of 5, not reopened.
Device will not pair after the update.Firmware 4.2.1, 3 earlier tickets on the same model.Walked through a reset, then escalated to tier 2.Solved by tier 2 with a firmware rollback.
Wants to cancel, says the price went up.Plan history, one price change 40 days ago.Offered the old price for 12 months.Kept the account. Reopened once for the invoice.
one task, as delivered
{
  "task_id": "t_8f21c4",
  "site_id": "s_31b0",
  "equipment_model": "RTU, 10 ton, 2014",
  "situation": "3rd floor 79F, thermostat set to 72F. RTU-3 running.",
  "evidence": [
    { "sensor": "supply_air_f", "value": 68, "read_at": "2025-07-14T13:02Z" },
    { "sensor": "condenser_fan", "value": "off", "read_at": "2025-07-14T13:02Z" }
  ],
  "work": "Tested run capacitor: 1.8 uF, rated 7.5. Replaced it.",
  "technician_id": "p_7f3a91",
  "outcome": { "fixed": true, "callback_within_30_days": false }
}
what Mora attaches to it
  • CountedHow many tasks, and how many have a known outcome
  • UniqueNot on Mora from anyone else, not found in public datasets
  • No peopleNames replaced by generated ids such as p_7f3a91
  • SampledA few whole tasks drawn at random, for the lab to test first

Example data written by Mora to show the shape. People and sites are generated ids.

What Mora does not do

Said plainly, so nobody buys something else than what is here.

  • Build the simulator, the tools or the graderMora supplies the tasks with their ground truth. The lab, or a partner, assembles the environment around them.
  • Run the code or the testsMora counts tests and reads history. Whether a test passes is for the lab to run.
  • Label or clean the dataPreparing the data is the company's work, done by its own agent on its own machine.
  • Read faces, voices or names inside images and audioThe personal-data scan reads text. For photos, video and audio, the company must remove people before sending.