Mora / Guides / RL environments

RL environments: what they are and what data they are built from

An RL environment is a simulated workplace where an AI agent practises a real job and gets scored. Here is what one contains, with a worked example, and where the data comes from.

Updated October 8, 2026

An RL environment (reinforcement learning environment) is a simulated workplace. An AI agent is dropped into a situation, given tools, allowed to act, and scored on whether it solved the problem. It is the difference between reading about a job and doing it.

A dataset says: here is a question, here is the answer. An environment says: here is a broken machine, here are fourteen things you can check, find out what is wrong.

What an environment contains

Part What it is Example, HVAC repair
Situation The starting state the agent sees "Third floor is 79°F, thermostat set to 72°F, rooftop unit running"
Tools Actions the agent can take Read sensors, open the wiring diagram, measure voltage, test a part, order a part
Hidden state The truth, which the agent has to discover The condenser fan capacitor has failed
Scoring A check on the outcome and how it was reached Correct part replaced, compressor not replaced, under 10 steps

The agent might ask for sensor readings, see that the condenser fan is off while the compressor runs, measure 240V at the motor, test the capacitor, read 1.8 µF against a rated 7.5, and replace it. The environment answers each step and scores the result.

Why labs buy them

Models already know the textbook. They can list every cause of high head pressure. What they lack is practice at investigating: choosing what to check first, recovering from a wrong guess, and stopping before the expensive mistake. That only comes from trying, failing and being scored, thousands of times.

Three things make an environment worth paying for:

  • A real outcome. The scoring comes from what happened, so nobody has to guess whether an answer was right.
  • Long tasks. Problems that take many steps, or come back over weeks, where the first three fixes treated symptoms.
  • Honest scoring. If an agent can score well without doing the job, it will learn to do that instead. Good environments check the final state of the world, as well as the answer the agent typed.

Where the data comes from

From ordinary businesses. A service company with ten years of work orders has, without meaning to, recorded hundreds of thousands of small experiments: a complaint, a diagnosis, a fix, and whether the customer called back. The same is true of aviation maintenance records, insurance claims files, support escalations and machine downtime logs.

A technician's note that reads "cust says AC not cooling. tech found bad cap. replaced 7.5 MFD. tested OK" already holds the situation, the hidden state and the outcome. Turning it into an environment means reconstructing what the technician could see and do at each step, then hiding the answer.

What makes records usable

  • The outcome is recorded: fixed, paid, reopened, called back.
  • Records about the same thing can be linked across time.
  • The evidence survives: readings, photos, notes, the parts used.
  • The company has the right to license them, and personal details can be removed.

Volume matters less than people expect. Fifty well-documented cases are enough to show a buyer what the full set would be like.

If your company has records like this

You don't need to build the environment. You need to know what you have. List it on Mora for free, and we handle verification, anonymization and the conversion into tasks. Read how to sell data to AI companies for the full process.