An AI eval (short for evaluation) is a test for an AI model: a fixed set of cases, each with an input and a way to tell whether the model's output was right, run the same way every time so results can be compared. LLM evaluation is the same idea applied to large language models (LLMs, the models behind chatbots and agents). It is how a team finds out whether a new model, a new prompt or a new tool made things better or worse.
Software has unit tests because code either returns the right value or it doesn't. A model gives a different answer each time and can be wrong in ways that read well. An eval is the unit test for that kind of system.
What an eval contains
| Part | What it is | Example, insurance claims |
|---|---|---|
| Test cases | The inputs the model is given | A claim file: the first report, the photos, the adjuster's notes |
| Known answer | What was correct, often called ground truth | What was finally paid, and whether the claim was disputed |
| Grader | The check that compares output to answer | Is the model's estimate within the margin the team chose |
| Metric | One score across all cases | Share of claims estimated within that margin |
The cases are kept out of training. A model that has already seen the answers is being tested on its memory, which is why labs call this a held-out set.
How answers get graded
| Grader | How it works | Good at | Weak at |
|---|---|---|---|
| Code check | Compare to the known answer, or run a test | Cheap, identical every run | Only works when there is one right answer |
| Human review | Someone who knows the job reads the output | Judgment, tone, safety | Slow, costly, and reviewers disagree |
| Model as judge | A second model scores the output against written criteria (often called LLM-as-judge) | Volume, open-ended answers | The judge makes its own mistakes and has to be checked against people |
Most teams mix the three: code checks where an answer is exact, a model judge for the rest, and people reviewing a sample of what the judge decided. All three depend on the same thing, which is knowing what a right answer looks like for each case.
Evaluating agents
An agent is a model that takes actions with tools over many steps: it looks things up, edits files, files a ticket. Grading the final message is not enough, so an agent eval checks two things: the state the agent left behind (was the refund issued, do the tests pass) and the path it took to get there (how many steps, which tools, any action it should not have taken).
SWE-bench is a well-known example. Its tasks are real issues from open-source code repositories, and the grader is the project's own tests. The cases are real and the answer was settled by what happened, which is why the benchmark is trusted.
An agent eval and an RL environment are nearly the same object: a situation, tools, a hidden answer and a score. The difference is what the score is used for. In an eval it measures the model. In an environment it trains it.
Evaluation frameworks
A framework is the harness around an eval. It loads the cases, calls the model, applies the graders and stores the results so two runs can be compared. Open-source ones include OpenAI Evals, Inspect, promptfoo and DeepEval. Hosted products add dashboards and logging on top.
Pick the one that fits the code your team already writes. None of them comes with test cases for your field, and that part decides whether the eval tells you anything.
The hard part is the test cases
Teams usually get cases from one of three places, and each has a limit.
| Source | The limit |
|---|---|
| Written by the team | Covers the situations the team thought of. Real work is stranger |
| Public benchmarks | The questions are on the internet, so they may be in the model's training data (labs call this contamination). Every lab has the same ones |
| Generated by a model | Tests what a model can already imagine, and the answer is still unverified |
What an eval needs is harder to find: real cases, never published, where the right answer became known afterwards. That is a description of ordinary business records.
- A support ticket records the problem, what the agent looked up, the fix, and whether the ticket reopened.
- A claims file records what was reported, what was estimated, and what was paid.
- A maintenance log records the fault, the repair, and whether the same fault came back.
In each one the outcome was decided by what happened next, not by somebody's opinion. Hide the outcome, give the model the rest, and the record is a test case with a grader built in.
What makes a record a good test case
- The outcome is written down: resolved, paid, reopened, returned.
- The evidence the person had at the time survives, so the model sees what they saw.
- It is hard. Routine cases that every model passes measure nothing.
- It has never been public, so no model has trained on it.
- The company has the right to license it, and personal details can be removed.
Where Mora fits
Mora does not run evals, and it does not label data. It is the market where companies list records like these and AI labs license them, after Mora has checked the rights, that no personal data is left, and that the records are not already public.
If your team builds evals, tell us what your model gets wrong and we source records against it. If your company holds records with outcomes, list them for free.