> For the complete documentation index, see [llms.txt](https://doc.duaer.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://doc.duaer.com/advanced-ai/evaluations/overview.md).

# Understand why to evaluate in Duaer

In Duaer, use evaluations against known test cases so AI digital organizations stay reliable across inputs.
## What evaluations are

Evaluation is how you check that an AI digital organization in Duaer is reliable. It is often what turns a flaky demo into something you can ship. You need it while building and after you go live.

The foundation is running a test dataset through the digital organization. The dataset has many cases. Each case has a sample input, and often the expected output too.

- Try a range of inputs so you see edge-case behavior
- Change prompts or models with more confidence, without breaking something else
- Compare performance across models or prompts

## Why evaluation is needed

Models are not like ordinary code you can reason about. They are black boxes. You measure them by running real inputs and reading the outputs.

You only gain confidence after enough cases that match the edge cases production will see.

## Two types of evaluation

Light evaluation (before go-live): a handful of hand-written examples is often enough to reach a releasable or proof-of-concept state. Compare outputs side by side without formal metrics yet. Steps are in [light evaluations](/advanced-ai/evaluations/light-evaluations.md).

Metric-based evaluation (after go-live): grow the set from production runs. When you find a bug, add that input. After a fix, rerun the whole set as a regression. When there are too many rows to read one by one, score quality with metrics and track them on the evaluations view. See [use metrics to measure quality](/advanced-ai/evaluations/metric-based-evaluations.md).

## How the two types compare

- Gain per iteration: large in light evaluation; smaller once you use metrics
- Dataset size: small for light; large for metrics
- Dataset sources: mostly hand-written (or lightly generated) for light; often production executions for metrics, plus optional generated rows
- Actual outputs: required for both
- Expected outputs: optional for light; usually required for metrics
- Evaluation metric: optional for light; required for metrics

When wiring fails or scores stay empty, start with [common issues](/advanced-ai/evaluations/tips-and-common-issues.md). The Evaluations tab “More info” link also points here.
## Questions

### What are evaluations in Duaer?

In Duaer, evaluations run an AI digital organization over a test set so you can see whether outputs stay reliable. Use them while building and after you go live.

### When should I use light vs metric-based evaluation in Duaer?

Use light evaluation while the set is small and you are still changing prompts—compare outputs by eye. Switch to metric-based evaluation when the set grows and you need regression scores on the evaluations view.

