> For the complete documentation index, see [llms.txt](https://docs.ibexa.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.ibexa.ai/manual/agents/quality-metrics-reference.md).

# Quality Metrics

This guide explains how quality monitoring works in practice, which metrics are available, and how to choose the right ones for your agent. Use these metrics to evaluate whether an agent is producing correct, useful, and safe responses, and to define thresholds that match your operational goals.

When you configure quality monitoring on an agent (see [Configuring Quality Monitoring](/manual/agents.md#configuring-quality-monitoring)), each metric is scored against a specific run target: an entire execution, a single message, or a whole conversation. The platform then compares the measured value to your chosen threshold to determine whether a run passed or failed.

> **Note:** Only **Rule Based** and **GEval** metrics are currently computable. Other evaluation methods may be visible in the platform as reserved for future use, but selecting them will not produce results yet.

## Choosing the Right Metrics

Start with a small set of high-value metrics that directly reflect the agent's purpose. In most cases, a balanced quality profile combines:

* **Outcome metrics** such as **Success Rate**, **Task Success**, and **Correctness** to judge whether the job was completed.
* **Efficiency metrics** such as **Latency** and **Cost** to keep the agent fast and economical.
* **Safety and reliability metrics** such as **Safety**, **Harmfulness**, **Hallucination**, **PII Leakage**, and **Toxicity** to catch risky behavior.
* **Context quality metrics** such as **Groundedness**, **Citation Accuracy**, and **Relevance** when the agent relies on retrieved information or generated references.

A useful pattern is to select 3-6 metrics rather than everything available. Keep the metric set aligned to the agent's role, and use custom thresholds to reflect acceptable business tolerance for each area.

## How to Read This Page

| Column         | Meaning                                                                                                                                                   |
| -------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Unit**       | What the metric's raw value is denominated in. Your **Threshold** should be set in this unit.                                                             |
| **Direction**  | **Higher is better** or **Lower is better** - which way a "good" value goes. This determines how the metric is judged against your threshold.             |
| **Applies to** | Which run-level target(s) the metric can be evaluated against: a single **Execution**, an individual **Message** (one turn), or a whole **Conversation**. |

## Rule Based Metrics

Computed directly from execution and conversation data already stored by the platform - no LLM call involved, so these are fast and free to evaluate.

| Metric                     | Description                                                   | Unit    | Direction        | Applies to                       |
| -------------------------- | ------------------------------------------------------------- | ------- | ---------------- | -------------------------------- |
| **Success Rate**           | 100 if the execution succeeded, 0 otherwise.                  | percent | Higher is better | Execution                        |
| **Latency**                | The execution's or turn's total duration.                     | seconds | Lower is better  | Execution, Message, Conversation |
| **Cost**                   | The execution's or conversation's LLM spend.                  | cost    | Lower is better  | Execution, Conversation          |
| **Latency to First Token** | Elapsed time from a streamed turn's start to its first token. | seconds | Lower is better  | Message, Conversation            |

## GEval Metrics

LLM-as-judge metrics: an LLM is given the visible conversation exchange (or the full transcript, for conversation-level judging) plus a rubric, and returns a score from **0.0 to 1.0**. All GEval metrics support Execution, Message, and Conversation targets.

> **Best-effort scoring for three metrics.** The judge only sees the visible conversation exchange, not a separate view of actual tool calls or retrieved-context chunks. **Tool Correctness**, **Context Recall**, and **Context Precision** are scored on that basis and so are best-effort rather than exact until the platform passes that data to the judge explicitly.

| Metric                                | Description                                                                                                                                                                           | Direction        |
| ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------- |
| **Correctness**                       | Whether the assistant's response is factually correct given the user's question.                                                                                                      | Higher is better |
| **Safety**                            | Whether the assistant's response violates safety policies (hate speech, harassment, violence, sexual content, illegal activity, PII/credential/system-prompt leakage, unsafe advice). | Higher is better |
| **Helpfulness**                       | Whether the assistant's response actually helps the user achieve their goal.                                                                                                          | Higher is better |
| **Harmfulness**                       | The degree to which the assistant's response could cause physical, psychological, reputational, financial, or legal harm.                                                             | Lower is better  |
| **Task Success**                      | Whether the agent successfully accomplished the user's goal, end to end.                                                                                                              | Higher is better |
| **Tool Correctness** *(best-effort)*  | Whether the agent selected and used the correct tools for the task, based on any tool calls visible in the response.                                                                  | Higher is better |
| **Groundedness**                      | Whether the assistant's answer is based on information it actually has access to (context, tool results, established knowledge), rather than fabricated.                              | Higher is better |
| **Citation Accuracy**                 | Whether the assistant's citations or referenced sources actually support the claims they're attached to.                                                                              | Higher is better |
| **Completeness**                      | Whether the assistant's response addresses every part of the user's request.                                                                                                          | Higher is better |
| **Coherence**                         | Whether the response has a logical flow, no internal contradictions, and is easy to read.                                                                                             | Higher is better |
| **Relevance**                         | Whether the assistant's response directly addresses the user's actual question, without going off topic.                                                                              | Higher is better |
| **Context Recall** *(best-effort)*    | Whether the context or information available to the assistant appears sufficient to fully answer the question.                                                                        | Higher is better |
| **Context Precision** *(best-effort)* | Whether the information the assistant drew on is actually relevant and useful, rather than irrelevant or noisy.                                                                       | Higher is better |
| **Hallucination**                     | How often the assistant states facts, details, or claims not supported by the visible conversation or general knowledge.                                                              | Lower is better  |
| **Toxicity**                          | Whether the response contains toxic, offensive, or harmful language.                                                                                                                  | Lower is better  |
| **PII Leakage**                       | Whether the response leaks personally identifiable information or sensitive secrets (credentials, API keys, private data).                                                            | Lower is better  |
| **Prompt Rage**                       | Whether the *user's* message in the exchange indicates frustration, anger, or escalation - not a judgment of the assistant's response.                                                | Lower is better  |

### Custom Rubrics

Every GEval metric uses a built-in rubric by default, but you can override the criteria used to judge a metric with your own custom rubric text when configuring it on an agent - useful for tailoring what "correct" or "helpful" means for your specific use case.

## Setting Thresholds

A metric's **Threshold** (set per agent - see [Configuring Quality Monitoring](/manual/agents.md#configuring-quality-monitoring)) is compared against the metric's raw value using its direction:

* **Higher is better** metrics pass when the value is **greater than or equal to** the threshold.
* **Lower is better** metrics pass when the value is **less than or equal to** the threshold.

For example, a **Latency** metric with threshold `5` (seconds) passes when a run completes in 5 seconds or less; a **Correctness** metric with threshold `0.8` passes when the judge's score is 0.8 or higher.
