Skip to main content
Open Evals under Agent settings (/settings/evals). Conversation Evals score your agent’s real logged conversations against a fixed set of quality criteria, so you can see not just whether conversations were resolved, but how well they were handled, and which part of the system to fix when they were not.

Run an evaluation

1

Pick a date range

Choose Last 7 days, Last 30 days, or a custom range. The page shows how many logged conversations in that range are eligible for scoring (a conversation needs at least one user turn and one agent turn).
2

Run evals

Click Run evals. Each eligible conversation is scored against every criterion, and the summary cards show the Overall score, the weakest criterion, and the run status.
3

Review low scores

Open individual conversations from the results table to see the per-criterion verdicts and the specific issues detected.

The criteria

Each criterion maps to a part of the system, which tells you where to make the fix:

Make it a habit

Evals are most useful as a routine, not a one-off:
  • Run them after any significant change: new knowledge sources, a new or edited Skill, or changed guardrails.
  • Compare the overall score across runs. A drop after a change points straight at that change.
  • Pair with the Topic Explorer: topics with negative sentiment and low eval scores are the same problem seen from two angles.
Conversation Evals score conversations that already happened. For pre-launch testing before real traffic exists, use Playground with a New test session and run through your expected journeys manually. Per-skill behavior can also be tested with skill simulations.

AI performance

Volume, resolutions, and CSAT trends.

Best practices

Fixes for the most common low-score causes.