/settings/evals). Conversation Evals score your agent’s real logged conversations against a fixed set of quality criteria, so you can see not just whether conversations were resolved, but how well they were handled, and which part of the system to fix when they were not.
Run an evaluation
1
Pick a date range
Choose Last 7 days, Last 30 days, or a custom range. The page shows how many logged conversations in that range are eligible for scoring (a conversation needs at least one user turn and one agent turn).
2
Run evals
Click Run evals. Each eligible conversation is scored against every criterion, and the summary cards show the Overall score, the weakest criterion, and the run status.
3
Review low scores
Open individual conversations from the results table to see the per-criterion verdicts and the specific issues detected.
The criteria
Each criterion maps to a part of the system, which tells you where to make the fix:Make it a habit
Evals are most useful as a routine, not a one-off:- Run them after any significant change: new knowledge sources, a new or edited Skill, or changed guardrails.
- Compare the overall score across runs. A drop after a change points straight at that change.
- Pair with the Topic Explorer: topics with negative sentiment and low eval scores are the same problem seen from two angles.
Conversation Evals score conversations that already happened. For pre-launch testing before real traffic exists, use Playground with a New test session and run through your expected journeys manually. Per-skill behavior can also be tested with skill simulations.
AI performance
Volume, resolutions, and CSAT trends.
Best practices
Fixes for the most common low-score causes.