Help the user understand what happened in a live streaming agent session.
If list_sessions or summarize_session fail with a connection error, the server
is not running. Tell the user:
uv run agentevals serve --dev
The --dev flag is required — plain serve does not enable the WebSocket endpoint
that sessions stream to. Also note: sessions are in-memory only with a 2-hour TTL.
If the server was restarted or the session is older than 2 hours, it is gone.
List sessions with
list_sessions. Show session ID, completion status, span count, and start time. Default to the most recent completed session unless the user specifies one.Summarize with
summarize_session. This converts raw OTLP spans into structured invocations with user messages, tool calls, and agent responses.Present as a readable narrative. For each turn:
Turn 1: User: [what the user asked] Tools: tool_name(arg=val, ...) → [one-line description of what this achieves] Response: [response text, truncated if long]Flag anything worth investigating:
- Turns with no tool calls when tools would be expected
- Empty or very short response after a long tool chain
- Same tool called repeatedly in one turn (possible loop)
- Abrupt stop mid-conversation
Offer next steps. If the user wants to score the session against a golden reference (pass/fail), suggest
/evalwith the session ID or a comparison run.
Source: agentevals-dev/agentevals — distributed by TomeVault.