Every call keeps its transcript and a record of what the agent did, so an unsuccessful call has evidence behind it.
From a poor call to the Builder
Open the Calls page and inspect the call in question.
The Builder opens in troubleshoot mode with that call's transcript and diagnostics already loaded. A banner confirms: the diagnostics are loaded, the fix will be applied to a draft, and callers stay on the live version.
Knowing why it failed is not needed. "It should have offered a morning slot" is enough. The Builder reads the evidence, explains what it finds, and proposes a change.
Success is measured, not guessed
Every agent the Builder designs carries its conversation goals: the outcomes agreed when it was built, in its description. A diagnosis opens with a verdict against them: each goal scored achieved, partial or missed, with evidence from the transcript and the record of what the agent's tools did. The call counts as a success only when every required goal was achieved and the agent followed its own instructions.
An agent with no recorded goals is assessed against its instructions alone, and the Builder offers to record goals as part of the fix, so the next diagnosis has a reference point.
Fixes are proposed, not imposed
Troubleshooting is a full editing session. The Builder follows its usual propose-and-confirm rhythm, every accepted fix autosaves into a draft, and the draft can be tested before anything reaches a caller. If the call ran on an older version, the Builder says so. The problem may already have been fixed.
Resuming a session
While a session is in progress, that team's card on the Agents page shows a Resume troubleshooting button. It reopens the same conversation with the diagnostics still in context.
One call, one session: each hand-off from the Calls page starts a fresh conversation with that call in context. Reviewing a different call means handing it off from Calls again.
When the agent is not in a team
For a call answered by a standalone agent, the Builder still reads the diagnostics and reports what it finds, but it advises rather than edits, and its suggestions are applied by hand in the agent editor.
Troubleshoot the pattern, not the single case. If three callers reach the same point of failure, send the clearest example to the Builder and mention the others.