Capture the failure with its context
When an agent gets something wrong — a bad tool call, a misread policy, a wrong escalation — LonsLabs records the full trace and the human judgment that flagged it. Similar failures are clustered so the team sees one problem, not forty tickets.