TRIAGE: Severity-Ranked Multi-Agent Failure Attribution
Xizhi Wang
Abstract
Failure attribution for multi-agent LLM systems is typically formulated as identifying a single decisive error in a failed execution trace. Yet failed traces often contain multiple interacting errors and admit several plausible intervention points. To examine whether an annotated decisive error is also the most promising point for repair, we introduce TriageBench, comprising LLM-council judgments over 1,184 Who&When intervention candidates. The council prioritizes a different step from the original label in 59.3% of traces. We therefore reframe failure attribution as severity-ranked diagnostic triage and instantiate it as TRIAGE, a zero-shot, cost-aware cascade combining LLM-as-judge suspicion scoring, structural impact, and LLM-simulated replay over generated repair patches. Under top-K evaluation, TRIAGE reaches 73.9% top-5 on Who&When (vs. 42.4% top-1) and 93.2% top-10 on multi-error TRAIL (vs. 56.8% top-1). In a separate evaluation of 326 traces, strict ratings with gold-answer access identify at least one plausible top-3 repair on 56.7% (human rater) and 81.6% (three-LLM panel majority) of traces. Our code and TriageBench are publicly available.