There are good surveys on LLMs for testing, and on multi-agent systems in general — but nobody has really put the two together and asked how reliable these agents are. There just isn’t much evidence on when, or why, their output goes wrong.
How reliable are LLM-based agents for software testing, and what architectural designs and mitigation strategies influence their dependability?
Kitchenham, B. & Charters, S. (2007). Guidelines for performing SLRs in software engineering (EBSE-2007-01). · Page, M. J. et al. (2021). PRISMA 2020. BMJ, 372, n71.
24 studies classified · 2 REST-API studies compare single- & multi-agent variants and count in both A1 & A2.
Run the same agent on the same task twice and you can get two different answers. That makes the results hard to reproduce, and hard to compare.
Counts include 2 REST-API studies (Agentic LLMs for REST API, 2025; Test Amplification for REST APIs, 2026) in both A1 & A2. · 2 secondary surveys (He et al.; Xia et al., 2025) excluded.