Multi-Agent Systems Burn 6.36x the Energy: A Sustainability Review of Agentic LLM Dev Tools
Across five SE tasks (code gen, tech debt, vulnerability detection, log parsing, log analysis), four-agent setups used 6.36x the energy and 6.07x the latency of non-agentic baselines, but improved accuracy on only one task.
- Versus non-agentic (NA), single-agent (SA) used 1.20x energy / 1.23x latency, dual-agent (DA) 3.06x / 3.32x, and four-agent (MA) 6.36x / 6.07x (averaged over 5 tasks, 6 models, 3 hardware platforms).
- Accuracy gains were task-specific — MA's 55.4% F1 beat NA (49.7%), SA (49.2%), and DA (35.6%) on vulnerability detection, but log-parsing exact-match fell from 34.5% (NA) to 26.1% (MA) as more agents were added.
- Of 66 three-objective (accuracy/energy/latency) Pareto-optimal configurations across 15 task-hardware pairs, 59 (89%) were non-agentic or single-agent, and only 1 was four-agent.

























































