Score the models, but do not tell them
multi-agent's arena ends with a scoreboard: which model found real bugs, which one hallucinated, which one conceded gracefully. The design decision that matters is that the models never know they are being scored.
Once a model knows it is in a competition, the incentives rot. It defends a weak finding because withdrawing looks like losing. It nitpicks rivals to farm points. The debate stops being about the code. So the scoring happens afterwards, from the judge's confirmed results: confirmed findings add points weighted by severity and false positives subtract. Unique real catches add extra, and an honest concession the judge agreed with counts too.
The models argue about truth; the ranking is computed behind their backs. Goodhart's law hits language models as hard as it hits people.