Loading...
The multi-agent communication graph is emitted by a trained autoregressive generator, and that generator is then fine-tuned RLHF-style against a learned reward model that jointly scores task correctness and structural compactness, so the topology designer itself is optimised to produce sparse graphs that keep accuracy while cutting token cost.
Source: https://arxiv.org/abs/2608.20099