- axis
- multi-agent
- verification
- det-gold
- whatItTests
- A dynamic sealed-bid, multi-attribute auction benchmark for agentic commerce where LLM agents set prices against hidden customer preferences, real-time-adapting competitors, and unannounced demand shocks; grounded in closed-form customer utilities so profit and acquisition outcomes can be scored exactly, run across 11 frontier LLMs from four providers.
- saturation
- open
- patternLabFit
- Its finding that the acquisition leader (Gemini 3.1 Pro), the profit leader (Opus 4.6), and the fastest-to-recover-after-a-shock agent are three different models is a concrete case for why Pattern Lab should track multiple distinct outcome metrics rather than a single passRate when agents compete rather than cooperate; the closed-form utility function gives an exact, non-judge scoring oracle for any competitive-market Pattern Lab scenario.
- notes
- Submitted 2026-07-30 (within the discovery window). Distinct from every tracked multi-agent entry (Sotopia, MAgIC, COMMA, LLM-Coordination, Alem, GPTNT, MultiAgentBench, BattleAgentBench): Bazaar is a competitive-market/pricing benchmark where independent LLM agents each optimize their own payoff under hidden preferences and adversarial competitor adaptation, scored with an exact closed-form utility function rather than a milestone KPI, social rubric, or cooperative-puzzle success rate. Not to be confused with the earlier 'Cattle Trade' bluffing/bidding/bargaining benchmark (arXiv:2605.14537, May 2026, outside this window and not evaluated for inclusion here since it predates the lens's discovery cutoff).