To evaluate agents, Arena randomizes both the orchestrator and the harness during real tasks, capturing interaction effects that model-only rankings miss. Listen
Angelopoulos says agents are far more heterogeneous than models: the harness can be a multi-component system with sub-models and sub-agents, and nobody can currently say which model-harness combination is best. Arena's agent mode gives the agent a computer for open-world, Cowork-style tasks and can connect a code base. The goal is for users not to have to choose a harness or orchestrator at all, with routing, which Arena has built for a year, handling the choice automatically.
“we can randomize both the orchestrator and the harness for the agent so we can collect all of those interaction effects that make it challenging.”
Listen to the episode Episode Product Link to this Report a problem
From The AI Race Has a Leaderboard | Arena CEO Anastasios Angelopoulos.