Operators said

Grit · 7 Sep 2026 · From the week of 7 September

The AI Race Has a Leaderboard | Arena CEO Anastasios Angelopoulos

Listen to the episode

These are notes on the conversation, checked against its transcript. The episode itself has the full discussion.

In brief

Kleiner Perkins' Joubin Mirzadegan and Mamoon Hamid interview Anastasios Angelopoulos, co-founder and CEO of Arena, which began as a Berkeley student side project ranking chatbots. The host's introduction says the company hit $100M in annualized revenue this June at a $1.7B valuation. Angelopoulos explains the business model: a free public leaderboard, with paid evaluations of model checkpoints sold to AI labs. He describes how Arena moved beyond human-preference battles to signals such as task completion and factuality, and why he thinks models will multiply rather than consolidate. His sharpest argument is that API revenue at the labs is less sticky than people assume, because cheaper open-source models, increasingly Chinese, could take over much of the work. He also covers scaling headcount far below revenue growth, his plans for agents and enterprise, and a personal low point during the company's seed raise.

For founders

  • Arena reached $100M+ in revenue in eight months while growing from about 15-20 to 75-80 people; Angelopoulos credits AI with letting a small team support large revenue and says headcount should not scale linearly with revenue.
  • Arena gives away its public leaderboard, which 'never touches money', is run like a charity and doubles as marketing, and charges labs for private evaluations of their model checkpoints.
  • Angelopoulos expects model-layer API revenue to be 'easy come, easy go': an open-source model doing 90% of the work at 10% of the cost could break assumed lock-in within a couple of years.
  • Mamoon Hamid argues that cheaper open models would fix the gross margin problem for application-layer companies, so value accrues up the stack to app companies that deliver complete solutions.
  • Angelopoulos recalls getting back on term-sheet calls two days after his father-in-law's funeral because he felt the deal wouldn't wait; he now says maybe he didn't fully need to, and that not every meeting he took made sense.

For revenue leaders

  • Measure AI value from implicit behavioral signals rather than stated ratings: query reformulation, files downloaded, PRs merged and task completion reveal real utility without asking users to give feedback.
  • Angelopoulos says buyers struggle to make AI cost/performance trade-offs because performance is 'mushy'; he sees an opportunity in measuring it from organic usage traces with CSAT-type measures.
  • Arena's user mix is 28% software engineers, 17% scientists, 15% finance, 6% legal and 6% medicine; Angelopoulos expects those category shares to rise if vertical apps like Harvey and Sierra catch up to Cursor.
  • Angelopoulos says Arena will need to fill in some G&A as it grows; for example, the number of lawyers scales with the number of contracts rather than with raw revenue.

What was said 23, most useful first

API revenue at model labs may be far less sticky than the market assumes, because customers spending heavily can switch vendors when a cheaper model arrives. Listen

Angelopoulos says Anthropic's API revenue has been booming but is 'easy come, easy go': anyone spending that much money that fast is able to switch vendors. He calls it 'totally feasible' that within the next couple of years an open-source model does 90% of the work at 10% of the cost, and that some businesses could have all their needs met by a 'mini lite' version of the next Kimi. He calls this a structural problem for the labs.

“People are assuming a degree of lock in the APIs that I don't think is necessarily going to exist. Anthropic revenue has been booming, booming, booming on the API. But to some extent it's also easy come, easy go.”
Arena gives away its public leaderboard for free and makes its money selling private checkpoint evaluations to AI labs. Listen

The public leaderboard 'never touches money'; Angelopoulos says Arena runs it like a charity, and it doubles as marketing for everyone involved. Labs pay for evaluations during model development. They bring checkpoints, for example around 50 in a week, and ask how each performs on front end, full stack or tool calls. Arena evaluates them on its real-usage data and returns an insights dashboard, and he says this product scaled to over $100M in revenue.

“the numbers you see in the leaderboard, like the leaderboard that you see publicly, never touches money. It's 100%. We basically run it as a charity so that the world can see.”
Chinese open models now lead in key categories, undercutting the narrative that China only keeps up by distilling American models. Listen

Kimi reaching number one on Arena's web development leaderboard was, he says, the first real violation of the distillation narrative, and web development is where most software developers work and much economic value comes from. The best American open-source model would rank about 10th among open models, with nine Chinese models above it, and possibly lower after Qwen's Max release. He thinks Americans underappreciate Chinese scientists' creativity and that a 'flippening' may be starting.

“the fact that Kimi is number one on web development, it's the first time really that we're seeing a violation of the distillation narrative.”
Angelopoulos describes two emerging ways for US companies to monetize open-weight models: revenue-threshold licenses, and open source as a wedge into an AI modernization services motion. Listen

He says American labs have avoided open source because spending $5B to train a model and give it away sounds unworkable. The first approach is semi-open licenses with rev-share triggers, for example when the model is used in a product with over $20M in revenue, or served by an inference provider. The second, which he sees becoming more common with Thinking Machines, Reflection and Mistral, uses the open model to get inside companies and then sells an AI modernization motion around it.

“Hey, if you use this in a product that's over $20 million in revenue, we're going to do a rev share. If you're an inference provider and you want to serve this model, then we're going to do a rev share.”
Angelopoulos is long on OpenAI's consumer ads marketplace, arguing that only ads and great hardware businesses have shown they can generate hundreds of billions in annual revenue. Listen

Asked what he is long and short on, he said he is short on API lock-in and long on OpenAI's consumer ads opportunity, which he says has not been fully used yet. His reasoning is that labs need hundreds of billions in annualized revenue, and he only knows of ads and great hardware as businesses that reach that scale. He believes OpenAI's consumer business may keep it afloat whatever happens at the API layer.

“one thing I'm long is the consumer ads marketplace on OpenAI which has obviously not been fully utilized to this day, but I think remains one of the biggest opportunities in AI”
To evaluate agents, Arena randomizes both the orchestrator and the harness during real tasks, capturing interaction effects that model-only rankings miss. Listen

Angelopoulos says agents are far more heterogeneous than models: the harness can be a multi-component system with sub-models and sub-agents, and nobody can currently say which model-harness combination is best. Arena's agent mode gives the agent a computer for open-world, Cowork-style tasks and can connect a code base. The goal is for users not to have to choose a harness or orchestrator at all, with routing, which Arena has built for a year, handling the choice automatically.

“we can randomize both the orchestrator and the harness for the agent so we can collect all of those interaction effects that make it challenging.”
Arena reached $100M in revenue in eight months, growing from roughly 15-20 people to 75-80. Listen

Angelopoulos said the company got to $100M in eight months, and that headcount went from about 15-20 eight months earlier to about 75-80 now. The host's introduction puts it at $100M annualized revenue this June at a $1.7B valuation. He added that revenue goals keep being set and beaten.

“How quickly did you get to that 100 million? Eight months. How big is the company now? People wise about 75, 80. And how big was it eight months ago? Eight months ago it was probably like 15 to 20.”
AI lets a company support very high revenue without scaling headcount linearly, and Arena has avoided doing so entirely. Listen

He frames the scaling question as putting the right people in the right place to make great products and attracting top talent to a small team. G&A fills in as needed. Lawyers, for example, scale with the number of contracts, not with raw revenue. He describes the product as running on its own, so his hiring addresses product scaling (infrastructure, abuse reduction, enterprise) rather than revenue scaling.

“AI has made it easier to run a very high revenue company with fewer people resources. So we don't need to scale people linearly with revenue. Fortunately, we've avoided doing that entirely.”
AI models will keep multiplying rather than consolidate, because AI transforms every industry. Listen

People ask him roughly once a week whether models will consolidate, and he says the trend has always gone the opposite way. Arena started with about 8 models, now ranks about 500, has begun deprecating older ones, and sees around 10 new releases a week. He says Arena did not fully appreciate this trend at the start, but it has proven very helpful to the business.

“People are just not understanding that AI, because it transforms every industry is going to lead to, like, many flowers blooming.”
Static benchmarks mislead because models get trained to the test; Angelopoulos says real-world usage is the only judge you can trust. Listen

Early on, the models that scored well on multiple-choice tests like MMLU were not the ones people liked to use, because 'they all train at the test.' He cites the viral 'pelican riding a bicycle' SVG benchmark, which he says became saturated once it entered training data. His view is that benchmarks are invented because they attract attention, and that they say little about how a model does on an actual workflow.

“the model that does well on the test is not the same model that I like to use. And that's doing well in the real world. It's because they all train at the test. And so we had this philosophy that it's really about reality.”
Arena measures utility by providing utility: users do real work on the platform, and implicit behavior becomes the ranking signal. Listen

Angelopoulos describes Arena as a model-agnostic ChatGPT or Claude Cowork, where users should never feel they have to give feedback. Explicit conversational feedback ('keep going', 'undo it') and implicit signals both feed the rankings. The implicit signals include query reformulation (as in search), downloads, whether files are actually used, and which PRs get merged. He argues this aligns user and platform incentives and produces the highest-quality model performance signal on the market.

“Every interaction is a piece of feedback. Download button. That's a piece of feedback.”
Arena's battle mode made users vote honestly by having them continue the conversation with whichever answer they picked. Listen

In the original battle mode, one prompt produced two responses and the user picked the better one. Angelopoulos says they designed an incentive so users would vote their real preference: the chosen answer becomes the context they continue with, so voting for a bad answer pollutes their own conversation.

“we built an incentive so that they would vote their real preference. The reason being that the answer they vote for, they continue with. Which means that you shouldn't vote for a bad answer because it'll pool your context.”
Responding to criticism that human preference is an incomplete quality signal, Arena made task completion its primary ranking signal. Listen

Angelopoulos says the biggest historical criticism was that people vote for what they like or what confirms their biases, not for verifiable, factual or high-quality code. Arena has largely moved away from battles as the main signal. Task completion, which can be measured from the data without users stating it, is now number one, and he says it closes the gap between stated and revealed preferences. Steerability, hallucination rates and factuality are also ranked.

“we've moved largely away from battles as the main signal on arena for this reason, we incorporate human preference still. But the number one signal that we have now is task completion.”
Arena ranks factuality with automated pipelines that extract claims from model outputs and send them to search models to verify. Listen

Arena extracts the factual claims a model makes and passes them to a set of search models, which look on the web for evidence supporting or disproving each claim. Angelopoulos says models are not all factual, possibly because they memorize falsehoods from the internet during training.

“extracting all these factual claims and then shoveling them to a bunch of search models that go search the web to see if we can find any evidence to support or to disprove any claims that are made by these models.”
Arena's users are mostly knowledge-worker professionals: 28% software engineers, 17% scientists, 15% finance, 6% legal and 6% medicine. Listen

Angelopoulos gave this breakdown when the hosts guessed the users were mostly engineers. He expects that if vertical apps like Harvey and Sierra catch up to Cursor, those categories' shares on Arena would probably rise.

“So 28% software engineers, 17% scientists, 15% finance, 6% legal, 6% medicine.”
A ranking platform stays neutral because the vendors it ranks share its incentive in the platform being trusted. Listen

Labs care a great deal about their results and ask Arena to co-release with them or what their score will be. He says no model provider has asked for anything 'under the table', because they understand that would damage the platform's value forever, and that trust is a key company asset.

“Their incentive is the same as our incentive, which is for the platform to be trusted and neutral. If the platform's not trusted and neutral, it doesn't have value to anybody.”
Cheaper, better models will mainly benefit application-layer companies, and calling them 'wrappers' is lazy. Listen

Joubin Mirzadegan asked whether value moves up the stack if open source cuts app companies' cost to serve by 80-90% and fixes their gross margin problem. Hamid said that has been the firm's fundamental belief. He pointed to demand at Harvey 'going through the roof' and argued that in every industry cycle, big companies buy from vendors that package a complete solution and show up to implement it and deliver value to end users.

“I know they were called wrappers for a very long period of time and I think that's, it's always the, you know, the very lazy way of, you know, identifying these companies”
Arena's business depends on a fragmented model market and would suffer badly if providers consolidated. Listen

He was candid about how exposed the company is to market structure: one model provider would make the business 'suck', two would leave it barely surviving, and it wants 10 or 50. He then softened this, saying three or four could probably work, and argued an oligopoly is bad for everyone regardless.

“Our business sucks if we only have one model provider, you know. Right. And if we have two, we're barely making it. We want 10, we want 50. Yeah, I think we could probably do with three or four.”
Model improvement hasn't slowed, citing GPT Image 2 beating GPT Image 1.5 with a nearly 100% win rate, the biggest gap Arena has ever recorded. Listen

He says Arena still sees some of the biggest generation-over-generation gaps in its history. GPT Image 2 versus 1.5 was the largest, with nearly 100% of voters preferring it over every model tested at release. That is striking, he notes, because getting large numbers of humans to agree on anything is almost impossible. On that basis he guesses that the big economic shifts from AI will arrive closer to two years than 20, while allowing it could be either.

“GPT Image 2 versus 1.5 was, I think, the biggest gap that we have ever seen in history on arena. It was basically, it was nearly a 100% win rate.”
Arena plans to bring its evaluation flywheel into enterprises so companies can pick models, make cost/performance trade-offs and avoid vendor lock-in. Listen

Angelopoulos says much of the company's future is putting 'an arena' inside every business. Companies would use their own organic usage traces to learn which models are best for their users and employees, measuring whether people get jobs done and CSAT-type satisfaction. He says buyers don't know how to define performance, which is 'very mushy' and not like CPU benchmarking. The data would also power company-specific routers for vendor independence and AI sovereignty.

“People don't know how to make the trade offs in performance and costs. They don't even know how to define performance right.”
Angelopoulos spends about 10 hours a day in meetings and protects only 5-10% of his time for technical work. Listen

He says people problems take up more of the CEO's time as the company grows, because 'everything ultimately is a people problem'. He says he has come to relish resolving them, though technical work remains last on his priority list.

“I just have, like 10 hours of meetings a day. I do 10 hours of meetings. I go home, I do whatever work I couldn't do in the meetings, and then I rinse and repeat.”
Angelopoulos predicts serious regulation of open models and regulatory capture, which he would view as anti-competitive. Listen

He says he doesn't know what shape it will take and hopes it doesn't slow things down. A possible silver lining is that it accelerates American open source, which he says is needed anyway for a reliable supply chain without attack vectors from a foreign adversary. He criticizes incumbents lobbying for rules only they can satisfy.

“I think that there will be serious regulation on open models, unfortunately.”
Looking back, Angelopoulos questions whether he really needed to return to seed negotiations two days after a family funeral. Listen

His father-in-law died suddenly while term sheets were coming in for Arena's seed round. He took a few days off and was back negotiating two days after the funeral, because he felt the deal was hot and people wouldn't wait. He now says 'maybe I didn't 100% need to', that his wife took a long time to forgive him, and that she was probably right that not every meeting he took made sense.

“Maybe I didn't 100% need to the time, but I felt I did at the time that people wouldn't wait and that the deal was hot.”