Arena Raises $200M Series B at $3.1B Valuation as AI Evaluation Demand Climbs
The UC Berkeley spinout nearly doubled its valuation in 10 months after scaling commercial evaluation services and adding an alignment leaderboard.

Arena, the artificial intelligence evaluation platform that originated in 2023 as a UC Berkeley research project, has raised $200 million in Series B funding at a $3.1 billion valuation, according to reporting by TechCrunch AI (https://techcrunch.com/2026/10/08/popular-ai-leaderboard-arena-nearly-doubles-valuation-to-3-1b-valuation-in-10-months/).
The financing was co-led by Lightspeed Venture Partners and Khosla Ventures, with participation from Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, Andreessen Horowitz (a16z), Felicis, and other backers. The transaction nearly doubles the company's valuation from January, when it announced a $150 million Series A round at a $1.7 billion post-money valuation.
The valuation increase follows rapid revenue growth. Arena reported reaching $100 million in annualized run-rate revenue in June, up from $30 million in annualized revenue disclosed during its Series A announcement ten months prior.
Arena operates a free crowdsourced platform where users input prompts, test requests, and vote on model outputs. The startup reports that the consumer site attracts tens of millions of monthly visitors. In September of last year, Arena introduced AI Evaluations, a commercial offering providing AI labs and enterprise customers with performance analytics derived from user interactions.
The enterprise push arrived as developers and corporate buyers confronted limitations with traditional static benchmarks, which AI models have increasingly been trained or optimized to pass. In its funding announcement, the company stated that static benchmarks break down once models recognize testing environments, arguing that independent third-party measurement is needed to assess how models perform with real-world users.
Alongside the capital raise, Arena introduced an alignment category to its leaderboard to measure specific failure modes, including unauthorized actions, false attribution of sources or statements, and deceptive completion, where a model misrepresents whether it finished an assigned task. In preliminary rankings for the alignment category, OpenAI models held the top positions, with Anthropic's Claude Opus 5.5 and Claude Fable ranking sixth and ninth, respectively.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.


