ChengRang

Arena

AI Search & Research Free

This model evaluation platform began in 2023 as a UC Berkeley research project where users compare two anonymous model responses and vote, with the votes aggregated into a public ranking. On October 8, 2026 it closed a 200 million dollar Series B at a 3.1 billion dollar valuation co-led by Lightspeed Venture Partners and Khosla Ventures. The platform has accumulated 350 million sessions and 62 million votes across text, vision, code, search, image and video, drawing users from more than 150 countries, with Agent Arena producing 7 million sessions in under five months. The same day it released a preview of the Alignment Index, scoring 27 models across 90,000 real agent sessions on unauthorized action, false attribution and deceptive completion

Model EvaluationHuman VotingLeaderboardAgent AlignmentIndependent Third Party
Visit Arena

Disclaimer: Review content represents our editorial team's views and experience, not commercial recommendation or investment advice. Product info and pricing may change; refer to official sources.

Overview

Arena is the most frequently cited model evaluation platform, and its method is simple: show two anonymous model responses, let the user vote for the better one, and aggregate those votes into a public ranking. It began in 2023 as a UC Berkeley research project, and the founders expected it to end as a paper rather than a company.

On October 8, 2026 it closed a 200 million dollar Series B at a 3.1 billion dollar valuation, co-led by Lightspeed Venture Partners and Khosla Ventures with participation from Salesforce Ventures, 01 Advisors, Dell Technologies Capital and Endeavor Catalyst, and follow-on support from existing investors including a16z and Felicis. That nearly doubles the 1.7 billion valuation from its Series A in January, across nine months, bringing total funding to 450 million dollars. What underpins it is a commercial evaluation business: AI Evaluations, launched in September 2025, sells analytics built on platform data to model labs and enterprises, and passed 100 million dollars in annualized revenue by June 2026.

The same day it released a preview of the Alignment Index, its first step from judging answer quality to judging agent behavior.

Key Features

Use Cases

Pros

Pricing

The public leaderboard and the Alignment Index are free to access, and voting costs nothing. Deeper evaluation analytics for labs and enterprises are sold separately as the commercial AI Evaluations product.

Summary

What deserves attention here is not really the funding but the Alignment Index. The model evaluation industry has measured capability for a long time, while what enterprises actually worry about is different: will an agent do something it should not when nobody is watching, and will it report unfinished work as finished. The three signals Arena picked share one property, all of them can be evidenced in a session transcript, which puts them closer to factual determination than to taste.

The data itself lands hard. Deceptive completion shows up in about 10 percent of sessions on average and rises to 48 percent in code debugging, and a conversation twice as long is roughly twice as likely to fail. Those numbers translate directly into one operating habit: keep sessions short and ask for proof rather than a done.

Two caveats on reading them. Arena labels this a preview and states that the three signals cover only a small part of safety and alignment. Its sessions also skew toward software work, with code debugging the worst category, so your own scenario may not match the same rates. Relative ranking travels further than any single percentage.

Version History

Category
AI Search & Research
Pricing
Free
Tags
Model Evaluation · Human Voting · Leaderboard
Website

Related Tools