Chatbot Arena Leaderboard
Free value-added services
Comprehensive List of AI Tools AI model evaluation

Chatbot Arena Leaderboard

Chatbot Arena Leaderboard, an intelligent tool focused on the evaluation of AI models

Tags:

What is Arena AI?

Arena AI is a platform that evaluates generative AI models by comparing anonymous responses from real users; it was formerly known as Chatbot Arena and LMArena. When users enter the same prompt, they see responses from two models whose identities are hidden, and they can vote to select the better outcome.

The project was initially developed by LMSYS and researchers from the University of California, Berkeley, and later evolved into an independent platform. The current official name of this platform is simply Arena; the old LMSYS blog page is only useful for understanding its history and should no longer be used as the main entry point to the platform.

The relationship between Chatbot Arena, LMArena, and Arena

NamePhases and meaningsCurrent usage recommendations
Chatbot ArenaNames of the original research projects that began in 2023Used for papers, historical introductions, and early open-source code
LMArenaBrand names for the independent site and company stagesIt is still used by a large number of media outlets and users.
ArenaCurrent unified brandThe official website, rankings, and product descriptions should be used first.
LMSYSAcademic organizations that initiated the early researchIt cannot be considered completely equivalent to the current Arena Company and all its products.

How Arena works

  1. Select text, images, videos, code, or other elements suitable for the task at hand.
  2. Enter a genuine prompt that can distinguish the capabilities of the model.
  3. The platform randomly provides two anonymous models and generates results simultaneously.
  4. Compare the two sides in terms of accuracy, usefulness, style, and task completion.
  5. Choose: A is better, B is better, it’s a draw, or both are poor.
  6. Check the true identity of the model after submitting a vote.
  7. Continue the conversation or start a new round, so that the preference data can be included in the public evaluation system.

Core functions

1. Anonymous model battles

In Battle mode, the name of the model is hidden before voting, in order to reduce the impact of brand recognition and users’ preconceptions. The answers themselves can still reveal the style of the model; therefore, the anonymization mechanism can only reduce bias, but not eliminate it entirely.

2. Humans prefer voting.

The platform does not rely on fixed standard answers; instead, it allows real users to decide which outcome is more useful. It is particularly suitable for evaluating tasks such as writing, dialogue, images, and open-ended generation, which are difficult to assess using a single score.

3. Multi-track rankings

Arena has evolved from plain-text chatting to include images, videos, code, visual content, and other specialized formats. Different rankings use distinct input methods, model pools, and evaluation criteria, so it is not possible to compare rankings across different categories directly.

4. Direct mode

In addition to anonymous battles, users can also choose to test specific known models directly. The Direct mode is suitable for verifying particular models, but only votes that comply with the rules of anonymous evaluation should affect the rankings.

5. Public research and datasets

Arena publishes research papers, analyses, and processed datasets of human preferences, which are used to study model evaluation, adversarial attacks, rejection bias, and search capabilities. The publicly available data comes with its own privacy handling and licensing rules, and it cannot be considered as the complete original dataset.

List of main functions

  • Have two anonymous AI models answer the same prompt and compare them side by side.
  • It supports voting options for win, draw, and both are incorrect.
  • Revealing the model’s identity after voting reduces brand bias prior to voting.
  • Provides continuously updated rankings for text, images, videos, and code.
  • View more detailed model performance by category, language, or task.
  • Select a specific model directly to experience it and verify the results.
  • Preview models that have not yet been officially released or appear under a code name.
  • Publicly available data on human preferences, papers, and evaluation methods.
  • The feedback from the community is compiled and used to help model developers improve the product.

How to understand the scores on the ranking list?

Arena aggregates a large number of pairwise win/loss relationships to produce a ranking based on relative ability. The methods used for this purpose typically involve statistical models for pairwise comparisons such as Bradley-Terry, with resampling employed to estimate uncertainty. Users often refer to these scores as Elo scores, but the actual calculation methods should be determined according to the guidelines provided by the ranking system.

Indicators or conceptsWhat does it represent?It doesn’t tell us anything.
Arena ScoreThe relative preference strength of the model in the current voting dataIt’s not the absolute accuracy rate or general intelligence score.
RankPosition under the specified rankings and statistical rulesIt cannot be guaranteed that it will perform better in real-world scenarios.
confidence intervalRange of uncertainty in rankings and scoresWhen intervals overlap, it is not appropriate to emphasize minor differences in rankings.
Number of votesThe size of the combat samples that can be used for estimationA large number does not mean that the sample is free from bias.
Classification rankingsThe relative performance of the model in specific tasks or languagesIt cannot replace the task testing that needs to be carried out by the organization itself.

How to design effective contrast cues

  1. Choose a specific issue from your actual tasks, rather than just asking about general knowledge.
  2. Provide the necessary background, input data, constraints, and the desired output format.
  3. Avoid including personal information, trade secrets, and confidential materials.
  4. Have the two models perform the same task without changing the evaluation criteria midway.
  5. First, examine the facts, reasoning, and adherence to instructions independently, and then compare the style of expression.
  6. When it’s difficult to decide, choose a draw or indicate that both options are unsatisfactory; there’s no need to force a winner.
  7. For critical tasks, multiple different samples are used repeatedly; the choice is not based on a single match.

Which users are it suitable for

  • Ordinary AI users who wish to experience different cutting-edge models.
  • Developers who need to initially compare model writing, reasoning, coding, and visual capabilities.
  • Academics who study human preference evaluation and model ranking methods.
  • Industry observers who track the relative performance of new models versus preview models.
  • A product and technical team that can provide external references is needed to select the appropriate model.
  • Laboratories and manufacturers that wish to improve the model using feedback from real users.
  • Researchers who evaluate ranking attacks, biases, and statistical reliability.

Different arenas and application scenarios

  • Text dialogue: Comparison of knowledge, reasoning, creation, and instruction following.
  • Code: Compare code generation, debugging, interface implementation, and tool usage.
  • Images: Comparison of text-to-image generation, editing, and adherence to prompts.
  • Video: Compares the quality and consistency of videos generated from text or images.
  • Visual understanding: Comparative image question answering, charts, OCR, and multimodal reasoning.
  • Search and research: Comparative retrieval, citations, timeliness, and comprehensive answers.
  • Specific languages: Observe the relative preference of the model for Chinese and other languages.

Free access and account requirements

Arena offers the public free access to model battles and leaderboards; there are no paid plans available for ordinary users. Different models may have waiting times, speed limits, geographical restrictions, login limits, or daily usage limits, and the platform may also make adjustments based on its capacity.

FunctionsPriceAccounts and restrictions
View the rankingsFreeIt can usually be viewed directly.
Anonymous model battlesFreeLogin may be required, and there may be restrictions on speed or capacity.
Experience the Direct modelFreeThe available models and the number of times they can be used may vary.
Research datasetAvailable on the corresponding pageIt is necessary to comply with the license and usage terms of the dataset.
Model manufacturer evaluation and collaborationThere is no unified, publicly disclosed price.Coordinated separately by Arena and its partners

How to use the ranking system to select models correctly

  • First, select the category, language, and ranking list that are most relevant to the business.
  • When looking at scores and confidence intervals, focus not only on the first place and the small differences in rankings.
  • Verify that the model version listed in the ranking matches the version that is actually available for purchase.
  • Filtering is done based on cost, latency, context, data policies, and deployment methods.
  • Create an internal evaluation set that includes real business inputs.
  • Tests are conducted separately for facts, code, security, tool calls, and structured output.
  • After going live, continuously monitor model updates and business feedback.

The advantages of the ranking list

  • Using open-ended prompts with real users brings a experience closer to everyday life than static exams.
  • Anonymous presentation reduces certain brand biases and the impact of model reputation.
  • Pairwise comparison makes it easier to make judgments than requiring users to assign absolute scores.
  • The models and voting systems are continuously updated, allowing for a quick assessment of the performance of new versions.
  • It covers text, images, videos, code, and multi-modal domains.
  • Making research, methods, and some data available facilitates external review.
  • Small laboratories are allowed to receive user evaluations under the same mechanism as large manufacturers.

Usage restrictions and methodological biases

  • Voters and prompt providers are selected by users themselves, and therefore do not represent all countries, industries, or demographics.
  • Users may prefer longer, more confident, or more stylish responses over more accurate ones.
  • Anonymous models can still be identified through wording, format, and the style of refusal.
  • The rankings are based on relative ordering, and changes in the model pool and combat distribution can affect the scores.
  • Small differences in rankings may not have statistical significance and should be interpreted in conjunction with confidence intervals.
  • The preview version of the model may differ from the publicly released version in terms of prompts, sampling, or configuration.
  • Malicious voting, coordinated attacks, and data corruption require ongoing protection.
  • General human preferences cannot replace assessments in the fields of medicine, law, security, and business tasks.

Privacy and content submission

Officials state that the suggestions submitted by users are collected in order to ensure fair and transparent evaluations, and that some of this feedback may be shared with the developers of the models. Users should consider Arena as a public research and evaluation environment, rather than a private messaging tool for handling confidential data.

  • Do not submit names, contact information, account details, passwords, or identity information.
  • Do not paste internal company codes, customer data, or unpublished documents.
  • For images and videos, only materials that you own or for which you have authorization can be used.
  • For issues related to sensitive industries, synthetic or anonymized samples should be used.
  • Before voting, check whether the answer has revealed any sensitive information from the input.
  • When using public datasets, comply with privacy filtering, licensing, and redistribution rules.

Open-source projects and datasets

The early version of Chatbot Arena was developed based on LMSYS FastChat; FastChat is licensed under the Apache-2.0 license and provides multi-model services, a web interface, as well as code related to evaluation tasks. The current version of the Arena platform offers more specialized competition categories and commercial functionalities, so the original repository cannot fully represent the current state of this platform.

  • FastChat includes a model service, a dialogue interface, and an early version of Arena.
  • Arena publishes the processed data on human preferences on platforms such as Hugging Face.
  • Relevant papers introduce anonymous paired voting and ranking methods.
  • Special research has been conducted on visual attacks, search attacks, denial-of-service attacks, and ranking attacks.
  • Open-source code, research data, and online platforms are three distinct categories.
  • The original StableLM GitHub link of the database is not related to the Arena project, and therefore should not be considered the official source code.

Why can’t we just look at the Arena rankings?

  • Business costs and delays are not included in the general preference score.
  • Data residency, privacy, compliance, and supplier terms need to be evaluated separately.
  • Structured output, tool invocation, and long-term stability may lack sufficient samples.
  • The number of hints available for specific Chinese industries may be lower than that for general English tasks.
  • After the model is updated, there may be a brief mismatch between the version of the ranking list and the actual API version.
  • High-risk tasks require more reproducible testing and expert review.

Platform and open-source status

ProjectCurrent situation
Current nameArena AI, formerly known as LMArena and Chatbot Arena
Product formatWeb-based model battles, Direct experience, and public rankings
PriceThe public functions are available free of charge; there are no paid plans for regular users.
Developer APIThere is no unified model inference API available for regular users.
Early open-source projectsLMSYS FastChat, Apache-2.0 license
Public dataThe processed dataset of human preferences along with the research paper
Is it fully open source?No, the current complete platform is not equivalent to the early FastChat repository.

Basic information

fieldContent
Tool nameArena AI
Historical nameLMArena, Chatbot Arena
OriginLMSYS and UC Berkeley research projects
Tool typeAI model blind testing, ranking systems, and platforms for assessing human bias in evaluations
Core mechanismAnonymous pair battles and user voting
Overlay modalText, images, videos, code, and multimodal content
Price patternFree
Whether a reasoning API is providedNo
Is it open source?FastChat was made open source in its early stages; the current platform is not fully open source.

Recommendation score

4.7 / 5. Arena AI serves as an important resource for understanding the actual preferences of real users and evaluating the performance of advanced models; anonymous battles and public research hold unique value. However, the rankings do not represent an objective measure of overall performance, and the choice of model must take into account internal requirements, costs, security aspects, and deployment needs.

Frequently Asked Questions

Are LMArena and Chatbot Arena the same platform?

They belong to the same stage of development for that project; the current brand name is Arena. Chatbot Arena was the original name, while LMArena is a separate brand that emerged later.

Is Arena AI free?

Public battles, Direct experiences, and leaderboards are currently free, but the availability of models, waiting times, and the number of uses may vary depending on capacity.

Can I see the model name before voting?

The Battle mode is not displayed; identities are revealed only after voting is submitted. The Direct mode allows users to choose known models directly.

Is the model that ranks first on the list the best one?

Not necessarily. The ranking reflects relative preferences under specific user and prompt distributions; actual choices also take into account task accuracy, cost, latency, privacy, and deployment considerations.

Does Arena provide model APIs?

There is no unified production inference API available for ordinary developers. It is primarily used for experimentation, competition, and evaluation, and it is not a platform for aggregating model calls.

Is Arena open source?

In its early stages, FastChat made some research data available, but currently the full Arena website, all competition categories, and the operational systems cannot be considered fully open source.

©️Copyright notice: Unless otherwise specified, all articles on this site are copyrighted bySharing of AI toolsAll content on this site is original; without permission, no individual, media outlet, website, or organization may reproduce, copy, or otherwise distribute it, nor may they create mirrors of it on servers that are not owned by this site. Otherwise, we reserve the right to take legal action against such parties in accordance with the law.

Tools similar to the Chatbot Arena Leaderboard