Scholé

Scholé answer engine

how do I learn LMArena fast?

Top answer found in 0.31 s 1 interactive lesson 114 tools indexed free, no signup

2 try.schole.ai › learn › lmarena › concepts

Blind, paired, and voted[2]

Benchmarks measure what a test-writer thought to ask. Arena votes measure what real people preferred on their own prompts, without knowing which model answered. Both are useful; they answer different questions.

Battle

Two answers, no labels

You type a prompt, two hidden models reply side by side, and you pick the better one. Only after voting do the names appear.

Elo

A rating from many votes

Each vote nudges a rating the way chess ratings move. A model near the top of the leaderboard won more blind comparisons, on prompts people cared about.

Categories

Good at what?

Rankings differ by task: coding, long prompts, hard questions, a given language. The overall number hides the one that matters to you.

3 try.schole.ai › learn › lmarena › inside-the-lesson

What you will do in the lesson[1]

You answer and try things. Nobody demonstrates at you for five minutes.

  1. Run a battle on your own prompt. You bring a task from your actual work, vote, and only then see which models you compared.
  2. Catch your bias. You run a prompt where the longer answer is worse and notice whether you still voted for it.
  3. Read the leaderboard by category. You compare the overall ranking with the coding ranking and find a model that swaps places.
  4. Judge a tie. You get two answers that are both fine and decide when to vote tie versus splitting hairs.
  5. Pick a model for a job. You end by choosing a model for a realistic task using arena data and one benchmark, and say why.

4 try.schole.ai › learn › lmarena › questions

People also ask

Is LMArena free?

Yes. You can run battles without an account, and voting is how the leaderboard gets its data.

Can the rankings be gamed?

Somewhat: style, length, and confidence sway voters, which is why the site also publishes style-controlled rankings. Read those alongside the raw ones.

Where did it come from?

It started as a research project at UC Berkeley, in the same lab community as some of the Scholé team, and grew into the standard for human-preference model rankings.

Should I trust it over benchmarks?

Trust it for how a model feels to use on open-ended tasks. Trust benchmarks for narrow, checkable skills. Use both before spending money.

Learners from these organizations are already learning on Scholé

Partner organization logo Partner organization logo Partner organization logo Partner organization logo Partner organization logo Partner organization logo Partner organization logo