External benchmark rankings

Text Arena

Pairwise human preference for text responses.

LMArena Leaderboard

Ranked among models offered by everyais · 3 models

Data source: LMArena LeaderboardLicense: CC-BY-4.0Last synced:
LMArena Leaderboard leaderboard · Ranked among models offered by everyais
RankModelScore
1Claude Fable 5
Evaluation target / conditions: claude-fable-5
View source
1,494 Elo
2GLM-5.3 Flash
Evaluation target / conditions: glm-5.3-flash
View source
1,471 Elo
3Claude Opus 4.8
Evaluation target / conditions: claude-opus-4-8
View source
1,452 Elo

About these rankings

Each source, unit and index version has its own ranking. This is not the source’s full leaderboard. Equal scores share a rank.

Results use the source’s evaluation settings; they do not represent everyais default API settings.

Models without evaluation data are not ranked; missing data does not mean a low score.

Sync time is separate from the measurement date.

Text Arena | everyais