Epoch AI Benchmarking Hub
Ranked among models offered by everyais · 2 models
| Rank | Model | Score |
|---|---|---|
| 1 | Claude Opus 5 Evaluation target / conditions: claude-opus-5 | 92.9% |
| 2 | MiniMax M3 Evaluation target / conditions: MiniMax-M3 | 90.9% |
External benchmark rankings
Graduate-level science questions designed to resist simple retrieval.
Ranked among models offered by everyais · 2 models
| Rank | Model | Score |
|---|---|---|
| 1 | Claude Opus 5 Evaluation target / conditions: claude-opus-5 | 92.9% |
| 2 | MiniMax M3 Evaluation target / conditions: MiniMax-M3 | 90.9% |
Each source, unit and index version has its own ranking. This is not the source’s full leaderboard. Equal scores share a rank.
Results use the source’s evaluation settings; they do not represent everyais default API settings.
Models without evaluation data are not ranked; missing data does not mean a low score.
Sync time is separate from the measurement date.