登录

EVALUATION / MODEL SELECTION

Benchmarks

Compare model performance by task, then open the original benchmark for methodology and full results.

5Boards
72Models
08-29Updated

SELECT A LENS

Editorial reference scores, refreshed with the catalog.

Overall

9/9

A blended view of chat, reasoning, coding, and multimodal performance.

SourceLiveBench 2026-06-25Release 2026-06-25Updated 2026-08-29View source
Editorial reference
01+3
GPT-5.6OpenAI
81/ 100
ReasoningCodingTool use
Platform pricing
Details
02+2
GPT-5.5OpenAI
80/ 100
General reasoningCodingMulti-turn
Platform pricing
Details
03+1
GPT-5.4OpenAI
78/ 100
VisionDocumentsReasoning
Platform pricing
Details
04+4
Qwen3.8-Flash-NextQwen · Open source
76/ 100
Long contextOpenCost
Platform pricing
Details
050
75/ 100
Live informationReasoningTools
USD 3.0 / 1M
Details
06+2
73/ 100
CodingAgentsLong context
Platform pricing
Details
07-4
73/ 100
Complex tasksCodingLong context
USD 5.0 / 1M
Details
080
Qwen3Qwen · Open source
72/ 100
ChineseOpenMultilingual
Platform pricing
Details
09+2
GLM-5.3 FlashZhipu · Open source
71/ 100
ChineseAgentsLow latency
Platform pricing
Details

EVALUATION SOURCES

Go deeper into the original platforms

Methodology, raw results, and community context live at the source.

A public evaluation platform comparing LLM output quality through real-user voting.

GlobalUpdated 2026-08-29

An independent model and API analysis site covering quality, speed, and price.

GlobalUpdated 2026-08-29
L

A continuously refreshed, contamination-resistant benchmark with more reliable results.

GlobalUpdated 2026-08-29
S

A benchmark that tests model coding ability on real GitHub issues.

GlobalUpdated 2026-08-29
A

Aider's code-editing leaderboard, designed around real development tasks.

GlobalUpdated 2026-08-29

Shanghai AI Lab's open evaluation suite and a respected Chinese-language leaderboard.

ChinaUpdated 2026-08-29
S

A comprehensive Chinese LLM benchmark with monthly leaderboards.

ChinaUpdated 2026-08-29

A model popularity leaderboard based on real API usage and adoption.

GlobalUpdated 2026-08-29