EVALUATION / MODEL SELECTION
Benchmarks
Compare model performance by task, then open the original benchmark for methodology and full results.
SELECT A LENS
Editorial reference scores, refreshed with the catalog.
Overall
9/9A blended view of chat, reasoning, coding, and multimodal performance.
EVALUATION SOURCES
Go deeper into the original platforms
Methodology, raw results, and community context live at the source.
A public evaluation platform comparing LLM output quality through real-user voting.
An independent model and API analysis site covering quality, speed, and price.
A continuously refreshed, contamination-resistant benchmark with more reliable results.
A benchmark that tests model coding ability on real GitHub issues.
Shanghai AI Lab's open evaluation suite and a respected Chinese-language leaderboard.