2509 / MO2026

DAILY AI MODEL INDEX

Frontier Model Ranking

One transparent yardstick for task performance, with the evidence and limitations kept in view.

30 models23 tasksBenchmark release LiveBench-2026-06-25Data checked 2026-09-25

中文EN

AIHOT Frontier Model Consensus Ranking Top 30

Cost: USD / million tokens

Token Index Method and sources
01
GPT-6 Astra Max EffortOpenAIPublic API
Strong at both everyday analysis and complex reasoning
Released2026-09-04
Configurationmax
Input$10
Output$50
94.0
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
82.66
Science and complex reasoning
94.73
Coding
80.36
Agent / real repository work
57.32
02
Claude Fable 5.1 Max EffortAnthropicPublic API
Balanced across analysis, reasoning, coding, and real repository work
Released2026-09-01
Configurationmax
Input$10
Output$50
93.9
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
80.92
Science and complex reasoning
94.35
Coding
86.38
Agent / real repository work
66.06
03
Claude Fable 5 Max EffortAnthropicPublic API
Strong at both understanding requests and writing code
Released2026-06-09
Configurationmax
Input$10
Output$50
93.2
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
82.33
Science and complex reasoning
92.82
Coding
85.99
Agent / real repository work
62.17
04
Claude Opus 5 Max EffortAnthropicPublic API
Strong at complex reasoning and real repository work
Released2026-07-24
Configurationmax
Input$5
Output$25
93.1
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
75.67
Science and complex reasoning
93.47
Coding
81.45
Agent / real repository work
65.20
05
Especially strong at following instructions and everyday analysis
Released2026-09-02
Configurationxhigh
Input$1.25
Output$4.25
90.9
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
80.12
Science and complex reasoning
92.80
Coding
81.06
Agent / real repository work
64.09
06
GPT-5.6 Sol Max EffortOpenAIPublic API
Strong at both complex reasoning and coding
Released2026-07-09
Configurationmax
Input$5
Output$30
88.7
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
79.79
Science and complex reasoning
93.93
Coding
83.94
Agent / real repository work
56.21
07
GLM-5.3Z.aiOpen weights
Especially strong at sustained work inside real code repositories
Released2026-08-18
Configurationdefault
Input$1.4
Output$4.4
85.8
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
73.13
Science and complex reasoning
86.85
Coding
78.95
Agent / real repository work
60.91
08
Grok 4.6xAIPublic API
Especially strong at mathematics, science, and multi-step reasoning
Released2026-08-12
Configurationdefault
Input$2
Output$6
84.1
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
76.47
Science and complex reasoning
91.54
Coding
76.78
Agent / real repository work
57.02
09
Strong at both complex reasoning and coding
Released2026-09-22
Configurationmax
Input$4
Output$20
83.3
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
77.44
Science and complex reasoning
94.62
Coding
89.25
Agent / real repository work
71.72
10
Kimi K3Moonshot AIOpen weights
Especially strong at following instructions and everyday analysis
Released2026-07-16
Configurationdefault
Input$3
Output$15
83.1
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
78.54
Science and complex reasoning
87.55
Coding
81.45
Agent / real repository work
62.17
11
Strong at coding and sustained work in real repositories
Released2026-09-22
Configurationxhigh
Input$4
Output$20
81.4
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
77.45
Science and complex reasoning
93.73
Coding
89.25
Agent / real repository work
65.35
12
GPT-5.6 Terra Max EffortOpenAIPublic API
Especially strong at mathematics, science, and multi-step reasoning
Released2026-07-09
Configurationmax
Input$2.5
Output$15
80.5
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
75.61
Science and complex reasoning
92.77
Coding
78.25
Agent / real repository work
54.95
13
Strong at both complex reasoning and coding
Released2026-04-16
Configurationmax
Input$5
Output$25
78.8
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
72.58
Science and complex reasoning
91.75
Coding
81.83
Agent / real repository work
50.50
14
Gemini 3.8 Flash HighGooglePublic API
Especially strong at mathematics, science, and multi-step reasoning
Released2026-09-02
Configurationhigh
Input$0.75
Output$3.75
78.1
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
74.40
Science and complex reasoning
90.43
Coding
72.49
Agent / real repository work
54.24
15
GLM-5.3 FlashZ.AIOpen weights
Especially strong at sustained work inside real code repositories
Released2026-08-26
Configurationdefault
Input$0.15
Output$0.5
76.9
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
68.84
Science and complex reasoning
79.44
Coding
78.95
Agent / real repository work
56.77
16
GPT-6 Sol Max EffortOpenAIPublic API
Strong at both everyday analysis and complex reasoning
Released2026-09-22
Configurationmax
Input$2
Output$10
76.4
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
78.35
Science and complex reasoning
92.50
Coding
81.77
Agent / real repository work
52.88
17
Especially strong at writing correct code and completing programs
Released2026-04-16
Configurationxhigh
Input$5
Output$25
75.9
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
74.31
Science and complex reasoning
90.02
Coding
82.09
Agent / real repository work
50.66
18
Union AlphaStealthPublic API
Especially strong at writing correct code and completing programs
Released2026-09-16
Configurationdefault
Input$0
Output$0
74.6
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
73.34
Science and complex reasoning
88.02
Coding
82.15
Agent / real repository work
54.70
19
Grok 4.7 xHighxAIPublic API
Especially strong at following instructions and everyday analysis
Released2026-09-21
Configurationxhigh
Input$2
Output$6
74.4
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
77.43
Science and complex reasoning
89.17
Coding
77.16
Agent / real repository work
53.99
20
Qwen 3.8 MaxAlibabaOpen weights
Especially strong at sustained work inside real code repositories
Released2026-07-19
Configurationmax
Input$2
Output$6
73.6
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
77.39
Science and complex reasoning
89.76
Coding
72.87
Agent / real repository work
64.65
21
GPT-6 Luna Max EffortOpenAIPublic API
Especially strong at writing correct code and completing programs
Released2026-09-22
Configurationmax
Input$0.1
Output$0.5
70.8
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
67.71
Science and complex reasoning
85.45
Coding
78.95
Agent / real repository work
51.21
22
DeepSeek V4.1 Flash Max EffortDeepSeekOpen weights
Especially strong at sustained work inside real code repositories
Released2026-09-10
Configurationmax
Input$0.15
Output$0.6
66.7
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
76.83
Science and complex reasoning
89.99
Coding
80.04
Agent / real repository work
77.27
23
Gemini 3.7 Flash HighGooglePublic API
Especially strong at following instructions and everyday analysis
Released2026-08-13
Configurationhigh
Input$0.75
Output$3.75
64.9
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
77.78
Science and complex reasoning
90.63
Coding
78.89
Agent / real repository work
58.28
24
Grok 4.5xAIPublic API
Strong at understanding requests and handling real repository work
ReleasedTo verify
Configurationdefault
Input$2
Output$6
58.9
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
75.79
Science and complex reasoning
89.00
Coding
68.59
Agent / real repository work
56.46
25
Especially strong at following instructions and everyday analysis
Released2026-04-23
Configurationxhigh
Input$5
Output$30
58.4
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
79.89
Science and complex reasoning
92.76
Coding
82.15
Agent / real repository work
53.99
26
Strong at both everyday analysis and complex reasoning
Released2026-03-05
Configurationxhigh
Input$2.5
Output$15
58.0
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
77.39
Science and complex reasoning
91.13
Coding
77.54
Agent / real repository work
53.84
27
GLM-5.2Z.AIOpen weights
Especially strong at writing correct code and completing programs
ReleasedTo verify
Configurationdefault
Input$1.4
Output$4.4
54.4
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
70.76
Science and complex reasoning
84.20
Coding
79.65
Agent / real repository work
51.77
28
Claude Sonnet 5 xHigh EffortAnthropicPublic API
Strong at complex reasoning and real repository work
ReleasedTo verify
Configurationxhigh
Input$3
Output$15
52.8
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
70.19
Science and complex reasoning
90.82
Coding
80.68
Agent / real repository work
59.39
29
GPT-5.6 Luna Max EffortOpenAIPublic API
Especially strong at writing correct code and completing programs
Released2026-07-09
Configurationmax
Input$1
Output$6
49.7
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
70.24
Science and complex reasoning
86.42
Coding
82.91
Agent / real repository work
48.43
30
DeepSeek V4 Pro 0813DeepSeekOpen weights
Especially strong at mathematics, science, and multi-step reasoning
Released2026-08-13
Configurationdefault
Input$0.435
Output$0.87
49.0
Consensus evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
76.34
Science and complex reasoning
90.46
Coding
77.16
Agent / real repository work
54.95

AIHOT Frontier Model Consensus Ranking · AIHOT Consensus v1

Aggregating 9 public benchmarks from 7 independent organizations with fixed topic budgets into a unified AIHOT consensus score.

General Evaluation 30% + Human Preference 10% + Specialist Benchmarks 60% (Coding 12% + Math & Reasoning 12% + Writing 9% + Knowledge 6% + Vision 6% + Tools & Office 6% + Multilingual 6% + Industry 3%)

Consensus is evaluated across models with at least 3 independent evidence families. Missing sources reduce completeness rather than penalizing scores directly.

Aggregated across 7 independent benchmark organizations. Pricing verified from official vendor documentation.

A state link opens the latest list. Long images record the data date visible when shared.