1008 / MO2026

DAILY AI MODEL INDEX

Frontier Model Ranking

One transparent yardstick for task performance, with the evidence and limitations kept in view.

30 models23 tasksBenchmark release LiveBench-2026-06-25Data checked 2026-08-10

中文EN

Data refresh is delayed. This is the last fully verified snapshot. Last complete check: Sources not completed in this check: official vendor catalogs

AIdaily Capability Ranking Top 30

Cost: USD / million tokens

01
Claude Fable 5 Max EffortAnthropicPublic API
Balanced across analysis, reasoning, coding, and real repository work
Released2026-06-09
Configurationmax
Input$10
Output$50
80.8
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
82.33
Science and complex reasoning
92.82
Coding
85.99
Agent / real repository work
62.17
02
Strong at complex reasoning and real repository work
Released2026-07-22
Configurationmax
Input$5
Output$25
78.9
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
75.67
Science and complex reasoning
93.47
Coding
81.45
Agent / real repository work
65.20
03
GPT-5.6 Sol Max EffortOpenAIPublic API
Strong at both complex reasoning and coding
Released2026-07-09
Configurationmax
Input$5
Output$30
78.5
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
79.79
Science and complex reasoning
93.93
Coding
83.94
Agent / real repository work
56.21
04
Kimi K3Moonshot AIOpen weights
Strong at understanding requests and handling real repository work
Released2026-07-16
Configurationdefault
Input$3
Output$15
77.4
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
78.54
Science and complex reasoning
87.55
Coding
81.45
Agent / real repository work
62.17
05
Especially strong at following instructions and everyday analysis
Released2026-04-23
Configurationxhigh
Input$5
Output$30
77.2
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
79.89
Science and complex reasoning
92.76
Coding
82.15
Agent / real repository work
53.99
06
Qwen 3.8 MaxAlibabaPublic API
Especially strong at sustained work inside real code repositories
Released2026-07-19
Configurationmax
Input$2
Output$6
76.2
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
77.39
Science and complex reasoning
89.76
Coding
72.87
Agent / real repository work
64.65
07
Strong at understanding requests and handling real repository work
Released2026-08-05
Configurationxhigh
Input$1.25
Output$4.25
75.5
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
76.45
Science and complex reasoning
90.60
Coding
77.54
Agent / real repository work
57.58
08
GPT-5.6 Terra Max EffortOpenAIPublic API
Especially strong at mathematics, science, and multi-step reasoning
Released2026-07-09
Configurationmax
Input$2.5
Output$15
75.4
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
75.61
Science and complex reasoning
92.77
Coding
78.25
Agent / real repository work
54.95
09
Claude Sonnet 5 xHigh EffortAnthropicPublic API
Especially strong at sustained work inside real code repositories
ReleasedTo verify
Configurationxhigh
Input$3
Output$15
75.3
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
70.19
Science and complex reasoning
90.82
Coding
80.68
Agent / real repository work
59.39
10
Strong at both everyday analysis and complex reasoning
Released2026-03-05
Configurationxhigh
Input$2.5
Output$15
75.0
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
77.39
Science and complex reasoning
91.13
Coding
77.54
Agent / real repository work
53.84
11
Especially strong at writing correct code and completing programs
Released2026-04-16
Configurationxhigh
Input$5
Output$25
74.3
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
74.31
Science and complex reasoning
90.02
Coding
82.09
Agent / real repository work
50.66
12
Strong at both complex reasoning and coding
Released2026-04-16
Configurationmax
Input$5
Output$25
74.2
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
72.58
Science and complex reasoning
91.75
Coding
81.83
Agent / real repository work
50.50
13
Especially strong at sustained work inside real code repositories
ReleasedTo verify
Configurationxhigh
Input$1.25
Output$4.25
73.8
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
72.18
Science and complex reasoning
87.44
Coding
77.16
Agent / real repository work
58.54
14
Grok 4.5xAIPublic API
Strong at understanding requests and handling real repository work
ReleasedTo verify
Configurationdefault
Input$2
Output$6
72.5
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
75.79
Science and complex reasoning
89.00
Coding
68.59
Agent / real repository work
56.46
15
Especially strong at following instructions and everyday analysis
ReleasedTo verify
Configurationhigh
Input$2
Output$12
72.3
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
81.01
Science and complex reasoning
87.52
Coding
76.45
Agent / real repository work
44.14
16
GPT-5.2 CodexOpenAIPublic API
Especially strong at writing correct code and completing programs
Released2026-01-14
Configurationdefault
Input$1.75
Output$14
72.3
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
72.78
Science and complex reasoning
83.24
Coding
83.62
Agent / real repository work
49.39
17
Especially strong at mathematics, science, and multi-step reasoning
Released2026-02-05
Configurationhigh
Input$5
Output$25
72.1
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
72.16
Science and complex reasoning
89.00
Coding
78.18
Agent / real repository work
48.99
18
GPT-5.6 Luna Max EffortOpenAIPublic API
Especially strong at writing correct code and completing programs
Released2026-07-09
Configurationmax
Input$1
Output$6
72.0
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
70.24
Science and complex reasoning
86.42
Coding
82.91
Agent / real repository work
48.43
19
GPT-5.2 HighOpenAIPublic API
Especially strong at mathematics, science, and multi-step reasoning
Released2025-12-11
Configurationhigh
Input$1.75
Output$14
71.9
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
73.25
Science and complex reasoning
88.19
Coding
76.07
Agent / real repository work
50.25
20
Gemini 3.5 Flash HighGooglePublic API
Especially strong at following instructions and everyday analysis
Released2026-05-19
Configurationhigh
Input$1.5
Output$9
71.8
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
75.01
Science and complex reasoning
85.12
Coding
78.18
Agent / real repository work
48.99
21
GLM-5.2Z.AIOpen weights
Strong at coding and sustained work in real repositories
ReleasedTo verify
Configurationdefault
Input$1.4
Output$4.4
71.6
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
70.76
Science and complex reasoning
84.20
Coding
79.65
Agent / real repository work
51.77
22
DeepSeek V4 Flash 0731DeepSeekOpen weights
Especially strong at following instructions and everyday analysis
Released2026-07-31
Configurationdefault
Input$0.14
Output$0.28
70.8
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
74.67
Science and complex reasoning
86.71
Coding
74.98
Agent / real repository work
46.77
23
Gemini 3.6 Flash HighGooglePublic API
Especially strong at following instructions and everyday analysis
ReleasedTo verify
Configurationhigh
Input$1.5
Output$7.5
70.3
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
74.09
Science and complex reasoning
85.78
Coding
77.86
Agent / real repository work
43.43
24
Especially strong at writing correct code and completing programs
Released2026-02-17
Configurationmedium
Input$3
Output$15
70.1
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
72.42
Science and complex reasoning
85.88
Coding
79.27
Agent / real repository work
42.63
25
Especially strong at writing correct code and completing programs
Released2025-11-01
Configurationhigh
Input$5
Output$25
69.3
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
72.75
Science and complex reasoning
85.24
Coding
79.65
Agent / real repository work
39.70
26
Qwen 3.7 MaxAlibabaPublic API
Especially strong at following instructions and everyday analysis
ReleasedTo verify
Configurationmax
Input$2.5
Output$7.5
69.3
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
75.19
Science and complex reasoning
84.29
Coding
74.22
Agent / real repository work
43.59
27
Inkling xHigh EffortThinking MachinesPublic API
Especially strong at sustained work inside real code repositories
ReleasedTo verify
Configurationxhigh
Input$1.87
Output$4.68
69.0
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
72.11
Science and complex reasoning
83.36
Coding
71.02
Agent / real repository work
49.39
28
Kimi K2.6 ThinkingMoonshot AIOpen weights
Especially strong at writing correct code and completing programs
ReleasedTo verify
Configurationdefault
Input$0.95
Output$4
68.9
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
68.21
Science and complex reasoning
81.83
Coding
78.57
Agent / real repository work
46.92
29
DeepSeek V4 ProDeepSeekOpen weights
Especially strong at mathematics, science, and multi-step reasoning
Released2026-04-24
Configurationdefault
Input$0.435
Output$0.87
67.7
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
71.67
Science and complex reasoning
86.68
Coding
69.99
Agent / real repository work
42.63
30
GPT-5.4 Nano xHighOpenAIPublic API
Especially strong at mathematics, science, and multi-step reasoning
Released2026-03-17
Configurationxhigh
Input$0.2
Output$1.25
67.4
Four-pillar evidenceCompleteness: 23/23 tasks · 7/7 categories · 4/4 pillars
General
65.78
Science and complex reasoning
86.04
Coding
70.84
Agent / real repository work
46.77

AIdaily Capability Reference · Composite v1

Each of four pillars carries 25%. The score covers tasks from one LiveBench release only. It does not represent Chinese ability, price, speed, creative preference, or every kind of agent work.

General 25% + science and complex reasoning 25% + coding 25% + agent / real repository work 25%

A small score gap may not be meaningful. A model missing any task is not ranked, and missing values are never set to zero or reweighted.

Data attributed to LiveBench and Abacus.AI. AIdaily re-aggregates the public results into four pillars; this is not an official LiveBench total. The one-line strength is not vendor marketing.

A state link opens the latest list. Long images record the data date visible when shared.