Data refresh is delayed. This is the last fully verified snapshot. Last complete check: 2026-08-10T09:00:00.000+08:00 Sources not completed in this check: official vendor catalogs
Rank Model / vendor Best at Released Configuration Input Output AIdaily Score
01
Balanced across analysis, reasoning, coding, and real repository work
Released 2026-06-09
Configuration max
Input $10
Output $50
80.8
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 82.33
Science and complex reasoning 92.82
Coding 85.99
Agent / real repository work 62.17
Official page LiveBench source data
02
Strong at complex reasoning and real repository work
Released 2026-07-22
Configuration max
Input $5
Output $25
78.9
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 75.67
Science and complex reasoning 93.47
Coding 81.45
Agent / real repository work 65.20
Official page LiveBench source data
03
Strong at both complex reasoning and coding
Released 2026-07-09
Configuration max
Input $5
Output $30
78.5
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 79.79
Science and complex reasoning 93.93
Coding 83.94
Agent / real repository work 56.21
Official page LiveBench source data
04
Strong at understanding requests and handling real repository work
Released 2026-07-16
Configuration default
Input $3
Output $15
77.4
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 78.54
Science and complex reasoning 87.55
Coding 81.45
Agent / real repository work 62.17
Official page LiveBench source data
05
Especially strong at following instructions and everyday analysis
Released 2026-04-23
Configuration xhigh
Input $5
Output $30
77.2
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 79.89
Science and complex reasoning 92.76
Coding 82.15
Agent / real repository work 53.99
Official page LiveBench source data
06
Especially strong at sustained work inside real code repositories
Released 2026-07-19
Configuration max
Input $2
Output $6
76.2
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 77.39
Science and complex reasoning 89.76
Coding 72.87
Agent / real repository work 64.65
Official page LiveBench source data
07
Strong at understanding requests and handling real repository work
Released 2026-08-05
Configuration xhigh
Input $1.25
Output $4.25
75.5
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 76.45
Science and complex reasoning 90.60
Coding 77.54
Agent / real repository work 57.58
Official page LiveBench source data
08
Especially strong at mathematics, science, and multi-step reasoning
Released 2026-07-09
Configuration max
Input $2.5
Output $15
75.4
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 75.61
Science and complex reasoning 92.77
Coding 78.25
Agent / real repository work 54.95
Official page LiveBench source data
09
Especially strong at sustained work inside real code repositories
Released To verify
Configuration xhigh
Input $3
Output $15
75.3
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 70.19
Science and complex reasoning 90.82
Coding 80.68
Agent / real repository work 59.39
Official page LiveBench source data
10
Strong at both everyday analysis and complex reasoning
Released 2026-03-05
Configuration xhigh
Input $2.5
Output $15
75.0
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 77.39
Science and complex reasoning 91.13
Coding 77.54
Agent / real repository work 53.84
Official page LiveBench source data
11
Especially strong at writing correct code and completing programs
Released 2026-04-16
Configuration xhigh
Input $5
Output $25
74.3
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 74.31
Science and complex reasoning 90.02
Coding 82.09
Agent / real repository work 50.66
Official page LiveBench source data
12
Strong at both complex reasoning and coding
Released 2026-04-16
Configuration max
Input $5
Output $25
74.2
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 72.58
Science and complex reasoning 91.75
Coding 81.83
Agent / real repository work 50.50
Official page LiveBench source data
13
Especially strong at sustained work inside real code repositories
Released To verify
Configuration xhigh
Input $1.25
Output $4.25
73.8
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 72.18
Science and complex reasoning 87.44
Coding 77.16
Agent / real repository work 58.54
Official page LiveBench source data
14
Strong at understanding requests and handling real repository work
Released To verify
Configuration default
Input $2
Output $6
72.5
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 75.79
Science and complex reasoning 89.00
Coding 68.59
Agent / real repository work 56.46
Official page LiveBench source data
15
Especially strong at following instructions and everyday analysis
Released To verify
Configuration high
Input $2
Output $12
72.3
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 81.01
Science and complex reasoning 87.52
Coding 76.45
Agent / real repository work 44.14
Official page LiveBench source data
16
Especially strong at writing correct code and completing programs
Released 2026-01-14
Configuration default
Input $1.75
Output $14
72.3
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 72.78
Science and complex reasoning 83.24
Coding 83.62
Agent / real repository work 49.39
Official page LiveBench source data
17
Especially strong at mathematics, science, and multi-step reasoning
Released 2026-02-05
Configuration high
Input $5
Output $25
72.1
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 72.16
Science and complex reasoning 89.00
Coding 78.18
Agent / real repository work 48.99
Official page LiveBench source data
18
Especially strong at writing correct code and completing programs
Released 2026-07-09
Configuration max
Input $1
Output $6
72.0
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 70.24
Science and complex reasoning 86.42
Coding 82.91
Agent / real repository work 48.43
Official page LiveBench source data
19
Especially strong at mathematics, science, and multi-step reasoning
Released 2025-12-11
Configuration high
Input $1.75
Output $14
71.9
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 73.25
Science and complex reasoning 88.19
Coding 76.07
Agent / real repository work 50.25
Official page LiveBench source data
20
Especially strong at following instructions and everyday analysis
Released 2026-05-19
Configuration high
Input $1.5
Output $9
71.8
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 75.01
Science and complex reasoning 85.12
Coding 78.18
Agent / real repository work 48.99
Official page LiveBench source data
21
Strong at coding and sustained work in real repositories
Released To verify
Configuration default
Input $1.4
Output $4.4
71.6
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 70.76
Science and complex reasoning 84.20
Coding 79.65
Agent / real repository work 51.77
Official page LiveBench source data
22
Especially strong at following instructions and everyday analysis
Released 2026-07-31
Configuration default
Input $0.14
Output $0.28
70.8
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 74.67
Science and complex reasoning 86.71
Coding 74.98
Agent / real repository work 46.77
Official page LiveBench source data
23
Especially strong at following instructions and everyday analysis
Released To verify
Configuration high
Input $1.5
Output $7.5
70.3
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 74.09
Science and complex reasoning 85.78
Coding 77.86
Agent / real repository work 43.43
Official page LiveBench source data
24
Especially strong at writing correct code and completing programs
Released 2026-02-17
Configuration medium
Input $3
Output $15
70.1
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 72.42
Science and complex reasoning 85.88
Coding 79.27
Agent / real repository work 42.63
Official page LiveBench source data
25
Especially strong at writing correct code and completing programs
Released 2025-11-01
Configuration high
Input $5
Output $25
69.3
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 72.75
Science and complex reasoning 85.24
Coding 79.65
Agent / real repository work 39.70
Official page LiveBench source data
26
Especially strong at following instructions and everyday analysis
Released To verify
Configuration max
Input $2.5
Output $7.5
69.3
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 75.19
Science and complex reasoning 84.29
Coding 74.22
Agent / real repository work 43.59
Official page LiveBench source data
27
Especially strong at sustained work inside real code repositories
Released To verify
Configuration xhigh
Input $1.87
Output $4.68
69.0
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 72.11
Science and complex reasoning 83.36
Coding 71.02
Agent / real repository work 49.39
Official page LiveBench source data
28
Especially strong at writing correct code and completing programs
Released To verify
Configuration default
Input $0.95
Output $4
68.9
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 68.21
Science and complex reasoning 81.83
Coding 78.57
Agent / real repository work 46.92
Official page LiveBench source data
29
Especially strong at mathematics, science, and multi-step reasoning
Released 2026-04-24
Configuration default
Input $0.435
Output $0.87
67.7
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 71.67
Science and complex reasoning 86.68
Coding 69.99
Agent / real repository work 42.63
Official page LiveBench source data
30
Especially strong at mathematics, science, and multi-step reasoning
Released 2026-03-17
Configuration xhigh
Input $0.2
Output $1.25
67.4
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 65.78
Science and complex reasoning 86.04
Coding 70.84
Agent / real repository work 46.77
Official page LiveBench source data
Brand overview
General models Several current models can remain under one vendor, so a whole brand is not judged by a single model.
Especially strong at writing correct code and completing programs
2026-07-09
Availability Public APIStage StableAccess apiPrice $1 / $6Source
Strong at both complex reasoning and coding
2026-07-09
Availability Public APIStage StableAccess apiPrice $5 / $30Source
Especially strong at mathematics, science, and multi-step reasoning
2026-07-09
Availability Public APIStage StableAccess apiPrice $2.5 / $15Source
Especially strong at following instructions and everyday analysis
2026-04-23
Availability Public APIStage StableAccess apiPrice $5 / $30Source
Especially strong at writing correct code and completing programs
2026-03-17
Availability Public APIStage StableAccess apiPrice $0.75 / $4.5Source
Especially strong at mathematics, science, and multi-step reasoning
2026-03-17
Availability Public APIStage StableAccess apiPrice $0.2 / $1.25Source
Strong at both everyday analysis and complex reasoning
2026-03-05
Availability Public APIStage StableAccess apiPrice $2.5 / $15Source
Especially strong at writing correct code and completing programs
2026-01-14
Availability Public APIStage StableAccess apiPrice $1.75 / $14Source
Especially strong at mathematics, science, and multi-step reasoning
2025-12-11
Availability Public APIStage StableAccess apiPrice $1.75 / $14Source
Strong at complex reasoning and real repository work
2026-07-22
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Balanced across analysis, reasoning, coding, and real repository work
2026-06-09
Availability Public APIStage StableAccess apiPrice $10 / $50Source
Especially strong at writing correct code and completing programs
2026-04-16
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Strong at both complex reasoning and coding
2026-04-16
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Especially strong at writing correct code and completing programs
2026-02-17
Availability Public APIStage StableAccess apiPrice $3 / $15Source
Especially strong at mathematics, science, and multi-step reasoning
2026-02-05
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Especially strong at writing correct code and completing programs
2025-11-01
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Especially strong at sustained work inside real code repositories
To verify
Availability Public APIStage StableAccess apiPrice $3 / $15Source
Especially strong at following instructions and everyday analysis
2026-05-19
Availability Public APIStage StableAccess apiPrice $1.5 / $9Source
Especially strong at writing correct code and completing programs
To verify
Availability Public APIStage StableAccess apiPrice $0.3 / $2.5Source
Especially strong at following instructions and everyday analysis
To verify
Availability Public APIStage StableAccess apiPrice $1.5 / $7.5Source
Especially strong at following instructions and everyday analysis
To verify
Availability Public APIStage PreviewAccess apiPrice $2 / $12Source
Strong at understanding requests and handling real repository work
2026-08-05
Availability Public APIStage StableAccess apiPrice $1.25 / $4.25Source
Especially strong at sustained work inside real code repositories
To verify
Availability Public APIStage StableAccess apiPrice $1.25 / $4.25Source
Especially strong at mathematics, science, and multi-step reasoning
To verify
Availability Public APIStage StableAccess apiPrice $1.25 / $2.5Source
Strong at understanding requests and handling real repository work
To verify
Availability Public APIStage StableAccess apiPrice $2 / $6Source
Especially strong at sustained work inside real code repositories
To verify
Availability Public APIStage StableAccess apiPrice $1 / $2Source
Best for general enterprise work with European deployment options
2025-05-07
Availability ConsumerStage StableAccess consumerPrice To verifySource
Especially strong at sustained work inside real code repositories
2026-07-19
Availability Public APIStage StableAccess apiPrice $2 / $6Source
Especially strong at writing correct code and completing programs
2026-04-02
Availability Public APIStage StableAccess apiPrice $0.5 / $3Source
Especially strong at following instructions and everyday analysis
To verify
Availability Public APIStage StableAccess apiPrice $2.5 / $7.5Source
Especially strong at writing correct code and completing programs
To verify
Availability Open weightsStage StableAccess api, weightsPrice $0.6 / $3.6Source
Best for Chinese everyday tasks inside Tencent Cloud
2025-03-13
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for Chinese Q&A and content work in Baidu AI Cloud
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Especially strong at following instructions and everyday analysis
2026-07-31
Availability Open weightsStage StableAccess api, weightsPrice $0.14 / $0.28Source
Especially strong at mathematics, science, and multi-step reasoning
2026-04-24
Availability Open weightsStage StableAccess api, weightsPrice $0.435 / $0.87Source
Balanced across analysis, reasoning, coding, and real repository work
To verify
Availability Open weightsStage StableAccess api, weightsPrice $0.14 / $0.28Source
Strong at understanding requests and handling real repository work
2026-07-16
Availability Open weightsStage StableAccess api, weightsPrice $3 / $15Source
Especially strong at writing correct code and completing programs
To verify
Availability Open weightsStage StableAccess api, weightsPrice $0.95 / $4Source
Especially strong at sustained work inside real code repositories
To verify
Availability Open weightsStage StableAccess api, weightsPrice $0.95 / $4Source
Especially strong at following instructions and everyday analysis
To verify
Availability Public APIStage StableAccess apiPrice $0.3 / $1.2Source
Strong at coding and sustained work in real repositories
To verify
Availability Open weightsStage StableAccess api, weightsPrice $1.4 / $4.4Source
Thinking Machines 1 models
Especially strong at sustained work inside real code repositories
To verify
Availability Public APIStage StableAccess apiPrice $1.87 / $4.68Source
Public picks
General models Only models available to consumers or through a public API are recommended, grouped by practical use.
Overall capability
Balanced across analysis, reasoning, coding, and real repository work
2026-06-09
Availability Public APIStage StableAccess apiPrice $10 / $50Source
Strong at complex reasoning and real repository work
2026-07-22
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Strong at both complex reasoning and coding
2026-07-09
Availability Public APIStage StableAccess apiPrice $5 / $30Source
Coding and real repository work
Balanced across analysis, reasoning, coding, and real repository work
2026-06-09
Availability Public APIStage StableAccess apiPrice $10 / $50Source
Strong at complex reasoning and real repository work
2026-07-22
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Strong at understanding requests and handling real repository work
2026-07-16
Availability Open weightsStage StableAccess api, weightsPrice $3 / $15Source
Complex reasoning
Strong at both complex reasoning and coding
2026-07-09
Availability Public APIStage StableAccess apiPrice $5 / $30Source
Strong at complex reasoning and real repository work
2026-07-22
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Balanced across analysis, reasoning, coding, and real repository work
2026-06-09
Availability Public APIStage StableAccess apiPrice $10 / $50Source
Price friendly
Strong at understanding requests and handling real repository work
2026-08-05
Availability Public APIStage StableAccess apiPrice $1.25 / $4.25Source
Especially strong at sustained work inside real code repositories
To verify
Availability Public APIStage StableAccess apiPrice $1.25 / $4.25Source
Especially strong at sustained work inside real code repositories
2026-07-19
Availability Public APIStage StableAccess apiPrice $2 / $6Source
Open weights
Strong at understanding requests and handling real repository work
2026-07-16
Availability Open weightsStage StableAccess api, weightsPrice $3 / $15Source
Strong at coding and sustained work in real repositories
To verify
Availability Open weightsStage StableAccess api, weightsPrice $1.4 / $4.4Source
Especially strong at following instructions and everyday analysis
2026-07-31
Availability Open weightsStage StableAccess api, weightsPrice $0.14 / $0.28Source
No comparable capability ranking is available yet. These are public specialist models, and AIdaily does not force different modalities into one score.
Best for turning detailed instructions into editable finished images
2025-04-23
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for images with clear text and realistic detail
2025-05-20
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for Chinese creative images and multi-reference editing
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for fast, visually distinctive concept art
2025-04-03
Availability ConsumerStage StableAccess consumerPrice To verifySource
Several current models can remain under one vendor, so a whole brand is not judged by a single model.
Best for turning detailed instructions into editable finished images
2025-04-23
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for images with clear text and realistic detail
2025-05-20
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for Chinese creative images and multi-reference editing
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for fast, visually distinctive concept art
2025-04-03
Availability ConsumerStage StableAccess consumerPrice To verifySource
Only models available to consumers or through a public API are recommended, grouped by practical use.
New and current models
Best for images with clear text and realistic detail
2025-05-20
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for turning detailed instructions into editable finished images
2025-04-23
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for fast, visually distinctive concept art
2025-04-03
Availability ConsumerStage StableAccess consumerPrice To verifySource
No comparable capability ranking is available yet. These are public specialist models, and AIdaily does not force different modalities into one score.
Best for coherent shots and short videos with sound
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for cinematic visuals with native audio
2025-05-20
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for long narratives controlled by several references
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for cost-conscious video iteration
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for video with coherent motion and stable shots
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Several current models can remain under one vendor, so a whole brand is not judged by a single model.
Best for coherent shots and short videos with sound
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for cinematic visuals with native audio
2025-05-20
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for long narratives controlled by several references
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for cost-conscious video iteration
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for video with coherent motion and stable shots
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Only models available to consumers or through a public API are recommended, grouped by practical use.
New and current models
Best for cinematic visuals with native audio
2025-05-20
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for video with coherent motion and stable shots
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for cost-conscious video iteration
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
No comparable capability ranking is available yet. These are public specialist models, and AIdaily does not force different modalities into one score.
Best for quickly turning text into natural speech
2025-03-20
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for multi-speaker dialogue and fine voice control
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for expressive multilingual voice work
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Several current models can remain under one vendor, so a whole brand is not judged by a single model.
Best for quickly turning text into natural speech
2025-03-20
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for multi-speaker dialogue and fine voice control
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for expressive multilingual voice work
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Only models available to consumers or through a public API are recommended, grouped by practical use.
New and current models
Best for quickly turning text into natural speech
2025-03-20
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for expressive multilingual voice work
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for multi-speaker dialogue and fine voice control
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Restricted model watchlist These models are not broadly public, so they are excluded from the Top 30 and public picks.
Claude Mythos · Anthropic Not broadly available to consumers or through a public API. It stays on the watchlist and is excluded from ranking and recommendations.
AIdaily Capability Reference · Composite v1 Each of four pillars carries 25%. The score covers tasks from one LiveBench release only. It does not represent Chinese ability, price, speed, creative preference, or every kind of agent work.
General 25% + science and complex reasoning 25% + coding 25% + agent / real repository work 25%
A small score gap may not be meaningful. A model missing any task is not ranked, and missing values are never set to zero or reweighted.
Data attributed to LiveBench and Abacus.AI. AIdaily re-aggregates the public results into four pillars; this is not an official LiveBench total. The one-line strength is not vendor marketing.
A state link opens the latest list. Long images record the data date visible when shared.
LiveBench LiveBench source data Datasheet Apache 2.0
Today’s AI Digest All issues