OpenDesign Benchmark: DeepSeek V4.1 Flash challenges industry benchmarks with highly competitive cost performance
This week, OpenDesign, the operator of the well-known design benchmark website OpenDesign Arena, conducted a horizontal evaluation of the same batch of design tasks on 13 mainstream AI models. In the end, OpenAI's GPT-6 Astra won the top spot with the highest score. However, DeepSeek's latest model, V4.1 Flash, followed closely, with a score of 98% of the top score and a cost of only about 1.4% of the top model.
Evaluation Standards and Core Goals
OpenDesign Arena focuses on evaluating daily design work scenarios, including building Web applications, dashboards, mobile interfaces, and landing pages, with a total score of 100 points. Among them, 30 points are used to assess whether the output content truly meets the demand briefing; the remaining 70 points are used to score the design quality in terms of layout, hierarchical structure, color matching and style fit.
Unlike most AI rankings that aim to showcase universal capabilities, this review aims to answer a more specific and practical question: Which model should front-line web designers really use tomorrow?
Comparison of performance and cost data
Under this evaluation system, the performance of each major model is as follows:
- GPT-6 Astra: Average score of 82.7 points, took 11.1 minutes, and cost a single design of $1.61.
- DeepSeek V4.1 Flash: scores 81.2 points, takes only 5.3 minutes, and costs as low as $0.023.
- Claude Fable 5.1: scores 80.3 points, takes 12.8 minutes, and costs up to $3.66.
Among the 13 models participating in the test, except for the GPT-6 Astra who won with a slight advantage (1.5 points lead), the remaining 11 models (including Grok 4.6, Qwen 3.8-Max, Kimi K3, GLM-5.3 Flash and Gemini 3.8 Flash, etc.) all scored lower than DeepSeek V4.1 Flash and had higher running costs.
Technical analysis: Why can we achieve low cost and high performance?
DeepSeek explained the source of its cost advantage in its technical report on V4.1 Flash. Although the model has a total of 552 billion parameters (i.e. internal settings that the model adjusts to store knowledge during training), only 8 billion of them are activated when processing input prompts, and only 16 billion are activated when generating responses. DeepSeek calls this architecture a "causal coder-decoder" design, which is a key technology for its ability to achieve extremely fast completion times.
This is not the first time DeepSeek has used low-cost strategies to close the capability gap. Previously, the company's V4 Pro model performed close to Claude Fable 5 in another benchmark (within 5%), at a fraction of the latter's cost. In addition, DeepSeek is recruiting engineers in Beijing to build its own Code Harness, aiming to master the complete Agent Stack, rather than just serving as an underlying model provider.
Analysis of evaluation limitations and delivery rate
It should be noted that OpenDesign's test setup limits the scope of proof of these data. A score will be awarded only if the results output by the model can be rendered into a working web page; any blank, damaged, or truncated output will be counted as zero and will not be retested. Therefore, this benchmark measures reliable and day-to-day design output capabilities rather than general reasoning or coding skills.
In terms of delivery rate (i.e., the proportion of usage that can be delivered directly without modification):
- GPT-6 Astra: Delivery rate is 60%.
- DeepSeek V4.1 Flash: Delivery rate is 57.7%.
- Claude Fable 5.1: Delivery rate is 56.7%.
Although the GPT-6 Astra, released by OpenAI on September 3, is renowned for its wide range of capabilities, such as circuit board layout, tax draft drafting, and 3D scene construction, early testers pointed out that it was weaker than its predecessor in terms of writing capabilities. In OpenDesign's chart, the GPT-6 Astra reflects the typical "all-around" characteristics: compared to DeepSeek's more cost-effective entry-level option, it is slower and more expensive, but it still has the highest overall score of the model.

Exchange Ranking
Top Exchanges
24h Volume Ranking
Popularity Ranking
Exchange BTC Balance
Proof of Reserves
Decentralized Exchanges
Funding Rate
Funding Heatmap
Liquidation Data
Max Pain
Long/Short Ratio
Whale L/S Ratio
Binance/Okex/Huobi L/S
Bitfinex Margin L/S
ETF Tracker
Solana ETF
XRP ETF
Hong Kong ETF
Bitcoin Treasuries
Crypto Reversal
Ethereum Reserves
HyperLiquid Wallet Analysis
Hyperliquid Whale Watch
Large Transactions
On-chain Movement
Bitcoin ROI
Stablecoin Market Cap
Options Analysis
News
Articles
Economic Calendar
Features
Wallet
Contract Calculator
Security
Collections
Watchlist
Following