DeepSeek's New Model Nearly Matches GPT-6 Astra on Design—at 1.4% of the Cost

Summary

OpenDesign Arena benchmarked 13 AI models on practical design tasks like web apps, dashboards, mobile screens, and landing pages, scoring both brief adherence and design quality. GPT-6 Astra ranked first with 82.7 points, but DeepSeek V4.1 Flash was nearly as strong at 81.2, while costing far less: about $0.023 per design versus $1.61 for GPT-6 Astra and $3.66 for Claude Fable 5.1. DeepSeek also finished faster at 5.3 minutes, compared with 11.1 and 12.8 minutes. Eleven of the 13 tested models scored below DeepSeek and cost more. DeepSeek’s efficiency comes from a sparse Causal Encoder-Decoder setup, activating only a small fraction of its 552B parameters per prompt. The benchmark measures reliable, renderable design output, not broad reasoning. DeepSeek’s delivery rate was 57.7%, close to GPT-6 Astra’s 60%.