DeepSeek has released DeepSeek-V4-Flash-0731 on Hugging Face, shipping the weights on the same day as the announcement rather than behind a countdown campaign. The model is distributed under a pure MIT license.

The release is notable for its efficiency: the model ships already quantized in FP4, with a total size of roughly 167GB. Community members report it performs on par with GLM-5.2 while requiring far less VRAM, and some benchmarks show it outperforming DeepSeek-V4-Pro. According to the model card, code-agent benchmarks were evaluated using the minimal mode of DeepSeek Harness, a new official agent framework the company says will be released soon, configured with maximum reasoning effort, temperature 1.0 and top_p 0.95.

The local-AI community has reacted strongly, with users praising the immediate open-weight release and discussing ways to run it on consumer hardware, including requests for GGUF builds and Q2–Q4 quantizations for setups like an M5 Max or a 64GB RAM machine paired with a 3090. Early testers describe the model's quality as very strong, with UI-generation taste near Opus level.