
Models · 1 September 2026
DeepSeek published full weights for DeepSeek-V4-Flash-Vision-Exp on Hugging Face on 31 August under an MIT license, ten days after the model appeared on the company’s API. Hugging Face lists the experimental multimodal system at about 305 billion parameters. It is the first vision model in the V4-Flash family.
The architecture adds a 32-layer vision encoder and aligner on top of the V4-Flash text backbone: a 43-layer mixture-of-experts language model with 256 routed experts (six selected per token), DFlash attention, Hyper-Connections, and the DSpark forward path. The model card specifies a 1-million-token context window and up to 384,000 output tokens, with text-plus-image input and text output.
On DeepSeek’s published table, Vision-Exp lifts ApexBench Pass@1 to 36.5 from 26.2 for text-only V4-Flash-0731, close to the 39.4 listed for Anthropic’s Opus 4.8. Agents’ Last Exam is 27.3 versus Opus 4.8’s 25.7. Text-agent scores stay in the same band as the prior Flash model, including 83.9 on Terminal Bench 2.1.
The repository includes tokenizer files, an encoding reference, a minimal PyTorch inference path, and launch notes for vLLM and SGLang. Because the license is MIT, labs and cloud hosts can fine-tune, quantize, and serve the weights without a custom commercial agreement.
Key takeaway A 305-billion-parameter multimodal agent model is now downloadable under MIT, with published agent scores approaching a closed frontier system and a 1-million-token context window.
Photo: whale motif standing in for the DeepSeek mark and the 305B release, via Unsplash.
Sources: Hugging Face model card · AI Weekly alert
