DeepSeek has released the weights for its multimodal model DeepSeek-V4-Flash-Vision-Exp, allowing developers to download and deploy the model locally for the first time. The model is built on the V4-Flash architecture and adds image understanding capabilities, including recognition of screenshots and charts, while supporting Agent tasks when paired with external tools. The model’s performance on pure-text Agent benchmarks such as Terminal Bench is reported to be roughly in line with V4-Flash. The weights can now be integrated through frameworks including Transformers, vLLM, and SGLang, expanding access beyond the model’s previous API-only availability.