You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
DeepSeek-V4 is DeepSeek's fourth-generation flagship model, released as open weights under the MIT license on Hugging Face. It includes V4-Pro (1.6T MoE, 49B activated) and V4-Flash (284B MoE, 13B activated), both supporting a 1-million-token context and native multimodality.
Core Capabilities
Mixture-of-Experts architecture: 1.6T total parameters with sparse activation for efficient inference.
Long context: 1M-token context window for long-document and repository-scale tasks.
Native multimodality: handles text, image, and video understanding and generation.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
DeepSeek-V4: Detailed Introduction
Table of Contents
What is DeepSeek-V4?
DeepSeek-V4 is DeepSeek's fourth-generation flagship model, released as open weights under the MIT license on Hugging Face. It includes V4-Pro (1.6T MoE, 49B activated) and V4-Flash (284B MoE, 13B activated), both supporting a 1-million-token context and native multimodality.
Core Capabilities
Access
Pricing
The model weights are free under MIT. API access is usage-based; V4-Flash offers a lower-cost option while V4-Pro delivers top-tier performance.
All reactions