DeepSeek targets AGI in $10bn funding round

DeepSeek's $10bn push for AGI could redefine the open-source landscape for developers.
30-Second TL;DR
What Changed
DeepSeek is pursuing AGI as its primary long-term objective.
Why It Matters
This funding round signals a major shift in the competitive landscape for open-source AI, potentially challenging the dominance of closed-source labs.
What To Do Next
Monitor DeepSeek's GitHub repository for new model architecture releases that could optimize your own training pipelines.
Key Points
- •DeepSeek is pursuing AGI as its primary long-term objective.
- •The lab is prioritizing frontier research over revenue generation.
- •DeepSeek will maintain its commitment to releasing open-source models.
Deep Insight
Background and context from public sources — not the original article. 26 sources cited.
Enhanced Key Takeaways
- •DeepSeek is currently in late-stage discussions to raise 70 billion yuan (approximately $10 billion) in its first major external funding round, which could value the company at around $45 billion.
- •The company was founded in 2023 by Liang Wenfeng, who previously established the quantitative hedge fund High-Flyer, which initially provided DeepSeek's funding.
- •DeepSeek's models, such as DeepSeek-V3 and DeepSeek-R1, are noted for achieving performance comparable to leading proprietary models like GPT-4o and Gemini-3.0-Pro, while reportedly incurring significantly lower training costs, with DeepSeek-V3 costing around $5.6 million.
- •DeepSeek's V4-Pro, a 1.6-trillion-parameter Mixture-of-Experts system, and V4-Flash models, released in April 2026, are optimized for compatibility with Huawei Ascend and Cambricon silicon, alongside Nvidia GPUs, indicating a strategic focus on the Chinese domestic market.
- •While DeepSeek promotes an open-source model strategy, some of its models are released under custom licenses that include usage limitations, distinguishing them from the conventional definition of open-source software.
Competitor Analysis
- DeepSeek
- Open-source (with custom licenses) & Proprietary API; MoE architecture
- OpenAI (ChatGPT)
- Proprietary; GPT-series (e.g., GPT-4, GPT-4o)
- Google (Gemini)
- Proprietary; Multimodal (e.g., Gemini Ultra, Pro, Nano)
- Anthropic (Claude)
- Proprietary; Focus on ethical AI (e.g., Claude 3)
- Mistral AI
- Open-source; Diverse model offerings (e.g., Mistral Large 2)
- DeepSeek
- Cost-effective, high-performance reasoning, coding, math, efficient training, support for Chinese NLP, agentic capabilities
- OpenAI (ChatGPT)
- Versatile, general-purpose conversation, creative content, robust code assistance, multimodal tasks
- Google (Gemini)
- Advanced multimodal capabilities (text, image, audio, video), integrated with Google ecosystem, strong for productivity and complex data analysis
- Anthropic (Claude)
- Excels in reasoning, long-form content, ethical safeguards, long-context handling, PhD-level science reasoning, agentic tasks
- Mistral AI
- Customizable, efficient, high-performance open-source models, privacy-focused
- DeepSeek
- DeepSeek-V3/R1 comparable to GPT-4o and Gemini-3.0-Pro; DeepSeek-V3-0324 outperforms GPT-4.5 in math/coding
- OpenAI (ChatGPT)
- GPT-4 and GPT-4o are leading general-purpose AI tools
- Google (Gemini)
- Gemini Ultra surpasses human experts on MMLU benchmark; Gemini 1.5 Flash offers extended context
- Anthropic (Claude)
- Claude 3 leads in coding, PhD-level science reasoning, and agentic task completion
- Mistral AI
- Mistral Large 2 is a top-tier reasoning model
- DeepSeek
- Free open-source models for self-hosting; competitive API pricing ($0.27-$0.55/M input tokens, $1.10-$2.19/M output tokens)
- OpenAI (ChatGPT)
- Paid tiers (e.g., ChatGPT Plus ~$20/month); API access
- Google (Gemini)
- Paid tiers (e.g., Gemini Advanced ~$20/month); API access
- Anthropic (Claude)
- Paid tiers (e.g., Claude Pro ~$20/month); API access
- Mistral AI
- Mostly free open-source models; hosting costs may apply
- DeepSeek
- DeepSeek-R1/V3: 128K tokens
- OpenAI (ChatGPT)
- Varies by model (e.g., GPT-4 Turbo has 128K)
- Google (Gemini)
- Gemini 1.5 Flash offers extended context windows
- Anthropic (Claude)
- Hundreds of pages of text
- Mistral AI
- Mistral Large 2: 128K tokens
| Feature/Model | DeepSeek | OpenAI (ChatGPT) | Google (Gemini) | Anthropic (Claude) | Mistral AI |
|---|---|---|---|---|---|
| Model Type | Open-source (with custom licenses) & Proprietary API; MoE architecture | Proprietary; GPT-series (e.g., GPT-4, GPT-4o) | Proprietary; Multimodal (e.g., Gemini Ultra, Pro, Nano) | Proprietary; Focus on ethical AI (e.g., Claude 3) | Open-source; Diverse model offerings (e.g., Mistral Large 2) |
| Key Strengths | Cost-effective, high-performance reasoning, coding, math, efficient training, support for Chinese NLP, agentic capabilities | Versatile, general-purpose conversation, creative content, robust code assistance, multimodal tasks | Advanced multimodal capabilities (text, image, audio, video), integrated with Google ecosystem, strong for productivity and complex data analysis | Excels in reasoning, long-form content, ethical safeguards, long-context handling, PhD-level science reasoning, agentic tasks | Customizable, efficient, high-performance open-source models, privacy-focused |
| Benchmarks/Performance | DeepSeek-V3/R1 comparable to GPT-4o and Gemini-3.0-Pro; DeepSeek-V3-0324 outperforms GPT-4.5 in math/coding | GPT-4 and GPT-4o are leading general-purpose AI tools | Gemini Ultra surpasses human experts on MMLU benchmark; Gemini 1.5 Flash offers extended context | Claude 3 leads in coding, PhD-level science reasoning, and agentic task completion | Mistral Large 2 is a top-tier reasoning model |
| Pricing/Accessibility | Free open-source models for self-hosting; competitive API pricing ($0.27-$0.55/M input tokens, $1.10-$2.19/M output tokens) | Paid tiers (e.g., ChatGPT Plus ~$20/month); API access | Paid tiers (e.g., Gemini Advanced ~$20/month); API access | Paid tiers (e.g., Claude Pro ~$20/month); API access | Mostly free open-source models; hosting costs may apply |
| Context Window | DeepSeek-R1/V3: 128K tokens | Varies by model (e.g., GPT-4 Turbo has 128K) | Gemini 1.5 Flash offers extended context windows | Hundreds of pages of text | Mistral Large 2: 128K tokens |
Technical Deep Dive
- Architecture Foundation: DeepSeek models are large-scale language models based on deep neural networks, primarily utilizing a decoder-only Transformer architecture.
- Mixture-of-Experts (MoE): DeepSeek-V3 and DeepSeek-R1 employ a Mixture-of-Experts framework, allowing dynamic activation of relevant sub-networks ("experts") for a given task, enhancing efficiency. DeepSeek-V3 has 671 billion total parameters, with only 37 billion activated per token during inference.
- Multi-head Latent Attention (MLA): This mechanism optimizes the traditional key-value cache (KV cache) by compressing attention matrices into smaller latent representations, significantly reducing memory usage and improving inference latency.
- Multi-Token Prediction (MTP): DeepSeek-V3 incorporates an MTP objective, enabling the model to predict multiple tokens at once, which densifies training signals and improves performance on complex benchmarks.
- Context Length: DeepSeek-R1 inherits a 128K context length from DeepSeek-V3-Base, achieved through a two-stage extension process utilizing the YaRN (Yet another RoPE extensioN method) technique.
- Training Efficiency: DeepSeek-V3 utilizes FP8 mixed precision training and an auxiliary-loss-free strategy for load balancing, contributing to its economical training costs (reportedly around $5.6 million for full training).
- DeepSeek-V3.2 Innovations: This version introduced DeepSeek Sparse Attention (DSA) for scalable long-context reasoning, large-scale reinforcement learning for specialist models, and a "specialize-first" approach using distilled data.
- DeepSeek-V4: The V4-Pro model is a 1.6-trillion-parameter Mixture-of-Experts system, with the V4 family optimized for Huawei Ascend and Cambricon silicon in addition to Nvidia GPUs.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2015-06Liang Wenfeng co-founds High-Flyer Quantitative Investment Management.
- 2023-05Liang Wenfeng announces High-Flyer's pursuit of AGI and launches DeepSeek.
- 2023-07DeepSeek is officially founded as an AI company.
- 2024-12DeepSeek unveils its V3 model, a 671 billion parameter MoE architecture.
- 2025-01DeepSeek releases the DeepSeek-R1 reasoning model.
- 2026-04DeepSeek releases its V4-Pro (1.6T parameters) and V4-Flash (284B parameters) models.
Sources (26)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: The Next Web (TNW) ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.

