DeepSeek targets AGI in $10bn funding round

๐กDeepSeek's $10bn push for AGI could redefine the open-source landscape for developers.
โก 30-Second TL;DR
What Changed
DeepSeek is pursuing AGI as its primary long-term objective.
Why It Matters
This funding round signals a major shift in the competitive landscape for open-source AI, potentially challenging the dominance of closed-source labs.
What To Do Next
Monitor DeepSeek's GitHub repository for new model architecture releases that could optimize your own training pipelines.
Key Points
- โขDeepSeek is pursuing AGI as its primary long-term objective.
- โขThe lab is prioritizing frontier research over revenue generation.
- โขDeepSeek will maintain its commitment to releasing open-source models.
๐ง Deep Insight
Web-grounded analysis with 26 cited sources.
๐ Enhanced Key Takeaways
- โขDeepSeek is currently in late-stage discussions to raise 70 billion yuan (approximately $10 billion) in its first major external funding round, which could value the company at around $45 billion.
- โขThe company was founded in 2023 by Liang Wenfeng, who previously established the quantitative hedge fund High-Flyer, which initially provided DeepSeek's funding.
- โขDeepSeek's models, such as DeepSeek-V3 and DeepSeek-R1, are noted for achieving performance comparable to leading proprietary models like GPT-4o and Gemini-3.0-Pro, while reportedly incurring significantly lower training costs, with DeepSeek-V3 costing around $5.6 million.
- โขDeepSeek's V4-Pro, a 1.6-trillion-parameter Mixture-of-Experts system, and V4-Flash models, released in April 2026, are optimized for compatibility with Huawei Ascend and Cambricon silicon, alongside Nvidia GPUs, indicating a strategic focus on the Chinese domestic market.
- โขWhile DeepSeek promotes an open-source model strategy, some of its models are released under custom licenses that include usage limitations, distinguishing them from the conventional definition of open-source software.
๐ Competitor Analysisโธ Show
| Feature/Model | DeepSeek | OpenAI (ChatGPT) | Google (Gemini) | Anthropic (Claude) | Mistral AI |
|---|---|---|---|---|---|
| Model Type | Open-source (with custom licenses) & Proprietary API; MoE architecture | Proprietary; GPT-series (e.g., GPT-4, GPT-4o) | Proprietary; Multimodal (e.g., Gemini Ultra, Pro, Nano) | Proprietary; Focus on ethical AI (e.g., Claude 3) | Open-source; Diverse model offerings (e.g., Mistral Large 2) |
| Key Strengths | Cost-effective, high-performance reasoning, coding, math, efficient training, support for Chinese NLP, agentic capabilities | Versatile, general-purpose conversation, creative content, robust code assistance, multimodal tasks | Advanced multimodal capabilities (text, image, audio, video), integrated with Google ecosystem, strong for productivity and complex data analysis | Excels in reasoning, long-form content, ethical safeguards, long-context handling, PhD-level science reasoning, agentic tasks | Customizable, efficient, high-performance open-source models, privacy-focused |
| Benchmarks/Performance | DeepSeek-V3/R1 comparable to GPT-4o and Gemini-3.0-Pro; DeepSeek-V3-0324 outperforms GPT-4.5 in math/coding | GPT-4 and GPT-4o are leading general-purpose AI tools | Gemini Ultra surpasses human experts on MMLU benchmark; Gemini 1.5 Flash offers extended context | Claude 3 leads in coding, PhD-level science reasoning, and agentic task completion | Mistral Large 2 is a top-tier reasoning model |
| Pricing/Accessibility | Free open-source models for self-hosting; competitive API pricing ($0.27-$0.55/M input tokens, $1.10-$2.19/M output tokens) | Paid tiers (e.g., ChatGPT Plus ~$20/month); API access | Paid tiers (e.g., Gemini Advanced ~$20/month); API access | Paid tiers (e.g., Claude Pro ~$20/month); API access | Mostly free open-source models; hosting costs may apply |
| Context Window | DeepSeek-R1/V3: 128K tokens | Varies by model (e.g., GPT-4 Turbo has 128K) | Gemini 1.5 Flash offers extended context windows | Hundreds of pages of text | Mistral Large 2: 128K tokens |
๐ ๏ธ Technical Deep Dive
- Architecture Foundation: DeepSeek models are large-scale language models based on deep neural networks, primarily utilizing a decoder-only Transformer architecture.
- Mixture-of-Experts (MoE): DeepSeek-V3 and DeepSeek-R1 employ a Mixture-of-Experts framework, allowing dynamic activation of relevant sub-networks ("experts") for a given task, enhancing efficiency. DeepSeek-V3 has 671 billion total parameters, with only 37 billion activated per token during inference.
- Multi-head Latent Attention (MLA): This mechanism optimizes the traditional key-value cache (KV cache) by compressing attention matrices into smaller latent representations, significantly reducing memory usage and improving inference latency.
- Multi-Token Prediction (MTP): DeepSeek-V3 incorporates an MTP objective, enabling the model to predict multiple tokens at once, which densifies training signals and improves performance on complex benchmarks.
- Context Length: DeepSeek-R1 inherits a 128K context length from DeepSeek-V3-Base, achieved through a two-stage extension process utilizing the YaRN (Yet another RoPE extensioN method) technique.
- Training Efficiency: DeepSeek-V3 utilizes FP8 mixed precision training and an auxiliary-loss-free strategy for load balancing, contributing to its economical training costs (reportedly around $5.6 million for full training).
- DeepSeek-V3.2 Innovations: This version introduced DeepSeek Sparse Attention (DSA) for scalable long-context reasoning, large-scale reinforcement learning for specialist models, and a "specialize-first" approach using distilled data.
- DeepSeek-V4: The V4-Pro model is a 1.6-trillion-parameter Mixture-of-Experts system, with the V4 family optimized for Huawei Ascend and Cambricon silicon in addition to Nvidia GPUs.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (26)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- techinasia.com
- businesstimes.com.sg
- theedgesingapore.com
- asiabusinessoutlook.com
- investing.com
- wikipedia.org
- fintechweekly.com
- ourchinastory.com
- frederick.ai
- entrepreneur.com
- towardsai.net
- medium.com
- baseten.co
- digitalocean.com
- arxiv.org
- github.com
- thenextweb.com
- deepseek.com
- sgu.ac.id
- sintra.ai
- g2.com
- medium.com
- techjarvisai.com
- bentoml.com
- fireworks.ai
- geeksforgeeks.org
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates

PACA: Open-source tool for ancient fossil coordinate mapping

The Hidden Costs of Solo Sales Handoffs

Hotel groups launch ChatGPT apps for travel booking

Nvidia may guarantee $250bn for OpenAI data center
AI-curated news aggregator. All content rights belong to original publishers.
Original source: The Next Web (TNW) โ