SourceStalecollected in 33m

DeepSeek Targets AGI in Massive $10 Billion Funding Round

Read original on Bloomberg Technology
#agi#funding#ai-strategy

DeepSeek's $10B pivot to AGI signals a major shift in the competitive landscape for foundational AI models.

30-Second TL;DR

What Changed

DeepSeek is seeking 70 billion yuan ($10 billion) in new capital.

Why It Matters

This strategic pivot suggests that DeepSeek intends to compete directly with top-tier labs like OpenAI and Anthropic by focusing on foundational model capabilities rather than immediate productization.

What To Do Next

Monitor DeepSeek's GitHub and research papers closely for new architectural breakthroughs that may challenge current SOTA models.

Who should care:Founders & Product Leaders

Key Points

  • DeepSeek is seeking 70 billion yuan ($10 billion) in new capital.
  • The company is shifting focus toward AGI research rather than short-term revenue.
  • Management is actively communicating this long-term vision to potential investors.
Key numbers$10 billion$300 million$45

Deep Insight

Background and context from public sources — not the original article. 19 sources cited.

Enhanced Key Takeaways

  • DeepSeek was founded in July 2023 by Liang Wenfeng, co-founder of the quantitative hedge fund High-Flyer Capital Management, which initially funded DeepSeek's operations.
  • Prior to this funding round, DeepSeek had largely been self-funded by High-Flyer's balance sheet for two years, having previously turned down external investments from leading Chinese venture capital firms and major tech companies.
  • The current funding round, initially targeting a $10 billion valuation for a $300 million raise, has seen its valuation rapidly escalate to a reported $45-50 billion, with China's state-backed semiconductor investment vehicle, the 'Big Fund,' potentially leading the investment.
  • DeepSeek's models, such as DeepSeek-R1 and DeepSeek V4, are recognized for their cost-efficiency in training and inference, often achieving performance comparable to leading models like OpenAI's GPT-4/o1 and Anthropic's Claude Opus at a significantly lower cost.
  • DeepSeek's commitment to open-weight models, released under the MIT License, distinguishes its approach in the competitive AI landscape, fostering transparency and community-driven innovation.

Competitor Analysis

Architecture
DeepSeek V4 Pro
Mixture-of-Experts (MoE), Multi-head Latent Attention (MLA), FP8 training
OpenAI (e.g., GPT-5.5)
Closed-source, likely Transformer-based, potentially MoE
Anthropic (e.g., Claude Opus 4.7)
Closed-source, likely Transformer-based
Openness
DeepSeek V4 Pro
Open-weight (MIT License)
OpenAI (e.g., GPT-5.5)
Closed-source
Anthropic (e.g., Claude Opus 4.7)
Closed-source
Context Window
DeepSeek V4 Pro
1 Million tokens
OpenAI (e.g., GPT-5.5)
124,000 tokens (ChatGPT-o1 Mini)
Anthropic (e.g., Claude Opus 4.7)
Generally long, but specific for Opus 4.7 not detailed
Pricing (Input/1M tokens)
DeepSeek V4 Pro
$0.435 (V4 Pro)
OpenAI (e.g., GPT-5.5)
$0.75 (GPT-5.4 mini)
Anthropic (e.g., Claude Opus 4.7)
Higher than DeepSeek V4 Pro
Pricing (Output/1M tokens)
DeepSeek V4 Pro
$0.87 (V4 Pro)
OpenAI (e.g., GPT-5.5)
$4.50 (GPT-5.4 mini)
Anthropic (e.g., Claude Opus 4.7)
Higher than DeepSeek V4 Pro
Key Benchmarks (May 2026 CAISI Evaluation vs. GPT-5.5)
DeepSeek V4 Pro
Mathematics: 96% (tied)
OpenAI (e.g., GPT-5.5)
Mathematics: 96%
Anthropic (e.g., Claude Opus 4.7)
N/A (not directly compared in this report)
DeepSeek V4 Pro
Cybersecurity: 32% (vs. 71%)
OpenAI (e.g., GPT-5.5)
Cybersecurity: 71%
Anthropic (e.g., Claude Opus 4.7)
N/A
DeepSeek V4 Pro
Software Engineering: 74% (vs. 81%)
OpenAI (e.g., GPT-5.5)
Software Engineering: 81%
Anthropic (e.g., Claude Opus 4.7)
N/A
Overall Performance (CAISI Elo Rating)
DeepSeek V4 Pro
800 (comparable to GPT-5.4 mini)
OpenAI (e.g., GPT-5.5)
1260 (GPT-5.5), 999 (Opus 4.6)
Anthropic (e.g., Claude Opus 4.7)
999 (Opus 4.6)
Security (Jailbreaking Susceptibility)
DeepSeek V4 Pro
94% compliance (R1-0528)
OpenAI (e.g., GPT-5.5)
8% compliance (US reference models)
Anthropic (e.g., Claude Opus 4.7)
8% compliance (US reference models)

Technical Deep Dive

  • DeepSeek models (e.g., V3, R1, V4) primarily utilize a Mixture-of-Experts (MoE) architecture for economical training and efficient inference.
  • DeepSeek V4 Pro features 1.6 trillion total parameters with 49 billion activated parameters, supporting a 1-million-token context window.
  • DeepSeek V3 has 671 billion total parameters with 37 billion activated for each token, pre-trained on 14.8 trillion tokens.
  • The architecture incorporates Multi-head Latent Attention (MLA) for efficient inference and DeepSeekMoE for cost-effective training.
  • DeepSeek has pioneered the use of FP8 precision for LLM pre-training, which doubles compute efficiency and halves memory usage compared to BF16.
  • DeepSeek-LLM models use an auto-regressive transformer decoder architecture, similar to LLaMA, incorporating SwiGLU, RoPE, and RMSNorm.
  • DeepSeek-R1 was notably trained using reinforcement learning, focusing on structured reasoning and the emergence of 'thinking brackets' for problem-solving.

Future ImplicationsAI analysis grounded in cited sources

DeepSeek's shift to external funding and explicit AGI focus will intensify the global AI race.
The substantial capital infusion and clear AGI goal position DeepSeek as a major contender, compelling other frontier AI labs to accelerate their own research and development efforts.
DeepSeek's open-weight model strategy will continue to influence the open-source AI community.
By releasing powerful models under permissive licenses and demonstrating cost-efficient training, DeepSeek encourages broader participation and innovation in AI development, potentially democratizing access to advanced AI capabilities.
China's state-backed investment in DeepSeek signals a strategic national priority for AI model capability.
The involvement of the 'Big Fund' indicates a shift in Beijing's strategy to prioritize AI model development as a response to semiconductor export controls, aiming for self-sufficiency in advanced AI.

Timeline

2023-07
DeepSeek founded by Liang Wenfeng, initially funded by High-Flyer Capital Management.
2023-11
DeepSeek released its first model, DeepSeek Coder, followed by the DeepSeek-LLM series.
2025-01
DeepSeek launched its DeepSeek-R1 model and an eponymous chatbot, noted for cost-efficiency and performance.
2025-08
DeepSeek V3.1 released, featuring a hybrid architecture and surpassing prior models on certain benchmarks.
2026-04
DeepSeek began seeking its first external funding round, initially targeting $300 million at a $10 billion valuation.
2026-04-24
DeepSeek released a preview of its V4 series (V4-Pro and V4-Flash), featuring 1.6 trillion total parameters and a 1-million-token context window.

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Bloomberg Technology

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.