🗾Stalecollected in 38m

Meta Unveils 4 New MTIA AI Chips

Meta Unveils 4 New MTIA AI Chips
PostLinkedIn
🗾Read original on ITmedia AI+ (日本)
#ai-chips#data-centers#custom-siliconmtiametamtianvidia

💡Meta's custom AI chips + 5GW data center plan challenge NVIDIA's infra dominance

⚡ 30-Second TL;DR

What Changed

Four new MTIA models specialized for AI inference

Why It Matters

Meta's in-house chips reduce dependency on third-party hardware like NVIDIA, potentially lowering AI training costs industry-wide. This accelerates custom AI infrastructure scaling for hyperscalers.

What To Do Next

Benchmark MTIA inference performance against NVIDIA GPUs for your custom AI workloads.

Who should care:Enterprise & Security Teams

Key Points

  • Four new MTIA models specialized for AI inference
  • Significant reduction in development cycles
  • Massive NVIDIA GPU cluster deployments underway
  • Future 5GW 'Hyperion' data center plan announced

🧠 Deep Insight

Background and context from public sources — not the original article. 9 sources cited.

🔑 Enhanced Key Takeaways

  • MTIA v1, the first-generation chip designed in 2020, uses TSMC 7nm process, operates at 800 MHz with 102.4 TOPS at INT8 and 25W TDP, featuring 64 PEs in an 8x8 grid and 128 MB on-chip SRAM.[1][3]
  • Meta announced MTIA v2i, the second-generation inference chip, at ISCA'23, which is now deployed at scale for production workloads.[9]
  • In April 2024, Meta unveiled MTIA v2 supporting generative AI inference for recommendation systems on Facebook and Instagram, with plans to ramp adoption throughout 2025.[5]

🛠️ Technical Deep Dive

  • MTIA v1 architecture includes 64 Processing Elements (PEs) in an 8x8 grid connected via mesh network, each with 128 KB local SRAM, supporting thread/data/instruction/memory-level parallelism.[1]
  • Memory subsystem: 128 MB on-chip SRAM shared among PEs, scalable to 128 GB off-chip LPDDR5 DRAM for high bandwidth and low latency.[1]
  • Fabricated on TSMC 7nm, clocked at 800 MHz, delivers 102.4 TOPS (INT8) / 51.2 TFLOPS (FP16), TDP 25W, includes RISC-V cores.[1][6]

🔮 Future ImplicationsAI analysis grounded in cited sources

Meta will reduce NVIDIA GPU dependency by 2027 through expanded MTIA deployments.
MTIA v2 ramp-up in 2025 and training chip extensions aim to cut costs on external GPUs after billions spent on NVIDIA hardware.[5][2]
Custom ASICs like MTIA will achieve 2x efficiency over GPUs for Meta's recommendation workloads.
MTIA v1 benchmarks show twice the efficiency of GPUs in Meta's tests for low/medium-complexity inference, with software optimizations planned.[1][4]

Timeline

2020-01
Designed first-generation MTIA v1 ASIC for recommendation inference workloads.
2023-05
Announced MTIA family and detailed v1 specs at AI Infra @ Scale event.
2023-06
Presented MTIA v1 paper at ISCA conference.
2023-07
Presented MTIA v2i, second-generation inference chip, at ISCA'23.
2024-04
Announced MTIA v2 for generative AI inference and recommendation systems.
2026-03
Unveiled four new MTIA models optimized for inference with shortened development cycles.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: ITmedia AI+ (日本)

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.