Fractile raises $220m for in-memory-compute inference chip production

๐กNew hardware architecture aiming to solve the memory bottleneck for LLM inference with major industry backing.
โก 30-Second TL;DR
What Changed
Secured $220 million in funding led by Accel.
Why It Matters
This funding signals a shift toward specialized hardware architectures that bypass the memory wall, potentially offering significant latency improvements for large-scale LLM inference.
What To Do Next
Monitor Fractile's public benchmarks against H100s to evaluate if their in-memory architecture fits your specific inference workload requirements.
Key Points
- โขSecured $220 million in funding led by Accel.
- โขPat Gelsinger joined as an angel investor.
- โขArchitecture integrates compute and memory on a single die.
- โขAnthropic is reportedly in early discussions to become a customer.
๐ง Deep Insight
Web-grounded analysis with 12 cited sources.
๐ Enhanced Key Takeaways
- โขFractile was founded in 2022 by Walter Goodwin, an Oxford PhD in robotics, who developed the concept while researching large language models (LLMs) for general-purpose robots.
- โขThe company projects its chips to deliver AI inference performance that is 100 times faster, 10 times cheaper, and 20 times more energy-efficient than current Nvidia GPUs, specifically for LLMs like Llama2-70B.
- โขFractile's in-memory compute architecture utilizes SRAM (Static Random-Access Memory) to integrate memory and compute directly on the same die, thereby mitigating the data transfer bottleneck prevalent in conventional GPU-DRAM systems.
- โขThe startup emerged from stealth in July 2024 with an initial $15 million seed funding round and has committed ยฃ100 million to expand its UK operations, including establishing a new hardware engineering facility in Bristol.
- โขThe recent $220 million funding round, co-led by Factorial Funds, Accel, and Peter Thiel's Founders Fund, reportedly values Fractile at over $1 billion.
๐ ๏ธ Technical Deep Dive
- Fractile's core technology is an in-memory compute architecture designed for AI inference, particularly for large language models (LLMs).
- The architecture integrates compute and memory directly on the same silicon die.
- It employs SRAM (Static Random-Access Memory) to co-locate memory and processing units, aiming to eliminate the 'memory wall' or 'data-shuttling bottleneck' associated with moving data between traditional GPUs and off-chip DRAM.
- This approach is projected to achieve a 100-fold increase in effective bandwidth and significantly higher energy efficiency.
- Fractile claims its accelerators can run LLMs like Llama2-70B 100 times faster, at one-tenth the system cost, and 20 times more energy-efficiently than Nvidia H100 GPUs for decode tokens per second.
- The company is developing custom multiply-accumulate (MAC) circuits that also store state.
- There is speculation that Fractile may optimize MAC arrays for general matrix-vector (GEMV) operations rather than general matrix-matrix (GEMM) operations to enhance efficiency.
- The Fractile team includes experienced engineers from companies such as Graphcore, Nvidia, and Imagination Technologies.
- Fractile is developing its own software stack in conjunction with its hardware.
- Commercial readiness for its chips is anticipated around 2027.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (12)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: The Next Web (TNW) โ