Google Launches Gemini 3.1 Flash-Lite at 1/8th Pro Cost

๐กCheapest Gemini yet: 2.5x faster TTFT, strong benchmarks at 1/8 Pro cost.
โก 30-Second TL;DR
What Changed
1/8th cost of Gemini 3.1 Pro for enterprises
Why It Matters
This enables scalable AI deployment for real-time apps like support and moderation at fraction of Pro cost, boosting enterprise adoption. Competitive benchmarks position it against larger models, pressuring rivals on efficiency.
What To Do Next
Test Gemini 3.1 Flash-Lite in Google AI Studio for low-latency inference benchmarks.
Key Points
- โข1/8th cost of Gemini 3.1 Pro for enterprises
- โข2.5X faster time to first token vs Gemini 2.5 Flash
- โข363 tokens/sec output speed, up 45%
- โขElo score 1432 on Arena.ai leaderboard
- โขAdjustable thinking levels for reasoning control
๐ง Deep Insight
Background and context from public sources โ not the original article. 7 sources cited.
๐ Enhanced Key Takeaways
- โขGemini 3.1 Flash-Lite supports multimodal inputs including text, up to 3,000 images (7MB inline), 10 videos (up to 1 hour), and 8.4 hours of audio, with a 1M input token context window.[1][2][4]
- โขPricing is $0.25 per 1M input tokens and $1.50 per 1M output tokens, optimized for high-volume agentic tasks and low-latency applications.[3][4]
- โขAchieves 60.1% on MRCR v2 (8-needle) long context benchmark at 128k tokens and 21.0% at 1M tokens, outperforming several peers in needle retrieval.[2]
- โขFeatures include function calling, structured outputs, thinking, caching, batch API, code execution, and search grounding, but lacks audio/image generation and computer use.[4]
๐ ๏ธ Technical Deep Dive
- โขMaximum input tokens: 1,048,576; maximum output tokens: 65,535 (or 65,536 in preview).[1][2][4]
- โขNatively multimodal: supports text, images (up to 3,000 per prompt, 7MB inline/30MB GCS), documents (3,000 files, 1,000 pages/file, 50MB), video (10 files, ~1hr no audio), audio (8.4hrs or 1M tokens).[1][2][4]
- โขArchitecture based on Gemini 3.1 Pro, inheriting its limitations and acceptable usage policies; suited for high-volume, low-latency tasks but less capable than Pro.[2]
- โขSupported MIME types include image/png/jpeg/webp/heic/heif, video/mp4/webm/etc., audio/mp3/wav/etc., and documents pdf/text/plain.[1]
- โขCapabilities: function calling, structured outputs, thinking levels, file search, URL context, search grounding; no image/audio generation, live API, or computer use.[4]
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (7)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- docs.cloud.google.com โ 3 1 Flash Lite
- Google DeepMind โ Gemini 3 1 Flash Lite
- androidcentral.com โ Gemini 3 1 Flash Lite Is the Fast Help You Need If Youre a Dev with Complex Data
- ai.google.dev โ Gemini 3.1 Flash Lite Preview
- Google DeepMind โ Gemini 3 1 Flash Lite
- Google Blog โ Gemini 3 1 Pro
- openrouter.ai โ Gemini 3.1 Flash Lite Preview
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: VentureBeat โ
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.


