EXAONE 4.5 33B Models Now on Hugging Face

New 33B open model in GGUF/FP8—quantized for your local GPU setup
30-Second TL;DR
What Changed
EXAONE 4.5-33B base model released
Why It Matters
Provides open-weight 33B model option for practitioners seeking alternatives to Western LLMs, with quantization support boosting local hardware accessibility.
What To Do Next
Download EXAONE-4.5-33B-GGUF from Hugging Face and test inference with llama.cpp.
Key Points
- •EXAONE 4.5-33B base model released
- •FP8 quantized version available
- •GGUF format for local inference
- •Hosted on Hugging Face by LGAI-EXAONE
Deep Insight
AI-generated analysis for this event — not the original article.
Enhanced Key Takeaways
- •EXAONE 4.5 is developed by LG AI Research, specifically designed to excel in both English and Korean bilingual capabilities, distinguishing it from general-purpose models.
- •The 33B parameter size is strategically chosen to balance high-level reasoning performance with the ability to run on consumer-grade hardware, such as high-end NVIDIA RTX GPUs.
- •The release emphasizes a 'multimodal' architecture, enabling the model to process and understand both text and visual inputs, a significant upgrade from previous text-only iterations.
Competitor Analysis
- EXAONE 4.5 33B
- Bilingual (EN/KO) Multimodal
- Llama 3.1 70B
- General Purpose
- Mistral Small 22B
- Efficiency/Reasoning
- EXAONE 4.5 33B
- Multimodal
- Llama 3.1 70B
- Text-only
- Mistral Small 22B
- Text-only
- EXAONE 4.5 33B
- Moderate (Consumer)
- Llama 3.1 70B
- High (Enterprise)
- Mistral Small 22B
- Low/Moderate
- EXAONE 4.5 33B
- Native FP8/GGUF support
- Llama 3.1 70B
- Community-driven
- Mistral Small 22B
- Community-driven
| Feature | EXAONE 4.5 33B | Llama 3.1 70B | Mistral Small 22B |
|---|---|---|---|
| Primary Focus | Bilingual (EN/KO) Multimodal | General Purpose | Efficiency/Reasoning |
| Architecture | Multimodal | Text-only | Text-only |
| Hardware Req. | Moderate (Consumer) | High (Enterprise) | Low/Moderate |
| Quantization | Native FP8/GGUF support | Community-driven | Community-driven |
Technical Deep Dive
- Architecture: Multimodal transformer-based architecture capable of joint text-image processing.
- Parameter Count: 33 Billion parameters, optimized for dense inference.
- Quantization Support: Native support for FP8 (Floating Point 8) to reduce VRAM footprint without significant perplexity degradation.
- Inference Optimization: GGUF format integration allows for seamless compatibility with llama.cpp and related local inference engines.
- Context Window: Optimized for long-context retrieval tasks compared to the 4.0 series.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2022-05LG AI Research unveils the first iteration of the EXAONE model.
- 2023-07Release of EXAONE 2.0, focusing on improved multimodal capabilities.
- 2024-08LG AI Research releases EXAONE 3.0, expanding the model's reasoning and coding benchmarks.
- 2025-03Introduction of EXAONE 4.0, featuring enhanced efficiency for enterprise deployment.
- 2026-04Release of EXAONE 4.5 33B on Hugging Face with native FP8 and GGUF support.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.