SourceStalecollected in 4h

EXAONE 4.5 33B Models Now on Hugging Face

Read original on Reddit r/LocalLLaMA
#model-release#quantization#open-weight

New 33B open model in GGUF/FP8—quantized for your local GPU setup

30-Second TL;DR

What Changed

EXAONE 4.5-33B base model released

Why It Matters

Provides open-weight 33B model option for practitioners seeking alternatives to Western LLMs, with quantization support boosting local hardware accessibility.

What To Do Next

Download EXAONE-4.5-33B-GGUF from Hugging Face and test inference with llama.cpp.

Who should care:Developers & AI Engineers

Key Points

  • •EXAONE 4.5-33B base model released
  • •FP8 quantized version available
  • •GGUF format for local inference
  • •Hosted on Hugging Face by LGAI-EXAONE

Deep Insight

AI-generated analysis for this event — not the original article.

Enhanced Key Takeaways

  • •EXAONE 4.5 is developed by LG AI Research, specifically designed to excel in both English and Korean bilingual capabilities, distinguishing it from general-purpose models.
  • •The 33B parameter size is strategically chosen to balance high-level reasoning performance with the ability to run on consumer-grade hardware, such as high-end NVIDIA RTX GPUs.
  • •The release emphasizes a 'multimodal' architecture, enabling the model to process and understand both text and visual inputs, a significant upgrade from previous text-only iterations.

Competitor Analysis

Primary Focus
EXAONE 4.5 33B
Bilingual (EN/KO) Multimodal
Llama 3.1 70B
General Purpose
Mistral Small 22B
Efficiency/Reasoning
Architecture
EXAONE 4.5 33B
Multimodal
Llama 3.1 70B
Text-only
Mistral Small 22B
Text-only
Hardware Req.
EXAONE 4.5 33B
Moderate (Consumer)
Llama 3.1 70B
High (Enterprise)
Mistral Small 22B
Low/Moderate
Quantization
EXAONE 4.5 33B
Native FP8/GGUF support
Llama 3.1 70B
Community-driven
Mistral Small 22B
Community-driven

Technical Deep Dive

  • Architecture: Multimodal transformer-based architecture capable of joint text-image processing.
  • Parameter Count: 33 Billion parameters, optimized for dense inference.
  • Quantization Support: Native support for FP8 (Floating Point 8) to reduce VRAM footprint without significant perplexity degradation.
  • Inference Optimization: GGUF format integration allows for seamless compatibility with llama.cpp and related local inference engines.
  • Context Window: Optimized for long-context retrieval tasks compared to the 4.0 series.

Future ImplicationsAI analysis grounded in cited sources

LG AI Research will likely integrate EXAONE 4.5 into their enterprise 'EXAONE Universe' platform.
The company has consistently used its open-weights releases as a foundation for its proprietary B2B service offerings.
The model will see rapid adoption in the Korean enterprise sector for localized RAG applications.
The combination of high-performance bilingual capabilities and local deployment options addresses critical data sovereignty concerns for Korean firms.

Timeline

2022-05
LG AI Research unveils the first iteration of the EXAONE model.
2023-07
Release of EXAONE 2.0, focusing on improved multimodal capabilities.
2024-08
LG AI Research releases EXAONE 3.0, expanding the model's reasoning and coding benchmarks.
2025-03
Introduction of EXAONE 4.0, featuring enhanced efficiency for enterprise deployment.
2026-04
Release of EXAONE 4.5 33B on Hugging Face with native FP8 and GGUF support.

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA ↗

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.