Uncensored Nemotron-3-Super-120B for MLX
๐กUncensored 120B model hits 94% HumanEval โ top open-weight for local coding.
โก 30-Second TL;DR
What Changed
Uncensored 4-bit MLX model with refusal vector removal
Why It Matters
Enables high-performance uncensored inference on Apple silicon via MLX, potentially rivaling closed models in coding and safety benchmarks for local deployments.
What To Do Next
Download from Hugging Face and run with provided custom.py in MLX Studio.
Key Points
- โขUncensored 4-bit MLX model with refusal vector removal
- โขCustom .py and chat template included for MLX Studio
- โขBenchmarks: HarmBench 97%, HumanEval 94%
- โขNo FP16 ablation due to unique attention mechanism
- โขq6 and q8 quantizations releasing tomorrow
๐ง Deep Insight
Background and context from public sources โ not the original article. 6 sources cited.
๐ Enhanced Key Takeaways
- โขNemotron 3 Super was officially released by NVIDIA with fully public weights, training data, recipes, and checkpoints available for download.
- โขThe model was pre-trained on over 10 trillion tokens and supports up to 1M token context length across English, French, German, Italian, Japanese, Spanish, and Chinese.
- โขIt achieves 85.6% on PinchBench, the highest score among open models for agentic performance in OpenClaw environments.
- โขOn NVIDIA Blackwell GPUs like B200, it delivers up to 5x higher throughput than prior Nemotron models using NVFP4 precision.
๐ Competitor Analysisโธ Show
| Feature/Benchmark | Nemotron 3 Super-120B | GPT-OSS-120B | Qwen3.5-122B A10B |
|---|---|---|---|
| Intelligence Index | 36 | 33 | 42 |
| Throughput (vs. baseline) | Up to 5x higher | Baseline | 7.5x lower than Nemotron |
| PinchBench Score | 85.6% | Not specified | Not specified |
๐ ๏ธ Technical Deep Dive
- โขHybrid architecture: LatentMoE with interleaved Mamba-2 layers (linear-time sequence efficiency), MoE layers (top-22 routing, 120.6B total / 12.7B active parameters), and select Transformer Attention layers for precision reasoning.
- โขMulti-Token Prediction (MTP): Predicts multiple future tokens per forward pass, enabling speculative decoding and up to 7.5x higher throughput on long outputs (e.g., 64k tokens) vs. Qwen3.5-122B.
- โขNVFP4 pre-training: First in Nemotron 3 family; uses NVFP4 for most linear layers (weights, activations, gradients) with BF16/MXFP8 for stability in key layers like QKV/attention and embeddings.
- โขPost-training: Multi-environment RL across 21 configurations with 1.2M rollouts using NeMo Gym/RL; optimized for Blackwell GPUs (4x faster inference vs. FP8 on H100).
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (6)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- aws.amazon.com โ Prodview H2ctgt5ws72e4
- alphasignalai.substack.com โ Nvidia Releases Nemotron 3 Super
- developer.nvidia.com โ Introducing Nemotron 3 Super an Open Hybrid Mamba Transformer Moe for Agentic Reasoning
- Hugging Face โ Nvidia Nemotron 3 Super 120b A12b Nvfp4
- research.nvidia.com โ Nvidia Nemotron 3 Super Technical Report
- youtube.com โ Watch
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.
