Llama 8B Matches 70B on Multi-Hop QA
Llama 3.1 8B with structured chain-of-thought and graph context compression matches or exceeds Llama 3.3 70B on multi-hop QA benchmarks without fine-tuning. Retrieval succeeds 77-91% of the time, but reasoning is the bottleneck. Tested on HotpotQA, MuSiQue, and 2WikiMultiHopQA with 12x cost savings on Groq.

