
Karpathy Pushes Lord of the Rings Benchmark
Andrej Karpathy is advocating The Lord of the Rings as a new benchmark for evaluating large language models. The proposal suggests testing models on their ability to process and reason over a long, complex narrative.
量子位 · 43d ago














