AI Models Compete in Beijing and Shanghai Gaokao Essays

๐กSee how top LLMs handle complex, human-centric essay prompts in high-stakes academic testing environments.
โก 30-Second TL;DR
What Changed
AI models evaluated on 2026 Gaokao essay prompts
Why It Matters
Using high-stakes academic examinations as a benchmark provides a unique look at how models handle cultural nuance and complex narrative structures. This helps researchers identify gaps in reasoning and creative writing capabilities compared to human students.
What To Do Next
Analyze the output of your models against standardized essay benchmarks to identify weaknesses in argumentative structure and cultural alignment.
Key Points
- โขAI models evaluated on 2026 Gaokao essay prompts
- โขFocus on Beijing and Shanghai regional exam topics
- โขTesting model reasoning on AI and technology development themes
๐ง Deep Insight
Web-grounded analysis with 26 cited sources.
๐ Enhanced Key Takeaways
- โขAI models have been participating in Gaokao evaluations for several years, initially focusing on mathematics and now increasingly on essay writing, with evaluations often involving human examiners unaware of the essays' AI origin.
- โขWhile AI models demonstrate strong performance in Chinese language and English, they continue to face challenges with complex reasoning in mathematics and often lack the creativity, emotional depth, and nuanced understanding of classical Chinese passages found in high-scoring human essays.
- โขChinese AI models, such as Alibaba's Qwen2-72B and Baidu's ERNIE Bot 5.0, have shown competitive performance in Chinese language tasks and specific benchmarks, sometimes surpassing international models due to localized fine-tuning with Chinese datasets and past exam archives.
- โขThe integration of AI into the Gaokao ecosystem extends beyond essay generation to include personalized revision, psychological support, AI-powered security monitoring during exams, and university application consulting.
- โขIn response to concerns about fairness and cheating, Chinese tech companies like Tencent, ByteDance (Doubao), and Moonshot AI (Kimi) disabled certain AI functions during the 2025 Gaokao, and the Ministry of Education issued warnings against false advertising of 'AI predicted exam questions' for the 2026 Gaokao.
๐ Competitor Analysisโธ Show
markdown
| AI Model / Platform | Primary Developer | Gaokao Essay/Language Performance (2024/2025) | Gaokao Math Performance (2024/2025) | General Chinese Language Reasoning (2025/2026) |
|---|---|---|---|---|
| Qwen2-72B | Alibaba | Top scorer in 2024 Gaokao (303/420 total), 72% accuracy in Chinese language & literature. | 36% average accuracy across LLMs. | Qwen3-Max achieved 75.2% accuracy on QualBench, edging past GPT-4o. |
| GPT-4o | OpenAI | 296/420 total in 2024 Gaokao, 67% accuracy in Chinese language & literature, 81% in English. | 73/150 (second highest among LLMs). | GPT-o3 topped basic logic, GPT-5 was close second in overall reasoning. |
| ERNIE Bot 5.0 | Baidu | Took on Gaokao essay prompts in 2023. | N/A | Ranked #1 Chinese AI model, #8 globally in text performance (early 2026). Outperformed ChatGPT-4 in Chinese interventional radiology questions. |
| InternLM 2.0 | Shanghai AI Lab | 295.5/420 total in 2024 Gaokao. | 75/150 (highest among LLMs). | N/A |
| DeepSeek | DeepSeek | Users tested on actual exam prompts (2025). | N/A | DeepSeek (R1) achieved highest overall performance in Chinese Medical Licensing Exam (454.8 mean score). |
| Doubao | ByteDance | Disabled exam-relevant features during 2025 Gaokao. | N/A | Below average human performance in Chinese Medical Licensing Exam (413.7 mean score). |
| AI-MATHS | Chengdu Zhunxingyunxue Technology | N/A | Scored 105/150 on Beijing math paper, 100/150 on national paper in 2017. | N/A |
๐ ๏ธ Technical Deep Dive
- AI-MATHS (2017): Utilized 11 servers, big data technology, and natural language recognition for solving math problems.
- Baidu ERNIE Bot 5.0 (2026): Employs a mixture-of-experts (MoE) architecture with 2.4 trillion parameters, activating less than 3% per query to enhance inference efficiency. It uses a unified autoregressive architecture for native full-modality understanding and generation (text, images, audio, video).
- AI Essay Grading Systems: These systems evaluate essays across multiple dimensions including idea development, organization, language use, and spelling/format. They apply distinct evaluation criteria for analytical writing (logic, argumentation) versus expressive writing (emotion, reflection) and can provide feedback aligned with official exam graders. Some advanced systems can even read handwritten essays.
- Gaokao Evaluation Methodology (e.g., GAOKAO-Eval): Involves using genuinely unseen data, ensuring temporal isolation, and operating in a closed-book environment to prevent data leakage. Subjective questions are graded by experienced human examiners who are unaware of the AI origin of the responses.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (26)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- scmp.com
- people.cn
- china.org.cn
- xinhuanet.com
- chinadaily.com.cn
- chinadaily.com.cn
- yicaiglobal.com
- arxiv.org
- chinadaily.com.cn
- chinadailyhk.com
- asianews.network
- youtube.com
- auntminnie.com
- nih.gov
- hku.hk
- aicerts.ai
- reddit.com
- youtube.com
- scio.gov.cn
- xinhuanet.com
- globaltimes.cn
- globaltimes.cn
- businessinsider.com
- chinamediaproject.org
- microsoft.com
- geniebench.ai
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
Same topic
Explore #benchmarking
Same product
More on llm-benchmarking
Same source
Latest from cnBeta (Full RSS)
Qwen Code Benchmark Infrastructure Validation

Improving SFT reasoning with Self-Distilled Reasoning

Meta testing StoryKit for AI-generated children's stories

WHO study confirms mobile phones do not cause brain cancer
AI-curated news aggregator. All content rights belong to original publishers.
Original source: cnBeta (Full RSS) โ