๐Ÿ‡จ๐Ÿ‡ณStalecollected in 75m

AI Models Compete in Beijing and Shanghai Gaokao Essays

AI Models Compete in Beijing and Shanghai Gaokao Essays
PostLinkedIn
๐Ÿ‡จ๐Ÿ‡ณRead original on cnBeta (Full RSS)

๐Ÿ’กSee how top LLMs handle complex, human-centric essay prompts in high-stakes academic testing environments.

โšก 30-Second TL;DR

What Changed

AI models evaluated on 2026 Gaokao essay prompts

Why It Matters

Using high-stakes academic examinations as a benchmark provides a unique look at how models handle cultural nuance and complex narrative structures. This helps researchers identify gaps in reasoning and creative writing capabilities compared to human students.

What To Do Next

Analyze the output of your models against standardized essay benchmarks to identify weaknesses in argumentative structure and cultural alignment.

Who should care:Researchers & Academics

Key Points

  • โ€ขAI models evaluated on 2026 Gaokao essay prompts
  • โ€ขFocus on Beijing and Shanghai regional exam topics
  • โ€ขTesting model reasoning on AI and technology development themes

๐Ÿง  Deep Insight

Web-grounded analysis with 26 cited sources.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขAI models have been participating in Gaokao evaluations for several years, initially focusing on mathematics and now increasingly on essay writing, with evaluations often involving human examiners unaware of the essays' AI origin.
  • โ€ขWhile AI models demonstrate strong performance in Chinese language and English, they continue to face challenges with complex reasoning in mathematics and often lack the creativity, emotional depth, and nuanced understanding of classical Chinese passages found in high-scoring human essays.
  • โ€ขChinese AI models, such as Alibaba's Qwen2-72B and Baidu's ERNIE Bot 5.0, have shown competitive performance in Chinese language tasks and specific benchmarks, sometimes surpassing international models due to localized fine-tuning with Chinese datasets and past exam archives.
  • โ€ขThe integration of AI into the Gaokao ecosystem extends beyond essay generation to include personalized revision, psychological support, AI-powered security monitoring during exams, and university application consulting.
  • โ€ขIn response to concerns about fairness and cheating, Chinese tech companies like Tencent, ByteDance (Doubao), and Moonshot AI (Kimi) disabled certain AI functions during the 2025 Gaokao, and the Ministry of Education issued warnings against false advertising of 'AI predicted exam questions' for the 2026 Gaokao.
๐Ÿ“Š Competitor Analysisโ–ธ Show

markdown

AI Model / PlatformPrimary DeveloperGaokao Essay/Language Performance (2024/2025)Gaokao Math Performance (2024/2025)General Chinese Language Reasoning (2025/2026)
Qwen2-72BAlibabaTop scorer in 2024 Gaokao (303/420 total), 72% accuracy in Chinese language & literature.36% average accuracy across LLMs.Qwen3-Max achieved 75.2% accuracy on QualBench, edging past GPT-4o.
GPT-4oOpenAI296/420 total in 2024 Gaokao, 67% accuracy in Chinese language & literature, 81% in English.73/150 (second highest among LLMs).GPT-o3 topped basic logic, GPT-5 was close second in overall reasoning.
ERNIE Bot 5.0BaiduTook on Gaokao essay prompts in 2023.N/ARanked #1 Chinese AI model, #8 globally in text performance (early 2026). Outperformed ChatGPT-4 in Chinese interventional radiology questions.
InternLM 2.0Shanghai AI Lab295.5/420 total in 2024 Gaokao.75/150 (highest among LLMs).N/A
DeepSeekDeepSeekUsers tested on actual exam prompts (2025).N/ADeepSeek (R1) achieved highest overall performance in Chinese Medical Licensing Exam (454.8 mean score).
DoubaoByteDanceDisabled exam-relevant features during 2025 Gaokao.N/ABelow average human performance in Chinese Medical Licensing Exam (413.7 mean score).
AI-MATHSChengdu Zhunxingyunxue TechnologyN/AScored 105/150 on Beijing math paper, 100/150 on national paper in 2017.N/A

๐Ÿ› ๏ธ Technical Deep Dive

  • AI-MATHS (2017): Utilized 11 servers, big data technology, and natural language recognition for solving math problems.
  • Baidu ERNIE Bot 5.0 (2026): Employs a mixture-of-experts (MoE) architecture with 2.4 trillion parameters, activating less than 3% per query to enhance inference efficiency. It uses a unified autoregressive architecture for native full-modality understanding and generation (text, images, audio, video).
  • AI Essay Grading Systems: These systems evaluate essays across multiple dimensions including idea development, organization, language use, and spelling/format. They apply distinct evaluation criteria for analytical writing (logic, argumentation) versus expressive writing (emotion, reflection) and can provide feedback aligned with official exam graders. Some advanced systems can even read handwritten essays.
  • Gaokao Evaluation Methodology (e.g., GAOKAO-Eval): Involves using genuinely unseen data, ensuring temporal isolation, and operating in a closed-book environment to prevent data leakage. Subjective questions are graded by experienced human examiners who are unaware of the AI origin of the responses.

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

AI will significantly transform educational assessment methods in China.
The increasing use of AI for essay grading, personalized feedback, and exam security suggests a shift towards AI-assisted evaluation, potentially improving efficiency, consistency, and accessibility in the Gaokao ecosystem.
Chinese AI models will continue to gain a competitive edge in Chinese language and culturally specific tasks.
Localized fine-tuning with curated Chinese datasets and past exam archives has already shown to boost performance in benchmarks like Gaokao and QualBench, indicating a strategic advantage for domestic models in their native language context.
The ethical and regulatory challenges surrounding AI in high-stakes exams will intensify.
The Ministry of Education's warnings against AI-predicted exam questions and tech companies disabling AI functions during Gaokao highlight ongoing concerns about cheating, fairness, and the integrity of the examination process, which will necessitate evolving policies and safeguards.

โณ Timeline

2017-06
AI-MATHS takes Gaokao math exam
2018-05
Chinese schools begin testing AI for essay grading
2023-06
ChatGPT and Ernie Bot publicly attempt Gaokao essays
2024-06
Shanghai AI Lab evaluates LLMs on Gaokao, noting strong language, weak math
2025-06
Chinese tech firms disable AI functions during Gaokao to prevent cheating
2026-06
Ministry of Education warns against 'AI predicted exam questions' for Gaokao
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: cnBeta (Full RSS) โ†—