EduEVAL-DB:AI家教評估資料集
💡New dataset for benchmarking AI tutors on bias, facts, and teaching quality—fine-tune lightweight models now.
⚡ 30-Second TL;DR
有什麼變化
854 個解釋對應 139 個 ScienceQA 問題,涵蓋科學、語言、社會科學
為什麼重要
此資料集透過評估 LLM 解釋中的教學風險,推進安全教育 AI 發展。它支援訓練輕量模型用於裝置端應用,民主化 AI 家教評估。研究人員現可依據真實教學標準基準測試教育 AI 系統。
下一步行動
Download EduEVAL-DB from arXiv and fine-tune Llama 3.1 8B for pedagogical risk detection.
關鍵要點
- •854 個解釋對應 139 個 ScienceQA 問題,涵蓋科學、語言、社會科學
- •一位真人教師與六個經提示工程的 LLM 模擬角色
- •五維度教學風險評分表與二元標籤
- •半自動標註並經專家教師審核
- •基準測試 Gemini 2.5 Pro 對比微調 Llama 3.1 8B 的可部署風險偵測
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 4 個來源。
🔑 增強重點摘要
- •EduEVAL-DB contains 854 explanations for 139 curated ScienceQA questions across K-12 science, language, and social science subjects[1][2].
- •Includes one human-teacher explanation per question and six LLM-simulated teacher roles created via prompt engineering, inspired by real educational styles and shortcomings[1][2][3].
- •Features a pedagogical risk rubric with five dimensions: factual correctness, explanatory depth/completeness, focus/relevance, student-level appropriateness, and ideological bias, using binary risk labels[1][2][3].
- •Annotations performed via semi-automatic process with expert teacher review; dataset is publicly released for training and evaluating LLM-based tutors and evaluators[1][2].
- •Benchmarks show Gemini 2.5 Pro outperforming fine-tuned Llama 3.1 8B in risk detection, with fine-tuning improving calibration, sensitivity, and deployability on consumer hardware[1].
📊 競品分析▸ Show
| Feature | EduEVAL-DB | ScienceQA |
|---|---|---|
| Explanations per Question | 7 (1 human + 6 LLM) | Primarily QA pairs with images/text |
| Focus | Pedagogical risk evaluation | Visual question answering benchmarks |
| Rubric Dimensions | 5 (correctness, depth, focus, appropriateness, bias) | Accuracy on science questions |
| Benchmarks | Gemini 2.5 Pro vs. Llama 3.1 8B fine-tuned | Various LLMs on QA accuracy |
| Hardware | Consumer-deployable models | Not specified |
🛠️ 技術深入
- •Dataset derived from curated subset of ScienceQA benchmark, covering K-12 levels[1][2].
- •LLM-simulated roles instantiated via prompt engineering to mimic instructional styles and common shortcomings[1][2][3].
- •Binary risk labels annotated semi-automatically with expert teacher review for all five rubric dimensions[1][2].
- •Fine-tuning Llama 3.1 8B on EduEVAL-DB improves MAE trends, confusion matrix sensitivity to risk-present cases, and reduces majority label bias despite class imbalance[1].
- •Gemini 2.5 Pro leverages broader factual knowledge for advantages in evaluation, while fine-tuned model supports local deployment[1].
🔮 前景展望AI analysis grounded in cited sources
EduEVAL-DB enables training of locally deployable pedagogical evaluators, advancing safer AI tutors by assessing beyond factual accuracy to include depth, focus, appropriateness, and bias, potentially standardizing K-12 AI education tools[1].
⏳ 時間線
📎 來源 (4)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI ↗
每週 AI 簡報
每週一封,可隨時退訂。
