📄較早收集於 20h

EduEVAL-DB:AI家教評估資料集

EduEVAL-DB:AI家教評估資料集
PostLinkedIn
📄閱讀原文: ArXiv AI

💡New dataset for benchmarking AI tutors on bias, facts, and teaching quality—fine-tune lightweight models now.

⚡ 30-Second TL;DR

有什麼變化

854 個解釋對應 139 個 ScienceQA 問題,涵蓋科學、語言、社會科學

為什麼重要

此資料集透過評估 LLM 解釋中的教學風險,推進安全教育 AI 發展。它支援訓練輕量模型用於裝置端應用,民主化 AI 家教評估。研究人員現可依據真實教學標準基準測試教育 AI 系統。

下一步行動

Download EduEVAL-DB from arXiv and fine-tune Llama 3.1 8B for pedagogical risk detection.

誰應關注:Researchers & Academics

關鍵要點

  • 854 個解釋對應 139 個 ScienceQA 問題,涵蓋科學、語言、社會科學
  • 一位真人教師與六個經提示工程的 LLM 模擬角色
  • 五維度教學風險評分表與二元標籤
  • 半自動標註並經專家教師審核
  • 基準測試 Gemini 2.5 Pro 對比微調 Llama 3.1 8B 的可部署風險偵測

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 4 個來源。

🔑 增強重點摘要

  • EduEVAL-DB contains 854 explanations for 139 curated ScienceQA questions across K-12 science, language, and social science subjects[1][2].
  • Includes one human-teacher explanation per question and six LLM-simulated teacher roles created via prompt engineering, inspired by real educational styles and shortcomings[1][2][3].
  • Features a pedagogical risk rubric with five dimensions: factual correctness, explanatory depth/completeness, focus/relevance, student-level appropriateness, and ideological bias, using binary risk labels[1][2][3].
  • Annotations performed via semi-automatic process with expert teacher review; dataset is publicly released for training and evaluating LLM-based tutors and evaluators[1][2].
  • Benchmarks show Gemini 2.5 Pro outperforming fine-tuned Llama 3.1 8B in risk detection, with fine-tuning improving calibration, sensitivity, and deployability on consumer hardware[1].
📊 競品分析▸ Show
FeatureEduEVAL-DBScienceQA
Explanations per Question7 (1 human + 6 LLM)Primarily QA pairs with images/text
FocusPedagogical risk evaluationVisual question answering benchmarks
Rubric Dimensions5 (correctness, depth, focus, appropriateness, bias)Accuracy on science questions
BenchmarksGemini 2.5 Pro vs. Llama 3.1 8B fine-tunedVarious LLMs on QA accuracy
HardwareConsumer-deployable modelsNot specified

🛠️ 技術深入

  • Dataset derived from curated subset of ScienceQA benchmark, covering K-12 levels[1][2].
  • LLM-simulated roles instantiated via prompt engineering to mimic instructional styles and common shortcomings[1][2][3].
  • Binary risk labels annotated semi-automatically with expert teacher review for all five rubric dimensions[1][2].
  • Fine-tuning Llama 3.1 8B on EduEVAL-DB improves MAE trends, confusion matrix sensitivity to risk-present cases, and reduces majority label bias despite class imbalance[1].
  • Gemini 2.5 Pro leverages broader factual knowledge for advantages in evaluation, while fine-tuned model supports local deployment[1].

🔮 前景展望AI analysis grounded in cited sources

EduEVAL-DB enables training of locally deployable pedagogical evaluators, advancing safer AI tutors by assessing beyond factual accuracy to include depth, focus, appropriateness, and bias, potentially standardizing K-12 AI education tools[1].

時間線

2026-02-17
EduEVAL-DB paper submitted to arXiv (arXiv:2602.15531v1)

📎 來源 (4)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. arXiv — 2602
  2. arXiv — 2602
  3. chatpaper.com — 238399
  4. slideshare.net — 273768967
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。