Search

Tag: #llm-evaluation69 results

快速合理性檢查:最常見AI研究建議

快速合理性檢查:最常見AI研究建議

作者建議初級AI研究人員進行快速合理性檢查,避免在有缺陷想法上浪費時間,例如驗證資料偏差、相關性和摘要統計。範例包括檢查LLM輸出中是否有「隱藏任務」等詞彙、工具呼叫成功率,以及推理鏈長度。檢視典型資料集範例有助辨識失敗是否來自能力缺口或其他問題。

AI Alignment ForumCommunityApr 2#sanity-checks#research-tips#llm-evaluation
PELLI Framework Boosts LLM Code Quality

PELLI Framework Boosts LLM Code Quality

PELLI is an iterative framework for integrating LLMs into software generation, evaluating code on maintainability, performance, and reliability. It tests five popular LLMs across three domains using Python standards. GPT-4T and Gemini outperform others, with prompt design impacting quality.

ArXiv AIResearchFeb 12#research#pelli#v1
Page 7 of 7