Search

直接匹配不多,已補上最新動態。

Tag: #none4 results

Exposing Ground Truth Illusion in Annotations

Exposing Ground Truth Illusion in Annotations

Literature review critiques 'ground truth' in ML data annotation as a positivistic fallacy ignoring human subjectivity. Analyzes 346 papers from top venues revealing biases like anchoring and geographic hegemony. Proposes roadmap for pluralistic infrastructures embracing disagreement.

ArXiv AIResearchFeb 13#research#arxiv#none
GitHub Tackles Eternal September for Maintainers

GitHub Tackles Eternal September for Maintainers

Open source enters 'Eternal September' with reduced contribution friction and surging activity. Maintainers adapt via trust signals, triage methods, and community solutions. GitHub announces plans to support maintainers amid this shift.

GitHub BlogOfficialFeb 12#other#github#none
Safely Deferring to Capable AIs

Safely Deferring to Capable AIs

The article explores strategies for safely deferring key decisions to advanced AIs, especially in rushed scenarios where control becomes infeasible. It emphasizes deferring only slightly above the capability needed for automating safety research, assuming scheming is handled separately. Prosaic methods and supervised AI labor are proposed to enhance alignment, wisdom, and effectiveness on complex tasks.

AI Alignment ForumCommunityFeb 12#research#ai-alignment-forum#none
Inference Scaling vs Larger Tasks Clarified

Inference Scaling vs Larger Tasks Clarified

Distinguishes rising LLM inference compute into larger tasks (human-like linear scaling) vs. true inefficiency beyond human cost fractions. Uses Pareto frontier of budget vs. 50% reliability time-horizon. Argues much progress is bigger tasks, not unsustainable scaling.

AI Alignment ForumCommunityFeb 11#research#inference-scaling#none
Ox Alpha:神秘模型公開亮相

Ox Alpha:神秘模型公開亮相

Ox Alpha 是一款透過 OpenRouter 與 OpenCode 提供的匿名模型,具備 100 萬 token 上下文、圖片與影片輸入、工具呼叫及免費使用等能力。早期測試顯示其推理與程式設計表現強勁,但開發者、參數量、訓練資料與正式基準排名仍未獲確認。