Search

Tag: #research297 results

Trace Length as LLM Uncertainty Signal

Trace Length as LLM Uncertainty Signal

Apple researchers demonstrate that reasoning trace length serves as a simple, effective confidence estimator in large reasoning models. It performs comparably to verbalized confidence across models, datasets, and prompts, acting complementarily. The work shows reasoning post-training alters the trace-confidence relationship.

Apple Machine LearningOfficialFeb 12#research#apple-ml#general
Real-World Tool Agent Evaluation

Real-World Tool Agent Evaluation

Hugging Face explores OpenEnv for evaluating tool-using AI agents in practical settings. The post details methodologies for real-world testing. It highlights performance insights and benchmarks for agent capabilities.

Hugging Face BlogOfficialFeb 12#research#hugging-face#openenv
Mapping UX Design for Computer Agents

Mapping UX Design for Computer Agents

Study maps UX design space for LLM-based computer use agents via two-phase research. Phase 1 reviewed systems and interviewed eight UX/AI practitioners to create taxonomy. Categories cover user prompts, explainability, user control, and more.

Apple Machine LearningOfficialFeb 12#research#apple-ml#ux-design
Apple Maps UX for LLM Computer Agents

Apple Maps UX for LLM Computer Agents

Apple's Machine Learning team conducted a two-phase study to explore user experience design for LLM-based computer use agents. Phase 1 reviewed existing systems and interviewed eight UX/AI practitioners to create a taxonomy covering user prompts, explainability, user control, and more. The work aims to understand optimal user interactions with these UI-interacting agents.

Apple Machine LearningOfficialFeb 12#research#apple#na
Ada Palmer Reinvents Renaissance Teaching

Ada Palmer Reinvents Renaissance Teaching

Ada Palmer developed a three-week immersive simulation of the 1492 papal election for 60 students role-playing historical figures. This experiential method provides deep context for Machiavelli's The Prince, making references visceral rather than abstract. Combined history and poli sci classes revealed interdisciplinary insights on hypothetical scenarios.

LessWrong AICommunityFeb 11#research#ada-palmer#no-version
Inference Scaling vs Larger Tasks

Inference Scaling vs Larger Tasks

Distinguishes inference scaling from natural compute increases for bigger tasks in LLMs. Uses Pareto frontiers of compute budget vs. task time-horizon to analyze efficiency. Argues true scaling concerns arise only when exceeding human-equivalent costs inefficiently.

AI Alignment ForumCommunityFeb 11#research#llms#ai
Page 29 of 30