All Updates

Page 709 of 1669

May 7, 2026

πŸ“„
ArXiv AIβ€’86d ago

Poly-Time Thiele Rules on Voter Intervals

Resolves open question: Thiele rules, including PAV, computable in polynomial time on voter interval (VI) domain via LP with guaranteed integral optimum and fast algorithm. Extends to VCI and LC domains, proving LC strictly contains VCI via graph theory. NP-hard on tree-based VCI generalization.

#social-choice#voting-rules#linear-programming
πŸ“„
ArXiv AIβ€’86d ago

LLMs Fail at Exact Computationβ€”PoT Wins

Researchers evaluate prompting methods like CoT, PoT, and Self-Consistency on deterministic tasks such as counting and arithmetic. PoT achieves perfect accuracy by generating executable code for external interpreters, outperforming others. Findings recommend tool integration or specialized models for reliable exact computation.

#prompting-strategies#deterministic-tasks#code-execution
πŸ“„
ArXiv AIβ€’86d ago

Framework Detects Team Mental Model Gaps

Proposes a framework to categorize four mental model discrepancies in task-based team dialogues: unsupported beliefs, false beliefs, belief contradictions, and omissions. Analyzes dialogues from 20 dyad teams in sequential collaborative object identification tasks. Demonstrates these patterns predict future misalignments using historical discrepancy counts as baseline.

#mental-models#team-dialogues#smm-discrepancies
πŸ“„
ArXiv AIβ€’86d ago

CreativityBench: LLM Creative Tool Benchmark

CreativityBench is a new benchmark evaluating LLMs' creative reasoning through affordance-based tool repurposing. It includes a 4K-entity affordance KB with 150K+ annotations and 14K constrained tasks. Tests on 10 SOTA LLMs reveal failures in identifying parts, affordances, and mechanisms, with scaling and CoT offering limited improvements.

#benchmark#affordances#tool-repurposing
πŸ“„
ArXiv AIβ€’86d ago

cotomi Act Learns Work by Watching

cotomi Act is a browser-based agent that automates multi-step tasks and learns organizational knowledge from user browsing behavior. It achieves 80.4% success on the WebArena benchmark, surpassing the 78.2% human baseline. The system includes a shared workspace with artifacts like task boards and wikis derived from observed behavior.

#web-agents
πŸ“„
ArXiv AIβ€’86d ago

Code Boosts LLM Symbolic Regression

A novel LLM-based evolutionary search framework introduces programmatic context augmentation for symbolic regression. It enables code-based dataset interactions to extract signals beyond MSE metrics. Outperforms baselines on LLM-SRBench in efficiency and accuracy.

#symbolic-regression#llm-evolution#data-analysis
🐯
θ™Žε—…β€’86d ago

Big Tech AI Mandates Fuel Formalism Farce

Major firms mandate AI usage with assessments, prompting employees to game systems by having AI write 'AIεΏƒεΎ—' reports or consume tokens on vast GitHub codebases. This creates a spectacle of faked productivity in the AI revolution. Questions if corporate AI pushes are performative rather than substantive.

#ai-adoption#enterprise#formailism
πŸ“„
ArXiv AIβ€’86d ago

AI Adoption Mismatch: Goals vs Worker Experiences

AI adoption fails as workers, invisible in design, resist due to usability issues and misaligned expectations. Interviews in healthcare, finance, and management reveal barriers like poor interoperability, limited control, and weak communication. Strategies at individual, task, and organizational levels are proposed to center workers.

#ai-adoption#worker-experience
πŸ“„
ArXiv AIβ€’86d ago

ADAPTS: AI Agents Track Mental Health Symptoms

ADAPTS framework uses mixture-of-agents LLMs to rate depression and anxiety from clinical interviews, decomposing them into symptom-specific tasks with auditable justifications. It outperforms human raters on high-discrepancy interviews (error 22 vs 26) and achieves ICC 0.877 with extended protocols across N=204 samples. Text-based but extensible to multimodal for scalable psychiatric assessment.

#mental-health-ai#affective-computing#llm-agents
πŸ“Š
Bloomberg Technologyβ€’86d ago

Alibaba Surges Past Tencent on Chip Buzz

Alibaba shares are outpacing Tencent amid a rally in Asian chipmakers. Investors show enthusiasm for Alibaba's ambitious semiconductor unit. This creates a divergence between China's two internet giants.

#semiconductors#stock-rally#china-big-tech
πŸ“Š
Bloomberg Technologyβ€’86d ago

NEXTDC CEO: AI Boom Beats Sleep

NEXTDC's boss is short on sleep but flush with new funds. The Australian data center operator warns investors: you snooze, you lose amid AI boom. Highlights surging demand for AI infrastructure.

#data-centers#ai-demand#funding
🐯
θ™Žε—…β€’86d ago

Unitree G1 Robot Flies on Commercial Plane

A Unitree G1 humanoid robot named Bebop boarded Southwest Airlines flight by purchasing a cabin seat due to its heavy transport case. It walked through the airport, danced with passengers, but its battery exceeding FAA limits was removed. The incident highlights regulatory challenges for treating robots as passengers versus cargo.

#humanoid-robot#robotics-logistics#aviation-regulations
πŸ€–
Reddit r/MachineLearningβ€’86d ago

Meta ProgramBench Tests AI Offline Program Recreation

Meta Superintelligence Lab launches ProgramBench, a benchmark challenging SOTA AI models to recreate complex executable programs like ffmpeg, SQLite, and ripgrep entirely from scratch without internet access. The Reddit post highlights this evaluation of AI's program synthesis capabilities. It questions if current top models can achieve this feat.

#benchmark#program-synthesis#offline-ai
πŸ”₯
36ζ°ͺβ€’86d ago

Yahua invests in embodied elderly care robots

Yahua Electronics strategically invests in Huaxi Tech, global embodied pension robot firm. They will combine medical interaction pipelines and full-stack robot tech for C/B-end synergies. Focus: accelerate R&D and mass deployment in care scenarios.

#robotics#funding#elderly-care
πŸ’°
ι’›εͺ’体‒86d ago

SOCAMM2 Explodes Memory Scene

SOCAMM2 is generating massive buzz in the memory industry. It is positioned as the next frontrunner in AI storage solutions.

#ai-storage#memory-hardware#industry-buzz
πŸ’°
ι’›εͺ’体‒86d ago

Users Fed Up with Lifeless AI Creations

Cyber residents are increasingly frustrated with AI-generated works lacking human essence. A quote highlights how cheap averages make unique exceptions valuable.

#ai-content#user-feedback#humanization
πŸ’°
ι’›εͺ’体‒86d ago

Silicon Valley Talent Flees with $18.8B Funding

Top talents from Silicon Valley giants are mass-exiting to launch startups. Capital has poured $18.8 billion into this entrepreneurship wave.

#talent-exodus#startup-funding#venture-capital
πŸ’°
ι’›εͺ’体‒86d ago

Samsung Dominates AI Memory Boom

AI training and inference demand surges for HBM, high-performance DRAM, and enterprise SSDs, not just GPUs. Samsung, a global memory leader, is perfectly positioned to capitalize. This shift bolsters Samsung's edge in the AI hardware ecosystem.

#memory-chips#ai-hardware#supply-chain
πŸ”₯
36ζ°ͺβ€’86d ago

A-shares rally as semiconductors surge

A-share indices rise at midday: Shanghai +0.25%, Shenzhen +0.71%, ChiNext +0.98%. Semiconductors, comm equipment, auto parts lead gains; Changfei Fiber +7%, Newquan Shares +6%, Hua Hong +4%. Coal/energy sectors lag.

#stock-market#a-shares#chip-sector
🐯
θ™Žε—…β€’86d ago

Hermes Surges: Stability Dethrones OpenClaw

Hermes gains massive popularity as developers prioritize stable execution over raw power. Ambitious 'full-brain' AI approaches are outpaced by conservative, reliable methods. OpenClaw reluctantly cedes ground to stability-focused models.

#model-stability#llm-comparison
Page 709 of 1669