All Updates
Page 709 of 1669
May 7, 2026
Poly-Time Thiele Rules on Voter Intervals
Resolves open question: Thiele rules, including PAV, computable in polynomial time on voter interval (VI) domain via LP with guaranteed integral optimum and fast algorithm. Extends to VCI and LC domains, proving LC strictly contains VCI via graph theory. NP-hard on tree-based VCI generalization.
LLMs Fail at Exact ComputationβPoT Wins
Researchers evaluate prompting methods like CoT, PoT, and Self-Consistency on deterministic tasks such as counting and arithmetic. PoT achieves perfect accuracy by generating executable code for external interpreters, outperforming others. Findings recommend tool integration or specialized models for reliable exact computation.
Framework Detects Team Mental Model Gaps
Proposes a framework to categorize four mental model discrepancies in task-based team dialogues: unsupported beliefs, false beliefs, belief contradictions, and omissions. Analyzes dialogues from 20 dyad teams in sequential collaborative object identification tasks. Demonstrates these patterns predict future misalignments using historical discrepancy counts as baseline.
CreativityBench: LLM Creative Tool Benchmark
CreativityBench is a new benchmark evaluating LLMs' creative reasoning through affordance-based tool repurposing. It includes a 4K-entity affordance KB with 150K+ annotations and 14K constrained tasks. Tests on 10 SOTA LLMs reveal failures in identifying parts, affordances, and mechanisms, with scaling and CoT offering limited improvements.
cotomi Act Learns Work by Watching
cotomi Act is a browser-based agent that automates multi-step tasks and learns organizational knowledge from user browsing behavior. It achieves 80.4% success on the WebArena benchmark, surpassing the 78.2% human baseline. The system includes a shared workspace with artifacts like task boards and wikis derived from observed behavior.
Code Boosts LLM Symbolic Regression
A novel LLM-based evolutionary search framework introduces programmatic context augmentation for symbolic regression. It enables code-based dataset interactions to extract signals beyond MSE metrics. Outperforms baselines on LLM-SRBench in efficiency and accuracy.
Big Tech AI Mandates Fuel Formalism Farce
Major firms mandate AI usage with assessments, prompting employees to game systems by having AI write 'AIεΏεΎ' reports or consume tokens on vast GitHub codebases. This creates a spectacle of faked productivity in the AI revolution. Questions if corporate AI pushes are performative rather than substantive.
AI Adoption Mismatch: Goals vs Worker Experiences
AI adoption fails as workers, invisible in design, resist due to usability issues and misaligned expectations. Interviews in healthcare, finance, and management reveal barriers like poor interoperability, limited control, and weak communication. Strategies at individual, task, and organizational levels are proposed to center workers.
ADAPTS: AI Agents Track Mental Health Symptoms
ADAPTS framework uses mixture-of-agents LLMs to rate depression and anxiety from clinical interviews, decomposing them into symptom-specific tasks with auditable justifications. It outperforms human raters on high-discrepancy interviews (error 22 vs 26) and achieves ICC 0.877 with extended protocols across N=204 samples. Text-based but extensible to multimodal for scalable psychiatric assessment.
Alibaba Surges Past Tencent on Chip Buzz
Alibaba shares are outpacing Tencent amid a rally in Asian chipmakers. Investors show enthusiasm for Alibaba's ambitious semiconductor unit. This creates a divergence between China's two internet giants.
NEXTDC CEO: AI Boom Beats Sleep
NEXTDC's boss is short on sleep but flush with new funds. The Australian data center operator warns investors: you snooze, you lose amid AI boom. Highlights surging demand for AI infrastructure.
Unitree G1 Robot Flies on Commercial Plane
A Unitree G1 humanoid robot named Bebop boarded Southwest Airlines flight by purchasing a cabin seat due to its heavy transport case. It walked through the airport, danced with passengers, but its battery exceeding FAA limits was removed. The incident highlights regulatory challenges for treating robots as passengers versus cargo.
Meta ProgramBench Tests AI Offline Program Recreation
Meta Superintelligence Lab launches ProgramBench, a benchmark challenging SOTA AI models to recreate complex executable programs like ffmpeg, SQLite, and ripgrep entirely from scratch without internet access. The Reddit post highlights this evaluation of AI's program synthesis capabilities. It questions if current top models can achieve this feat.
Yahua invests in embodied elderly care robots
Yahua Electronics strategically invests in Huaxi Tech, global embodied pension robot firm. They will combine medical interaction pipelines and full-stack robot tech for C/B-end synergies. Focus: accelerate R&D and mass deployment in care scenarios.
SOCAMM2 Explodes Memory Scene
SOCAMM2 is generating massive buzz in the memory industry. It is positioned as the next frontrunner in AI storage solutions.
Users Fed Up with Lifeless AI Creations
Cyber residents are increasingly frustrated with AI-generated works lacking human essence. A quote highlights how cheap averages make unique exceptions valuable.
Silicon Valley Talent Flees with $18.8B Funding
Top talents from Silicon Valley giants are mass-exiting to launch startups. Capital has poured $18.8 billion into this entrepreneurship wave.
Samsung Dominates AI Memory Boom
AI training and inference demand surges for HBM, high-performance DRAM, and enterprise SSDs, not just GPUs. Samsung, a global memory leader, is perfectly positioned to capitalize. This shift bolsters Samsung's edge in the AI hardware ecosystem.
A-shares rally as semiconductors surge
A-share indices rise at midday: Shanghai +0.25%, Shenzhen +0.71%, ChiNext +0.98%. Semiconductors, comm equipment, auto parts lead gains; Changfei Fiber +7%, Newquan Shares +6%, Hua Hong +4%. Coal/energy sectors lag.
Hermes Surges: Stability Dethrones OpenClaw
Hermes gains massive popularity as developers prioritize stable execution over raw power. Ambitious 'full-brain' AI approaches are outpaced by conservative, reliable methods. OpenClaw reluctantly cedes ground to stability-focused models.