Search

Tag: #v1120 results

MetaphorStar Masters Image Metaphor Reasoning

MetaphorStar Masters Image Metaphor Reasoning

MetaphorStar uses end-to-end visual RL for image metaphor understanding, featuring TFQ-Data dataset, TFQ-GRPO method, and TFQ-Bench. MetaphorStar-32B sets SOTA on implication benchmarks, outperforming 20+ MLLMs including Gemini-3.0-pro. Improves general visual reasoning via scaling analyses.

ArXiv AIResearchFeb 12#research#metaphorstar#v1
MERIT Boosts LLM Negotiation Skills

MERIT Boosts LLM Negotiation Skills

AgoraBench tests LLMs in nine bargaining scenarios like deception; utility metrics measure human alignment. MERIT feedback via prompting/finetuning elicits deeper strategy and opponent awareness. Outperforms baselines in negotiation power and acquisition.

ArXiv AIResearchFeb 12#research#agorabench#v1
MEL Boosts LLM Reasoning via Meta-Experience

MEL Boosts LLM Reasoning via Meta-Experience

Meta-Experience Learning (MEL) enhances RLVR by internalizing error-derived meta-experience into LLM memory. Uses self-verification for contrastive analysis of trajectories. Achieves 3.92%-4.73% Pass@1 gains across model sizes.

ArXiv AIResearchFeb 12#research#mel#v1
LRMs Fail to Transfer Reasoning to ToM

LRMs Fail to Transfer Reasoning to ToM

Study compares reasoning vs non-reasoning LLMs on ToM benchmarks, finding no consistent gains and sometimes worse performance. Insights reveal slow thinking collapse, need for adaptive reasoning, and option-matching shortcuts. Interventions like S2F and T2M mitigate issues.

ArXiv AIResearchFeb 12#research#tom-study#v1
LOREN: Low-Rank Adaptation for Neural Receivers

LOREN: Low-Rank Adaptation for Neural Receivers

LOREN introduces low-rank adapters to enable code-rate adaptation in neural receivers without storing separate weights. It freezes a shared base network and trains lightweight adapters per code rate. Achieves comparable performance with major hardware savings.

ArXiv AIResearchFeb 12#research#loren#v1
LoRA Enables Modular Chemistry Prediction

LoRA Enables Modular Chemistry Prediction

Evaluates LoRA for parameter-efficient fine-tuning of LLMs on organic reaction datasets like USPTO and C-H functionalisation. Matches full fine-tuning accuracy while preserving multi-task performance and mitigating forgetting. Reveals distinct reactivity patterns for better adaptation.

ArXiv AIResearchFeb 12#research#lora#v1
Locomo-Plus Tests LLM Cognitive Memory

Locomo-Plus Tests LLM Cognitive Memory

Locomo-Plus benchmarks cognitive memory in LLM agents under cue-trigger disconnects, focusing on latent conversational constraints. It proposes constraint consistency evaluation over string-matching. Reveals gaps in existing memory systems.

ArXiv AIResearchFeb 12#research#locomo-plus#v1
LLMs Outstrategize Humans in Games

LLMs Outstrategize Humans in Games

Uses AlphaEvolve to discover interpretable models of human and LLM strategic behavior from data. Analysis on iterated rock-paper-scissors shows frontier LLMs capable of deeper strategy than humans. Provides foundation for understanding behavioral differences in interactions.

ArXiv AIResearchFeb 12#research#alphacevolve#v1
LLMs Generate Planning Abstractions

LLMs Generate Planning Abstractions

Prompts pretrained LLMs to create QNP abstractions for generalized planning from domains and tasks. Automated debugging detects/fixes errors iteratively. Guided LLMs produce useful abstractions for qualitative numerical planning.

ArXiv AIResearchFeb 12#research#qnp-generator#v1
Page 6 of 12