Search

Tag: #research297 results

V2G Transforms CAD into Auditable Graphs

V2G Transforms CAD into Auditable Graphs

V2G pipeline converts CAD diagrams into property graphs, capturing component topology and connectivity overlooked by pixel-based MLLMs. It delivers major accuracy gains on electrical schematic compliance benchmarks where top MLLMs fail. Benchmark and code released on GitHub for further research.

TSR Boosts Multi-Turn RL for LLM Agents

TSR Boosts Multi-Turn RL for LLM Agents

TSR introduces trajectory-search rollouts to enhance multi-turn reinforcement learning for LLM agents. It uses lightweight tree-style search for high-quality trajectories, improving rollout generation and stabilizing training. Achieves up to 15% performance gains on tasks like Sokoban and WebShop.

ArXiv AIResearchFeb 13#research#tsr#llm-agents
TRACER Aggregates Risks in Agent Trajectories

TRACER Aggregates Risks in Agent Trajectories

TRACER is a trajectory-level uncertainty metric for tool-using agents, combining surprisal, repetition, and coherence signals with tail-focused aggregation. Improves AUROC by 37% and AUARC by 55% on tau^2-bench for failure prediction. Code and benchmark on GitHub.

ArXiv AIResearchFeb 13#research#tracer#ai-agents
ThinkRouter Boosts Reasoning Efficiency

ThinkRouter Boosts Reasoning Efficiency

ThinkRouter introduces confidence-aware routing between latent and discrete spaces for efficient AI reasoning. It switches to discrete tokens during low-confidence steps to reduce noise from latent embeddings. Experiments show major accuracy gains on STEM and coding tasks while shortening outputs.

ArXiv AIResearchFeb 13#research#thinkrouter#ai-reasoning
Text2GQL-Bench: New Graph Query Benchmark

Text2GQL-Bench: New Graph Query Benchmark

Text2GQL-Bench introduces a unified benchmark for Text-to-Graph-Query-Language systems with 178,184 question-query pairs across 13 domains and multiple GQLs. It features a scalable dataset generation framework and a multi-metric evaluation including grammatical validity, similarity, semantic alignment, and execution accuracy. Evaluations show LLMs struggle with ISO-GQL, achieving only 4% zero-shot execution accuracy, improving to 50% with 3-shot prompting and 45.1% with fine-tuning.

Surveying Multi-Agent Communication Paradigms

Surveying Multi-Agent Communication Paradigms

This survey frames multi-agent communication via the Five Ws, tracing evolution from MARL's hand-designed protocols to emergent language and LLM-based systems. It highlights trade-offs in interpretability, scalability, and generalization across paradigms. Practical design patterns and open challenges are distilled for hybrid systems.

ArXiv AIResearchFeb 13#research#arxiv#multi-agent
SemaPop: Semantic Population Synthesis

SemaPop: Semantic Population Synthesis

SemaPop uses LLMs for semantic-conditioned population synthesis, deriving personas from surveys. Integrates with WGAN-GP for statistical alignment and behavioral realism. Achieves better marginal/joint distribution matches with diversity.

ArXiv AIResearchFeb 13#research#semapop#llms
scPilot Automates Single-Cell Analysis

scPilot Automates Single-Cell Analysis

scPilot enables LLMs to reason over single-cell RNA-seq data using natural language and on-demand tools for annotation, trajectories, and TF targeting. Paired with scBench benchmark, it shows gains like 11% accuracy lift via iterative reasoning. Transparent traces explain biological insights.

ArXiv AIResearchFeb 13#research#scpilot#scbench
SCF-RKL Advances Model Merging

SCF-RKL Advances Model Merging

SCF-RKL introduces sparse, distribution-aware model merging using reverse KL divergence to minimize interference. It selectively fuses complementary parameters, preserving stable representations and integrating new capabilities. Evaluations on 24 benchmarks show superior performance in reasoning, instruction following, and safety.

ArXiv AIResearchFeb 13#research#scf-rkl#ai
Quark Medical Alignment Paradigm Launched

Quark Medical Alignment Paradigm Launched

Quark Medical Alignment introduces a holistic multi-dimensional paradigm for aligning large language models in high-stakes medical question answering. It decomposes objectives into four categories with closed-loop optimization using observable metrics, diagnosis, and rewards. A unified mechanism with Reference-Frozen Normalization and Tri-Factor Adaptive Dynamic Weighting resolves scale mismatches and optimization conflicts.

ArXiv AIResearchFeb 13#research#quark#medical-alignment
Page 9 of 30