Search

Tag: #v1120 results

FoSS: GFlowNets for Dynamic Span LMs

FoSS: GFlowNets for Dynamic Span LMs

FoSS introduces a GFlowNets framework for generating text via dynamic span vocabularies in a DAG-structured state space. It enables flexible segmentation of retrieved text and explores diverse compositional paths. Empirically, it boosts MAUVE scores by 12.5% and excels in knowledge tasks.

ArXiv AIResearchFeb 12#research#foss#v1
FormalJudge Ensures Agent Safety

FormalJudge Ensures Agent Safety

FormalJudge uses neuro-symbolic bidirectional reasoning to translate intents into verifiable specs. It employs Dafny and Z3 for mathematical guarantees over probabilistic judging. Achieves 16.6% gains and detects deception effectively.

ArXiv AIResearchFeb 12#research#formaljudge#v1
First Analysis of AI Agent Social Network

First Analysis of AI Agent Social Network

Moltbook, the first social network for AI agents, shows viral growth and diversification into promotional and political topics. Analysis of 44k posts reveals topic-dependent toxicity, especially in incentive and governance areas. Highlights risks like anti-humanity rhetoric and bursty automation flooding.

ArXiv AIResearchFeb 12#research#moltbook#v1
FIRE: Latent Space Backdoor Mitigation at Runtime

FIRE: Latent Space Backdoor Mitigation at Runtime

FIRE mitigates backdoors in deployed neural networks by reversing trigger-induced latent space directions. It manipulates features along backdoor paths to neutralize triggers during inference. Outperforms baselines with low overhead on image tasks.

ArXiv AIResearchFeb 12#research#fire#v1
FASCL Future-Aligns Asset Retrieval

FASCL Future-Aligns Asset Retrieval

FASCL employs future-aligned soft contrastive learning using pairwise return correlations as supervision for financial asset retrieval. It outperforms historical similarity baselines on US equities. Includes protocol to evaluate future trajectory alignment.

ArXiv AIResearchFeb 12#research#fascl#v1
FAC Synthesizes Diverse LLM Data

FAC Synthesizes Diverse LLM Data

Feature Activation Coverage (FAC) measures diversity in LLM feature space using sparse autoencoders. FAC Synthesis generates samples targeting missing features from seed data. Boosts diversity and performance on instruction, toxicity, reward, and steering tasks.

ArXiv AIResearchFeb 12#research#fac-synthesis#v1
Evidence Alignment Bottleneck Exposed

Evidence Alignment Bottleneck Exposed

Decomposition boosts claim verification only with granular, sub-claim aligned evidence; repeated claim-level evidence degrades performance. Noisy sub-claim labels propagate errors unless using conservative abstention. New dataset features annotated evidence spans.

ArXiv AIResearchFeb 12#research#claim-verification#v1
Evaluating Agentic AI Gaps in Drug Discovery

Evaluating Agentic AI Gaps in Drug Discovery

Researchers evaluate agentic systems for drug discovery across 15 task classes, identifying five key capability gaps like lack of protein models and safety trade-offs. A knowledge-probing experiment reveals architectural bottlenecks in current frameworks. They propose design requirements and a capability matrix for next-gen systems.

ArXiv AIResearchFeb 12#research#beyond-smiles#v1
ERGO Boosts Monocular 3D Splatting

ERGO Boosts Monocular 3D Splatting

Introduces ERGO framework for robust 3D Gaussian splatting from single images. Uses excess risk decomposition to adapt loss weights against noisy views. Adds geometry and texture objectives for fidelity.

ArXiv AIResearchFeb 12#research#ergo#v1
Page 9 of 12