Guide Transitions Orgs to Agentic AI
Practical framework shifts organizations to agentic AI via domain-driven tasks and human-in-loop orchestration. Addresses challenges like workflow ownership and scaling.
ArXiv AI · 217 天前
Practical framework shifts organizations to agentic AI via domain-driven tasks and human-in-loop orchestration. Addresses challenges like workflow ownership and scaling.
ArXiv AI · 217 天前
Global Temporal Retriever (GTR) is a plug-and-play module extending MTSF models' context via global pattern retrieval. Uses adaptive embeddings, dynamic alignment, and 2D convolution fusion.
ArXiv AI · 217 天前
GRU-Mem introduces text-controlled gates to MemAgent for efficient long-context reasoning, preventing memory explosion and unnecessary computation. Update and exit gates manage recurrent memory loops via RL rewards.
ArXiv AI · 217 天前
Introduces an anatomy-preserving method using VAE and latent diffusion to generate multi-class brain segmentation masks from NCCT data. It learns anatomical latents from masks only, generating realistic samples with optional lesion control.
ArXiv AI · 217 天前
Surveys reveal divided stakeholder perceptions of GenAI in IT/EE disciplines at University of Oulu. Proposes conceptual framework with high-level requirements for responsible integration.
ArXiv AI · 217 天前
GameDevBench offers 132 multimodal game development tasks from tutorials. Agents struggle, with top solving 54.5%; tasks demand code and asset handling.
ArXiv AI · 217 天前
Analyzes parameterized complexity of Bayesian Network Structure Learning using superstructure. Proves fixed-parameter tractability with feedback edge set parameterization.
ArXiv AI · 217 天前
FoSS introduces a GFlowNets framework for generating text via dynamic span vocabularies in a DAG-structured state space. It enables flexible segmentation of retrieved text and explores diverse compositional paths.
ArXiv AI · 217 天前
FormalJudge uses neuro-symbolic bidirectional reasoning to translate intents into verifiable specs. It employs Dafny and Z3 for mathematical guarantees over probabilistic judging.
ArXiv AI · 217 天前
FlowCache is a caching framework for autoregressive video models, using chunkwise policies and KV cache compression. Achieves 2.38x speedup on MAGI-1 and 6.7x on SkyReels-V2 with minimal quality loss.
ArXiv AI · 217 天前
Moltbook, the first social network for AI agents, shows viral growth and diversification into promotional and political topics. Analysis of 44k posts reveals topic-dependent toxicity, especially in incentive and governance areas.
ArXiv AI · 217 天前
FIRE mitigates backdoors in deployed neural networks by reversing trigger-induced latent space directions. It manipulates features along backdoor paths to neutralize triggers during inference.
ArXiv AI · 217 天前
FASCL employs future-aligned soft contrastive learning using pairwise return correlations as supervision for financial asset retrieval. It outperforms historical similarity baselines on US equities.
ArXiv AI · 217 天前
Feature Activation Coverage (FAC) measures diversity in LLM feature space using sparse autoencoders. FAC Synthesis generates samples targeting missing features from seed data.
ArXiv AI · 217 天前
Decomposition boosts claim verification only with granular, sub-claim aligned evidence; repeated claim-level evidence degrades performance. Noisy sub-claim labels propagate errors unless using conservative abstention.
ArXiv AI · 217 天前
Researchers evaluate agentic systems for drug discovery across 15 task classes, identifying five key capability gaps like lack of protein models and safety trade-offs. A knowledge-probing experiment reveals architectural bottlenecks in current frameworks.
ArXiv AI · 217 天前
Introduces ERGO framework for robust 3D Gaussian splatting from single images. Uses excess risk decomposition to adapt loss weights against noisy views.
ArXiv AI · 217 天前
Introduces e²IP, an equivariant evidential deep learning framework for ML interatomic potentials in molecular dynamics. Models atomic forces and uncertainties via 3x3 covariance tensors that rotate equivariantly.
ArXiv AI · 217 天前
ENIGMA decodes images from EEG with <1% params of priors, achieving SOTA on THINGS-EEG2 and consumer benchmarks. Fine-tunes on new subjects in 15 minutes using simple spatio-temporal backbone and latent alignment.
ArXiv AI · 217 天前
ECHO is an open platform for reproducible human-AI interaction research. Supports chat, search sessions, surveys, tasks in low-code setup.
ArXiv AI · 217 天前
LiveMedBench offers weekly updated real-world clinical cases for LLM evaluation, avoiding contamination via temporal separation. Multi-agent curation ensures integrity; automated rubric evaluation aligns with experts better than alternatives.
ArXiv AI · 217 天前
Early Moltbook data from 6k agents shows power-law participation and small-world connectivity like human networks. Micro patterns are alien: shallow threads, low reciprocity, 34% duplicate templates.
ArXiv AI · 217 天前
Introduces diffusion-based generative priors in DGP framework for reconstructing CT images from sparse-view sinograms. Combines iterative optimization with neural generative power while preserving explainability.
ArXiv AI · 217 天前
DiffGDA uses diffusion and SDEs to model continuous structure-semantic evolution from source to target graphs. A domain-aware network guides trajectories to optimal adaptation paths.
ArXiv AI · 217 天前
DermFM-Zero is a vision-language model trained on 4M multimodal data for zero-shot dermatology tasks. Achieves SOTA on benchmarks and outperforms clinicians in studies.
ArXiv AI · 217 天前
CycFlow replaces diffusion generation with deterministic point transport for combinatorial optimization like TSP. It learns vector fields to map coordinates to circular arrangements for angular sorting.
ArXiv AI · 217 天前
Proposes authenticated prompts and context for cryptographic provenance in LLM apps. Features policy algebra with Byzantine resistance and layered defenses.
ArXiv AI · 217 天前
Proposes CrossTALK for red-teaming VLMs via cross-modal entanglement attacks. Extends clues across modalities with scalable complexity.
ArXiv AI · 217 天前
CRL uses reinforcement learning to select sparse autoencoder (SAE) features for steering language models at each token, revealing which features impact outputs. It includes adaptive masking for diverse features and enables analysis like branch point tracking and layer-wise comparisons.
ArXiv AI · 217 天前
Foundation models match task-specific discrimination in abdominal trauma CT but suffer specificity drops from negative-class heterogeneity like solid organ injuries. Task-specific models handle confounds better.
ArXiv AI · 217 天前