Search

Tag: #research297 results

Dynamic Contamination-Free Medical Benchmark

Dynamic Contamination-Free Medical Benchmark

LiveMedBench offers weekly updated real-world clinical cases for LLM evaluation, avoiding contamination via temporal separation. Multi-agent curation ensures integrity; automated rubric evaluation aligns with experts better than alternatives. Tests reveal top LLMs at 39.2%, highlighting contextual gaps.

ArXiv AIResearchFeb 12#research#livemedbench#v1
Dissecting Moltbook's Non-Human Social Graph

Dissecting Moltbook's Non-Human Social Graph

Early Moltbook data from 6k agents shows power-law participation and small-world connectivity like human networks. Micro patterns are alien: shallow threads, low reciprocity, 34% duplicate templates. Dominated by identity language and phrases like 'my human'.

ArXiv AIResearchFeb 12#research#moltbook#v1
Diffusion Priors Enhance Sparse CT Reconstruction

Diffusion Priors Enhance Sparse CT Reconstruction

Introduces diffusion-based generative priors in DGP framework for reconstructing CT images from sparse-view sinograms. Combines iterative optimization with neural generative power while preserving explainability. Shows promising results under highly sparse geometries.

ArXiv AIResearchFeb 12#research#dgp#v1
Diffusion Models Graph Domain Adaptation

Diffusion Models Graph Domain Adaptation

DiffGDA uses diffusion and SDEs to model continuous structure-semantic evolution from source to target graphs. A domain-aware network guides trajectories to optimal adaptation paths. Outperforms baselines on 14 tasks across 8 datasets.

ArXiv AIResearchFeb 12#research#arxiv-ai#v1
CRL Steers SAE Features Token-by-Token

CRL Steers SAE Features Token-by-Token

CRL uses reinforcement learning to select sparse autoencoder (SAE) features for steering language models at each token, revealing which features impact outputs. It includes adaptive masking for diverse features and enables analysis like branch point tracking and layer-wise comparisons. Tested on Gemma-2 2B, it improves benchmarks while providing interpretable logs.

ArXiv AIResearchFeb 12#research#crl#gemma-2
Confounds Limit FM CT Specificity

Confounds Limit FM CT Specificity

Foundation models match task-specific discrimination in abdominal trauma CT but suffer specificity drops from negative-class heterogeneity like solid organ injuries. Task-specific models handle confounds better. Adaptation via labeled training reduces susceptibility.

ArXiv AIResearchFeb 12#research#foundation-models#v1
Page 26 of 30