開源工具自動化 arXiv 論文篩選與摘要
Research Radar 是一款開源工具,透過根據使用者定義的研究興趣對 arXiv 論文進行評分,實現每日論文監控自動化。它利用多階段 LLM 流程篩選相關論文並生成深度摘要,大幅節省研究人員的時間。
Tag: #arxiv10 results
Research Radar 是一款開源工具,透過根據使用者定義的研究興趣對 arXiv 論文進行評分,實現每日論文監控自動化。它利用多階段 LLM 流程篩選相關論文並生成深度摘要,大幅節省研究人員的時間。
Arxiv cs.LG 分類每日上傳 100-200 篇論文。另有 cs.AI、math.OC 等類別更多 ML 論文。社群討論應對海量論文的策略。
用戶計劃將 ArXiv 論文拆成兩部分,將第一部分改標題與框架後投稿 NeurIPS 2026。擔心審稿人視為增量研究。尋求 r/MachineLearning 社群經驗。
This survey frames multi-agent communication via the Five Ws, tracing evolution from MARL's hand-designed protocols to emergent language and LLM-based systems. It highlights trade-offs in interpretability, scalability, and generalization across paradigms. Practical design patterns and open challenges are distilled for hybrid systems.
NMIPS introduces a unified neuro-symbolic framework for solving PDE families with shared structures but varying parameters. It discovers interpretable analytical solutions via multifactorial optimization and affine transfer for efficiency. Experiments show up to 35.7% accuracy gains over baselines.
INTENT is an inference-time planner for budget-constrained LLM agents using costly tools. Leverages hierarchical world model for intention-aware cost anticipation. Outperforms baselines on StableToolBench under budgets and price shifts.
Literature review critiques 'ground truth' in ML data annotation as a positivistic fallacy ignoring human subjectivity. Analyzes 346 papers from top venues revealing biases like anchoring and geographic hegemony. Proposes roadmap for pluralistic infrastructures embracing disagreement.
DrIGM introduces distributionally robust IGM for MARL, ensuring decentralized actions align under uncertainties via robust value factorization. Compatible with VDN/QMIX/QTRAN without reward shaping. Boosts OOD performance in SustainGym and StarCraft.
CausalAgent is a multi-agent system automating end-to-end causal inference via natural language. Integrates MAS, RAG, and MCP for data cleaning to report generation. Lowers barriers with interactive visualizations for non-experts.
AgentNoiseBench evaluates tool-using LLM agents' robustness in noisy real-world environments. Categorizes noise into user-noise and tool-noise; injects controllable perturbations into benchmarks. Reveals performance drops across models under perturbations.