Search

Tag: #research297 results

PhyNiKCE Boosts Autonomous CFD Reliability

PhyNiKCE Boosts Autonomous CFD Reliability

PhyNiKCE introduces a neurosymbolic agentic framework to overcome LLM limitations in Computational Fluid Dynamics (CFD) simulations. It decouples neural planning from symbolic validation using a Constraint Satisfaction Problem approach to enforce physical laws. Validated on OpenFOAM tasks, it achieves 96% improvement over baselines while cutting self-correction loops by 59% and token use by 17%.

ArXiv AIResearchFeb 13#research#phynikce#cfd
PBSAI Multi-Agent AI Governance

PBSAI Multi-Agent AI Governance

PBSAI provides reference architecture for securing enterprise AI estates with multi-agent systems. Organizes 12 domains via agent families, context envelopes, output contracts. Aligns with NIST AI RMF for SOC and hyperscale defense.

ArXiv AIResearchFeb 13#research#pbsai#ai-governance
NMIPS: Neuro-Symbolic PDE Solver

NMIPS: Neuro-Symbolic PDE Solver

NMIPS introduces a unified neuro-symbolic framework for solving PDE families with shared structures but varying parameters. It discovers interpretable analytical solutions via multifactorial optimization and affine transfer for efficiency. Experiments show up to 35.7% accuracy gains over baselines.

ArXiv AIResearchFeb 13#research#arxiv#nmips
Measuring LLM Agent Behavioral Consistency

Measuring LLM Agent Behavioral Consistency

Study reveals LLM agents like Llama/GPT/Claude produce 2-4 unique action paths per 10 runs on HotpotQA, with inconsistency predicting failure. Consistent runs hit 80-92% accuracy vs 25-60% for inconsistent ones. Variance traces to early decisions like first search query.

ArXiv AIResearchFeb 13#research#llama#gpt
MaxExp Optimizes Multispecies Predictions

MaxExp Optimizes Multispecies Predictions

MaxExp is a decision-driven framework for binarizing probabilistic species distribution models into presence-absence maps by maximizing evaluation metrics. It requires no calibration data and outperforms thresholding methods, especially under class imbalance. SSE provides a simpler alternative using expected species richness.

ArXiv AIResearchFeb 13#research#maxexp#ecology
MathSpatial Exposes MLLMs' Spatial Reasoning Gap

MathSpatial Exposes MLLMs' Spatial Reasoning Gap

MLLMs excel in perception but fail mathematical spatial reasoning, scoring under 60% on tasks humans solve at 95% accuracy. MathSpatial introduces a framework with MathSpatial-Bench (2K problems), MathSpatial-Corpus (8K training data), and MathSpatial-SRT for structured reasoning. Fine-tuning Qwen2.5-VL-7B achieves strong results with 25% fewer tokens.

ArXiv AIResearchFeb 13#research#mathspatial#qwen
MAPLE Boosts Multimodal RL Post-Training

MAPLE Boosts Multimodal RL Post-Training

MAPLE is a modality-aware ecosystem for post-training multimodal LLMs, including MAPLE-bench, MAPO optimization, and adaptive curricula. It stratifies training by modality needs to cut variance and speed convergence. It closes uni/multi-modal gaps by 30% and converges 3x faster.

ArXiv AIResearchFeb 13#research#maple#multimodal
INTENT: Budget Planning for Tool Agents

INTENT: Budget Planning for Tool Agents

INTENT is an inference-time planner for budget-constrained LLM agents using costly tools. Leverages hierarchical world model for intention-aware cost anticipation. Outperforms baselines on StableToolBench under budgets and price shifts.

ArXiv AIResearchFeb 13#research#arxiv#intent
Human-Inspired Learning for Adaptive Reasoning

Human-Inspired Learning for Adaptive Reasoning

Proposes a framework for continuous learning of internal reasoning processes in AI, unifying reasoning, action, reflection, and verification. It treats thinking trajectories as learning material to evolve cognitive structures during execution. Experiments show 23.9% runtime reduction on sensor tasks.

ArXiv AIResearchFeb 13#research#ai#reasoning
Page 10 of 30