
OpenAI Launches Fast Codex-Spark Variant
OpenAI released GPT-5.3-Codex-Spark, a lightweight version of its Codex programming tool. It targets rapid iteration with extreme inference speed.
cnBeta (Full RSS) · 214d ago
Every story we have kept, newest first.
Looking for the daily editions? → Past editions
Page 1359 of 1372

OpenAI released GPT-5.3-Codex-Spark, a lightweight version of its Codex programming tool. It targets rapid iteration with extreme inference speed.
cnBeta (Full RSS) · 214d ago

OpenAI released GPT-5.3-Codex-Spark, a lightweight version of its Codex intelligent programming tool. Designed for extreme inference speed, it targets rapid iteration scenarios.
cnBeta (Full RSS) · 214d ago
Voxtral Realtime achieves Whisper-quality transcription at 480ms latency via end-to-end streaming training. Features causal audio encoder and Ada RMS-Norm.
ArXiv AI · 214d ago
V2G pipeline converts CAD diagrams into property graphs, capturing component topology and connectivity overlooked by pixel-based MLLMs. It delivers major accuracy gains on electrical schematic compliance benchmarks where top MLLMs fail.
ArXiv AI · 214d ago
TSR introduces trajectory-search rollouts to enhance multi-turn reinforcement learning for LLM agents. It uses lightweight tree-style search for high-quality trajectories, improving rollout generation and stabilizing training.
ArXiv AI · 214d ago
TRACER is a trajectory-level uncertainty metric for tool-using agents, combining surprisal, repetition, and coherence signals with tail-focused aggregation. Improves AUROC by 37% and AUARC by 55% on tau^2-bench for failure prediction.
ArXiv AI · 214d ago
ThinkRouter introduces confidence-aware routing between latent and discrete spaces for efficient AI reasoning. It switches to discrete tokens during low-confidence steps to reduce noise from latent embeddings.
ArXiv AI · 214d ago
Text2GQL-Bench introduces a unified benchmark for Text-to-Graph-Query-Language systems with 178,184 question-query pairs across 13 domains and multiple GQLs. It features a scalable dataset generation framework and a multi-metric evaluation including grammatical validity, similarity, semantic alignment, and execution accuracy.
ArXiv AI · 214d ago
This survey frames multi-agent communication via the Five Ws, tracing evolution from MARL's hand-designed protocols to emergent language and LLM-based systems. It highlights trade-offs in interpretability, scalability, and generalization across paradigms.
ArXiv AI · 214d ago
SemaPop uses LLMs for semantic-conditioned population synthesis, deriving personas from surveys. Integrates with WGAN-GP for statistical alignment and behavioral realism.
ArXiv AI · 214d ago
scPilot enables LLMs to reason over single-cell RNA-seq data using natural language and on-demand tools for annotation, trajectories, and TF targeting. Paired with scBench benchmark, it shows gains like 11% accuracy lift via iterative reasoning.
ArXiv AI · 214d ago
SCF-RKL introduces sparse, distribution-aware model merging using reverse KL divergence to minimize interference. It selectively fuses complementary parameters, preserving stable representations and integrating new capabilities.
ArXiv AI · 214d ago
Quark Medical Alignment introduces a holistic multi-dimensional paradigm for aligning large language models in high-stakes medical question answering. It decomposes objectives into four categories with closed-loop optimization using observable metrics, diagnosis, and rewards.
ArXiv AI · 214d ago
PhyNiKCE introduces a neurosymbolic agentic framework to overcome LLM limitations in Computational Fluid Dynamics (CFD) simulations. It decouples neural planning from symbolic validation using a Constraint Satisfaction Problem approach to enforce physical laws.
ArXiv AI · 214d ago
PBSAI provides reference architecture for securing enterprise AI estates with multi-agent systems. Organizes 12 domains via agent families, context envelopes, output contracts.
ArXiv AI · 214d ago
NMIPS introduces a unified neuro-symbolic framework for solving PDE families with shared structures but varying parameters. It discovers interpretable analytical solutions via multifactorial optimization and affine transfer for efficiency.
ArXiv AI · 214d ago
Study reveals LLM agents like Llama/GPT/Claude produce 2-4 unique action paths per 10 runs on HotpotQA, with inconsistency predicting failure. Consistent runs hit 80-92% accuracy vs 25-60% for inconsistent ones.
ArXiv AI · 214d ago
MaxExp is a decision-driven framework for binarizing probabilistic species distribution models into presence-absence maps by maximizing evaluation metrics. It requires no calibration data and outperforms thresholding methods, especially under class imbalance.
ArXiv AI · 214d ago
MLLMs excel in perception but fail mathematical spatial reasoning, scoring under 60% on tasks humans solve at 95% accuracy. MathSpatial introduces a framework with MathSpatial-Bench (2K problems), MathSpatial-Corpus (8K training data), and MathSpatial-SRT for structured reasoning.
ArXiv AI · 214d ago
MAPLE is a modality-aware ecosystem for post-training multimodal LLMs, including MAPLE-bench, MAPO optimization, and adaptive curricula. It stratifies training by modality needs to cut variance and speed convergence.
ArXiv AI · 214d ago
LGS uses VAE latent space and Transformer dynamics for generalizable PDE simulation. Uncertainty knob and flow forcing stabilize long-horizon predictions.
ArXiv AI · 214d ago
INTENT is an inference-time planner for budget-constrained LLM agents using costly tools. Leverages hierarchical world model for intention-aware cost anticipation.
ArXiv AI · 214d ago
Proposes a framework for continuous learning of internal reasoning processes in AI, unifying reasoning, action, reflection, and verification. It treats thinking trajectories as learning material to evolve cognitive structures during execution.
ArXiv AI · 214d ago
GHOST applies structured pruning to Mamba2 using forward-pass controllability and observability metrics, avoiding backpropagation. Achieves 50% state reduction with ~1 PPL rise on WikiText-2 across 130M-2.7B models.
ArXiv AI · 214d ago
Literature review critiques 'ground truth' in ML data annotation as a positivistic fallacy ignoring human subjectivity. Analyzes 346 papers from top venues revealing biases like anchoring and geographic hegemony.
ArXiv AI · 214d ago
New research identifies 'rung collapse' in LLMs, where models confuse associations with causal interventions, leading to flawed reasoning under distributional shifts. It proposes Epistemic Regret Minimization (ERM), a belief revision method that penalizes causal errors independently of task success.
ArXiv AI · 214d ago
DrIGM introduces distributionally robust IGM for MARL, ensuring decentralized actions align under uncertainties via robust value factorization. Compatible with VDN/QMIX/QTRAN without reward shaping.
ArXiv AI · 214d ago
Formalizes decision-valued maps tracking representation impacts on outcomes. DecisionDB logs, replays, audits using content-based IDs and write-once storage.
ArXiv AI · 214d ago
DashAI introduces a human-centered XAI module integrating PDP, PFI, and KernelSHAP for no-code ML users. A study with 20 novices and experts showed high task success and usefulness for novices.
ArXiv AI · 214d ago
Researchers apply Crosscoders for the first time to compare LLMs across different architectures, introducing Dedicated Feature Crosscoders (DFCs) to isolate unique model features. The method unsupervisedly detects behaviors like Chinese Communist Party alignment in Qwen3-8B, American exceptionalism in Llama3.1-8B-Instruct, and copyright refusals in GPT-OSS-20B.
ArXiv AI · 214d ago