
GPT-5.3-Codex-Spark: Real-Time Coding Model
OpenAI introduces GPT-5.3-Codex-Spark, its first real-time coding model. It delivers 15x faster generation and a 128k context window.
OpenAI Blog · 217d ago
Every story we have kept, newest first.
Looking for the daily editions? → Past editions
Page 1365 of 1372

OpenAI introduces GPT-5.3-Codex-Spark, its first real-time coding model. It delivers 15x faster generation and a 128k context window.
OpenAI Blog · 217d ago

OpenAI introduced GPT-5.3-Codex-Spark, the first real-time coding model. It offers 15x faster generation and 128k context length.
OpenAI News · 217d ago

DeepSeek launched its R1 reasoning model in January 2025. This marked a pivotal shift for Chinese AI development.
MIT Technology Review · 217d ago

DeepSeek released its R1 reasoning model in January 2025. This marked a turning point for Chinese AI development.
MIT Technology Review · 217d ago

DeepSeek launched R1 reasoning model in January 2025, signaling a turning point for Chinese AI. Companies are rapidly advancing open-source models.
MIT Technology Review · 217d ago

Microsoft is developing powerful in-house AI models to achieve self-sufficiency and reduce dependence on OpenAI. This strategic shift follows a relationship reorganization in October last year.
cnBeta (Full RSS) · 217d ago

Anthropic is opening premium features to free Claude users, including file creation/editing, third-party connectors, and Skills. This counters OpenAI's introduction of ads in free and low-tier ChatGPT.
cnBeta (Full RSS) · 217d ago

Google appears set to release Gemini 3.1 Pro soon, with model references already spotted in related arenas. This follows recent launches like Zhipu's open-source GLM-5 and DeepSeek's upgraded model with larger context window.
cnBeta (Full RSS) · 217d ago

Google Gemini and related tools now refuse Disney character generation requests after Disney's IP infringement notice. The update rolled out about two months after Disney's December cease-and-desist letter.
cnBeta (Full RSS) · 217d ago
Cosmo3DFlow uses 3D wavelet transform and flow matching for efficient cosmological inference from N-body simulations. Addresses sparsity via spectral compression, enabling 50x faster sampling than diffusion models.
ArXiv AI · 217d ago
VulReaD uses a security knowledge graph and teacher LLM for CWE-consistent vulnerability detection beyond binary classification. Student models are fine-tuned with ORPO for taxonomy-aligned reasoning.
ArXiv AI · 217d ago
Found-RL integrates foundation models into RL for end-to-end driving via async batch inference to cut latency. Distills VLM guidance using VMR, AWAG; CLIP rewards shaped by conditional alignment.
ArXiv AI · 217d ago
Vision-Centric Jailbreak Attack (VJA) uses visual inputs to bypass safety in image editing models. IESBench benchmark tests vulnerabilities with up to 80.9% success rates.
ArXiv AI · 217d ago
VESPO introduces variational sequence-level soft policy optimization to tackle training instability in RL for LLMs caused by policy staleness and async execution. It derives a closed-form reshaping kernel for importance weights without length normalization.
ArXiv AI · 217d ago
Versor uses Conformal Geometric Algebra (CGA) for sequence modeling with SE(3)-equivariance. Outperforms Transformers on N-body dynamics, topology, and benchmarks with fewer parameters.
ArXiv AI · 217d ago
V-STAR addresses probability-reward mismatch in generative recsys via value-guided decoding and sibling-relative RL. VED efficiently explores high-potential prefixes; Sibling-GRPO focuses on decisive branches.
ArXiv AI · 217d ago
EVA is a cross-species, multimodal foundation model harmonizing transcriptomics and histology for immunology. It shows scaling laws and SOTA on 39 tasks from discovery to clinical trials.
ArXiv AI · 217d ago
Develops theory for random projections in computing influence functions, covering unregularized, regularized, and factorized cases. Shows exact preservation conditions and handles out-of-range gradients via leakage term.
ArXiv AI · 217d ago
TwiFF-2.7M dataset and model advance VCoT for videos via future frame generation. TwiFF-Bench evaluates reasoning trajectories.
ArXiv AI · 217d ago
Transformer training on modular arithmetic tasks collapses high-dimensional parameters to 3-4D execution manifolds. This structure explains attention concentration, SGD integrability, and sparse autoencoder limits.
ArXiv AI · 217d ago
NMRTrans uses set transformers on experimental NMR spectra for molecular structure elucidation, trained on NMRSpec corpus from literature. It models spectra as unordered peak sets aligning with NMR physics.
ArXiv AI · 217d ago
Integrates neural networks, topological data analysis, and Bayesian methods for AI in military domains. Covers image, time-series, graph applications like fraud detection.
ArXiv AI · 217d ago
Inference-time scaling in language models leads to adaptive resource rationality without explicit cost rewards. Models shift from brute-force to analytic strategies as task complexity rises.
ArXiv AI · 217d ago
TokaMark standardizes AI evaluation on MAST tokamak data with unified multi-modal access and 14 tasks. Harmonizes formats, metadata, and protocols for reproducible comparisons.
ArXiv AI · 217d ago
Text-guided framework enhances weakly supervised multimodal video anomaly detection. Employs in-context learning for anomaly text augmentation and multi-scale bottleneck Transformer for fusion.
ArXiv AI · 217d ago
Introduces δ_TCB metric to quantify LLM internal state robustness against perturbations, beyond traditional accuracy. Linked to output embedding geometry, it reveals prediction instabilities missed by perplexity.
ArXiv AI · 217d ago
LHAW generates controllable underspecified long-horizon tasks by removing info across goals, constraints, inputs, context. Validates via agent trials, classifying ambiguity impacts.
ArXiv AI · 217d ago
SynergyKGC fuses entity semantics with heterogeneous topologies via cross-modal synergy. Uses density-dependent anchoring and double-tower consistency.
ArXiv AI · 217d ago
Step 3.5 Flash is a 196B MoE model with 11B active params for agentic tasks. Optimized with sliding-window attention and MTP-3 for low-latency inference.
ArXiv AI · 217d ago
McNemar's test framework detects post-optimization LLM degradations via per-sample comparisons. Aggregates across benchmarks with controlled false positives.
ArXiv AI · 217d ago