
Chrome Auto Browse: Wins and Crashes
Chrome's Auto Browse agent was tested for web surfing tasks. It demonstrated impressive capabilities in some areas. However, it also experienced spectacular failures.
Tag: #research297 results

Chrome's Auto Browse agent was tested for web surfing tasks. It demonstrated impressive capabilities in some areas. However, it also experienced spectacular failures.

OpenAI introduces GPT-5.3-Codex-Spark, its first real-time coding model. It delivers 15x faster generation and a 128k context window. The model is now available in research preview for ChatGPT Pro users.

Microsoft is developing powerful in-house AI models to achieve self-sufficiency and reduce dependence on OpenAI. This strategic shift follows a relationship reorganization in October last year. The company is now independently building its most advanced AI technology.
Cosmo3DFlow uses 3D wavelet transform and flow matching for efficient cosmological inference from N-body simulations. Addresses sparsity via spectral compression, enabling 50x faster sampling than diffusion models. Samples initial conditions in seconds at 128^3 resolution.
VulReaD uses a security knowledge graph and teacher LLM for CWE-consistent vulnerability detection beyond binary classification. Student models are fine-tuned with ORPO for taxonomy-aligned reasoning. Boosts F1 scores significantly on real datasets.
Found-RL integrates foundation models into RL for end-to-end driving via async batch inference to cut latency. Distills VLM guidance using VMR, AWAG; CLIP rewards shaped by conditional alignment. Lightweight policy matches VLM perf at 500 FPS.
VESPO introduces variational sequence-level soft policy optimization to tackle training instability in RL for LLMs caused by policy staleness and async execution. It derives a closed-form reshaping kernel for importance weights without length normalization. Experiments demonstrate stable training up to 64x staleness on math benchmarks.
Versor uses Conformal Geometric Algebra (CGA) for sequence modeling with SE(3)-equivariance. Outperforms Transformers on N-body dynamics, topology, and benchmarks with fewer parameters. Offers linear complexity and interpretability via rotors.
V-STAR addresses probability-reward mismatch in generative recsys via value-guided decoding and sibling-relative RL. VED efficiently explores high-potential prefixes; Sibling-GRPO focuses on decisive branches. Outperforms baselines in accuracy and diversity.
EVA is a cross-species, multimodal foundation model harmonizing transcriptomics and histology for immunology. It shows scaling laws and SOTA on 39 tasks from discovery to clinical trials. Open version released for transcriptomics research.