Latest Computer Vision News & Updates
Detection, segmentation, OCR and visual understanding — the perception layer powering robotics, autonomy and content tools.
341 articles
AI Gives Ukrainian Kamikaze Drones Autonomous Target Tracking
US company Auterion is providing AI capabilities that allow Ukraine’s low-cost kamikaze drones to track targets autonomously. A $100 million deal will equip 50,000 Ukrainian drones with the US-developed technology.
Xunying Deploys 500 AI Cameras at EWC Paris
Xunying Tail 2 became the official partner of the Esports World Cup for a second consecutive year, deploying 500 AI cameras at the Paris event. Its AI tracking and PTZR gimbals reportedly allow one operator to control 60–70 devices, reducing reliance on large camera teams.
Building a Pipeline for Editable Textbook Figure Extraction
A developer is seeking a cost-effective, human-in-the-loop pipeline to convert static academic textbook figures into structured, editable digital assets. The goal is to detect figures, remove existing labels, and maintain underlying artwork for frontend rendering.
Top 5 smartphone makers to adopt 1:1 front sensors
Major smartphone manufacturers are reportedly adopting 1:1 square front camera sensors in 2025 to improve digital zoom and image quality during orientation switching. This shift follows Apple's implementation of the technology in the iPhone 17 series.
Photography is becoming a new frontier for smart hardware
As camera hardware costs drop, the focus of innovation is shifting toward 'autonomous photography' hardware like robots and drones. These devices aim to automate the entire process of framing, tracking, and content creation.
SenseNova-Vision: Unified Open-Source Visual Foundation Model
SenseTime has released SenseNova-Vision, a unified visual foundation model that integrates tasks like detection, segmentation, and 3D reconstruction. It is fully open-source, including a 50 million sample visual instruction corpus.
High schooler uses AI to screen autism via retina scans
A 17-year-old student developed 'RetinaMind', an AI tool using convolutional neural networks to analyze retinal images for autism and ADHD screening. The project achieved an 89% accuracy rate in tests, demonstrating the potential of AI in neurodevelopmental diagnostics.
Ant Group's LingBot-Vision Outperforms Meta's DINOv3
Ant Group's LingBot-Vision foundation model, with 1.1B parameters, has surpassed Meta's 7B DINOv3 in performance. The model is part of the LingBot-Depth 2.0 spatial perception suite, achieving 12 world-first benchmarks.
Worldmodeldata raises £7M to turn gameplay into AI data
Cambridge-based startup Worldmodeldata has raised £7 million to develop technology that converts video game footage into high-quality AI training data. The company aims to help AI models better understand physical world interactions.
Tesla Cybercab adds camera cleaning system for FSD
Tesla has integrated a spray-based cleaning system for the side cameras on its Cybercab prototype to ensure clear vision for autonomous driving. This hardware solution addresses a critical pain point where dirt or weather conditions obscure sensors.
NVIDIA releases LocateAnything for high-speed object detection
NVIDIA, in collaboration with universities, introduced LocateAnything, a model optimized for rapid object detection in robotics and AI agents. It utilizes a novel Parallel Box Decoding framework to achieve high-speed, precise visual-language localization.
TikTok Tests AI Portrait Detection to Combat Deepfakes
TikTok is testing a new tool that allows creators to identify if their likeness has been used by AI without authorization. The feature is currently in limited testing for US creators to help them report deepfake content.
Study Reveals Key Differences Between AI and Human Vision
Researchers at the University of York found that while artificial neural networks (ANNs) excel at predicting object recognition, their internal processing mechanisms differ significantly from primate brains. This suggests that current AI models may be achieving results through methods that do not mirror biological reality.
DialogueVPR: Interactive Reasoning for Visual Place Recognition
DialogueVPR introduces a dialogue-driven reasoning paradigm for geo-localization, moving beyond static retrieval. It includes the DlgQuest-Cities benchmark and the DQ-pilot framework to handle ambiguous natural language queries through interactive questioning.
SenseTime Releases Unified Vision Model SenseNova-Vision
SenseTime has open-sourced SenseNova-Vision, a unified model that integrates detection, segmentation, depth prediction, and 3D reconstruction. It currently leads the HuggingFace Any-to-Any leaderboard.
Build Multi-Camera 3D Tracking with NVIDIA DeepStream 9.1
NVIDIA DeepStream 9.1 introduces advanced capabilities for multi-camera 3D object tracking, addressing the limitations of traditional 2D tracking. This update enables developers to maintain object identity across different camera views in complex environments like warehouses and smart buildings.
Building visual intelligence with Amazon Bedrock and MCP servers
This guide introduces the Computer Vision MCP Server, a standardized interface for integrating visual processing into AI agents. It simplifies the integration of complex computer vision capabilities into broader application architectures.
CIA: AI Drones Drastically Reduce Battlefield Survival Time
CIA Director John Ratcliffe reports that AI-powered attack drones in Ukraine have reduced the average survival time for Russian soldiers to 20-30 minutes.
SenseTime open-sources SenseNova-Vision unified vision model
SenseTime has released SenseNova-Vision, a unified foundation model capable of handling multiple vision tasks like object detection and 3D reconstruction. It replaces the need for separate specialist models by integrating these capabilities into one system.
IMGNet: Face Verification via Sign Patterns Instead of Cosine
IMGNet is a novel face verification model that replaces traditional cosine similarity with sliding window sign pattern matching. It achieves competitive accuracy on LFW benchmarks with a lightweight 10.58 MB footprint.
Mistral AI introduces Robostral Navigate for single-camera navigation
Mistral AI has unveiled Robostral Navigate, a new solution focused on AI-powered navigation using only a single camera. This development signals Mistral's expansion into embodied AI and robotics applications.
Ant Group unveils breakthrough robot vision for transparent objects
Ant Group's embodied AI unit, Robbyant, has launched LingBot-Depth 2.0 and LingBot-Vision. These models are designed to help robots accurately perceive and navigate around glass, mirrors, and other transparent surfaces.
LingBot-Depth 2.0 achieves SOTA on masked depth benchmarks
LingBot-Depth 2.0 introduces sensor-validity masking, training models on the specific failure distributions of RGB-D cameras. It outperforms existing methods on 7 of 8 benchmarks by focusing on real-world sensor limitations like transparent surfaces and textureless areas.
NovoViz Develops High-Speed Low-Power SPAD Sensor
Swiss firm NovoViz has developed a Single-Photon Avalanche Diode (SPAD) edge imaging sensor that significantly reduces power consumption. The technology enables an equivalent frame rate of up to 100 million frames per second.
First open-source spatial-native embodied vision model released
Ant Lingbo has released the first spatial-native embodied vision foundation model. This model enhances the ability of robots to perceive and understand 3D spatial environments.
Ant Group's LingBot-Depth 2.0 Launches for Embodied AI
Ant Group's Lingbo Technology has released LingBot-Depth 2.0, a spatial perception model trained on 150 million data points. It is accompanied by the LingBot-Vision foundation model to enhance robot spatial awareness and precision.
LingBot-Vision: New Masked Boundary Modeling for Self-Supervised Pretraining
LingBot-Vision introduces a novel self-supervised pretraining approach that focuses on boundary-bearing tokens to improve visual representation learning. It achieves competitive results on NYUv2 depth completion benchmarks while using significantly less data than DINOv3.
iOS 27 code hints at new Apple wearable
Leaked iOS 27 source code references a new device model B790. The device is explicitly linked to 'Visual Intelligence' capabilities.
VideoFlexTok: Flexible-Length Video Tokenization
VideoFlexTok introduces a coarse-to-fine tokenization method that moves away from rigid 3D grid representations. This approach allows for variable-length tokenization based on video complexity, improving efficiency and information preservation.
ETH Zurich’s bidirectional pixel turns screens into cameras
Researchers at ETH Zurich have developed a bidirectional pixel capable of both emitting light and recording it. This innovation could allow screens to function as cameras without external sensors.