Latest Computer Vision News & Updates

Detection, segmentation, OCR and visual understanding — the perception layer powering robotics, autonomy and content tools.

341 articles

⚛️
Ars Technica AI17h ago

AI Gives Ukrainian Kamikaze Drones Autonomous Target Tracking

US company Auterion is providing AI capabilities that allow Ukraine’s low-cost kamikaze drones to track targets autonomously. A $100 million deal will equip 50,000 Ukrainian drones with the US-developed technology.

🐼
Pandaily12h ago

Xunying Deploys 500 AI Cameras at EWC Paris

Xunying Tail 2 became the official partner of the Esports World Cup for a second consecutive year, deploying 500 AI cameras at the Paris event. Its AI tracking and PTZR gimbals reportedly allow one operator to control 60–70 devices, reducing reliance on large camera teams.

🤖
Reddit r/MachineLearningYesterday

Building a Pipeline for Editable Textbook Figure Extraction

A developer is seeking a cost-effective, human-in-the-loop pipeline to convert static academic textbook figures into structured, editable digital assets. The goal is to detect figures, remove existing labels, and maintain underlying artwork for frontend rendering.

🏠
IT之家Yesterday

Top 5 smartphone makers to adopt 1:1 front sensors

Major smartphone manufacturers are reportedly adopting 1:1 square front camera sensors in 2025 to improve digital zoom and image quality during orientation switching. This shift follows Apple's implementation of the technology in the iPhone 17 series.

🐯
虎嗅20d ago

Photography is becoming a new frontier for smart hardware

As camera hardware costs drop, the focus of innovation is shifting toward 'autonomous photography' hardware like robots and drones. These devices aim to automate the entire process of framing, tracking, and content creation.

🏠
IT之家22d ago

SenseNova-Vision: Unified Open-Source Visual Foundation Model

SenseTime has released SenseNova-Vision, a unified visual foundation model that integrates tasks like detection, segmentation, and 3D reconstruction. It is fully open-source, including a 50 million sample visual instruction corpus.

🐯
虎嗅27d ago

High schooler uses AI to screen autism via retina scans

A 17-year-old student developed 'RetinaMind', an AI tool using convolutional neural networks to analyze retinal images for autism and ADHD screening. The project achieved an 89% accuracy rate in tests, demonstrating the potential of AI in neurodevelopmental diagnostics.

🐼
Pandaily28d ago

Ant Group's LingBot-Vision Outperforms Meta's DINOv3

Ant Group's LingBot-Vision foundation model, with 1.1B parameters, has surpassed Meta's 7B DINOv3 in performance. The model is part of the LingBot-Depth 2.0 spatial perception suite, achieving 12 world-first benchmarks.

🌍
The Next Web (TNW)29d ago

Worldmodeldata raises £7M to turn gameplay into AI data

Cambridge-based startup Worldmodeldata has raised £7 million to develop technology that converts video game footage into high-quality AI training data. The company aims to help AI models better understand physical world interactions.

🏠
IT之家44d ago

Tesla Cybercab adds camera cleaning system for FSD

Tesla has integrated a spray-based cleaning system for the side cameras on its Cybercab prototype to ensure clear vision for autonomous driving. This hardware solution addresses a critical pain point where dirt or weather conditions obscure sensors.

🏠
IT之家66d ago

NVIDIA releases LocateAnything for high-speed object detection

NVIDIA, in collaboration with universities, introduced LocateAnything, a model optimized for rapid object detection in robotics and AI agents. It utilizes a novel Parallel Box Decoding framework to achieve high-speed, precise visual-language localization.

🇨🇳
cnBeta (Full RSS)17d ago

TikTok Tests AI Portrait Detection to Combat Deepfakes

TikTok is testing a new tool that allows creators to identify if their likeness has been used by AI without authorization. The feature is currently in limited testing for US creators to help them report deepfake content.

🇨🇳
cnBeta (Full RSS)17d ago

Study Reveals Key Differences Between AI and Human Vision

Researchers at the University of York found that while artificial neural networks (ANNs) excel at predicting object recognition, their internal processing mechanisms differ significantly from primate brains. This suggests that current AI models may be achieving results through methods that do not mirror biological reality.

📄
ArXiv AI18d ago

DialogueVPR: Interactive Reasoning for Visual Place Recognition

DialogueVPR introduces a dialogue-driven reasoning paradigm for geo-localization, moving beyond static retrieval. It includes the DlgQuest-Cities benchmark and the DQ-pilot framework to handle ambiguous natural language queries through interactive questioning.

🐼
Pandaily18d ago

SenseTime Releases Unified Vision Model SenseNova-Vision

SenseTime has open-sourced SenseNova-Vision, a unified model that integrates detection, segmentation, depth prediction, and 3D reconstruction. It currently leads the HuggingFace Any-to-Any leaderboard.

🟩
NVIDIA Developer Blog19d ago

Build Multi-Camera 3D Tracking with NVIDIA DeepStream 9.1

NVIDIA DeepStream 9.1 introduces advanced capabilities for multi-camera 3D object tracking, addressing the limitations of traditional 2D tracking. This update enables developers to maintain object identity across different camera views in complex environments like warehouses and smart buildings.

☁️
AWS Machine Learning Blog19d ago

Building visual intelligence with Amazon Bedrock and MCP servers

This guide introduces the Computer Vision MCP Server, a standardized interface for integrating visual processing into AI agents. It simplifies the integration of complex computer vision capabilities into broader application architectures.

📊
Bloomberg Technology19d ago

CIA: AI Drones Drastically Reduce Battlefield Survival Time

CIA Director John Ratcliffe reports that AI-powered attack drones in Ukraine have reduced the average survival time for Russian soldiers to 20-30 minutes.

🇨🇳
TechNode21d ago

SenseTime open-sources SenseNova-Vision unified vision model

SenseTime has released SenseNova-Vision, a unified foundation model capable of handling multiple vision tasks like object detection and 3D reconstruction. It replaces the need for separate specialist models by integrating these capabilities into one system.

🤖
Reddit r/MachineLearning25d ago

IMGNet: Face Verification via Sign Patterns Instead of Cosine

IMGNet is a novel face verification model that replaces traditional cosine similarity with sliding window sign pattern matching. It achieves competitive accuracy on LFW benchmarks with a lightweight 10.58 MB footprint.

🦙
Reddit r/LocalLLaMA27d ago

Mistral AI introduces Robostral Navigate for single-camera navigation

Mistral AI has unveiled Robostral Navigate, a new solution focused on AI-powered navigation using only a single camera. This development signals Mistral's expansion into embodied AI and robotics applications.

🇭🇰
SCMP Technology28d ago

Ant Group unveils breakthrough robot vision for transparent objects

Ant Group's embodied AI unit, Robbyant, has launched LingBot-Depth 2.0 and LingBot-Vision. These models are designed to help robots accurately perceive and navigate around glass, mirrors, and other transparent surfaces.

🤖
Reddit r/MachineLearning28d ago

LingBot-Depth 2.0 achieves SOTA on masked depth benchmarks

LingBot-Depth 2.0 introduces sensor-validity masking, training models on the specific failure distributions of RGB-D cameras. It outperforms existing methods on 7 of 8 benchmarks by focusing on real-world sensor limitations like transparent surfaces and textureless areas.

💰
钛媒体28d ago

NovoViz Develops High-Speed Low-Power SPAD Sensor

Swiss firm NovoViz has developed a Single-Photon Avalanche Diode (SPAD) edge imaging sensor that significantly reduces power consumption. The technology enables an equivalent frame rate of up to 100 million frames per second.

⚛️
量子位28d ago

First open-source spatial-native embodied vision model released

Ant Lingbo has released the first spatial-native embodied vision foundation model. This model enhances the ability of robots to perceive and understand 3D spatial environments.

🔥
36氪28d ago

Ant Group's LingBot-Depth 2.0 Launches for Embodied AI

Ant Group's Lingbo Technology has released LingBot-Depth 2.0, a spatial perception model trained on 150 million data points. It is accompanied by the LingBot-Vision foundation model to enhance robot spatial awareness and precision.

🤖
Reddit r/MachineLearning28d ago

LingBot-Vision: New Masked Boundary Modeling for Self-Supervised Pretraining

LingBot-Vision introduces a novel self-supervised pretraining approach that focuses on boundary-bearing tokens to improve visual representation learning. It achieves competitive results on NYUv2 depth completion benchmarks while using significantly less data than DINOv3.

🇨🇳
cnBeta (Full RSS)30d ago

iOS 27 code hints at new Apple wearable

Leaked iOS 27 source code references a new device model B790. The device is explicitly linked to 'Visual Intelligence' capabilities.

🍎
Apple Machine Learning33d ago

VideoFlexTok: Flexible-Length Video Tokenization

VideoFlexTok introduces a coarse-to-fine tokenization method that moves away from rigid 3D grid representations. This approach allows for variable-length tokenization based on video complexity, improving efficiency and information preservation.

🌍
The Next Web (TNW)36d ago

ETH Zurich’s bidirectional pixel turns screens into cameras

Researchers at ETH Zurich have developed a bidirectional pixel capable of both emitting light and recording it. This innovation could allow screens to function as cameras without external sensors.