A new research paper evaluates Large Vision-Language Models (LVLMs) for 2D object detection in automated driving systems, focusing on Safety Of The Intended Functionality (SOTIF) conditions. The study used the PeSOTIF dataset to compare ten LVLMs, including Gemini 3, against specialized detectors like YOLOv5 and RT-DETRv4. Results indicate that top LVLMs show superior recall and comparable performance to specialized detectors under natural degradations, though geometric precision remains an advantage for the latter under specific perturbations. AI
IMPACT LVLMs show potential for enhancing safety in autonomous driving systems by improving object detection recall.
RANK_REASON The cluster contains a research paper evaluating AI models on a specific benchmark.
- EfficientDet: Scalable and Efficient Object Detection
- Faster R-CNN
- Fast R-CNN
- Mask R-CNN
- R-CNN
- Region proposal networks for automated bounding box detection and text segmentation
- RetinaNet
- YOLO
- Gemini 3
- Large Vision-Language Models
- PeSOTIF
- RT-DETRv4
- Safety Of The Intended Functionality
- YOLOv5
- Zhao Yongqi
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →