Researchers have developed new datasets and pipelines to improve the understanding of surgical videos by vision-language models (VLMs). SurgAtlas, a large-scale dataset with over 2,391 hours of surgical videos, includes both open and minimally invasive procedures and offers diverse annotations for training foundation models. Additionally, the SurgSTU-Pipeline generates fine-grained spatial-temporal question-answer samples for surgical videos, addressing the challenge of creating such datasets manually. When applied to existing surgical video data, this pipeline creates the SurgSTU dataset, which has been shown to enhance the spatial-temporal understanding capabilities of VLMs in surgical contexts. AI
IMPACT These advancements could lead to more sophisticated AI tools for computer-assisted surgery and medical training.
RANK_REASON The cluster describes new research papers introducing datasets and pipelines for surgical video understanding.
Read on Hugging Face Daily Papers →
- arXiv
- computer-assisted surgery
- Hugging Face
- large-language models
- Lennart Maack
- SurgSTU
- SurgSTU-Pipeline
- vision-language model
- Qwen3-VL-8B
- SurgAtlas
- YouTube
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →