PulseAugur
实时 12:18:30
English(EN) C3-Bench: A Context-Aware Change Captioning Benchmark

新的基准和模型推动图像变化字幕和分割的进步

研究人员正在开发新的图像变化字幕和分割方法,旨在提高配对图像描述的准确性和细节。引入了几个新框架和基准,包括用于联合语义推理和空间分割的CCRC,使用文本引导对比损失的DFM,以及用于验证和定位字幕错误的GAVEL。此外,C3-Bench为上下文感知变化字幕提供了一个全面的基准,揭示了包括GPT-5.2等最先进的LLM在内的当前模型的局限性。RSICCLLM被提出为第一个专门用于遥感图像变化字幕的大型视觉语言模型后训练框架。 AI

影响 图像变化字幕和分割的进步可以改进监控、图像编辑和遥感分析等应用。

排序理由 多篇研究论文介绍了用于图像变化字幕和分割任务的新基准和模型。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 13 个来源。 我们如何撰写摘要 →

新的基准和模型推动图像变化字幕和分割的进步

报道来源 [13]

  1. arXiv cs.CL TIER_1 English(EN) · Licheng Zhang, Bach Le, Pengtao Zhao, Naveed Akhtar ·

    超越像素差异:为 Web UI 可视回归测试进行图像变化字幕生成基准测试

    arXiv:2607.01728v1 Announce Type: cross Abstract: Visual regression testing (VRT) is a standard quality assurance step in modern software release pipelines. On every change, it re-renders user interface (UI) screenshots, compares each one against an approved baseline image, and r…

  2. arXiv cs.CL TIER_1 English(EN) · Naveed Akhtar ·

    超越像素差异:为 Web UI 可视回归测试进行图像变化字幕基准测试

    Visual regression testing (VRT) is a standard quality assurance step in modern software release pipelines. On every change, it re-renders user interface (UI) screenshots, compares each one against an approved baseline image, and routes any detected difference to a human reviewer …

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越像素差异:为 Web UI 可视回归测试进行图像变化字幕的基准测试

    Visual regression testing (VRT) is a standard quality assurance step in modern software release pipelines. On every change, it re-renders user interface (UI) screenshots, compares each one against an approved baseline image, and routes any detected difference to a human reviewer …

  4. arXiv cs.AI TIER_1 English(EN) · Jinhong Hu, Xiaoping Wang, Shuyin Huang, Guojin Zhong, Kaitai Liu, Kai Lu ·

    CCRC:一种面向图像变化字幕和分割的变化感知字幕和推理链

    arXiv:2606.28724v1 Announce Type: cross Abstract: Understanding and localizing subtle changes between paired images is critical for tasks such as surveillance and image editing. However, traditional Image Change Captioning (ICC) methods lack spatial grounding, limiting their prec…

  5. arXiv cs.LG TIER_1 English(EN) · Yelin Wang, Zijia Song, Chuanguang Yang, Miaoyu Wang, Zhulin An, Libo Huang, Yongjun Xu ·

    DFM:用于遥感影像变化字幕生成的文本引导门控对比损失的差异特征建模

    arXiv:2606.27410v1 Announce Type: cross Abstract: The primary goal of Remote Sensing Image Change Captioning (RSICC) is to automatically generate descriptions of changes between remote sensing images captured at different time points. Existing models still rely on a single autore…

  6. arXiv cs.CL TIER_1 English(EN) · Zixian Gao, Atsushi Hashimoto, Kuniaki Saito ·

    GAVEL:基于事实的字幕错误验证与定位

    arXiv:2606.26923v1 Announce Type: new Abstract: Vision-language models (VLMs) often produce hallucinated or inconsistent outputs, where text and images are not properly aligned. Addressing this issue requires not only detecting misalignment but also explaining the discrepancy and…

  7. arXiv cs.CL TIER_1 English(EN) · Kuniaki Saito ·

    GAVEL:基于事实的字幕错误验证与定位

    Vision-language models (VLMs) often produce hallucinated or inconsistent outputs, where text and images are not properly aligned. Addressing this issue requires not only detecting misalignment but also explaining the discrepancy and localizing its visual evidence. We introduce GA…

  8. Hugging Face Daily Papers TIER_1 English(EN) ·

    C3-Bench:一个上下文感知的变化字幕基准

    While Change Captioning systems have garnered substantial attention to respond to our evolving world, their true performance on diverse real-world change contexts remains largely unexplored due to the lack of comprehensive evaluation frameworks. To fill this gap, we propose C3-Be…

  9. arXiv cs.AI TIER_1 English(EN) · Ue-Hwan Kim ·

    C3-Bench:一个上下文感知的变化描述基准

    While Change Captioning systems have garnered substantial attention to respond to our evolving world, their true performance on diverse real-world change contexts remains largely unexplored due to the lack of comprehensive evaluation frameworks. To fill this gap, we propose C3-Be…

  10. arXiv cs.CV TIER_1 English(EN) · Yelin Wang, Zijia Song, Shuo Ye, Chuanguang Yang, Miaoyu Wang, Yong Xu, Zhulin An, Yongjun Xu, Zitong Yu ·

    RSICCLLM:用于遥感图像变化字幕的多模态大语言模型

    arXiv:2606.28266v1 Announce Type: new Abstract: Remote Sensing Image Change Captioning (RSICC) aims to describe changes between bi-temporal remote sensing images and holds significant research and application value. However, most existing methods rely on conventional deep learnin…

  11. arXiv cs.CV TIER_1 English(EN) · Zitong Yu ·

    RSICCLLM:一种用于遥感图像变化字幕的多模态大语言模型

    Remote Sensing Image Change Captioning (RSICC) aims to describe changes between bi-temporal remote sensing images and holds significant research and application value. However, most existing methods rely on conventional deep learning architectures, and the limited model capacity …

  12. arXiv cs.CV TIER_1 English(EN) · Jae-Woo Kim, Hyeongbeom Kim, Ue-Hwan Kim ·

    C3-Bench:一个上下文感知型变化描述基准

    arXiv:2606.25445v1 Announce Type: new Abstract: While Change Captioning systems have garnered substantial attention to respond to our evolving world, their true performance on diverse real-world change contexts remains largely unexplored due to the lack of comprehensive evaluatio…

  13. arXiv cs.CV TIER_1 English(EN) · Phuc-Tan Nguyen, Hieu Nguyen, Minh-Triet Tran, Trung-Nghia Le ·

    VisChronos:通过现实生活事件彻底改变图像字幕生成

    arXiv:2606.24058v1 Announce Type: new Abstract: This paper aims to bridge the semantic gap between visual content and natural language understanding by leveraging historical events in the real world as a source of knowledge for caption generation. We propose VisChronos, a novel f…