PulseAugur
中
实时 07:50:09
English(EN) Mid-Training Language Models on Raw Video

原始视频中期训练提升LLM在视觉任务上的性能

研究人员探索了在中期训练大型语言模型时使用原始视频数据(无字幕或文本监督)的有效性。通过将视频帧编码为视觉标记,并训练语言模型来预测下一个视觉标记,他们发现在视频和图像基准测试中都有显著的改进。与未经此中期训练的模型相比,中期训练的Qwen3-1.7B模型在视频任务上提高了2.9分,在图像任务上提高了5.1分。值得注意的是,文本性能保持稳定,而视觉任务上的提升在训练早期就出现了。 AI

影响 这项研究提出了一种新颖的、自监督的方法,利用原始视频来增强多模态LLM,有可能减少对昂贵的基于文本的标注的依赖。

排序理由 该集群描述了一篇在arXiv上发表的研究论文,其中详细介绍了一种新的语言模型训练方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

原始视频中期训练提升LLM在视觉任务上的性能

本文如何被排名

Signal score
20 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一篇在arXiv上发表的研究论文,其中详细介绍了一种新的语言模型训练方法。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Jaedong Hwang, Xiaoqian Shen, Ernie Chang, Changsheng Zhao, Chong Zhou, Saksham Suri, Qi Qian, Zechun Liu, Lemeng Wu, Qinsi Wang, Raghuraman Krishnamoorthi, Wei Wen ·

    在原始视频上进行模型中期训练

    arXiv:2610.11019v1 Announce Type: cross Abstract: Multimodal large language models learn mostly from paired image-text data or annotated video, and raw web video is rarely used to further train an existing language model. We study whether raw video, with no captions and no text l…