PulseAugur
实时 21:11:30
English(EN) How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems

新的基准VQABench分析了云VLM VQA系统的成本-质量

一项名为VQABench的新基准已被开发出来,用于评估基于云的视觉语言模型(VLM)在视觉问答(VQA)系统中的成本-质量权衡。研究强调,客户端输入预处理技术虽然提供了潜在的优化,但并非普遍适用于所有模型或任务。该研究分析了四种商用VLM和三个VQA数据集上的12种预处理方法,涉及超过95,000次API调用。研究结果表明,预处理的有效性高度依赖于特定的VLM、API提供商和任务表述,选择不当的策略可能会增加成本并降低准确性。 AI

影响 为优化VQA系统提供了关键见解,在基于云的VLM部署中平衡答案质量与成本和延迟。

排序理由 该集群包含一篇研究论文,详细介绍了用于评估基于VLM的VQA系统的新基准。

在 arXiv cs.IR (Information Retrieval) 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的基准VQABench分析了云VLM VQA系统的成本-质量

报道来源 [2]

  1. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Guohao Lan ·

    回答我的问题需要多少成本?对基于云的VLM视觉问答系统进行基准测试

    Vision-language models (VLMs) are becoming a practical backend for mobile visual question answering (VQA) systems, enabling smartphones and smart glasses to answer users' questions about the physical world. Since modern VLMs remain difficult to run on mobile and edge devices, VQA…

  2. arXiv cs.CV TIER_1 English(EN) · Henri Vanhuynegem, Weitao Xu, Yiran Shen, Guohao Lan ·

    回答我的问题需要多少成本?对基于云的VLM视觉问答系统进行基准测试

    arXiv:2608.07861v1 Announce Type: new Abstract: Vision-language models (VLMs) are becoming a practical backend for mobile visual question answering (VQA) systems, enabling smartphones and smart glasses to answer users' questions about the physical world. Since modern VLMs remain …