A new benchmark called VQABench has been developed to evaluate the cost-quality trade-offs of cloud-based Vision-Language Models (VLMs) for visual question answering (VQA) systems. The research highlights that client-side input preprocessing techniques, while offering potential optimizations, do not universally benefit all models or tasks. The study analyzed 12 preprocessing methods across four commercial VLMs and three VQA datasets, involving over 95,000 API calls. Findings indicate that the effectiveness of preprocessing is highly dependent on the specific VLM, API provider, and task formulation, with poorly chosen strategies potentially increasing costs and reducing accuracy. AI
IMPACT Provides critical insights for optimizing VQA systems, balancing answer quality with cost and latency in cloud-based VLM deployments.
RANK_REASON The cluster contains a research paper detailing a new benchmark for evaluating VLM-based VQA systems.
Read on arXiv cs.IR (Information Retrieval) →
- alphaXiv
- arXiv
- CatalyzeX
- Cloud VLM-based VQA Systems
- DagsHub
- Gotit.pub
- Hugging Face
- ScienceCast
- Vision--Language Models
- VQABench
- Connected Papers
- CORE Recommender
- Litmaps
- scite Smart Citations
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →