This paper introduces a novel approach to testing deep learning-based image retrieval systems by utilizing image augmentation techniques as test generators. The research categorizes 50 augmentation methods and empirically evaluates their effectiveness in generating diverse test cases. Experiments conducted on CIFAR-10, ImageNet-1K, and March Networks datasets, using Amazon Titan and OpenCLIP models, reveal that weather simulation and SaSPA techniques yield the highest embedding uncertainty and failure rates while maintaining realism. Conversely, GAN-based augmentations showed lower realism due to synthetic artifacts. AI
IMPACT Provides practical guidelines for selecting augmentation techniques to improve test diversity and realism for image retrieval systems.
RANK_REASON The item is a research paper detailing a new methodology for testing deep learning systems. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- Amazon Titan
- CatalyX
- CIFAR-10
- DagsHub
- generative adversarial network
- Hugging Face
- ImageNet-1K
- LLaVA
- March Networks
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →