Researchers have developed PetQA, a new benchmark designed to evaluate the veterinary knowledge and clinical reasoning capabilities of large language models (LLMs) and large vision-language models (LVLMs). The benchmark includes over 10,000 question-answer pairs derived from real-world queries about dogs and cats, with answers provided by veterinary experts. Initial evaluations of eighteen models using metrics like ROUGE and BERTScore revealed limitations in current AI systems for veterinary care, highlighting the need for improved adaptation methods. AI
IMPACT This benchmark could drive the development of more reliable AI systems for veterinary diagnostics and care.
RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- BERTScore
- Cats
- dog
- LLMs
- LVLMs
- PetQA
- PetQA-Bench
- retrieval-augmented generation
- ROUGE
- supervised fine-tuning
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →