Researchers have introduced LCSHBench, a new benchmark dataset designed to evaluate automated subject cataloging for Library of Congress Subject Headings (LCSH). The dataset comprises 22,346 books in 15 languages, sourced from Harvard, Columbia, and Princeton catalogs, with records selected only when at least two agencies agreed on the LCSH assignment. LCSHBench accounts for both exact and conceptual matches, addressing the common discrepancy where libraries agree on topics but differ in precise heading expression. AI
IMPACT Provides a standardized evaluation for multilingual subject cataloging, potentially improving AI's ability to organize and retrieve information across languages.
RANK_REASON The cluster describes a new benchmark dataset for a specific NLP task.
Read on arXiv cs.IR (Information Retrieval) →
- Harvard University
- LCSHBench
- Library of Congress Subject Heading Assignment
- Columbia University
- Princeton University
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →