Researchers have introduced the APEX-VW Corpus, a new dataset designed for document-level English-Spanish post-editing in the healthcare domain. This corpus is built from recent NHS virtual-ward documents and incorporates professional post-editing performed within Trados Studio, ensuring controlled machine translation, terminology, and quality assurance. Unlike previous sentence-level datasets, APEX-VW preserves document order and CAT-tool context, making it suitable for studying terminology normalization and correction propagation in realistic translation workflows. AI
IMPACT This dataset could advance research in machine translation post-editing and terminology normalization for specialized domains.
RANK_REASON The cluster contains a research paper introducing a new dataset. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →