PulseAugur
EN
LIVE 11:41:54

New CzechDocs dataset aids format-preserving machine translation research

Researchers have introduced CzechDocs, a new dataset designed for evaluating machine translation systems that preserve document formatting. This multiway parallel dataset includes documents in Czech and several minority languages such as Ukrainian, English, and Vietnamese, formatted in HTML, DOCX, and PDF. The dataset aims to advance research in document-level translation, with a validation split and evaluation toolkit already released, and a test split planned for a future shared task. AI

IMPACT Facilitates the development of machine translation systems capable of maintaining document formatting, crucial for accurate cross-lingual information retrieval.

RANK_REASON The cluster describes a new dataset released on arXiv for research purposes.

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New CzechDocs dataset aids format-preserving machine translation research

COVERAGE [2]

  1. arXiv cs.CL TIER_1 English(EN) · Josef Jon, Ond\v{r}ej Bojar ·

    CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia

    arXiv:2606.20212v1 Announce Type: new Abstract: We present CzechDocs, a multiway parallel dataset of formatted documents (HTML, DOCX, and PDF) covering Czech and minority languages used in Czechia-primarily Ukrainian and English, with smaller portions of Vietnamese, Russian and o…

  2. arXiv cs.CL TIER_1 English(EN) · Ondřej Bojar ·

    CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia

    We present CzechDocs, a multiway parallel dataset of formatted documents (HTML, DOCX, and PDF) covering Czech and minority languages used in Czechia-primarily Ukrainian and English, with smaller portions of Vietnamese, Russian and other languages. The dataset is designed to suppo…