Researchers have developed two unsupervised learning strategies, TSDAE and contrastive sentence embedding (CSE), to create effective sentence embeddings for ancient languages. These methods adapt existing language models using only raw text, outperforming multilingual and specialized baselines on tasks like reuse identification and correspondence retrieval in Latin and Ancient Greek texts. The adapted encoders are particularly effective even with noisy, automatically transcribed text, and a tool called Paraphrasis is available for non-specialists to use the full pipeline. AI
IMPACT Enables semantic analysis of historical texts, potentially unlocking new insights in digital humanities and linguistics.
RANK_REASON Academic paper detailing new unsupervised learning methods for NLP tasks on ancient languages. [lever_c_demoted from research: ic=1 ai=1.0]
Read on arXiv cs.IR (Information Retrieval) →
- Ancient Greek
- Athanasius
- Augustine
- Controllable Spoken Dialogue Generation
- Jerome
- Latin
- Paraphrasis poetica in historiam Ionae prophetae
- TSDAE
- Uniform Manifold Approximation and Projection
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →