PulseAugur
EN
LIVE 07:58:41

LLMs improve biomedical translation for low-resource Arabic-script languages

Researchers have explored cross-lingual transfer learning to improve machine translation for low-resource Arabic-script languages in the biomedical domain. By using Arabic and Persian as pivot languages, they fine-tuned small decoder-only LLMs with LoRA adapters. The study evaluated three transfer strategies: few-shot learning, minimal supervised adaptation, and zero-data LoRA adapter merging, finding that supervised adaptation with only 500 sentences yielded significant improvements for Dari and Urdu. Adapter merging proved effective for closely related languages without requiring target-language biomedical data, though Pashto and Sorani Kurdish showed limitations due to structural distance from the pivot languages. AI

IMPACT This research could significantly improve access to biomedical information for speakers of low-resource languages, potentially aiding global health initiatives.

RANK_REASON The cluster contains an academic paper detailing a new methodology for machine translation. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

LLMs improve biomedical translation for low-resource Arabic-script languages

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Abdullah Alabdullah, Arash Eslamighayour, Sarp Harbalioglu, Lifeng Han ·

    Biomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter Merging

    arXiv:2607.22300v1 Announce Type: new Abstract: We present a systematic study of healthcare-domain cross-lingual transfer to address the scarcity of biomedical NMT resources for Arabic-script languages. We use Arabic and Persian as higher-resource pivots to improve translation fo…