Researchers have released GLAN-QnA-KR, a large-scale Korean instruction-QA corpus containing over 300,000 rows. This corpus was generated using Microsoft's Phi-3.5-MoE-instruct model and a seedless taxonomy-driven synthesis pipeline. Notably, the data exhibits a low rate of duplicate questions and has undergone contamination audits against several Korean benchmarks, showing minimal overlap with test sets. AI
IMPACT Provides a large, high-quality dataset for training Korean language models, potentially improving their performance on various tasks.
RANK_REASON Release of a new academic paper detailing a synthetic instruction corpus. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- GLAN-QnA-KR
- HAE-RAE-Bench
- Hugging Face Hub
- KMMLU
- Kobestelia
- Korean
- Microsoft
- OpenRail Association
- Phi-3.5-MoE-instruct
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →