PulseAugur
EN
LIVE 07:23:42

New Korean instruction corpus released, generated by Phi-3.5-MoE-instruct

Researchers have released GLAN-QnA-KR, a large-scale Korean instruction-QA corpus containing over 300,000 rows. This corpus was generated using Microsoft's Phi-3.5-MoE-instruct model and a seedless taxonomy-driven synthesis pipeline. Notably, the data exhibits a low rate of duplicate questions and has undergone contamination audits against several Korean benchmarks, showing minimal overlap with test sets. AI

IMPACT Provides a large, high-quality dataset for training Korean language models, potentially improving their performance on various tasks.

RANK_REASON Release of a new academic paper detailing a synthetic instruction corpus. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New Korean instruction corpus released, generated by Phi-3.5-MoE-instruct

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Daekeun Kim ·

    GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction Corpus

    arXiv:2607.20443v1 Announce Type: new Abstract: We release GLAN-QnA-KR, a 303,581-row openly redistributable Korean instruction-QA corpus produced via the seedless taxonomy-driven GLAN synthesis pipeline with Microsoft's Phi-3.5-MoE-instruct as the producer model (generation: 202…