Researchers have introduced the Chinese Data-Text Pair (CDTP), a large-scale dataset designed to evaluate Chinese-language knowledge-grounded Large Language Models (LLMs). The dataset contains over 7 million instances, pairing Chinese text with KG triples, and is intended to support tasks such as Knowledge Graph Completion (KGC), Question Answering (QA), and Triple-to-Text Generation (T2T). CDTP specifically addresses Chinese linguistic nuances like polysemy and word-segmentation ambiguity, aiming to improve structured reasoning and factual understanding in LLMs. Experiments indicate that while model scale alone is insufficient, fine-tuning on CDTP enhances performance and robustness on these knowledge-intensive tasks. AI
IMPACT Provides a specialized benchmark for evaluating and improving Chinese LLMs in knowledge-grounded tasks.
RANK_REASON The cluster describes a new academic dataset and benchmark for evaluating LLMs, published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Chinese Data-Text Pair (CDTP)
- Hugging Face
- Jiapu Wang
- Knowledge Graph Completion (KGC)
- Knowledge Graph (KG)
- Large Language Models (LLMs)
- Question Answering (QA)
- Triple-to-Text Generation (T2T)
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →