PulseAugur
实时 07:11:54
English(EN) Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer

新的Arkios语言模型在英语-尼泊尔语文本上进行训练

研究人员推出了Arkios,一个拥有10.4亿参数的语言模型,该模型在1500亿个英语和尼泊尔语的token上进行了训练。该模型使用了自定义训练堆栈和梵文感知分词器。评估显示,Arkios在ARC-Easy和ARC-Challenge基准测试中表现优于同类开放模型,但这可能是由于与基准测试的数据重叠,而非整体能力。该研究还强调了评估低资源语言模型所面临的挑战,因为标准的单项选择题提示可能导致尼泊尔语理解的随机猜测结果。 AI

影响 推出一个新的双语模型,并强调了低资源语言的评估挑战。

排序理由 该集群描述了一篇关于创建和评估新型语言模型的新研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的Arkios语言模型在英语-尼泊尔语文本上进行训练

本文如何被排名

Signal score
24 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一篇关于创建和评估新型语言模型的新研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun ·

    Arkios:一个从头开始训练的开放式双语英语-尼泊尔语语言模型,配备了德瓦那加里文感知分词器

    arXiv:2608.30092v1 Announce Type: cross Abstract: We present Arkios, a 1.04B-parameter dense transformer pretrained from scratch on 150B tokens of bilingual English-Nepali text, using a custom single-file C/CUDA training stack and a Devanagari-aware byte-level BPE tokenizer built…