PulseAugur
中
实时 21:36:00
English(EN) No wonder Qwen and Gemma are so different

Qwen和Gemma的tokenization差异影响代码和语言任务

一位r/LocalLLaMA用户观察到Qwen 35B A3B和Gemma 26B A4B在tokenizing代码方面存在显著差异。Qwen将一段330行的HTML/JS代码片段处理成1609个token,而Gemma将相同的输入tokenized为4258个token。这种差异可能解释了Qwen在代码任务上的优势以及Gemma在语言处理方面的优势,因为Qwen似乎将代码视为一种不同的输入类型,而Gemma则像处理自然语言一样将其分解。用户还指出,对于一个较短的指令文档,tokenization计数几乎相同。 AI

影响 强调了tokenization策略如何影响模型在特定任务(如编码与自然语言处理)上的性能。

排序理由 用户观察和现有模型分析,并非新发布或研究论文。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Qwen和Gemma的tokenization差异影响代码和语言任务

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Commentary
用户观察和现有模型分析,并非新发布或研究论文。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
54 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/WhoRoger ·

    难怪 Qwen 和 Gemma 如此不同

    <!-- SC_OFF --><div class="md"><p>Pasted the same HTML/JS code (330 lines) into Qwen 35B A3B and Gemma 26B A4B.</p> <p>Qwen: tokenized the input to 1609 tokens</p> <p>Gemma: tokenized the input to 4258 tokens.</p> <p>Damn. I've never noticed this before and I haven't seen people …