PulseAugur
中
实时 02:56:45
English(EN) Learning Latent Protein Languages for Autoregressive Generation

新的潜在蛋白质语言提升自回归生成模型

研究人员开发了两种新颖的潜在蛋白质语言:蛋白质潜在语言(PLL)和结构潜在语言(SLL),旨在改进用于蛋白质序列和结构生成的自回归Transformer模型。PLL将序列映射到源自ESM-2编码器的上下文字母表,而SLL则调整VQ-VAE模型来表示骨架几何。当作为自回归模型(PLLM和SLLM)进行训练时,与传统的氨基酸分词相比,这些语言在计算扩展指数方面有所提高,并显著减少了低复杂度生成。SLL在序列到结构预测的验证困惑度方面也有显著降低,并生成了多样化、新颖的骨架结构。 AI

影响 这些潜在语言可以通过提高生成模型的效率和质量来加速蛋白质的设计和发现。

排序理由 该集群描述了一篇介绍蛋白质生成新方法的论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的潜在蛋白质语言提升自回归生成模型

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一篇介绍蛋白质生成新方法的论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
5 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    学习潜在蛋白质语言以进行自回归生成

    Autoregressive transformers remain comparatively weak for protein sequence and structure generation. We study the role of target representation: amino acid tokens encode residue identities without explicit contextual semantics, while backbone coordinates require a discrete repres…