PulseAugur
EN
LIVE 13:51:08
ENTITY Gemma 2-2B

Gemma 2-2B

PulseAugur coverage of Gemma 2-2B — every cluster mentioning Gemma 2-2B across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
8
18 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
6
15 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

5 day(s) with sentiment data

RECENT · PAGE 1/2 · 36 TOTAL
  1. TOOL · CL_255061 ·

    New LLMs Developed for Moroccan Arabic and Arabizi Dialects

    A developer has fine-tuned two open-source large language models, SILMA-9B-Darija and SILMA-2B-Darija, to better understand and generate Moroccan Arabic (Darija) and its informal Arabizi script. These models were traine…

  2. TOOL · CL_254408 ·

    New framework offers formal guarantees for LLM interpretability

    A new formal verification framework has been developed to address the fragility of mechanistic interpretability in large language models. Researchers demonstrated that minor input changes can drastically alter the inter…

  3. RESEARCH · CL_245206 ·

    New AI alignment methods improve efficiency and multi-dimensional control · 3 sources tracked

    Researchers are developing new methods for aligning AI models with human preferences, aiming to improve efficiency and performance. One approach, DSPA, uses inference-time steering to condition alignment on prompts, sho…

  4. TOOL · CL_241951 ·

    Small Qwen3 LLM on Old Phone Controls Desktop Browser

    A demonstration showcases the Qwen3-0.6B language model, running on a 2017 Samsung Note 8, successfully controlling a desktop Google Chrome browser. The model processed structured page representations to perform tasks l…

  5. TOOL · CL_232567 ·

    Cross-model KV cache sharing promises to speed up multi-model AI inference

    Two research papers propose a method called cross-model KV cache sharing to improve the efficiency of multi-model AI inference pipelines. This technique allows the key-value states computed by one model during its initi…

  6. TOOL · CL_231597 ·

    New 'patterning' technique debiases AI reward models, shows cross-model transfer

    Researchers have developed a new technique called "patterning" to debias reward models used in AI training. This method reweights preference pairs based on their impact on benchmark losses, effectively reducing stylisti…

  7. RESEARCH · CL_229094 ·

    Hugging Face unveils efficient multimodal encoder NeoMME, study favors encoders for Indic NER

    Hugging Face has introduced NeoMME, a new family of multilingual multimodal encoders designed for efficiency. Unlike many generative models, NeoMME uses a single bidirectional Transformer to process both text and image …

  8. TOOL · CL_228940 ·

    New method enables cross-model KV state sharing for LLMs

    Researchers have developed a novel "universal context-reuse layer" that enables KV (key-value) state sharing between different large language models, even those with varying architectures, tokenizers, and scales. This c…

  9. TOOL · CL_221057 ·

    New research explores grammar's geometry in Transformer layers

    A new research paper explores the geometric properties of language representations within Transformer models. The study investigates how the intrinsic dimensionality (ID) of these representations changes across layers a…

  10. TOOL · CL_215939 ·

    LLMs fine-tuned for malaria drug discovery outperform proprietary models

    A new study introduces Malaria-Instruct, a dataset designed for malaria drug discovery using large language models (LLMs). The research evaluated several open-source LLMs, finding that fine-tuned models significantly ou…

  11. TOOL · CL_198057 ·

    LLMs exhibit congruency effects similar to human cognition in conflict tasks

    Researchers have developed a novel verbal conflict task to investigate congruency effects in large language models, drawing parallels to psychological and neuroscience studies. The task involves prompts that elicit a de…

  12. TOOL · CL_196113 ·

    New method uses Koopman operator for model interpretability

    Researchers have developed a new method for mechanistic interpretability called "Intrinsic Structure" that uses the Koopman operator to analyze the spectral properties of a model's internal dynamics. This approach aims …

  13. TOOL · CL_195168 ·

    HyperSAE uses Poincaré geometry to boost Sparse Autoencoder performance

    A new PyTorch library called HyperSAE has been developed to improve the efficiency of Sparse Autoencoders (SAEs) by employing Poincaré hyperbolic geometry. This approach addresses the limitations of standard SAEs, which…

  14. TOOL · CL_193516 ·

    New LLM fine-tuning method targets performance and carbon emission break-even

    Researchers have developed a new fine-tuning method that incorporates a differentiable energy surrogate to optimize for both performance and carbon emissions in Large Language Models (LLMs). This approach aims to achiev…

  15. TOOL · CL_191133 ·

    New Tiled SVD Method Extracts Network Mechanisms Directly From Weights

    Researchers have developed a new method called column-tiled SVD to extract usable weight mechanisms directly from linear sites within neural networks. This approach identifies concepts within the network's weights thems…

  16. TOOL · CL_165006 ·

    New training method enhances LLM interpretability by reducing signal loss

    Researchers have developed a new method called replacement-aware training to improve the interpretability of large language models. This technique trains sparse auto-encoders (SAEs) to be robust to errors introduced by …

  17. TOOL · CL_160861 ·

    Gemma 2-2B research finds active feature planes have less holonomy

    A new research paper published on arXiv investigates the concentration of holonomy within specific feature planes of the Gemma 2-2B model. The study preregistered its methodology and analysis rules before inspecting the…

  18. TOOL · CL_151945 ·

    New 'prolepsis' phenomenon identified in small transformer models

    Researchers have identified a phenomenon called 'prolepsis' in small transformer models, where the model commits to a decision early in its processing and cannot correct it. This commitment is sustained by task-specific…

  19. SIGNIFICANT · CL_100834 ·

    Google's Gemma 2 models achieve high performance with efficient architecture

    Google's new Gemma 2 models, particularly the 27B parameter version, are demonstrating significant performance gains through architectural innovations rather than just increased size. These models utilize a hybrid atten…

  20. RESEARCH · CL_99632 ·

    New research identifies actionable directions to mitigate AI model misalignment

    Researchers have identified a method to detect and mitigate emergent misalignment in language models by analyzing activation directions. This approach, tested across four model families including Qwen2.5-1.5B, Gemma-2-2…