PulseAugur
EN
LIVE 10:44:08
ENTITY GPT-2 small

GPT-2 small

PulseAugur coverage of GPT-2 small — every cluster mentioning GPT-2 small across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
31
31 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
29
29 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

5 day(s) with sentiment data

RECENT · PAGE 1/2 · 31 TOTAL
  1. TOOL · CL_277298 ·

    New theory explains language model self-repair via counterweights

    Researchers have proposed a new framework for understanding self-repair in language models, suggesting that interventions on model components can be viewed as points on a coordinate axis representing a counterfactual co…

  2. TOOL · CL_277203 ·

    New research questions causal claims from language model attention-head ablations

    A new research paper published on arXiv investigates the reliability of attention-head ablations in language models for making causal claims about component functions. The study, using GPT-2 small and DistilGPT2, demons…

  3. RESEARCH · CL_268263 ·

    New research advances causal inference methods for AI models · 6 sources tracked

    Researchers are developing new methods to improve causal inference in machine learning models, particularly when dealing with unmeasured confounding factors. One approach, Hidden-Pathway Contribution (HPC), aims to dist…

  4. TOOL · CL_254408 ·

    New framework offers formal guarantees for LLM interpretability

    A new formal verification framework has been developed to address the fragility of mechanistic interpretability in large language models. Researchers demonstrated that minor input changes can drastically alter the inter…

  5. TOOL · CL_251599 ·

    AuroraAI Research releases Aurora1.0-150M language model

    AuroraAI Research has released Aurora1.0-150M, a new 150 million parameter language model. The model's performance is comparable to GPT-2 Small, with benchmark scores including 62.24% on PIQA and 32.20% on Hellaswag. It…

  6. TOOL · CL_244910 ·

    New PAC Privacy Method Enhances Secure Autoregressive Generation

    Researchers have developed a new method for ensuring privacy in autoregressive language generation, a technique previously limited to classification tasks. This approach, called PAC-Private Autoregressive Generation, ca…

  7. RESEARCH · CL_235425 ·

    New ObserverBench framework evaluates AI interpretability for interventions

    Researchers have introduced ObserverBench, a new framework designed to evaluate the effectiveness of internal estimators, or "observers," in guiding AI interventions and safety monitoring. The benchmark distinguishes be…

  8. TOOL · CL_229319 ·

    New method traces distinguishability in transformers using stochastic LayerNorm

    Researchers have developed a new method to analyze the internal workings of transformer models by introducing stochastic Layer Normalization. This modification allows for the statistical distinguishability of representa…

  9. TOOL · CL_226011 ·

    User creates interactive game to explore GPT-2 Small model internals

    A user has created an interactive "game" that allows others to explore the internal workings of the GPT-2 Small model. Participants can suggest neuron IDs, and the user will report back on what they find within the mode…

  10. TOOL · CL_221392 ·

    Gated Recurrent Transformer offers depth and efficiency over standard models

    Researchers have introduced a novel architecture called the Gated Recurrent Transformer, which aims to improve the expressivity and memory efficiency of transformer models. This new design reuses a shared core across mu…

  11. TOOL · CL_213067 ·

    Weight tying in language models: Gradient analysis and performance impact

    Researchers explored the implications of weight tying in language models, specifically how tying the input embedding and output projection matrices affects gradient calculations and model performance. They found that de…

  12. TOOL · CL_211995 ·

    Mechanistic Tomography framework unifies AI model interpretability methods

    Researchers have introduced "Mechanistic Tomography," a framework for interpretability in AI models. This approach unifies various measurement techniques like patching and Hessian-vector products under a shared mathemat…

  13. TOOL · CL_208541 ·

    FishBack method improves transformer activation steering using non-Euclidean geometry

    Researchers have developed a new method called FishBack to improve activation steering in transformers, a technique for modifying language model behavior without updating parameters. Existing methods are often unstable …

  14. TOOL · CL_206280 ·

    RecurrentGPT introduces recurrent modulation for transformer efficiency

    Researchers have introduced RecurrentGPT, a novel transformer architecture designed to enhance expressivity and memory efficiency in large language models. This model utilizes recurrent modulation, allowing a shared cor…

  15. TOOL · CL_203878 ·

    AI Interpretability Evidence Unreliable for Regulatory Compliance, Study Finds

    A new research paper argues that evidence derived from mechanistic interpretability, a method used to understand AI decision-making, is not reliable enough to meet regulatory requirements. The study found that even with…

  16. TOOL · CL_196113 ·

    New method uses Koopman operator for model interpretability

    Researchers have developed a new method for mechanistic interpretability called "Intrinsic Structure" that uses the Koopman operator to analyze the spectral properties of a model's internal dynamics. This approach aims …

  17. RESEARCH · CL_185428 ·

    New research probes MUON optimizer's convergence and proposes MALT extension

    Two new research papers explore the MUON optimization algorithm, a method used in training large language models. The first paper introduces MALT, an extension of MUON that incorporates lightweight diagonal precondition…

  18. RESEARCH · CL_154389 ·

    Researchers explore adaptive depth and cyclic folding for Transformer optimization

    Two new research papers explore novel approaches to optimizing Transformer models by dynamically adjusting their depth. The first paper, "Adaptive Depth in Looped Transformers," investigates learned halting gates and tr…

  19. TOOL · CL_150221 ·

    GPT-2 Small embedding geometry around "Trump" analyzed

    Researchers explored the embedding geometry of the token "Trump" within the GPT-2 Small model's static embedding table. By analyzing nearest neighbors under both discretized and continuous representations of the token's…

  20. RESEARCH · CL_135237 ·

    New framework enhances statistical rigor for AI model interpretability

    Researchers have developed Certified Interventional Fidelity (CIF), a new statistical framework designed to rigorously evaluate causal claims in mechanistic interpretability. CIF treats evaluation metrics as causal esti…