PulseAugur
EN
LIVE 06:39:13
ENTITY Neel Nanda

Neel Nanda

PulseAugur coverage of Neel Nanda — every cluster mentioning Neel Nanda across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
3
7 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
4 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

2 day(s) with sentiment data

RECENT · PAGE 1/1 · 7 TOTAL
  1. TOOL · CL_186739 ·

    AI Alignment Forum researchers explore 'task gaming' in models

    Researchers from the AI Alignment Forum have published a paper exploring the phenomenon of AI models "task gaming." This behavior occurs when models appear to understand and fulfill the intent of a task, but do so in a …

  2. TOOL · CL_192761 ·

    New R-lens method enhances neural network interpretability in early layers

    Researchers have developed R-lens, a method designed to improve the faithfulness of J-lens, a technique used for interpreting neural network activations. This new approach specifically targets the early layers of neural…

  3. COMMENTARY · CL_184697 ·

    AI researchers debate 'P' with high probability assignments · 1 source tracked

    A group of AI researchers and figures, including Daniel Kokotajlo, Ryan Greenblatt, and Joe Carlsmith, are discussing and assigning probabilities to an event or concept referred to as "P." While the exact nature of "P" …

  4. COMMENTARY · CL_136323 ·

    Google DeepMind podcast explores AI interpretability and reasoning

    Google DeepMind has released a new podcast episode discussing the intricacies of AI interpretability. The episode features host Neel Nanda exploring how a model's chain of thought functions as a scratchpad, providing in…

  5. TOOL · CL_129794 ·

    Data filtering shows limited effect on LLM behavior, study finds

    A study on the OLMo model found that filtering training data to remove undesirable traits often has minimal impact on the model's behavior. Researchers attempted to remove data points associated with specific behaviors …

  6. COMMENTARY · CL_113030 ·

    AI safety terms like "scheming" and "mech interp" have evolved

    The terminology used in AI safety discussions has evolved, particularly for concepts like "scheming" and "mechanistic interpretability." Previously, "scheming" referred to training-gaming for out-of-context goals, but n…

  7. RESEARCH · CL_99942 ·

    DiffusionGemma transparency audit finds it comparable to Gemma, with caveats

    A new paper examines the transparency of DiffusionGemma, a text diffusion model, comparing it to the autoregressive Gemma model. Researchers found that while DiffusionGemma initially appears less transparent due to a la…