Logit Lens
PulseAugur coverage of Logit Lens — every cluster mentioning Logit Lens across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
New LLM-Microscope tool reveals punctuation's hidden role in transformer context
Researchers have developed LLM-Microscope, a toolkit designed to analyze how large language models process and retain contextual information. The tool reveals that seemingly minor tokens like punctuation and determiners…
-
New Jacobian Lens Tool Tests if AI Models Use Internal Signals
Researchers have developed a new interpretability tool called the Jacobian lens, designed to determine if a model's internal signals are actively used in its decision-making process. Unlike previous methods like the log…
-
New method reveals stable readout features in language models
Researchers have introduced Sparse Readout Prism (SRP), a novel method to analyze language model internal states by decomposing the readout matrix into sparse features. This approach aims to decouple the analysis of hid…
-
Anthropic's Jacobian Lens Offers New Insight into AI Model Hidden Layers
Researchers have detailed a new interpretability tool called the Jacobian lens (J-lens), developed by Anthropic. This tool addresses limitations of previous methods like the logit lens, which struggled to accurately int…
-
New research questions superposition in language models
A new paper analyzes the phenomenon of "superposition" in language models, where multiple solutions might be maintained simultaneously within a single representation. Researchers investigated this using three different …
-
LLMs struggle to decode mythological knowledge beyond dominant traditions
A new research paper investigates how 18 open-source large language models represent and decode mythological knowledge. The study found that while models can represent knowledge about deities like Zeus, Jupiter, and Tho…
-
Anthropic unveils J-lens for debugging LLM internal states
Anthropic has developed a new interpretability technique called the Jacobian Lens (J-lens) to better understand the internal workings of large language models. This tool provides insights into intermediate concepts and …
-
Anthropic unveils J-Lens to visualize LLM internal thought processes
Anthropic has introduced a new interpretability technique called the Jacobian Lens (J-Lens) to visualize the internal thought processes of its large language models, specifically Claude. This J-Lens reveals a hidden "J-…
-
Anthropic paper introduces J-space as LLM 'global workspace'
Anthropic has released a paper detailing a new interpretability technique called the Jacobian Lens, which identifies a 'J-space' within language models. This J-space appears to function as a global workspace, holding ve…
-
Research: LLM uncertainty doesn't alter inference dynamics
A new research paper explores how large language models (LLMs) process uncertainty during inference. The study, utilizing a variant of the Logit Lens called Tuned Lens, analyzed layer-wise probability trajectories acros…
-
Speech-language models implicitly transcribe spoken words, study finds
A new research paper published on arXiv explores the internal workings of interleaved speech-language models (SLMs). The study reveals that these models, even when not explicitly trained for speech recognition, undergo …
-
New Query Lens method enhances AI model interpretability
Researchers have introduced Query Lens, a new method designed to improve the interpretability of sparse features in AI models. This technique extends existing approaches by analyzing both the input features that activat…