Researchers have identified a specific type of neuron in large language models called "weakening neurons" that play a significant role in the models' output. These neurons, primarily found in later layers of transformer models, exhibit a negative cosine similarity between their input and output weight vectors. Despite their relative scarcity, these weakening neurons activate frequently and exert a substantial influence on the model's behavior, particularly when gate values are negative. AI
IMPACT This research offers a new method for analyzing LLM internals, potentially leading to better model interpretability and control.
RANK_REASON The cluster contains a research paper detailing findings about the internal workings of transformer models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- IArxiv
- ScienceCast
- Sebastian Gerstner
- transformers
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →