Researchers are exploring the intersection of psychology and large language models (LLMs) to understand their internal representations and behaviors. One study adapted Funder's personality triad framework to analyze LLMs, using SAE decomposition to identify and control trait-like internal features that influence behavior across different situations. Another study applied psychological methods, specifically prompt-based adaptations of the Implicit Association Test, to investigate sentiment differences in LLMs like ChatGPT across racial conditions, though findings on bias were weak and analysis-dependent. AI
IMPACT These studies suggest that psychological frameworks can be used to probe and potentially control LLM behavior, offering insights into their internal workings and potential biases.
RANK_REASON The cluster contains two academic papers exploring LLM behavior and internal representations using psychological frameworks and methods.
Read on Hugging Face Daily Papers →
- alphaXiv
- CatalyzeX
- Connected Papers
- DagsHub
- Funder
- Gotit.pub
- Hugging Face
- Litmaps
- LLMs
- SAE decomposition
- ScienceCast
- scite Smart Citations
- ChatGPT
- GPT-3.5T
- GPT-4
- GPT-4T
- Implicit Association Test
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →