Researchers have introduced MAxBench, a new evaluation framework designed to assess multinomial concept representations in language models. This framework is geometry-agnostic and samples from recovered concept representations to compare different localization methods. The study found that affine subspaces are more reliable for steering and have better recall than rank-one or linear subspaces, with non-zero offsets contributing significantly to this advantage. Manifold steering also proved competitive, though no method consistently outperformed prompting. AI
IMPACT Introduces a new benchmark for evaluating fine-grained control and steerability in language models, potentially advancing interpretability research.
RANK_REASON The cluster contains a research paper detailing a new evaluation framework for language models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →