linear probes
PulseAugur coverage of linear probes — every cluster mentioning linear probes across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Paper argues detached linear probes won't improve AI interpretability
A recent paper proposes using detached linear probes within an RL optimization process to prevent models from outmaneuvering interpretability tools. However, the author argues this approach is flawed, as RL itself is de…
-
New theory links Mahalanobis Cosine Similarity to probe performance
Researchers have theoretically and empirically demonstrated that Mahalanobis Cosine Similarity (MCS) is a strong predictor of a linear probe's Out-of-Distribution AUROC. This relationship holds across various models, la…
-
New paper proposes Mahalanobis cosine similarity for probe comparison
A new paper introduces the Mahalanobis cosine similarity (MCS) as a theoretically grounded method for comparing linear probes, which are commonly used in interpretability research. Unlike standard cosine similarity, MCS…