R-Judge
PulseAugur coverage of R-Judge — every cluster mentioning R-Judge across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New FGLGuard system enhances LLM multi-agent safety via federated graph learning
Researchers have developed FGLGuard, a novel system for enhancing the safety of multi-agent systems (MAS) powered by large language models (LLMs). This approach utilizes federated graph learning to train a graph neural …
-
New TRACE method improves safety detection for long-horizon LLM agents
Researchers have introduced TRACE, a novel method for enhancing the safety of long-horizon Large Language Model (LLM) agents. TRACE addresses the challenge of detecting sparse and delayed safety risks that are often mis…
-
New LLM safety judge test reveals unreliability in current evaluation methods
Researchers have introduced a new method called policy invariance to assess the reliability of LLM-based safety judges. This approach tests whether an LLM's safety verdicts are consistent regardless of how the evaluatio…