MultiWOZ
PulseAugur coverage of MultiWOZ — every cluster mentioning MultiWOZ across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New benchmark TPBench evaluates dialogue compression methods
Researchers have introduced TPBench, a new benchmark designed to evaluate dialogue compression methods by focusing on specific information targets rather than a single retention score. TPBench assesses a model's ability…
-
New MAPS framework models subjective perspectives in multi-agent dialogue
Researchers have introduced MAPS (Multi-Agent Perspective Spaces), a new framework designed to model dialogue between AI agents with distinct cognitive styles. Unlike current systems that enforce semantic uniformity, MA…
-
New "Reclaim Evaluation" reveals language models' "brittle memory" problem
Researchers have introduced "Reclaim Evaluation" to assess language models' memory capabilities, finding that a memory retaining incorrect conclusions is more detrimental than an empty one. This "brittle memory" phenome…