LongMemEval-S
PulseAugur coverage of LongMemEval-S — every cluster mentioning LongMemEval-S across labs, papers, and developer communities, ranked by signal.
7 day(s) with sentiment data
-
MemStrata achieves 95% accuracy on long-context benchmarks with local Qwen reader
Researchers have developed MemStrata, a system that achieves high source-aware accuracy on long-context evaluation benchmarks. Using a local Qwen 3.8 27B Q4_K_M reader, MemStrata CL1 reached 95% accuracy on LongMemEval-…
-
TAGGRAPH framework evaluates LLM agent memory systems · 3 sources tracked
Researchers have developed TAGGRAPH, a novel framework for evaluating LLM agent memory systems. The system uses a controlled evaluation framework with shared conversational memories, employing localized graph configurat…
-
New research questions long-term memory evaluation methods in LLMs
A new paper on arXiv details an audit of long-term memory evaluation methods for retrieval chains. The study found inconsistencies in scoring due to reader variation and the need for repeated judging, with scores fluctu…
-
AI memory systems achieve high scores on benchmarks, one for LLMs, one for robotics
A new research paper introduces an auditable long-term memory system for AI that achieved high scores on the LongMemEval-S benchmark. Using Claude Opus as a reader, the system scored 479/500 and 475/500, performing comp…
-
Mnemon agent uses dual-system approach for LLM long-term memory
Researchers have introduced Mnemon, a novel memory agent designed for long-term context in large language models. Mnemon distinguishes between fast, judgment-based tasks (System 1) and slow, reasoning-based tasks (Syste…
-
New research tackles LLM agent memory for long-horizon tasks · 8 sources tracked
Multiple research papers explore advanced memory management techniques for large language model (LLM) agents to improve their performance on long-horizon tasks. These studies introduce novel frameworks and methods such …
-
AI agent memory frameworks detail security and reproducibility measures
Two distinct AI agent memory frameworks, skillmem and nautilus-compass, have detailed their approaches to security and reproducibility. Skillmem addressed a vulnerability where external text could be mistaken for agent …
-
Nautilus-Compass agent memory layer outperforms Mem0 on retrieval benchmarks
A new open-source memory layer for AI agents, named Nautilus-Compass, has demonstrated superior performance compared to Mem0 Agent Memory Framework on the LongMemEval-S retrieval benchmark. The Nautilus-Compass system a…
-
Fortunate Recall enhances LLM memory management with ontology-driven policies
Researchers have introduced Fortunate Recall (FR), a novel policy layer designed to improve Large Language Model (LLM) memory management. FR addresses the issue of unbounded memory growth and degrading retrieval precisi…
-
Nautilus-Compass launches LLM-free AI agent memory layer
Nautilus-Compass has released an open-source memory and reliability layer for AI agents that aims to improve long-term memory retrieval without relying on LLM extraction at write time. This approach stores raw text embe…
-
Uteke memory engine shows high consistency across CPU architectures
Codecora has published benchmark results for its open-source memory engine, Uteke, demonstrating high recall rates on the LongMemEval-S benchmark. A key finding is the engine's consistency across different CPU architect…
-
Wontopos Tablet 2 shows strong performance in multilingual and multimodal memory retrieval
Researchers have evaluated Wontopos Tablet 2, a long-term memory engine for language models, on various text retrieval benchmarks. The system achieved high scores on LongMemEval-S (95.7%) and BEAM-1M (67.5%), though the…
-
Wontopos Tablet 2 advances LLM memory retrieval without lexical matching · 2 sources tracked
A new research paper introduces Wontopos Tablet 2, a long-term memory engine for language models designed for multilingual and multimodal retrieval without relying on lexical matching. The system demonstrated strong per…
-
New framework enhances AI conversational memory with user-aware recall
Researchers have developed a new framework called Profile-guided Personalized Retrieval Optimization (PPRO) to enhance the long-term memory recall capabilities of conversational AI agents. This system creates user profi…
-
RAG compression evaluation flawed, hides model performance differences
A new research paper published on arXiv highlights a critical flaw in how Retrieval-Augmented Generation (RAG) compression is evaluated. The study demonstrates that fixed compression methods can mask significant perform…
-
Verbatim LLM conversation chunks outperform extracted facts in memory retrieval
A new research paper challenges the common practice of distilling long LLM conversations into structured artifacts like facts or events. The study found that using verbatim conversation chunks significantly outperformed…
-
DeferMem framework enhances LLM long-term memory QA with RL
Researchers have developed DeferMem, a new framework designed to improve question answering for large language model agents dealing with long-term conversational memory. This system separates the process into initial br…
-
Nautilus Compass detects LLM agent persona drift without model access
Researchers have developed Nautilus Compass, a novel system designed to detect persona drift in large language model (LLM) agents operating in production environments. This black-box method functions solely at the promp…