LongMemEval
PulseAugur coverage of LongMemEval — every cluster mentioning LongMemEval across labs, papers, and developer communities, ranked by signal.
- used by Long Context Modeling 90%
- instance of Gotit.pub 90%
- instance of Mem0 Agent Memory Framework 90%
- instance of alphaXiv 90%
- instance of ScienceCast 90%
- instance of locomotive class 90%
- instance of Long Context Modeling 70%
- used by DagsHub 70%
- instance of CatalyzeX 70%
- instance of DagsHub 70%
- competes with Mem0 Agent Memory Framework 60%
6 day(s) with sentiment data
-
New LLM memory systems tackle long-term conversation challenges · 5 sources tracked
Researchers are developing novel methods to manage and retrieve information from long-term conversational memory in Large Language Models (LLMs). These approaches aim to overcome the limitations of full-context injectio…
-
KVMem virtualizes million-token AI agent workspaces on consumer GPUs
Researchers have developed KVMem, a system designed to manage large context windows for AI agents, enabling them to operate with up to one million tokens on consumer-grade GPUs. This virtualization technique stores over…
-
Developer creates benchmark for AI companion apps to combat affiliate spam
A developer has created a benchmark called companion-bench to evaluate AI companion applications, addressing the lack of objective testing for these apps. Unlike benchmarks that test raw language models, companion-bench…
-
New methods enable LLMs to compress context efficiently
Researchers have developed new methods for compressing context in large language models, allowing them to process more information efficiently. FlexComp, a framework from arXiv, enables a single model to handle variable…
-
New method supervises LLM agent memory using audit trails
Researchers have introduced Hindsight Memory-PRM, a novel method for supervising memory management in long-horizon Large Language Model (LLM) agents. This approach leverages the audit trail of retrieval hits and answer-…
-
New PostgreSQL-Native Graph RAG Engine Improves Temporal Accuracy
Researchers have developed post-graph-rag, an open-source engine designed to improve the efficiency and accuracy of graph-based Retrieval Augmented Generation (RAG) systems. This new engine integrates embeddings, a cano…
-
New SearchWiki framework learns to navigate knowledge wikis for active information seeking
Researchers have developed SearchWiki, a framework designed to synthesize a corpus into a structured, navigable knowledge wiki. This system trains an agent, WikiResearcher-9B, to actively seek information through multi-…
-
New methods compress long contexts for LLMs beyond text and vision · 2 sources tracked
Two new research papers introduce novel methods for compressing long contexts in large language models. LatentPress utilizes continuous memory tokens, bypassing text reconstruction for faster and more efficient processi…
-
AI agents face memory loss due to rapid model churn; Uteke offers persistent memory solution
The rapid release cycle of AI models, exemplified by Qwen's five releases in 36 days, creates a significant maintenance burden for AI agents that rely on model-specific context for memory. This "churn tax" means prompts…
-
AI context windows expand, but dedicated memory systems remain crucial
AI models' context windows are expanding, but this does not equate to true memory systems. While larger windows offer more immediate workspace, they do not inherently solve issues of information persistence, retrieval, …
-
ArborMem framework enhances LLM memory for complex conversations
Researchers have introduced ArborMem, a novel memory framework designed for large language models to manage complex conversational states. ArborMem represents conversations as a navigable forest of interaction states, a…
-
LLM context window research shows more is not better for agents
Recent research indicates that increasing the context window size for LLM agents does not necessarily improve performance and can, in fact, degrade it. Studies show that models struggle to effectively utilize vast amoun…
-
Oracle agent memory system slashes token use, boosts LongMemEval accuracy
Researchers have developed a new memory system for Oracle agents that significantly reduces token usage. This system achieved 93.8% accuracy on the LongMemEval benchmark, using 10.7 times fewer tokens compared to tradit…
-
Memora memory system balances abstraction and specificity for AI agents · 2 sources tracked
Researchers have introduced Memora, a novel memory representation system designed to balance abstraction and specificity for AI agents. This system organizes information by abstracting primary concepts that index concre…
-
Microsoft unveils Memora memory system for AI agents
Microsoft Research has introduced Memora, a novel memory system designed to enhance the capabilities of AI agents in long-horizon tasks. Memora addresses the stateless nature of current AI models by decoupling memory co…
-
NLL-Guided Layer Selection Optimizes LLM Long-Context Efficiency
Researchers have developed a novel training-free method called NLL-guided layer selection to optimize the efficiency of long-context LLMs. This technique identifies which layers of a hybrid attention model should retain…
-
VEKTOR Slipstream beats GPT-4 on local memory benchmark
VEKTOR Slipstream, a local agent memory framework, achieved a 79% score on the LongMemEval benchmark, outperforming full-context GPT-4 by 12 points. This benchmark specifically tests real-world memory retrieval failures…
-
LLM Memory Systems Outperform Full Context on Long Histories
A new benchmark, LongMemEval, has demonstrated that retrieval-based memory systems outperform full-context baselines for LLM agents dealing with long conversation histories. While full context remains competitive for sh…
-
Regimes system improves AI agent reliability with auditable improvement loops
Researchers have developed a new system called Regimes that enhances the trustworthiness of autonomous AI improvement loops. This system uses an event-sourced agent runtime to log all changes, allowing for auditable dia…
-
AI Memory Systems Can Harm Performance, Research Finds
New research indicates that AI memory systems, while intended to improve user experience and task completion, can paradoxically degrade model performance and foster sycophantic tendencies. Studies show that these system…