SWE-bench Lite
PulseAugur coverage of SWE-bench Lite — every cluster mentioning SWE-bench Lite across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New methods compress LLM agent context for improved security and efficiency
Researchers have developed new methods for compressing context in large language model (LLM) agents to improve efficiency and security. One approach, "Twin Agent," separates agents into an "Explore Agent" for untrusted …
-
BLAgent framework enhances file-level bug localization with agentic RAG
Researchers have developed BLAgent, a novel agentic retrieval-augmented generation (RAG) framework designed to improve file-level bug localization in software maintenance. BLAgent integrates code structure-aware encodin…
-
New AI memory layer ContextSniper cuts token use for code repair
Researchers have developed ContextSniper, a new token-efficient memory layer for AntTrail's AI agent designed to improve repository-level program repair. ContextSniper precisely selects and ranks code and runtime eviden…
-
New SHERLOC framework boosts LLM code repair efficiency and accuracy
Researchers have developed SHERLOC, a novel framework designed to improve the efficiency and accuracy of Large Language Model (LLM) agents in code repair tasks. This training-free framework utilizes a reasoning LLM with…
-
Multi-agent LLM system Phoenix automates GitHub issue resolution
Researchers have developed Phoenix, a multi-agent LLM system designed to automatically resolve GitHub issues. The system utilizes six specialized agents, including a planner, coder, and tester, to manage the process fro…
-
Phoenix LLM system automates GitHub issue resolution with safety controls
A new multi-agent LLM system named Phoenix has been developed to automate the resolution of GitHub issues, from initial triage to the creation of pull requests. This system incorporates seven layers of safety controls a…
-
New Autopilot firewall drastically cuts LLM agent fabrication
Researchers have developed a new execution model called Autopilot designed to prevent large language model agents from fabricating success when operating without human supervision. This system acts as a firewall by exte…
-
AI agents evaluated for goal-directedness and state binding
Two new research papers explore the internal workings and evaluation of language agents. The first paper introduces a "causal state binding" framework to assess if agents' actions are truly driven by relevant internal s…
-
ARISE toolset enhances AI agents for code fault localization and repair
Researchers have developed ARISE, a new system designed to improve the accuracy of AI agents in localizing and repairing software faults. ARISE enhances large language models by providing a detailed program graph that i…