PulseAugur
EN
LIVE 07:34:33
ENTITY Large Language Model Agents

Large Language Model Agents

PulseAugur coverage of Large Language Model Agents — every cluster mentioning Large Language Model Agents across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
8
17 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
8
17 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

7 day(s) with sentiment data

RECENT · PAGE 1/1 · 17 TOTAL
  1. TOOL · CL_206003 ·

    New benchmark tests LLM agents for aviation copilot roles

    Researchers have introduced AeroCopilotBench, a novel two-tier benchmark designed to evaluate Large Language Model (LLM) agents in aviation scenarios. This benchmark includes a virtual cockpit environment called the Aer…

  2. TOOL · CL_194695 ·

    AI agents tested for evading biological tool detection

    Researchers have explored the capabilities of Large Language Model Agents in identifying and potentially evading biological tools used in nucleic acid synthesis. The study focused on screening evasion techniques, specif…

  3. TOOL · CL_195686 ·

    New AI framework aims to value biotech assets beyond cash flows

    A new multi-agent AI framework is proposed for valuing clinical-stage, cross-border biotechnology companies, addressing the limitation of traditional cash-flow based valuation methods. This framework incorporates a valu…

  4. RESEARCH · CL_191120 ·

    New research tackles credit assignment for LLM agents · 2 sources tracked

    Two new research papers from arXiv explore advanced credit assignment techniques for large language model agents. The first paper, "From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Langua…

  5. TOOL · CL_185314 ·

    MemFly framework optimizes LLM memory using information bottleneck principles

    Researchers have introduced MemFly, a novel framework designed to optimize the long-term memory capabilities of large language models (LLMs). This system utilizes information bottleneck principles to balance efficient c…

  6. TOOL · CL_174030 ·

    LLM agents automate PID tuning for chemical processes

    Researchers have developed a novel framework that leverages Large Language Models (LLMs) to automate the tuning of PID controllers in chemical processes. This approach mimics the iterative workflow of human engineers, u…

  7. TOOL · CL_174368 ·

    MemHarness framework enables LLM agents to reconstruct past experiences

    Researchers have introduced MemHarness, a novel framework designed to enhance large language model agents by enabling them to reconstruct past experiences rather than simply replaying them. This approach, inspired by hu…

  8. TOOL · CL_167225 ·

    LLM agents fail stress tests in robotic chemistry lab

    A new study published on arXiv details the use of a robotic chemistry laboratory to stress-test large language model (LLM) agents. Researchers found that LLM agents struggled with reliable physical action and adaptation…

  9. TOOL · CL_117633 ·

    LLM agents enable interpretable inverse design of MOFs

    Researchers have developed LLM4MOF, a framework that uses large language model agents for the inverse design of metal-organic frameworks (MOFs). This system autonomously reasons about chemistry, generates candidate MOFs…

  10. RESEARCH · CL_115208 ·

    Agentic AI poses scalable threat to mobility data privacy, study finds

    A new study published on arXiv demonstrates how agentic AI, specifically large language models, can automate the re-identification of individuals from mobility microdata. The research presents a pipeline where AI agents…

  11. TOOL · CL_96097 ·

    New Benchmark Evaluates AI Map Agents' Satisfaction-Aware Decision-Making

    Researchers have introduced MapSatisfyBench, a new benchmark designed to evaluate map agents' ability to understand and satisfy users' implicit needs beyond explicit task completion. The benchmark reconstructs complete …

  12. RESEARCH · CL_84355 ·

    LLM agents guide evolutionary molecular design for drug discovery

    Researchers have developed "My Chemical Harness," a novel framework for molecular design that integrates large language models (LLMs) with evolutionary algorithms. This system uses LLMs as high-level strategy controller…

  13. TOOL · CL_63379 ·

    CUHK team introduces SLIM for dynamic LLM agent skill management

    Researchers from the Chinese University of Hong Kong have developed SLIM, a novel framework for managing the lifecycle of skills used by large language model agents. SLIM dynamically assesses the contribution of each ex…

  14. RESEARCH · CL_48703 ·

    MemAudit framework audits poisoned LLM agent memory

    Researchers have developed MemAudit, a new framework designed to identify and audit malicious data within the memory of large language model agents. This post-hoc auditing system addresses the security vulnerability whe…

  15. TOOL · CL_44858 ·

    New framework improves LLM agent performance via execution alignment

    Researchers have developed a new framework called "harnesses" to improve the performance of large language model agents during inference. This approach focuses on aligning execution trajectories by separating harness fu…

  16. RESEARCH · CL_50824 ·

    New Benchmark Tests LLM Agents' Skill Formation From Experience

    A new benchmark called SkillEvolBench has been introduced to evaluate the ability of large language model (LLM) agents to distill episodic experience into reusable procedural skills. The benchmark consists of 180 tasks …

  17. RESEARCH · CL_43964 ·

    DeferMem framework enhances LLM long-term memory QA with RL

    Researchers have developed DeferMem, a new framework designed to improve question answering for large language model agents dealing with long-term conversational memory. This system separates the process into initial br…