PulseAugur
EN
LIVE 01:54:34
ENTITY HotpotQA

HotpotQA

PulseAugur coverage of HotpotQA — every cluster mentioning HotpotQA across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
23
66 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
23
64 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

14 day(s) with sentiment data

RECENT · PAGE 1/4 · 66 TOTAL
  1. TOOL · CL_195933 ·

    New RLMOpt method uses recursive language models for adaptive prompt optimization

    Researchers have developed RLMOpt, a novel prompt optimization method that utilizes a recursive language model (RLM) to drive the search policy. This RLM agent operates within a tool-based environment, analyzing task in…

  2. RESEARCH · CL_195687 ·

    New method optimizes LLM prompts by using cheaper models for most tasks

    Researchers have developed a novel method for optimizing large language model (LLM) prompts and agentic programs by decoupling the LLM's roles and utilizing cross-tier transfer. This approach involves running the high-v…

  3. TOOL · CL_193280 ·

    New EL-DGR framework improves LLM judge performance in reasoning pipelines

    A new research paper introduces Evidence-Locked Derive-Gate-Repair (EL-DGR), a novel decision-making framework for LLM judges within reasoning pipelines. The study demonstrates that EL-DGR significantly improves perform…

  4. RESEARCH · CL_193013 ·

    SAGE system optimizes RAG retrieval for latency and cost · 2 sources tracked

    Researchers have developed SAGE, a new adaptive retrieval policy for production Retrieval-Augmented Generation (RAG) systems. SAGE dynamically adjusts the number of passages retrieved per query based on estimated query …

  5. TOOL · CL_187240 ·

    New HERALD system audits AI search agent rewards for manipulation

    Researchers have developed HERALD, a new offline audit system designed to evaluate and improve the reward mechanisms for search agents. HERALD uses counterfactual interventions to distinguish between candidate-visible a…

  6. TOOL · CL_185362 ·

    New research reveals "referential dangling" failure in LLM prompt compression

    A new research paper identifies a significant failure mode in hard prompt compression techniques used for large language models, termed "referential dangling." This occurs when the compression process retains text conta…

  7. TOOL · CL_183528 ·

    LLM hallucination benchmarks misleading, study finds

    A recent analysis of four open-weight large language models—Phi-4 Mini, Mistral 7B Instruct v0.3, Qwen2.5-7B-Instruct, and Llama-3.1–8B-Instruct—reveals that hallucination benchmarks may be misleading. The study found t…

  8. TOOL · CL_183247 ·

    FLARE framework optimizes LLM instructions, outperforming GEPA

    Researchers have introduced FLARE, a new framework designed to optimize instructions for large language models. FLARE utilizes advanced reflective mechanisms and a limited set of few-shot reference examples to enhance p…

  9. RESEARCH · CL_190127 ·

    Referential Dangling: A New Failure Mode in LLM Prompt Compression

    A new paper identifies a significant failure mode in hard prompt compression techniques, termed "referential dangling." This occurs when methods designed to reduce context length by selecting high-scoring text segments …

  10. RESEARCH · CL_180475 ·

    New datasets and frameworks aim to improve LLM question answering completeness

    Researchers have developed new methods to improve the completeness and quality of answers generated by large language models (LLMs) for complex questions. Apple's research introduces DeepAmbigQA, a dataset and generatio…

  11. TOOL · CL_191661 ·

    FLARE framework outperforms GEPA in optimizing LLM instructions

    Researchers have introduced FLARE, a new framework for optimizing instructions in large language models. FLARE utilizes reflective mechanisms and a small set of few-shot examples to improve performance across various be…

  12. RESEARCH · CL_180205 ·

    New RAG verification method improves multi-hop question answering

    A new research paper proposes a novel approach to improve verification in retrieval-augmented generation (RAG) systems, particularly for multi-hop question answering. The study demonstrates that traditional per-chunk fi…

  13. TOOL · CL_181225 ·

    New BANDMAS framework slashes multi-agent communication costs

    Researchers have developed BANDMAS, a novel framework for multi-agent collaboration that optimizes communication by intelligently scheduling data packets. This system analyzes semantic features of messages to determine …

  14. TOOL · CL_167175 ·

    HyCE-RAG framework uses hypergraphs for explainable multi-hop question answering

    Researchers have introduced HyCE-RAG, a novel framework for explainable multi-hop question answering that utilizes hypergraphs to model complex relationships between entities and evidence. Unlike traditional RAG methods…

  15. TOOL · CL_160823 ·

    New TopoGuard defense tackles split-knowledge attacks in RAG systems

    A new defense mechanism called TopoGuard has been developed to combat split-knowledge attacks targeting retrieval-augmented generation (RAG) systems. These attacks involve injecting seemingly benign documents that, when…

  16. RESEARCH · CL_147776 ·

    AI agents' "bridge documents" offer causal utility beyond static relevance

    A new research paper titled "Bridge Evidence" explores the discrepancy between static and causal utility in retrieval systems used by multi-step AI agents. The study found that documents deemed useful by static evaluati…

  17. RESEARCH · CL_147773 ·

    SmartRAG enables LLMs on mobile devices with graph-based RAG

    Researchers have developed SmartRAG, a novel on-device framework designed to enable large language models (LLMs) to function as personal assistants on mobile devices. This system decomposes intelligence into four module…

  18. TOOL · CL_145667 ·

    New DeepStress framework stress-tests AI search agents on unreliable data

    Researchers have introduced DeepStress, a novel framework designed to stress-test the robustness of deep search agents. This framework controls the frequency of challenging evidence by replacing the retrieval module of …

  19. RESEARCH · CL_143672 ·

    QUBO optimization tackles evidence selection in RAG question answering

    Researchers have developed a novel method for evidence selection in retrieval-augmented question answering (RAG) by formulating it as a Quadratic Unconstrained Binary Optimization (QUBO) problem. This approach aims to i…

  20. TOOL · CL_141532 ·

    New framework CAFE streamlines AI system evaluation and component analysis

    Researchers have developed CAFE (Compound-AI Factorial Evaluation), an open-source platform designed to bring design of experiments principles to the evaluation of compound AI systems. CAFE allows practitioners to regis…