information retrieval
PulseAugur coverage of information retrieval — every cluster mentioning information retrieval across labs, papers, and developer communities, ranked by signal.
9 day(s) with sentiment data
-
New framework converts AI research agent trajectories into auditable evidence
Researchers have developed a new framework for converting experimental trajectories into auditable evidence for industrial research agents. This system verifies artifacts, qualifies claims, and consolidates evidence acr…
-
New dataset PROSLEX enhances AI's legal reasoning for statute prediction
Researchers have introduced PROSLEX, a new dataset designed to improve legal statute prediction by incorporating expert-annotated legal reasoning. This dataset, containing 1,623 Indian legal documents with statute predi…
-
RAG's roots traced to early 2000s IR research, not LLMs
A new paper argues that Retrieval-Augmented Generation (RAG), often seen as a novel LLM paradigm, has deep roots in earlier information retrieval and question answering research. The authors trace RAG's core concepts, s…
-
New method improves reliability of AI relevance decisions
Researchers have developed a new method called Label-wise Monotone Reliability Projection (MRP) to improve the reliability of relevance decisions in information retrieval systems. Unlike standard calibration techniques …
-
Music recommender calibration fails to impress users in new study
A new user study on music recommendation systems reveals that while users can perceive differences in the popularity composition of recommended track lists, they do not consistently prefer calibrated lists. The research…
-
Research paper argues classic IR metrics fail LLM agents
A research paper contends that traditional metrics for information retrieval are inadequate for evaluating the performance of Large Language Model (LLM) agents. The authors argue that these established metrics fail to c…
-
New Legal Information Retrieval Dataset Launched for Paragraph-Level Citations
Researchers have introduced LegalPincite, a new dataset designed to improve legal information retrieval by providing paragraph-level citation annotations. This dataset, derived from Court of Justice of the European Unio…
-
RecHarness automates recommender model optimization with bandit-routed agents
Researchers have developed RecHarness, a novel system designed to automate the optimization of recommender models. This system employs a bandit-routed agentic harness that separates the process into selecting modificati…
-
New e-commerce search system boosts item discoverability using LLMs · 3 sources tracked
Researchers have developed a new system for e-commerce search that enhances item discoverability by generating related user intents. This two-stage architecture uses large language models for common queries and a fine-t…
-
MediaWiki Code2Code Search improves semantic code discovery
Researchers have developed MediaWiki Code2Code Search, a novel neural retrieval system designed to improve semantic code discovery within large software ecosystems. This system indexes over 1.29 million structural entit…
-
arXiv paper questions cosine similarity's effectiveness in AI representations
A new arXiv paper titled "Semantics at an Angle: When Cosine Similarity Works Until It Doesn't" critically examines the widespread use of cosine similarity in machine learning representations. The paper, authored by Kis…
-
Research explores user control's impact on news filter bubbles
A new research paper explores the impact of user control over news recommendation systems on the formation of filter bubbles. The study designed a system that exposes inferred political and topical interests, allowing u…
-
New CwA method optimizes vector search by jointly learning partitions and probing functions
A new research paper introduces CwA (Cluster with Auctions), a method that jointly learns a balanced database partition and a neural probing function for large-scale approximate nearest neighbor search. This approach op…
-
Paper explores pre-trained word embeddings for improved answer selection
This paper explores the use of pre-trained word embeddings to enhance answer selection methods in information retrieval. The research demonstrates that integrating these embeddings can capture semantic relationships bet…
-
LLM-powered agentic system enhances Connected TV content discovery
Researchers have developed an LLM-powered agentic recommendation system for Connected TV (CTV) content discovery. This system aims to overcome limitations in traditional recommendation models by using LLMs to process di…
-
LLMs for RDF Dataset Search: Balancing Retrieval Effectiveness and Faithfulness
A new research paper explores the use of Large Language Models (LLMs) to generate metadata for RDF datasets, aiming to improve dataset searchability. The study evaluated six different metadata generation approaches, ass…
-
New Relevance-Based Embeddings Improve Candidate Retrieval in ML
Researchers have introduced a novel method for candidate retrieval in machine learning applications, termed Relevance-Based Embeddings. This approach aims to improve the efficiency of retrieving relevant items for a que…
-
New hybrid LLM approach optimizes query augmentation for information retrieval
A new research paper introduces a hybrid approach to query augmentation for information retrieval, merging prompting-based and reinforcement learning (RL) methods. The study found that simple, training-free query augmen…
-
New attack can identify hidden embedding models in AI systems
Researchers have developed a new method called an Embedding Inference Attack (EIA) that can identify the specific embedding model used by a black-box information retrieval system. This attack is effective even when the …
-
New methods enhance unsupervised cross-modal retrieval with limited data · 4 sources tracked
Researchers are developing new methods for unsupervised cross-modal retrieval, aiming to improve efficiency and reduce reliance on large, manually annotated datasets. Papers propose techniques like Attribute-Prompted Ke…