PulseAugur
EN
LIVE 17:40:28
ENTITY MS MARCO

MS MARCO

PulseAugur coverage of MS MARCO — every cluster mentioning MS MARCO across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
6
15 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
5
14 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

4 day(s) with sentiment data

RECENT · PAGE 1/2 · 28 TOTAL
  1. TOOL · CL_253683 ·

    DIY DSSM improves search with translation tables

    A developer has created a "poor man's" Deep Structured Semantic Model (DSSM) that uses count-based translation tables to enhance full-text search capabilities. This method enriches the inverted index by associating docu…

  2. RESEARCH · CL_242938 ·

    New Matryoshka Hash Representations improve RAG retrieval efficiency

    Researchers have developed Matryoshka Hash Representations (MHR), a novel method for compact semantic retrieval in retrieval-augmented generation (RAG) systems. MHR addresses the challenge of storing large vector indexe…

  3. TOOL · CL_239614 ·

    New 'Embedding Surgery' Technique Enhances Dense Retrieval Systems

    Researchers have developed a new technique called "embedding surgery" to improve the performance of dense retrieval systems. This method allows for localized, minimal updates to document embeddings at query time, guided…

  4. RESEARCH · CL_231512 ·

    New TRIS defense system combats knowledge poisoning in RAG models

    Researchers have developed TRIS, a Tri-Layer Retrieval Integrity Sieve, to combat knowledge poisoning in retrieval-augmented generation (RAG) systems. This middleware defense system sanitizes retrieved evidence by emplo…

  5. TOOL · CL_222857 ·

    Supervised fine-tuning impacts LLM instruction sensitivity differently by model scale

    A new study published on arXiv investigates how supervised fine-tuning (SFT) affects the instruction sensitivity of large language models. Researchers found that SFT consistently reduces instruction sensitivity in small…

  6. TOOL · CL_221080 ·

    New Turkish LLM MoganBert-TR trained with CLM-to-MLM curriculum

    Researchers have developed MoganBert-TR, a new Turkish encoder foundation model, and its accompanying embedding model, MoganBert-Embed. Trained from scratch on a filtered Turkish corpus using a novel CLM-to-MLM curricul…

  7. RESEARCH · CL_218084 ·

    New arXiv papers explore privacy, efficiency, and LLM integration in dense retrieval

    Four new arXiv papers explore advancements in dense retrieval, a key component for large language models in information retrieval tasks. The first paper introduces a privacy-preserving method using learned deep hashing …

  8. TOOL · CL_205636 ·

    Vector database admission control tackles retrieval hubs under workload drift

    Researchers have developed a new admission control mechanism for vector databases to mitigate the impact of retrieval hubs, where a single document dominates search results. This system maintains a set of sentinel queri…

  9. RESEARCH · CL_205628 ·

    New research explores static pruning and LLM-based query expansion for retrieval systems

    Two new research papers explore methods to improve information retrieval systems. The first paper, "Static Pruning Across Sparse Retrieval Regimes," investigates how static pruning techniques can be applied across diffe…

  10. TOOL · CL_199759 ·

    Vector database performance benchmarked across seven systems

    A new research paper provides a comprehensive empirical evaluation of seven prominent vector database systems, including FAISS, Qdrant, Milvus, Weaviate, Chroma, pgvector, and LanceDB. The study, which analyzed over 4 m…

  11. RESEARCH · CL_183200 ·

    New study evaluates RAG pipeline for scientific question answering · 2 sources tracked

    Researchers have introduced SciRet, a study examining retrieval-augmented generation (RAG) for scientific question answering using the CORD-19 dataset. The study evaluates a fixed RAG pipeline across three different cor…

  12. RESEARCH · CL_156459 ·

    PLAID-PRF enhances dense retrieval with centroid-aware pseudo-relevance feedback

    Researchers have introduced PLAID-PRF, a novel method for enhancing multi-vector dense retrieval models like ColBERT. This technique leverages centroid-like tokens within the PLAID framework to perform Pseudo-Relevance …

  13. RESEARCH · CL_128512 ·

    New benchmarks evaluate Portuguese text embedding models, revealing performance gaps

    Two new benchmarks, MTEB-PT and MTEB-PT (Brazilian Portuguese), have been released to evaluate text embedding models specifically for the Portuguese language. These benchmarks address the underrepresentation of Portugue…

  14. RESEARCH · CL_117090 ·

    New RAG research enhances LLM retrieval, unlearning, and faithfulness

    Multiple research papers are exploring advancements in retrieval-augmented generation (RAG) to improve the performance and efficiency of large language models. Apple's CLaRa framework unifies retrieval and generation in…

  15. TOOL · CL_115148 ·

    New method enhances explainability for dense embedding rankers

    Researchers have developed a new method called ChunkGroupSHAP to improve the explainability of dense embedding rankers used in information retrieval. This technique clusters semantically related text chunks across docum…

  16. TOOL · CL_111510 ·

    GPUSparse system accelerates learned sparse retrieval using GPU parallelization

    Researchers have developed GPUSparse, a novel system designed to accelerate learned sparse retrieval models by leveraging GPU parallelization. This system addresses the CPU-bound bottleneck in current sparse retrieval m…

  17. TOOL · CL_111511 ·

    TileMaxSim kernel boosts GPU retrieval model speed by 220x

    Researchers have developed TileMaxSim, a new IO-aware kernel for GPUs designed to significantly accelerate the MaxSim scoring process used in multi-vector retrieval models like ColBERT. Existing implementations are inef…

  18. COMMENTARY · CL_103119 ·

    AI agents fail due to flawed search index distribution, not prompting

    A common issue in AI agents is that their search results appear correct but lead to factually wrong answers due to problems with the underlying search index. This is not a prompting issue but a distribution problem, whe…

  19. TOOL · CL_93461 ·

    New indexing framework SPI boosts RAG performance in vector databases

    Researchers have introduced Semantic Pyramid Indexing (SPI), a novel indexing framework for vector databases designed to enhance retrieval-augmented generation (RAG) pipelines. SPI adapts the retrieval depth based on qu…

  20. RESEARCH · CL_72662 ·

    New research tackles noise and efficiency in full-duplex dialogue systems

    Two new research papers explore advancements in full-duplex spoken dialogue systems, which allow for simultaneous listening and speaking. One paper introduces Interference-Resilient Adaptive Fusion (IRAF) to improve rob…