PulseAugur
EN
LIVE 13:14:49
ENTITY TruthfulQA

TruthfulQA

PulseAugur coverage of TruthfulQA — every cluster mentioning TruthfulQA across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
4
19 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
4
19 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

2 day(s) with sentiment data

RECENT · PAGE 1/2 · 31 TOTAL
  1. RESEARCH · CL_256972 ·

    New CoSQ framework helps LLMs decide when to abstain from answering

    Researchers have developed a new framework called Chain-of-Self-Questioning (CoSQ) to help large language models (LLMs) determine when to abstain from answering questions if their factual support is weak. CoSQ variants,…

  2. TOOL · CL_254356 ·

    New ALTAS method improves LLM reliability in clinical question answering

    Researchers have developed ALTAS, a novel method for improving the reliability of Large Language Models (LLMs) in clinical question answering. ALTAS utilizes a trajectory-gated router that analyzes terminal entropy and …

  3. RESEARCH · CL_233440 ·

    New methods enhance LLM steering for behavior control · 2 sources tracked

    Two new research papers introduce novel methods for steering large language models to suppress undesired behaviors. GAPS (Gated Activation steering via Posterior and Separability) employs dimension-level gates to select…

  4. RESEARCH · CL_218205 ·

    New research tackles AI hallucinations in video and language models

    Researchers are developing new methods to combat hallucinations in AI models, particularly in video-language and large language models. One approach, CounterVid, uses counterfactual video generation to create synthetic …

  5. TOOL · CL_217837 ·

    LLM Leaderboards Skewed by Config-Fragile Items, Study Finds

    A new paper from arXiv reveals that modern large language model (LLM) leaderboards are significantly influenced by "config-fragile" items, meaning the way questions are presented and answers are evaluated can drasticall…

  6. TOOL · CL_215931 ·

    New VA-DPO method enables controllable emotion generation in language models

    Researchers have developed a new method called VA-DPO to enable language models to generate text with controllable emotions. Unlike previous methods that use discrete labels, VA-DPO specifies desired affect as a continu…

  7. RESEARCH · CL_215727 ·

    New agent detects misinformation in RAG systems

    Researchers have developed an "Evaluation Agent" to address the security and reliability gap in Retrieval-Augmented Generation (RAG) systems. This agent acts as middleware to detect misinformation and knowledge poisonin…

  8. TOOL · CL_205909 ·

    New HASSUM framework guides multi-agent AI with semantic uncertainty

    Researchers have developed a new framework called HASSUM to improve coordination in multi-agent AI systems by guiding orchestration decisions with semantic uncertainty. This method uses semantic entropy and density to a…

  9. RESEARCH · CL_196080 ·

    New Power Law Graph Attention generalizes SDPA with learned operator

    Researchers have introduced a novel attention mechanism called Power Law Graph Attention (PLGA) that generalizes scaled dot-product attention (SDPA) by using a learned, input-generated bilinear operator. This new archit…

  10. TOOL · CL_183528 ·

    LLM hallucination benchmarks misleading, study finds

    A recent analysis of four open-weight large language models—Phi-4 Mini, Mistral 7B Instruct v0.3, Qwen2.5-7B-Instruct, and Llama-3.1–8B-Instruct—reveals that hallucination benchmarks may be misleading. The study found t…

  11. TOOL · CL_177381 ·

    DeepSeek V4 flash version shows strong performance on MMLU-Pro, GPQA Diamond

    DeepSeek V4 has released a new "flash" version, reportedly achieving impressive scores on benchmarks like MMLU-Pro, GPQA Diamond, and TruthfulQA. The model is noted for its strong performance relative to its size, with …

  12. TOOL · CL_174040 ·

    New framework uses psychometric tests to evaluate LLM behavioral consistency

    Researchers have developed a new framework for evaluating the behavioral consistency of large language models (LLMs) using situational judgment tests (SJTs) and multidimensional item response theory (MIRT). This approac…

  13. TOOL · CL_160746 ·

    LLM diversity metrics may not measure diversity, study finds

    A new research paper published on arXiv questions the effectiveness of common diversity metrics used in Large Language Model (LLM) ensembles. The study found that these metrics often correlate more with the models' over…

  14. TOOL · CL_154354 ·

    New Benchmark Suite Evaluates LLMs on Kyrgyz Language Understanding

    Researchers have developed KyrgyzLLM-Bench, a new benchmark suite designed to evaluate large language models (LLMs) on the Kyrgyz language. This suite includes natively authored datasets like KyrgyzMMLU and KyrgyzRC, al…

  15. TOOL · CL_154351 ·

    New Research: Language Models' Self-Judgement Overrides Objective Correctness

    A new research paper published on arXiv explores the phenomenon of self-judgment confounding in language models, where models' own assessments of their output's correctness can override objective correctness. The study …

  16. TOOL · CL_139854 ·

    J-space entropy shows mixed results as an error predictor in Qwen3-4B

    A recent study explored using "J-space entropy," an internal metric within language models, to predict errors, particularly hallucinations. The research tested this hypothesis on the Qwen3-4B model across seven diverse …

  17. RESEARCH · CL_129095 ·

    AI hallucinations: new research probes reasoning and cross-lingual generalization

    Two new research papers explore the phenomenon of "hallucinations" in AI models, focusing on how these errors influence downstream reasoning and whether detection signals generalize across languages and domains. The fir…

  18. TOOL · CL_123138 ·

    New approach quantifies neural network uncertainty using gradient norms

    Researchers have developed a novel method for quantifying uncertainty in neural networks, particularly large language models, by approximating predictive uncertainty using gradient norms and an isotropy assumption. This…

  19. RESEARCH · CL_117343 ·

    New SEVA agent tackles LLM hallucination with detailed verification

    Researchers have developed SEVA, a novel self-evolving verification agent designed to combat hallucination in LLM-based systems. Unlike traditional verifiers that provide opaque binary labels, SEVA offers detailed evide…

  20. RESEARCH · CL_111612 ·

    New metric ConflictScore measures LLMs' handling of conflicting evidence

    Researchers have introduced ConflictScore, a new metric designed to evaluate how well language models handle conflicting information within their grounding documents. Unlike existing metrics that only check for support …