PulseAugur
实时 06:00:58
English(EN) BIRD-History: A Benchmark for History-Driven Text-to-SQL with Fine-Grained Knowledge Annotations

新的BIRD-History基准使用查询日志评估文本到SQL系统

研究人员推出了BIRD-History,这是一个新的基准,旨在评估文本到SQL系统利用历史查询日志理解不明确的自然语言问题的能力。该基准包含11个数据库中的1,393个任务,并带有注解以识别过去SQL脚本中的相关知识。提出的插件式检索器从这些历史日志中提取和重新排序外部知识,并在与现有文本到SQL系统集成时展示了持续的性能改进。 AI

影响 该基准可以通过使文本到SQL系统更好地理解历史数据中的上下文来提高其准确性。

排序理由 该条目描述了一个用于评估AI系统的新基准和相关方法论,发布在arXiv上。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的BIRD-History基准使用查询日志评估文本到SQL系统

本文如何被排名

Signal score
36 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该条目描述了一个用于评估AI系统的新基准和相关方法论,发布在arXiv上。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Yunfan Zhou, Qiming Shi, Yizhou Yang, Di Weng, Yingcai Wu ·

    BIRD-History:一个用于历史驱动的、具有细粒度知识标注的Text-to-SQL基准测试

    arXiv:2608.29345v1 Announce Type: new Abstract: While recent Large Language Model (LLM)-based text-to-SQL systems achieve impressive performance on standard benchmarks, they struggle when user queries implicitly rely on domain-specific knowledge, such as business logic, data conv…