PulseAugur
中
实时 23:24:10

新框架用发现曲线分析网络爬取数据

研究人员开发了一个新的框架,用于分析纵向网络爬取数据,即演变中的URL群体的顺序样本。该框架引入了“发现曲线”来衡量爬取窗口内的累积URL足迹。通过将发现曲线与成对包含分析进行比较,研究人员在应用于Common Crawl和德国学术网络档案时,发现了预测上的分歧。这些差异表明网络存在一个由持久核心和动态外壳组成的双组分模型,该模型能更好地调和观测到的数据。 AI

排序理由 该集群包含一篇学术论文,详细介绍了一种分析网络爬取数据的新方法。[lever_c_demoted from research: ic=1 ai=0.1]

在 arXiv cs.IR (Information Retrieval) 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新框架用发现曲线分析网络爬取数据

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇学术论文,详细介绍了一种分析网络爬取数据的新方法。[lever_c_demoted from research: ic=1 ai=0.1]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
Low
Off-topic or adjacent — cluster remains reachable but doesn't surface in AI-industry rankings.
Story freshness
85 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Luca Foppiano ·

    爬虫所见之度量:纵向网络爬取中的发现曲线、核心持久性与外壳动态

    A longitudinal web crawl is a sequence of partial samples of an evolving URL population. Pairwise containment between two crawls is the standard probe; under a simple \emph{urn} model of the crawl -- each round samples a fraction of the URLs and replaces a fraction -- it recovers…