PulseAugur
中
实时 16:46:30

Scrapy 2.19 新增默认 RemoteControl 扩展,支持爬取过程中执行 Python 代码

Scrapy 的最新版本 2.19 引入了一个新的 RemoteControl 扩展,该扩展默认启用,允许用户在实时爬取过程中执行 Python 代码。此扩展会在一个随机端口上启动一个 HTTP 服务器,并通过 bearer token 进行保护,并创建一个包含爬取详情的作业文件。服务器提供两个端点:/status 用于查看爬取信息,/execute 用于执行任意 Python 代码,并返回其输出或遇到的任何错误。执行的代码可以访问实时 Crawler 实例和一个持久化的 stash 字典,从而实现对正在进行的爬取的动态控制和检查。 AI

影响 增强了网络抓取任务的控制和调试能力,可能提高了 AI 数据收集管道的效率。

排序理由 现有软件工具的新功能发布。

在 dev.to — MCP tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Scrapy 2.19 新增默认 RemoteControl 扩展,支持爬取过程中执行 Python 代码

本文如何被排名

Signal score
17 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
现有软件工具的新功能发布。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. dev.to — MCP tag TIER_1 English(EN) · John Rooney ·

    Scrapy RemoteControl 扩展内部解析

    <p>Scrapy 2.19 added an extension called <a href="https://github.com/scrapy/scrapy/blob/master/scrapy/extensions/remote_control.py" rel="noopener noreferrer"><code>RemoteControl</code></a>, and it's on by default, which means every crawl you start on the asyncio reactor is alread…