PulseAugur
实时 18:06:25
English(EN) We recovered 575k crop labels from a decade of manual Photoshop work to automate book digitization - more data, ResNet-50, and higher resolution all failed; ten operator clicks per book beat them [P]

人类裁剪决策在书籍数字化方面优于 AI 扩展

研究人员开发了一种新颖的方法,通过利用操作员十年来的手动裁剪决策来自动化稀有书籍的数字化。这种方法涉及恢复 575,729 个 Photoshop 工作中的裁剪标签,比扩展训练数据、使用 ResNet-50 模型或增加输入分辨率更有效。关键的见解是,操作员对页边距嵌入的一致(尽管是隐形的)偏好是关键因素,而从每本书的十个操作员校正的裁剪中获得的简单中值残差显著提高了性能。 AI

影响 该方法强调了利用人类生成的隐式数据进行 AI 任务的潜力,表明从纯像素学习转向整合人类偏好模型。

排序理由 该项目描述了一种使用源自人力的数据集来自动化任务的新颖方法,并将其有效性与标准机器学习方法进行了比较。[lever_c_demoted from research: ic=1 ai=0.7]

在 r/MachineLearning 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

人类裁剪决策在书籍数字化方面优于 AI 扩展

本文如何被排名

Signal score
9 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目描述了一种使用源自人力的数据集来自动化任务的新颖方法,并将其有效性与标准机器学习方法进行了比较。[lever_c_demoted from research: ic=1 ai=0.7]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. r/MachineLearning TIER_1 English(EN) · /u/laamaleph ·

    我们从十年的手动Photoshop工作中恢复了57.5万个农作物标签以实现图书数字化自动化——更多数据、ResNet-50和更高分辨率均告失败;每本书十次操作员点击胜过它们[P]

    <!-- SC_OFF --><div class="md"><p>Author here. Ibteda Digital Library is a private community archive in Pakistan — for ten years we digitized rare Urdu books (lithographs, dictionaries, periodicals) on a DIY camera rig, finishing every page by hand in Photoshop. When we wound dow…