PulseAugur
中
实时 13:00:28

AEGIS 系统优化共享 GPU 上的深度学习训练

研究人员开发了 AEGIS,一个运行时调度系统,旨在提高共享多 GPU 服务器上深度学习训练的效率。AEGIS 通过集成内存可行性检查、放置后观察和运行时压力过滤来管理多个深度学习工作负载的共置。与独占分配相比,这种方法旨在减少资源利用不足和排队时间,同时还能减轻不太复杂的共置方法可能出现的性能下降和内存不足故障。使用各种工作负载进行的评估表明,与独占分配相比,AEGIS 可将训练完成时间最多缩短 27%,与 Lucid 和 Horus 等其他共置系统相比,可缩短 16-21%。 AI

影响 优化深度学习训练的 GPU 利用率,可能降低成本并加速开发周期。

排序理由 该集群描述了一篇研究论文,其中详细介绍了一个用于优化深度学习训练基础设施的新系统。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

AEGIS 系统优化共享 GPU 上的深度学习训练

本文如何被排名

Signal score
7 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一篇研究论文,其中详细介绍了一个用于优化深度学习训练基础设施的新系统。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Ehsan Yousefzadeh-Asl-Miandoab, B\"u\c{s}ra Karatay Demiray, Florina M. Ciorba, Pamela Delgado, P{\i}nar T\"oz\"un ·

    AEGIS:运行时引导的多租户深度学习训练GPU协同

    arXiv:2508.19073v4 Announce Type: replace-cross Abstract: Deep learning training commonly runs on shared multi-tenant GPU servers, where exclusive allocation provides isolation but can leave resources underutilized and increase queueing time. Collocation can improve efficiency, b…