PulseAugur
EN
LIVE 08:13:21

High-res video generation with LingBot-Video requires substantial GPU power

A user on Reddit shared details about running LingBot-Video at a high resolution of 1088x1920. The process required four RTX PRO 6000 Max-Q GPUs, each with 96GB of VRAM, and took approximately 20 minutes to generate just 3 seconds of video. The setup utilized FSDP2 and context parallelism across the GPUs, with each card peaking at 57GB of VRAM usage. The user also noted the complexity of managing memory and the need for a separate 27B model to generate structured JSON captions for the video. AI

IMPACT Demonstrates the significant hardware demands for high-resolution AI video generation.

RANK_REASON User-generated content detailing the hardware requirements for a specific AI video generation model.

Read on r/StableDiffusion →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

High-res video generation with LingBot-Video requires substantial GPU power

COVERAGE [1]

  1. r/StableDiffusion TIER_2 English(EN) · /u/NewVeterinarian5384 ·

    LingBot-Video at 1088x1920 on 4x RTX PRO 6000 Max-Q (57 GB a card, just under 20 minutes for 3 seconds)

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1v94i6d/lingbotvideo_at_1088x1920_on_4x_rtx_pro_6000_maxq/"> <img alt="LingBot-Video at 1088x1920 on 4x RTX PRO 6000 Max-Q (57 GB a card, just under 20 minutes for 3 seconds)" src="https://external-previe…