Verse-Bench
PulseAugur coverage of Verse-Bench — every cluster mentioning Verse-Bench across labs, papers, and developer communities, ranked by signal.
-
NAVA model generates synchronized audio and video from text prompts
A new 6.3 billion parameter model named NAVA has been released, capable of generating synchronized audio and video from a single text prompt. It features multi-speaker speech control and image-conditioned continuations.…
-
Baidu releases NAVA, a 6.3B parameter audio-visual generation model
Baidu has released NAVA, a 6.3 billion parameter model capable of generating synchronized audio and video from a single text prompt. This model utilizes an Align-then-Fuse MMDiT architecture to achieve state-of-the-art …
-
New frameworks and benchmarks advance audio-visual generation
Researchers have introduced OmniCustom, a framework for customizing both video identity and audio timbre simultaneously from reference images and audio. This DiT-based model uses separate LoRA modules for identity and t…