PulseAugur
EN
LIVE 08:15:05

New benchmark SlidesGen-Bench evaluates LLM slide generation

Researchers have introduced SlidesGen-Bench, a new benchmark designed to evaluate the performance of large language models in generating presentation slides. This benchmark focuses on universality, quantification, and reliability, treating slide outputs as visual renderings to remain agnostic to the underlying generation method. It quantitatively assesses slides across content, aesthetics, and editability, and includes the Slides-Align1.5k dataset to ensure alignment with human preferences. AI

IMPACT Provides a standardized method for evaluating LLM slide generation capabilities, potentially driving improvements in content, aesthetics, and editability.

RANK_REASON The item describes a new academic paper introducing a benchmark for evaluating LLM slide generation. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark SlidesGen-Bench evaluates LLM slide generation

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Yunqiao Yang, Wenbo Li, Houxing Ren, Zimu Lu, Ke Wang, Zhiyuan Huang, Zhuofan Zong, Mingjie Zhan, Hongsheng Li ·

    SlidesGen-Bench: Evaluating Slides Generation via Computational and Quantitative Metrics

    arXiv:2601.09487v2 Announce Type: replace Abstract: The rapid evolution of Large Language Models (LLMs) has fostered diverse paradigms for automated slide generation, ranging from code-driven layouts to image-centric synthesis. However, evaluating these heterogeneous systems rema…