This paper introduces a black-box evaluation framework designed to assess how well Large Language Models (LLMs) can generate Design Structure Matrices (DSMs) from technical documentation. The framework uses a reproducible methodology to compare LLM-generated DSMs against manually validated ground-truth matrices. It incorporates structural, classification, and stability metrics, culminating in a Composite Quality Score (Q). Experiments reveal that while LLMs can produce plausible DSMs with good reproducibility on well-structured inputs, they struggle with ambiguity, inconsistent definitions, and prompt variations, highlighting current limitations in LLM-driven automation for model-based systems engineering. AI
IMPACT This framework could accelerate the integration of LLMs into model-based systems engineering by providing a standardized way to audit their performance in generating complex technical matrices.
RANK_REASON The cluster contains an academic paper detailing a new evaluation framework for LLM capabilities.
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →