A preliminary study explored the effectiveness of Large Language Models (LLMs) in extracting architectural design decisions from source code commits. Researchers tested four LLMs—Gemini 3-Pro, DeepSeek-R1, Kimi K2, and Qwen3—using zero-shot and few-shot prompting on 30 developer-written decisions. While all models achieved a BERT-F1 score above 0.81, few-shot prompting showed slight improvements. However, the generated decisions were often too lengthy, implementation-focused, and lacked the underlying rationale, indicating a need for more architecture-aware LLM systems. AI
IMPACT Highlights opportunities for architecture-aware LLM systems in software engineering and automated knowledge management.
RANK_REASON The cluster contains an academic paper detailing a study on LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →