A preliminary study explored the ability of four Large Language Models (LLMs) to extract Architectural Design Decisions (ADDs) from source code commits. Models including Gemini 3-Pro, DeepSeek-R1, Kimi K2, and Qwen3 were tested using zero-shot and few-shot prompting on 30 ADDs. While all models achieved a BERT-F1 score above 0.81, the generated decisions were often too lengthy, implementation-focused, and lacked the underlying rationale, indicating a need for more architecture-aware LLM systems. AI
IMPACT Highlights opportunities for architecture-aware LLM systems and automated Architectural Knowledge Management.
RANK_REASON The cluster is about an academic paper presenting a study on LLM capabilities for a specific task.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →