PulseAugur
EN
LIVE 09:02:57

4B open-weight models achieve near GPT-4 medical QA performance in Swedish

Researchers have found that smaller, open-weight language models, specifically 4B parameter models, are approaching the performance of larger models like GPT-4 on medical question-answering tasks in Swedish. Models such as Qwen3.5-4B and Gemma4-E4B have demonstrated significant accuracy on Swedish medical licensing exams without extensive post-training, with Qwen3.5-4B reaching 87% accuracy. Interestingly, Qwen3.5-4B performs this reasoning in English despite Swedish prompts, suggesting language barriers are less of an obstacle than previously thought for these models. AI

IMPACT Demonstrates that smaller, open-weight models can achieve high performance on specialized tasks, potentially lowering the barrier for AI adoption in niche domains.

RANK_REASON Research paper detailing performance of smaller LLMs on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/MachineLearning →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

4B open-weight models achieve near GPT-4 medical QA performance in Swedish

COVERAGE [1]

  1. r/MachineLearning TIER_1 English(EN) · /u/AccomplishedCat4770 ·

    Open-weight 4B models approach o3-level medical question answering in Swedish [P]

    <!-- SC_OFF --><div class="md"><p>I have been running some experiments with smaller open-weight LLMs on multiple-choice questions of Swedish medical licensing exams. On a dataset called MedQA-SWE, GPT-4 scored 84% accuracy in 2024 and o3 scored 88% in 2025 on a smaller, overlappi…