A new study published on arXiv evaluated the performance of 16 large language models (LLMs) against 60 practicing traditional Chinese medicine (TCM) physicians. The LLMs demonstrated superior performance in areas like medical advice and diagnostic tasks, scoring higher than human physicians in evaluations by senior TCM experts. However, the models exhibited limitations in prescription accuracy, including herb selection and dosage, and produced concerning outputs such as hallucinations and template-driven responses, necessitating physician oversight. AI
IMPACT LLMs show potential for decision support in specialized medical fields, but require human oversight for safety and accuracy.
RANK_REASON The cluster contains an academic paper evaluating LLM performance on a specific domain. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →