Researchers have developed a new benchmark to evaluate how effectively smaller language models utilize legal texts provided in their context, particularly in the domain of Bangladeshi law. The study found that while fine-tuning can improve accuracy on legal question-answering tasks, it does not necessarily enhance the models' ability to rely on supplied legal provisions. The methodology involved creating a bilingual statutory corpus and fine-tuning examples, then employing techniques like constrained scoring and controlled removal of governing provisions to differentiate between scorer, retriever, and model effects. Results indicated that some models showed significant accuracy gains from fine-tuning, but this did not translate to increased reliance on the provided legal context. AI
IMPACT Highlights limitations in current small LLMs' ability to leverage provided context, suggesting a need for more robust evaluation methods in specialized domains.
RANK_REASON Academic paper detailing a new benchmark and evaluation of small language models. [lever_c_demoted from research: ic=1 ai=1.0]
- Bangladesh
- bar council
- Gemma 4-E2B
- Hugging Face
- Llama 3.2:1b
- Llama 3.2:3b
- LoRA+
- Qwen3.5-0.8B
- Qwen3.5 2B
- Qwen3.5 4B
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →