A new paper evaluates the performance of seven open-source large language models (LLMs) for retrieval-augmented generation (RAG) tasks specifically within the environmental, social, and governance (ESG) domain. The study utilized 498 real-world ESG reports from EU-listed companies and a set of 100 synthetic QA pairs to assess models like GLM 4.7 Flash, Nemotron-3-nano:4b, and Qwen3:4b-instruct. While retrieval performance was generally strong across models, generation metrics like faithfulness and factual correctness showed significant variation, indicating a need for domain-specific fine-tuning. AI
IMPACT Provides data-driven guidance for selecting and fine-tuning open-source LLMs for specialized ESG reporting tasks.
RANK_REASON The cluster contains an academic paper evaluating open-source LLMs on a specific domain task. [lever_c_demoted from research: ic=1 ai=1.0]
- environmental, social and corporate governance
- EU
- gemma3:4b
- gemma4:e2b
- gemma4:e4b
- GLM 4.7 Flash
- Hugging Face
- large-language models
- Ragas
- retrieval-augmented generation
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →