Researchers have investigated how instruction-tuned Transformer models, specifically LLaMA and Mistral, encode discourse relations in English, with a focus on causation and antithesis. Using interpretability techniques, the study found that early model layers make predictions at mid-sequence tokens, while mid-level layers finalize decisions near the end. Some layers showed a preference for specific answers, indicating an asymmetric representation of discourse-based reasoning. AI
IMPACT Provides insights into how LLMs process complex linguistic structures, potentially improving their reasoning and ethical behavior.
RANK_REASON The cluster contains a research paper detailing findings on how specific LLMs encode discourse relations. [lever_c_demoted from research: ic=1 ai=1.0]
- Abhidip Bhattacharyya
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- llama
- ScienceCast
- Transformer++
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →