A recent benchmark by Vals AI reveals that while AI tools excel at summarizing lengthy legal documents and answering questions with citations, they struggle with comparing different versions of contracts. Human lawyers outperformed AI in identifying subtle discrepancies between document revisions, a task requiring precise positional context retention. This highlights a limitation in current AI models, where long context windows do not always translate to accurate performance on complex tasks like redlining, as demonstrated by the CLAUSE benchmark and research from Stanford RegLab. AI
IMPACT Highlights the need for specialized AI tools for complex legal tasks like contract redlining, beyond general summarization capabilities.
RANK_REASON The item details findings from a new benchmark comparing AI legal tools against human lawyers on specific tasks. [lever_c_demoted from research: ic=1 ai=1.0]
- Claude 3.5 Sonnet
- CLAUSE
- CoCounsel
- Daniel E. Ho
- Eacles
- GigaChat 2 Max
- GPT-4
- GPT-4o
- Harvey
- Harvey Assistant
- Lexis+ AI
- NoLiMa
- Stanford RegLab
- Vals AI
- Vincent Aizebeoje Balogun
- VLAIR
- Westlaw AI-Assisted Research
- YandexGPT
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →