A benchmark test evaluated three AI models—Claude Sonnet 5.5, Claude Opus 5.5, and Gemini 3.7 Flash—on their ability to correctly update spreadsheet formulas when copied to new cells. Claude Opus 5.5 and Gemini 3.7 Flash achieved perfect scores, accurately returning the updated formulas in all tested cases. Claude Sonnet 5.5 performed poorly, scoring only 0.25, often failing to provide a usable formula even when demonstrating partial understanding of reference movement. AI
IMPACT Highlights differences in LLM capabilities for practical spreadsheet assistance, indicating that model choice significantly impacts usability for formula manipulation.
RANK_REASON Benchmark of AI model performance on a specific task. [lever_c_demoted from research: ic=1 ai=0.7]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →