SpreadsheetBench
PulseAugur coverage of SpreadsheetBench — every cluster mentioning SpreadsheetBench across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New BREW framework enables LLM agents to learn from experience
Researchers have developed BREW, a novel framework designed to enable Large Language Model (LLM)-based agents to learn from past experiences. Unlike current agents that restart learning with each session, BREW distills …
-
New ProCUA-SFT dataset boosts AI agent desktop performance
Researchers have developed ProCUA-SFT, a new dataset designed to improve the training of computer-use agents (CUAs) that interact with graphical desktop environments. Existing datasets like AgentNet have shown negative …
-
AI models benchmarked for Excel accuracy; specialized tools lead
A new benchmark called SpreadsheetBench evaluates AI models on their accuracy in handling Excel documents. The benchmark uses real-world tasks from Excel forums, requiring exact cell-by-cell accuracy and testing complex…
-
SkillOpt optimizes AI agent skills using validated parameter edits
A new paper introduces SkillOpt, a method for optimizing AI agent skills by treating markdown skill files as trainable parameters. The approach uses a frontier model to propose bounded edits, which are then validated ag…