PulseAugur
EN
LIVE 07:41:43
ENTITY SpreadsheetBench

SpreadsheetBench

PulseAugur coverage of SpreadsheetBench — every cluster mentioning SpreadsheetBench across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
1
6 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
4 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

1 day(s) with sentiment data

RECENT · PAGE 1/1 · 6 TOTAL
  1. SIGNIFICANT · CL_242890 ·

    Ant Group's Bailing launches finance-focused AI model and evaluation benchmark

    Ant Group's Bailing has released Ling-3.0-flash-Fin, its first finance-enhanced open-source model designed to handle complex financial workflows beyond simple Q&A. This model aims to move AI from answering single questi…

  2. COMMENTARY · CL_185732 ·

    ChatGPT's spreadsheet accuracy questioned by independent tests

    A recent analysis has revealed significant discrepancies in ChatGPT's performance when handling financial spreadsheets, despite OpenAI's claims of high accuracy. While OpenAI reported an improvement from 43.7% to 87.3% …

  3. TOOL · CL_139589 ·

    New BREW framework enables LLM agents to learn from experience

    Researchers have developed BREW, a novel framework designed to enable Large Language Model (LLM)-based agents to learn from past experiences. Unlike current agents that restart learning with each session, BREW distills …

  4. RESEARCH · CL_95769 ·

    New ProCUA-SFT dataset boosts AI agent desktop performance

    Researchers have developed ProCUA-SFT, a new dataset designed to improve the training of computer-use agents (CUAs) that interact with graphical desktop environments. Existing datasets like AgentNet have shown negative …

  5. TOOL · CL_61569 ·

    AI models benchmarked for Excel accuracy; specialized tools lead

    A new benchmark called SpreadsheetBench evaluates AI models on their accuracy in handling Excel documents. The benchmark uses real-world tasks from Excel forums, requiring exact cell-by-cell accuracy and testing complex…

  6. TOOL · CL_52060 ·

    SkillOpt optimizes AI agent skills using validated parameter edits

    A new paper introduces SkillOpt, a method for optimizing AI agent skills by treating markdown skill files as trainable parameters. The approach uses a frontier model to propose bounded edits, which are then validated ag…