EvalBench
PulseAugur coverage of EvalBench — every cluster mentioning EvalBench across labs, papers, and developer communities, ranked by signal.
-
Chunking strategies for dense retrieval evaluated for effectiveness and cost · 2 sources tracked
A new paper evaluates eight different chunking strategies for dense retrieval systems, considering not only retrieval effectiveness but also system-level costs like throughput and latency. The research indicates that co…
-
EvalBench simplifies LLM evaluation schema with flexible data model
The author describes a refactoring of the EvalBench evaluation schema to simplify the addition of new benchmark suites. Previously, each new suite required database schema migrations, aggregation query changes, and fron…
-
LLM evaluation tool EvalBench reveals critical statistical bug
The creator of EvalBench, a platform for evaluating large language models, discovered a critical bug in their statistical calculations. The platform incorrectly reported a negative retry count for one of the models, ind…