Toolkit Labs
PulseAugur coverage of Toolkit Labs — every cluster mentioning Toolkit Labs across labs, papers, and developer communities, ranked by signal.
-
AI API schema rejections plague Gemini, Groq; models fail before generation
A recent analysis of API calls revealed that a significant number of structured output requests failed not due to model errors, but because the API providers rejected the JSON schema itself. Toolkit Labs found that 28 o…
-
json-repair excels at extracting LLM-wrapped JSON, outperforming jsonshim
A new benchmark, MALFORMED-300, evaluates how well parsers can extract JSON from malformed LLM outputs, particularly when JSON is wrapped in other formats like XML, HTML, or SQL. The json-repair tool achieved a perfect …
-
Developer releases Unicode string corpus to expose text handling flaws
A developer has created a corpus of 300 Unicode strings designed to expose flaws in naive text handling functions, particularly `len()`. The corpus includes various complex Unicode features like ZWJ sequences and astral…
-
Nostr NIP-90 job market sees 1,860 requests but zero bids
An analysis of the Nostr NIP-90 job market over 30 days revealed a significant imbalance between job requests and bids. Out of 1,860 job requests, none included a bid tag specifying a maximum payment, indicating that cu…
-
LLM JSON parsers struggle with truncated output; Toolkit Labs releases benchmark
Toolkit Labs has released findings from a benchmark of JSON parsers designed to handle malformed output from large language models. The study focused on 25 cases of truncated JSON, where a model's output is cut off mid-…