Phind
PulseAugur coverage of Phind — every cluster mentioning Phind across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New benchmark SWE-sweep tests LLMs on proactive bug fixing
Researchers from Meta, Stanford, Harvard, and UW have developed SWE-sweep, a new benchmark designed to evaluate large language models' ability to proactively identify and fix bugs in large codebases before they impact u…
-
New AI coding benchmarks test deep software engineering capabilities
New coding benchmarks are emerging that aim to test deeper AI capabilities in software engineering beyond traditional metrics. Program-Bench requires agents to reconstruct code from a compiled binary and documentation, …
-
Cursor AI coding assistant integrates multiple LLMs, including OpenAI and Gemini
A new AI coding assistant named Cursor has been released, aiming to improve the developer experience. It integrates with various large language models, including OpenAI's models, Claude, Gemini, Code Llama, Starcoder, M…
-
Open-source coding LLMs now rival proprietary leaders, shifting focus to workflow fit
The landscape of open-source coding LLMs has rapidly advanced, with several models now rivaling proprietary leaders on practical software engineering tasks. This shift means the focus has moved from whether open-source …