SWE-Bench Multilingual
PulseAugur coverage of SWE-Bench Multilingual — every cluster mentioning SWE-Bench Multilingual across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New tool PAIChecker identifies PR-issue misalignment in LLM benchmarks
Researchers have developed PAIChecker, a multi-agent system designed to identify and correct misalignments between pull requests (PRs) and their associated issues in benchmarks used to evaluate large language models (LL…
-
MindForge pipeline trains small LLMs for full software engineering lifecycle
Researchers have developed MindForge, an automated pipeline designed to train smaller language models in comprehensive software engineering tasks. This system converts open-source command-line programs into source-free …
-
Poolside AI releases Laguna S 2.1, a compact coding model with 1M context
Poolside AI has released Laguna S 2.1, an 118B parameter Mixture-of-Experts model with 8B activated parameters and a 1M token context window. The model was developed in under nine weeks and demonstrates strong performan…
-
Xiaomi's MiMo-V2-Flash leads open-source coding benchmarks with efficient MoE architecture
Xiaomi has developed MiMo-V2-Flash, a 309-billion-parameter Mixture-of-Experts model that leads open-source options on SWE-Bench for coding tasks. This model achieves high performance with significantly less computation…
-
Tencent releases Hy3, an open 295B MoE model with 256K context
Tencent has released Hy3, an open-source 295 billion parameter Mixture-of-Experts (MoE) model designed for complex reasoning, agentic workflows, and long-context tasks. The model activates only 21 billion parameters per…
-
Dockerless verifies AI coding agent patches without running tests
Researchers have introduced Dockerless, a novel method for verifying code patches generated by AI coding agents without the need to execute repository-specific tests. This approach bypasses the costly and time-consuming…
-
Poolside releases Laguna M.1, a 225B MoE model for agentic coding
Poolside has released Laguna M.1, a 225 billion parameter Mixture-of-Experts model optimized for agentic coding tasks. The model features a large sparse MoE architecture with 256 experts and global attention, enabling i…
-
Fireworks AI enables training of trillion-parameter MoE models
Fireworks AI has developed a new training infrastructure that enables the fine-tuning of trillion-parameter Mixture-of-Experts (MoE) models, overcoming previous memory and orchestration bottlenecks. This platform was in…