English
PulseAugur coverage of English — every cluster mentioning English across labs, papers, and developer communities, ranked by signal.
- instance of German 90%
- instance of ScienceCast 90%
- instance of Portuguese 90%
- instance of Korean 90%
- instance of Russian 90%
- instance of languages of Africa 90%
- developed by Tamil 90%
- developed Romanian 90%
- developed Tamil 90%
- used by machine translation 75%
- instance of Standard Chinese 70%
- instance of DagsHub 70%
23 day(s) with sentiment data
-
AI models fail African languages, losing 90% of safety signal
AI models demonstrate a significant degradation in safety alignment when tested with African languages, retaining less than 10% of the safety signal observed in English. This deficiency leaves speakers of these language…
-
Expensive translation model proved most confidently wrong, glossary fixes all
A developer discovered that a high-priced, flagship translation model produced more fluent but dangerously incorrect translations for specialized domain terminology compared to a cheaper model. When a glossary of domain…
-
GFlowNets used to generate novel LLM attacks in English and Turkish
Researchers have developed a novel method using GFlowNets to automatically generate adversarial attacks against Large Language Models (LLMs). This approach trains an attacker model to identify vulnerabilities in a victi…
-
AssemblyAI launches Universal-3.5 Pro with expanded multilingual transcription
AssemblyAI has released its Universal-3.5 Pro model, enhancing its multilingual transcription capabilities. This new model supports automatic language detection and native code-switching across 18 languages, a significa…
-
BLEU and ROUGE metrics explained for language model evaluation
BLEU and ROUGE are key metrics used to evaluate the performance of language models, particularly in tasks like machine translation and text summarization. BLEU focuses on precision of n-grams and includes a penalty for …
-
New R3S framework boosts multilingual LLM reasoning without external data
Researchers have developed R3S, a novel reinforcement learning framework designed to improve multilingual understanding and reasoning in large language models. This framework addresses bottlenecks in processing non-Engl…
-
New Polish Vision-Language Benchmark PoVisLE Introduced
Researchers have introduced PoVisLE, a new benchmark designed to evaluate Polish vision-language models (VLMs). Unlike existing benchmarks that are primarily English-centric and focus on surface-level recognition, PoVis…
-
New method improves low-resource language translation in NMT models
Researchers have developed a new method for initializing embeddings in multilingual neural machine translation models for low-resource languages. This approach involves averaging the embeddings of typologically related …
-
New MCIF benchmark tests multimodal and crosslingual LLM instruction following
Researchers have introduced MCIF, a new benchmark designed to evaluate multimodal and crosslingual instruction-following capabilities in large language models. This benchmark is unique in its use of scientific talks as …
-
New pipeline tackles gender bias in English-Romanian machine translation
Researchers have developed a novel pipeline to address gender bias in English-to-Romanian machine translation. Their method employs a fine-tuned large language model to identify gender in English sentences and insert ge…
-
New metrics needed for Classical Chinese to English AI translation
Researchers have investigated the effectiveness of current automatic evaluation metrics for translating Classical Chinese to English, a task where large language models show surprising proficiency but lack reliable asse…
-
Researchers pinpoint cross-lingual refusal circuit in multilingual MoE model
Researchers have identified a specific circuit within a multilingual Mixture-of-Experts (MoE) model, named sarvam, that is responsible for refusing harmful requests. This circuit's ability to refuse is language-invarian…
-
New Claude Code plugins aim to simplify AI output into plain English
Two distinct plugins have been developed for Claude Code to enhance the clarity of its output. One plugin, based on ISO 24495 standards, aims to ensure Claude's responses are in plain language, offering skills for vario…
-
Multilingual LLMs vulnerable to attacks bypassing English safety guardrails
Current AI safety alignment methods, which primarily focus on English, create significant vulnerabilities in multilingual large language models. These models can be exploited through attacks in less common languages or …
-
New AfriNLLB models offer efficient translation for 15 African languages
Researchers have developed AfriNLLB, a suite of lightweight translation models designed for African languages. These models are derived from the NLLB-200 600M architecture, which has been compressed through layer prunin…
-
Tokenization premiums create AI cost barriers for non-English languages · arXiv cs.CL
A new study published on arXiv introduces the Tokenization Equity Audit (TEA), a benchmark designed to measure disparities in how large language models tokenize different languages. The research found that semantically …
-
AI agent's prompt injection detector fails on non-English attacks
A security audit of an open-source agent framework revealed a significant vulnerability in its prompt injection detection system. The scanner, which inspects context files, memory writes, and tool outputs, failed to det…
-
AI Community Innovates with MiniMax Distillation, New ASR Model, and Gemini Rumors
The AI community has rapidly developed a distilled LoRA model based on MiniMax AI's recently released weights, showcasing rapid innovation. Separately, a new open-source model, Audio8-ASR-0.1B, has been released, offeri…
-
Multilingual RAG systems pose privacy risks, study finds
A new study published on arXiv investigates privacy risks in multilingual Retrieval-Augmented Generation (RAG) systems. Researchers tested an English-source synthetic dataset with queries in five languages, using a Qwen…
-
New distillation technique boosts multilingual math reasoning in LLMs
Researchers have explored On-Policy Delta Distillation (OPD^2), an advancement over On-Policy Distillation (OPD), for multilingual mathematical reasoning. Experiments using the Qwen3 model demonstrated that OPD^2 signif…