Turkish
PulseAugur coverage of Turkish — every cluster mentioning Turkish across labs, papers, and developer communities, ranked by signal.
6 day(s) with sentiment data
-
GFlowNets used to generate novel LLM attacks in English and Turkish
Researchers have developed a novel method using GFlowNets to automatically generate adversarial attacks against Large Language Models (LLMs). This approach trains an attacker model to identify vulnerabilities in a victi…
-
AssemblyAI launches Universal-3.5 Pro with expanded multilingual transcription
AssemblyAI has released its Universal-3.5 Pro model, enhancing its multilingual transcription capabilities. This new model supports automatic language detection and native code-switching across 18 languages, a significa…
-
AI agent's prompt injection detector fails on non-English attacks
A security audit of an open-source agent framework revealed a significant vulnerability in its prompt injection detection system. The scanner, which inspects context files, memory writes, and tool outputs, failed to det…
-
HukukBERT: New Language Model for Turkish Legal Texts
Researchers have developed HukukBERT, a specialized language model designed to process Turkish legal texts. This model was trained on a substantial corpus of Turkish legal documents using a hybrid pre-training approach …
-
Cross-lingual transfer in Turkic languages shows strong pair-specific performance
Researchers have investigated cross-lingual transfer techniques for machine translation within the Turkic language family, focusing on Turkish, Azerbaijani, Uzbek, Kazakh, and Kyrgyz. Their findings indicate that transf…
-
Research paper flags structural flaws in LLM-as-Judge synthetic corpora
A new research paper highlights a critical issue in creating synthetic datasets for evaluating Large Language Models (LLMs) used as judges. The study reveals that the process of generating 'hallucinated' answers for the…
-
Healthcare LLMs show significant cross-lingual factual disparities, paper finds
A new arXiv paper highlights significant disparities in the factual accuracy of Large Language Models (LLMs) when answering healthcare-related questions across different languages. Researchers developed a multilingual d…
-
LLMs' grasp of sarcasm questioned in Turkish language study
Large Language Models (LLMs) may struggle with understanding sarcasm due to its reliance on context and linguistic nuances, unlike simpler tasks like sentiment analysis. A project at Acıbadem University, in collaboratio…
-
Cross-lingual transfer learning shows mixed results for low-resource ASR
Researchers have explored cross-lingual transfer learning to improve automatic speech recognition (ASR) for low-resource languages. One study successfully used Sinhala to enhance Dhivehi ASR, achieving a 12.89% Word Err…
-
AI agent evaluation rankings shift based on judge language, study finds
A new research paper explores how the language used in evaluating AI agents can significantly impact their performance rankings. The study localized prompts for an Agent-as-a-Judge framework into five diverse languages,…
-
New BERT models tackle hate speech in Turkish and Arabic
Researchers have developed advanced BERT-based models for hate speech detection in Turkish and Arabic languages. The study introduces a new dataset covering five topics in Turkish, including refugees, the Israel-Palesti…
-
New ToxiREX dataset tackles implicit toxicity across six languages
Researchers have introduced ToxiREX, a new multilingual dataset designed to capture implicit and context-dependent toxicity in online conversations. The dataset comprises Reddit comment threads, annotated using a struct…
-
LLMs tested for Turkish scam detection using new audio-transcript dataset
Researchers have explored the effectiveness of large language models (LLMs) in detecting phone call scams in Turkish, a low-resource language. They introduced a new dataset of 100 aligned audio-transcript pairs of scam …
-
Morpheus: New Turkish Language Model Achieves Superior Morphological Alignment
Researchers have developed Morpheus, a novel neural tokenizer and word embedder specifically designed for the Turkish language. Unlike traditional subword tokenizers that can fragment Turkish's agglutinative structure, …
-
New SindBERT model advances Turkish NLP capabilities
Researchers have developed SindBERT, a new large-scale RoBERTa-based language model specifically for Turkish. Trained on over 300 GB of Turkish text, SindBERT is available in base and large configurations, marking the f…
-
New Turkish Embedding Model Achieves 8K Context Window
Researchers have developed embeddingmagibu-200m, a new Turkish-focused sentence embedding model that significantly enhances semantic search and related tasks. This model boasts a 768-dimensional vector output and an 8,1…
-
New Turkish embedding model achieves SOTA with efficient adaptation
Researchers have developed a new Turkish-focused sentence embedding model called embeddingmagibu-200m, which significantly outperforms larger teacher models while requiring fewer computational resources. The model was c…
-
Model collapse explained by cultural evolution theory
Researchers have reframed the phenomenon of model collapse, where large language models degrade when trained on their own outputs, as a cultural evolution process. By applying iterated learning theory, they derived and …
-
LLMs quantify syncretism's effect on language agreement errors
Researchers have investigated how morphological syncretism influences agreement attraction errors in verbs across different languages. Using large language models to measure processing proxies like surprisal and attenti…
-
The macOS Natural Language framework and Nalaprop https:// web.brid.gy/r/https://eclectic light.co/2026/04/22/the-macos-natural-language-framework-and-nalaprop/
The macOS Natural Language framework offers robust support for analyzing text in various languages, enabling applications to deploy custom machine learning models. While major Large Language Models are predominantly tra…