Brazilian Portuguese
PulseAugur coverage of Brazilian Portuguese — every cluster mentioning Brazilian Portuguese across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New pipeline curates 41B token European Portuguese web corpus for LLMs
Researchers have developed a new pipeline for curating high-quality web corpora specifically for European Portuguese (PT-PT). This pipeline addresses challenges like dialectal overlap with Brazilian Portuguese and the s…
-
New benchmark evaluates vehicle attribute classification in surveillance
Researchers have introduced the Unconstrained Vehicle Identification Benchmark (UVIB) to evaluate vehicle attribute classification in diverse surveillance scenarios. This benchmark, comprising 84,835 images from seven B…
-
New Brazilian face recognition benchmark addresses demographic bias
Researchers have introduced UFPR-PEs, a new benchmark for evaluating face recognition bias. This dataset utilizes public videos of Brazilian politicians, annotated with self-declared race and color labels that include t…
-
LLMs show promise for Portuguese legal IR relevance assessment, despite bias
Researchers have developed NormasTCU, a new dataset for Brazilian Portuguese Information Retrieval (IR) that includes 14,469 legal documents and human relevance judgments. The study evaluated the effectiveness of using …
-
Plaintiff hides AI prompt injection in court filing
A plaintiff in a Connecticut court case attempted to use prompt injection by hiding instructions within a court filing using white text on a white background. The hidden text directed any AI model processing the documen…
-
NVIDIA Magpie TTS adds 3 languages, now supports 12 total
NVIDIA has expanded its Magpie Text-to-Speech (TTS) model to support 12 languages, including the recent additions of Arabic, Korean, and Brazilian Portuguese. This open-weights model, featuring 364 million parameters, i…
-
New AI framework enhances hate speech detection with moral rationales
Researchers have developed a novel framework called Supervised Moral Rationale Attention (SMRA) to improve the interpretability and robustness of hate speech detection models. Unlike previous methods that relied on surf…
-
DharmaOCR achieves superior performance on Brazilian Portuguese OCR
DharmaOCR, an OCR model specialized for Brazilian Portuguese, has demonstrated superior performance compared to models like Mistral OCR4 and Unlimited-OCR. This advantage stems from a two-stage training process: initial…
-
New Whisper-based system improves prosodic boundary detection in Brazilian Portuguese
Researchers have developed SAMPA, a new system for automatically segmenting prosodic boundaries in Brazilian Portuguese speech. This system is based on fine-tuning the Whisper large-v3 model, a significant advancement o…
-
New benchmarks evaluate Portuguese text embedding models, revealing performance gaps
Two new benchmarks, MTEB-PT and MTEB-PT (Brazilian Portuguese), have been released to evaluate text embedding models specifically for the Portuguese language. These benchmarks address the underrepresentation of Portugue…
-
New framework TOTEN improves tokenization of technical notation
Researchers have developed TOTEN, a knowledge-based ontological tokenization framework designed to improve the semantic understanding of technical notation in Brazilian Portuguese. Unlike traditional byte-pair encoding,…
-
New benchmark reveals LLM bias towards Brazilian Portuguese
A new benchmark called P3B3 has been developed to assess how large language models (LLMs) handle variations in Portuguese, specifically European Portuguese (pt-PT) and Brazilian Portuguese (pt-BR). The benchmark aims to…
-
New benchmark tests clinical LLMs in Brazilian Portuguese
Researchers have developed ClinicalBr, a new bilingual benchmark for evaluating clinical Large Language Models in Brazilian Portuguese and English. The benchmark, derived from real Brazilian medical case reports, covers…
-
User seeks help fine-tuning Kokoro for Brazilian Portuguese
A user is seeking advice on locally installing and fine-tuning the Kokoro language model, specifically for Brazilian Portuguese. They are experiencing poor performance with non-English languages when using the Open Rout…
-
New method extracts accent features from Portuguese speech using acoustic labels
Researchers have developed a new method to extract accent features from spoken Brazilian Portuguese without relying on sociolinguistic labels. This approach uses acoustic labels and a phoneme-based forced aligner to iso…
-
FalAR corpus boosts European Portuguese ASR with 5,800 hours of parliamentary data
Researchers have introduced FalAR, a new large-scale speech corpus for European Portuguese parliamentary sessions, aiming to improve Automatic Speech Recognition (ASR) for the language. The corpus contains approximately…
-
New benchmark 'Prosa' evaluates LLMs on Brazilian Portuguese chats
Researchers have introduced Prosa, a new benchmark designed to evaluate Large Language Models (LLMs) using real user conversations in Brazilian Portuguese. This benchmark utilizes a rubric-based scoring system with mult…
-
New LLM bias benchmark measures opinion and sycophancy in AI assistants
Researchers have developed a new open-source method called llm-bias-bench to uncover the hidden opinions of large language models on contentious subjects. The technique employs two distinct probing strategies: direct qu…