Donut
PulseAugur coverage of Donut — every cluster mentioning Donut across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
OpenAI reportedly developing AI donut speaker with Jony Ive
OpenAI is reportedly developing its first piece of hardware, an AI-powered speaker shaped like a donut, codenamed "Donut." Designed in collaboration with Jony Ive, the device is intended to follow users around their hom…
-
Small model extracts text from white backgrounds, inspired by DONUT
A user on Reddit's r/MachineLearning subreddit has developed a small model capable of extracting text from images with a white background. Inspired by the DONUT model, the project initially aimed to extract information …
-
New synthetic dataset boosts Persian OCR capabilities
Researchers have introduced Persian Pixel, a large-scale synthetic dataset designed to improve Optical Character Recognition (OCR) for the Persian language. The dataset contains over 343,000 image-text pairs, generated …
-
New framework decouples trajectory forecasting from benchmark metrics
Researchers have proposed a new framework for trajectory forecasting in autonomous driving that decouples the training objective from specific benchmark metrics. This approach, called Trajectory Distribution Evaluation …
-
New 'Counterfeit Answers' attack targets OCR-free DocVQA models
Researchers have developed a novel adversarial attack method called "Counterfeit Answers" that can forge document content to manipulate OCR-free Document Visual Question Answering (DocVQA) models. This attack can induce…
-
Research compares multimodal models for document classification
A new research paper analyzes multimodal approaches for classifying visually-rich documents, comparing transformer and LLM-based architectures. The study evaluated LayoutLMv3, Donut, Qwen3-VL-32B-Instruct, and Qwen3-32B…
-
FILTR framework extracts topological features from 3D models using transformers
Researchers have developed FILTR, a novel framework designed to extract topological features from pretrained 3D models. This approach adapts a transformer decoder to generate persistence diagrams, which summarize a shap…