Vít
PulseAugur coverage of Vít — every cluster mentioning Vít across labs, papers, and developer communities, ranked by signal.
- instance of Swin Transformer 90%
- instance of Data Efficient Image Transformers 90%
- instance of Vision Transformers 90%
- used by DagsHub 70%
- used by ScienceCast 70%
- instance of alphaXiv 70%
- used by alphaXiv 70%
- used by Gotit.pub 70%
- used by residual neural network 70%
- used by CatalyzeX 70%
- used by Swin Transformer 70%
- instance of residual neural network 70%
16 day(s) with sentiment data
-
Deep learning models for breast cancer detection benchmarked for performance and emissions
A new paper benchmarks seven deep learning models for breast cancer detection, evaluating their performance and environmental impact. The study found that while EfficientNet and ResNet offer strong accuracy, they also p…
-
New AI model estimates eye health from fundus photos
Researchers have developed SpecF2M, a novel multi-task network designed to estimate axial length and refractive error from pediatric fundus photographs. This spectral-aware network integrates an anatomy-guided enhanceme…
-
Brain MRI Foundation Models Primarily Encode Acquisition Site, Not Anatomy
Researchers have discovered that frozen foundation models, when used to represent brain MRI data, primarily encode the site where the MRI was acquired rather than anatomical or clinical information. This effect was obse…
-
New iBKD framework transfers CNN inductive biases to Vision Transformers under data scarcity
Researchers have developed a new knowledge distillation framework called iBKD, designed to improve the performance of Vision Transformers (ViTs) when training data is limited. This method effectively transfers the induc…
-
New methods tackle LLM and VLM hallucinations with internal analysis · 2 sources tracked
Researchers have developed new methods to detect hallucinations in large language and vision-language models. UniProbe, a technique for Large VLMs, uses a graph neural network, a Vision Transformer, and a gated recurren…
-
New AQUA20 dataset targets challenging underwater species classification
Researchers have introduced AQUA20, a new benchmark dataset designed to improve underwater species classification. The dataset contains 8,171 images of 20 marine species, specifically curated to address challenges like …
-
New Visual Token Codec Boosts ViT Feature Compression Efficiency
Researchers have developed a new method called the Visual Token Codec (VTC) to compress intermediate features in Vision Transformer (ViT) models. VTC effectively utilizes the spatial correlations inherent in ViT patch t…
-
New AI Model XEns-CKD Improves Chronic Kidney Disease Detection Accuracy
Researchers have developed XEns-CKD, a new ensemble vision transformer model for detecting chronic kidney disease (CKD) stages from ultrasound images. This model, trained on a private dataset, achieved an 86.36% classif…
-
AI training efficiency: Gradient optimization methods benchmarked
A new research paper benchmarks five gradient optimizers and three memory strategies for AI training on constrained hardware. The study found that gradient accumulation is the most effective strategy for reducing traini…
-
PolypSteer framework generates synthetic medical images for AI training
Researchers have developed PolypSteer, a novel framework designed to generate synthetic medical imaging data for training AI models. This training-free approach uses activation steering within diffusion transformers to …
-
AI advances radiology report generation with controllable outputs and efficient processing
Researchers have developed new frameworks for generating radiology reports from medical images, addressing limitations in current AI models. One approach, RadFusion, integrates a classifier with a vision-language model …
-
UniSpace introduces unified visual representation for AI generation and editing
Researchers have developed UniSpace, a novel approach to visual representation that unifies understanding, generation, and editing tasks within a single model. By introducing "Patch Reparameterization," UniSpace modifie…
-
Meta releases open-source multimodal model Muse Glimmer
Meta has released Muse Glimmer, an open-source, multimodal, and agentic large language model. The model features a 30 billion parameter architecture that includes a 2 billion parameter vision encoder and a 28 billion pa…
-
SEED system offers explainable detection for AI-generated text forgeries
Researchers have developed SEED, a system designed to detect and explain AI-generated text forgeries. This system, which ranked third in the GenText-Forensics Challenge at ACM MM 2026, utilizes a Vision Transformer (ViT…
-
Huawei IJCAI 2026 papers highlight efficiency gains in AI models · 1 source tracked
Huawei presented four papers at IJCAI-ECAI 2026, shifting focus from scaling model size to optimizing efficiency and design. One paper details a hierarchical Vision Transformer (ViT) scaled to 30 billion parameters, ach…
-
PADFormer uses Vision Transformer for pose-agnostic anomaly detection
Researchers have introduced PADFormer, a new approach for detecting anomalies in images that can handle significant pose variations without relying on complex 3D reconstruction. This method utilizes a Vision Transformer…
-
New Transformer Architecture Optimizes Self-Supervised Learning
Researchers have developed an attention-only white-box Transformer model by integrating the LeJEPA self-supervised learning framework with optimization algorithms. This approach optimizes the sparse rate reduction objec…
-
New APQF framework automates AI model compression with LLM guidance
Researchers have developed APQF, an automated framework designed to optimize deep neural networks for efficiency on edge devices. This system uses an agentic approach, guided by LLM planners and profiling data, to deter…
-
New research explores VLM vs. vision-only models for autonomous driving
Researchers have developed a new approach to end-to-end driving systems by comparing vision-language models (VLMs) with traditional vision-only encoders. Their study found that while both types of models share significa…
-
AI framework extracts rich embeddings from microscopy images
Researchers have developed an AI framework to extract semantically rich image embeddings from optical microscopy images of particles and fibers. This system uses a multimodal teacher that combines visual embeddings with…