vision transformer
PulseAugur coverage of vision transformer — every cluster mentioning vision transformer across labs, papers, and developer communities, ranked by signal.
- competes with convolutional neural network 80%
- instance of Vít 70%
- used by Vít 70%
- used by Mae 70%
- used by residual neural network 70%
- used by ImageNet ILSVRC-2012 70%
- instance of residual neural network 70%
- used by ResNet50 70%
- used by Gotit.pub 60%
- used by alphaXiv 60%
- used by ScienceCast 60%
- uses convolutional neural network 60%
18 day(s) with sentiment data
-
New model drastically cuts facial emotion recognition computation
Researchers have developed a new model called Sparse Attention to Emotion (SAE) for facial emotion recognition. This model significantly reduces computational complexity by discarding up to 90% of image tokens, focusing…
-
AI model screens Parkinson's disease using face and voice without labels
Researchers have developed a novel method for screening Parkinson's disease using only facial expressions and voice analysis, eliminating the need for direct PD labels. This approach leverages frozen pretrained encoders…
-
New Vision Transformer framework enhances sparse dToF depth completion
Researchers have developed a new framework for dense metric depth completion from sparse direct Time-of-Flight (dToF) sensors. This method utilizes a depth-guided dual-branch Vision Transformer encoder that processes RG…
-
PADFormer uses Vision Transformer for pose-agnostic anomaly detection
Researchers have introduced PADFormer, a new approach for detecting anomalies in images that can handle significant pose variations without relying on complex 3D reconstruction. This method utilizes a Vision Transformer…
-
New AI framework mimics radiologist reasoning for chest X-ray analysis
Researchers have developed CHASE (Classification with Hierarchical Analysis and Structured Enforcement), a novel framework designed to improve the accuracy and consistency of chest X-ray interpretation by AI. Unlike pre…
-
Deep learning models enhance anomaly detection in zero-trust networks
Researchers have developed two deep learning models, a Vision Transformer (ViT) and a 1D Convolutional Neural Network (1D-CNN), to enhance anomaly detection within zero-trust software-defined networks. These models anal…
-
DeVIT method accelerates vision transformers using delta computation
Researchers have developed DeVIT, a novel method to accelerate vision transformers by utilizing delta computation for multiplier-less matrix multiplication. This approach aims to reduce the computational complexity and …
-
New SSR method refines object-centric masks without retraining
Researchers have developed a new training-free method called Similarity-Shift Refinement (SSR) to improve object-centric masks generated by vision transformers. SSR analyzes changes in patch similarity within the self-a…
-
Foundation model pretraining strategies impact retinal imaging transferability
A new arXiv paper explores how different pretraining strategies for foundation models impact their effectiveness when transferred to ultra-widefield retinal imaging tasks. Researchers compared Vision Transformer encoder…
-
AI models for COVID-19 CT scan classification show promising results · 2 papers
Two research papers submitted to arXiv propose novel methods for classifying COVID-19 cases from chest CT scans. The first paper introduces a hybrid 2D/3D Convolutional Neural Network (CNN) that extracts features from m…
-
BladeYOLO framework enhances wind turbine defect detection with limited data
Researchers have developed BladeYOLO, a new framework designed to improve the detection of defects on wind turbine blades, particularly in scenarios with limited annotated data. The system integrates a Vision Transforme…
-
New SAFViT module enhances nucleus segmentation in digital pathology
Researchers have introduced SAFViT, a novel Spatial Attention Fusion Gating module designed to enhance nucleus segmentation and classification in digital pathology. This module improves upon existing encoder-decoder arc…
-
JEPADepth framework enhances self-supervised monocular depth estimation
Researchers have developed JEPADepth, a novel self-supervised framework for monocular depth estimation that integrates a masked predictive representation learning objective inspired by Image Joint-Embedding Predictive A…
-
Panda framework enables real-time unsupervised anomaly detection in pelvic MRI
Researchers have developed a novel unsupervised anomaly detection framework called Panda, designed for real-time pelvic MRI imaging. This system utilizes a frozen DINOv3 Vision Transformer encoder and a noisy MLP bottle…
-
New DDVT Network Enhances Visual Question Answering Accuracy
Researchers have developed a new network architecture called the Dynamic Dual-level Vision Transformer Fusion Network (DDVT) for answer grounding in visual question answering. This approach aims to precisely locate imag…
-
MAViE encoder boosts vision-language model efficiency by 80%
Researchers have introduced MAViE, a Multi-scale Adaptive Vision Encoder designed to improve the efficiency and effectiveness of vision-language models. MAViE utilizes position-dependent gates to integrate features from…
-
Foundation models make multi-branch fusion less effective for vehicle re-ID
A new research paper questions the effectiveness of multi-branch architectures and the fusion of different backbone models for vehicle re-identification tasks in the era of foundation models. The study found that a sing…
-
New AI Vulnerability: Physical Objects Can Hijack Vision-Language Models
Researchers have identified a new class of vulnerabilities in multimodal Vision-Language Models (VLMs) called Physical Prompt Injection Attacks (PPIA). These attacks exploit the way VLMs process visual information, allo…
-
Vision Transformer and FFT-ReLU integrated for enhanced image deblurring
Researchers have developed a novel dual-domain architecture for image deblurring that integrates Vision Transformers (ViTs) with a frequency-domain FFT-ReLU module. This approach aims to enhance the recovery of sharp im…
-
New multimodal EEG foundation model achieves state-of-the-art in epilepsy detection
Researchers have developed a multimodal foundation model for electroencephalography (EEG) data, aiming to improve generalizability in epilepsy detection. The model integrates a Mamba-based raw signal encoder, a Vision T…