PulseAugur
EN
LIVE 18:10:21
ENTITY vision transformer

vision transformer

PulseAugur coverage of vision transformer — every cluster mentioning vision transformer across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
42
143 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
41
141 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

18 day(s) with sentiment data

RECENT · PAGE 1/8 · 143 TOTAL
  1. TOOL · CL_193888 ·

    New model drastically cuts facial emotion recognition computation

    Researchers have developed a new model called Sparse Attention to Emotion (SAE) for facial emotion recognition. This model significantly reduces computational complexity by discarding up to 90% of image tokens, focusing…

  2. TOOL · CL_193817 ·

    AI model screens Parkinson's disease using face and voice without labels

    Researchers have developed a novel method for screening Parkinson's disease using only facial expressions and voice analysis, eliminating the need for direct PD labels. This approach leverages frozen pretrained encoders…

  3. TOOL · CL_185499 ·

    New Vision Transformer framework enhances sparse dToF depth completion

    Researchers have developed a new framework for dense metric depth completion from sparse direct Time-of-Flight (dToF) sensors. This method utilizes a depth-guided dual-branch Vision Transformer encoder that processes RG…

  4. TOOL · CL_185476 ·

    PADFormer uses Vision Transformer for pose-agnostic anomaly detection

    Researchers have introduced PADFormer, a new approach for detecting anomalies in images that can handle significant pose variations without relying on complex 3D reconstruction. This method utilizes a Vision Transformer…

  5. TOOL · CL_183399 ·

    New AI framework mimics radiologist reasoning for chest X-ray analysis

    Researchers have developed CHASE (Classification with Hierarchical Analysis and Structured Enforcement), a novel framework designed to improve the accuracy and consistency of chest X-ray interpretation by AI. Unlike pre…

  6. TOOL · CL_183351 ·

    Deep learning models enhance anomaly detection in zero-trust networks

    Researchers have developed two deep learning models, a Vision Transformer (ViT) and a 1D Convolutional Neural Network (1D-CNN), to enhance anomaly detection within zero-trust software-defined networks. These models anal…

  7. TOOL · CL_180986 ·

    DeVIT method accelerates vision transformers using delta computation

    Researchers have developed DeVIT, a novel method to accelerate vision transformers by utilizing delta computation for multiplier-less matrix multiplication. This approach aims to reduce the computational complexity and …

  8. TOOL · CL_180968 ·

    New SSR method refines object-centric masks without retraining

    Researchers have developed a new training-free method called Similarity-Shift Refinement (SSR) to improve object-centric masks generated by vision transformers. SSR analyzes changes in patch similarity within the self-a…

  9. TOOL · CL_180934 ·

    Foundation model pretraining strategies impact retinal imaging transferability

    A new arXiv paper explores how different pretraining strategies for foundation models impact their effectiveness when transferred to ultra-widefield retinal imaging tasks. Researchers compared Vision Transformer encoder…

  10. RESEARCH · CL_178517 ·

    AI models for COVID-19 CT scan classification show promising results · 2 papers

    Two research papers submitted to arXiv propose novel methods for classifying COVID-19 cases from chest CT scans. The first paper introduces a hybrid 2D/3D Convolutional Neural Network (CNN) that extracts features from m…

  11. TOOL · CL_174296 ·

    BladeYOLO framework enhances wind turbine defect detection with limited data

    Researchers have developed BladeYOLO, a new framework designed to improve the detection of defects on wind turbine blades, particularly in scenarios with limited annotated data. The system integrates a Vision Transforme…

  12. TOOL · CL_174280 ·

    New SAFViT module enhances nucleus segmentation in digital pathology

    Researchers have introduced SAFViT, a novel Spatial Attention Fusion Gating module designed to enhance nucleus segmentation and classification in digital pathology. This module improves upon existing encoder-decoder arc…

  13. TOOL · CL_172011 ·

    JEPADepth framework enhances self-supervised monocular depth estimation

    Researchers have developed JEPADepth, a novel self-supervised framework for monocular depth estimation that integrates a masked predictive representation learning objective inspired by Image Joint-Embedding Predictive A…

  14. TOOL · CL_167862 ·

    Panda framework enables real-time unsupervised anomaly detection in pelvic MRI

    Researchers have developed a novel unsupervised anomaly detection framework called Panda, designed for real-time pelvic MRI imaging. This system utilizes a frozen DINOv3 Vision Transformer encoder and a noisy MLP bottle…

  15. TOOL · CL_167821 ·

    New DDVT Network Enhances Visual Question Answering Accuracy

    Researchers have developed a new network architecture called the Dynamic Dual-level Vision Transformer Fusion Network (DDVT) for answer grounding in visual question answering. This approach aims to precisely locate imag…

  16. TOOL · CL_179288 ·

    MAViE encoder boosts vision-language model efficiency by 80%

    Researchers have introduced MAViE, a Multi-scale Adaptive Vision Encoder designed to improve the efficiency and effectiveness of vision-language models. MAViE utilizes position-dependent gates to integrate features from…

  17. TOOL · CL_165133 ·

    Foundation models make multi-branch fusion less effective for vehicle re-ID

    A new research paper questions the effectiveness of multi-branch architectures and the fusion of different backbone models for vehicle re-identification tasks in the era of foundation models. The study found that a sing…

  18. TOOL · CL_164251 ·

    New AI Vulnerability: Physical Objects Can Hijack Vision-Language Models

    Researchers have identified a new class of vulnerabilities in multimodal Vision-Language Models (VLMs) called Physical Prompt Injection Attacks (PPIA). These attacks exploit the way VLMs process visual information, allo…

  19. TOOL · CL_160703 ·

    Vision Transformer and FFT-ReLU integrated for enhanced image deblurring

    Researchers have developed a novel dual-domain architecture for image deblurring that integrates Vision Transformers (ViTs) with a frequency-domain FFT-ReLU module. This approach aims to enhance the recovery of sharp im…

  20. TOOL · CL_160694 ·

    New multimodal EEG foundation model achieves state-of-the-art in epilepsy detection

    Researchers have developed a multimodal foundation model for electroencephalography (EEG) data, aiming to improve generalizability in epilepsy detection. The model integrates a Mamba-based raw signal encoder, a Vision T…