Vision Transformer Large
PulseAugur coverage of Vision Transformer Large — every cluster mentioning Vision Transformer Large across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
DCM-SAM system improves defect segmentation in metal additive manufacturing
Researchers have developed DCM-SAM, a novel system for segmenting defects in metal additive manufacturing parts using X-ray computed tomography. This system employs a defect-conditioned adaptive mixture of LoRA experts,…
-
Hyper-RED framework uses hypergraphs for scalable event camera pre-training
Researchers have developed Hyper-RED, a novel pre-training framework designed to improve event camera representation learning. This method utilizes semantic hypergraphs to transfer high-order semantic structures from im…
-
Bird ID models: Resolution vs. Architecture trade-offs on edge devices
A new study investigates the optimal input resolution for bird species identification models, particularly for edge devices like the NVIDIA Jetson Orin Nano. Researchers found that model architecture significantly impac…
-
New ARC-Bench protocol reveals critical flaws in frozen JEPA world models
Researchers have developed ARC-Bench, a new evaluation protocol designed to assess the action ranking capabilities of frozen latent world models. The study found that these models, which plan by scoring candidate action…
-
New ViTAMINS method enhances vision transformer training with synthetic negatives
Researchers have developed ViTAMINS, a novel method for training self-supervised vision transformers by incorporating synthetic hard negatives. This approach enhances representation quality, leading to significant impro…
-
New MemMTL framework enhances multi-task dense prediction with prototype memory
Researchers have developed MemMTL, a new framework for multi-task dense prediction that utilizes a learnable task-state prototype memory. This memory refines a compact task state derived from global visual context, whic…
-
New framework enhances cross-domain object detection with DINOv2
Researchers have developed a new framework called Semantic Localization-Enhanced Teacher (SLE-T) to improve cross-domain object detection using vision foundation models (VFMs). This method addresses issues like spatial-…
-
Ant Group's Lingbo releases suite of embodied AI models, including world action and video generation
Ant Group's Lingbo Technology has released a suite of new models aimed at advancing embodied AI and robotics. LingBot-VA 2.0 is presented as the first embodiment-native world action model, designed from the ground up fo…
-
New framework enables linear merging of billion-parameter transformers
Researchers have developed a new framework for merging large pretrained transformers, specifically those with billions of parameters. This method addresses limitations of previous approaches by optimizing interpolation …
-
New research explores merging large transformers and improving looped model stability
Two new research papers explore novel techniques for enhancing the capabilities and stability of large transformer models. The first paper introduces a scalable framework for linear mode connectivity (LMC) that allows f…
-
New Anatomy-Slot Method Enhances Retinal Diagnosis Accuracy
Researchers have developed a new unsupervised method called Anatomy-Slot for analyzing retinal images, which improves diagnostic accuracy by explicitly comparing homologous anatomical structures between the left and rig…
-
Anatomy-Slot method improves retinal diagnosis by modeling bilateral structures
Researchers have developed a new unsupervised method called Anatomy-Slot for retinal diagnosis that explicitly models the bilateral nature of the condition. By decomposing image patches into distinct slots and aligning …