MediaPipe
PulseAugur coverage of MediaPipe — every cluster mentioning MediaPipe across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
Low-cost reservoir computing model advances sign language recognition
Researchers have developed a novel, low-cost hybrid reservoir computing model for recognizing isolated sign language videos. This model utilizes MediaPipe to extract key body and hand points, which are then processed by…
-
New RSC-GestureNet system enhances traffic gesture recognition for autonomous driving
Researchers have developed RSC-GestureNet, a new system designed to reliably recognize Chinese traffic police gestures for autonomous driving applications. This model incorporates pose confidence as a key factor, down-w…
-
OpenAI's GPT Image 2 simplifies image editing, replacing complex tools
OpenAI's GPT Image 2, also known as ChatGPT Images 2.0, offers a streamlined approach to image editing, replacing the need for complex local setups like ADetailer with Stable Diffusion. Users can now make single-frame e…
-
Teenager builds fully local AI ecosystem CODA OS
A 17-year-old developer has created CODA OS, a comprehensive AI ecosystem designed to run entirely on local hardware without cloud dependencies. The system includes a 3D reconstruction tool, a custom 2.2B parameter lang…
-
New Transformer Model Enhances Face Recognition with Masked Faces
Researchers have developed PLGSA-Transformer, a novel framework for face recognition that addresses the challenges posed by facial masks. This system utilizes periocular landmark-guided spatial attention to focus on vis…
-
Google's webcam hand-scan reCAPTCHA quickly bypassed by testers · 3 sources tracked
Google is testing a new reCAPTCHA system that uses a webcam to scan users' hands, mapping 21 points to verify human identity. This experimental feature, part of Google Cloud Fraud Defense, aims to combat bots more effec…
-
New Gated Affect Transformer improves human motion prediction
Researchers have developed a new method called the Gated Affect Transformer (GAT) to improve human motion prediction by integrating facial affect cues with body pose data. The study found that simply combining these mod…
-
AI touch detection framework struggles with real-world mobile typing reconstruction
A new research paper details a multi-modal framework for detecting touch events on mobile keypads using video surveillance. The system integrates hand landmark detection, skin color filtering, motion detection, and edge…
-
Virtual ring try-on system uses AI for realistic placement
Researchers have developed a novel virtual ring try-on system that allows users to see how a ring would look on their hand. The system uses MediaPipe for hand point detection and YOLO-V8 for ring object detection to acc…
-
New AI tool Envisage visualizes rhinoplasty outcomes with novel evaluation metric
Researchers have developed Envisage, a new pipeline for visualizing the intended outcomes of rhinoplasty surgery using diffusion-based generative editing. This system is designed to provide localized edits and includes …
-
Withdrawn research paper details CPU-based fall detection system
A research paper, now withdrawn, proposed a fall detection system for elderly care that operates efficiently on standard CPUs. The system utilizes pose estimation via the MediaPipe framework to analyze motion and body p…
-
AI framework digitizes athlete profiling with VLM and RAG
Researchers have developed a novel LLM-based framework for holistic athlete profiling, designed to overcome the limitations of traditional manual or basic computer vision assessment methods. This agentic system, orchest…
-
AR hand pose estimation accurate for impaired hands
A new study published on arXiv investigates the accuracy of hand pose estimation in augmented reality (AR) applications, particularly for individuals with hand impairments. Researchers compared the HoloLens 2 HMD with s…
-
Open-source EyeTheia toolbox offers webcam-based gaze estimation
Researchers have developed EyeTheia, an open-source, lightweight deep learning pipeline for gaze estimation using standard webcams. The system combines landmark extraction with a convolutional neural network, offering r…
-
AI system offers real-time athletic performance analysis
Researchers have developed a lightweight prototype for real-time athletic performance analysis using markerless deep learning. The system integrates Human Pose Estimation (HPE) with exercise-specific logic to provide AI…
-
Open-source AI meeting platform Hoovik faces real-time inference challenges
Anupam Kumar, the creator of the open-source AI meeting platform Hoovik, found that the most challenging aspect of development was not the core WebRTC technology but managing real-time multimodal AI inference. This invo…
-
AI game "Hand Gesture at Doc Yang" uses MediaPipe for live hand tracking
A new browser-based game called "Hand Gesture at Doc Yang" utilizes AI, specifically MediaPipe, to detect and interpret live hand gestures. Players can score points by displaying two hand gestures simultaneously, with e…
-
ProxyFace adds local, emotional avatars to AI chats
ProxyFace is an open-source project that adds a local, expressive avatar to AI interactions. It utilizes a small, on-device emotion model and eye-tracking to make the avatar react to AI output and the user's gaze. The p…
-
AI research targets efficient, accessible sign language translation
Two new research papers explore advancements in sign language translation (SLT) technology, focusing on making systems more efficient and accessible for low-resource languages. One paper proposes a data-centric approach…
-
Tamaththul3D creates high-fidelity 3D Saudi Sign Language avatars from video
Researchers have developed Tamaththul3D, a novel pipeline for generating high-fidelity 3D avatars of Saudi Sign Language (SSL). This system addresses a significant gap in resources for Arabic Sign Language (ArSL), which…