ONNX
PulseAugur coverage of ONNX — every cluster mentioning ONNX across labs, papers, and developer communities, ranked by signal.
- used by WebGPU 90%
- used by DJ software 90%
- used by tensorrt 80%
- used by qdrant 70%
- used by graphics processing unit 70%
- used by Raspberry Pi 5 70%
- used by NVIDIA Jetson Orin Nano 8GB 70%
- used by Audio Developer Conference 70%
- used by Anmol Mishra 70%
- affiliated with tensorrt 60%
- used by Int8 60%
- instance of Apache Software License 2.0 60%
10 day(s) with sentiment data
-
Guide: Run AI text embeddings on CPUs, not expensive GPUs
A guide suggests that running text embedding models on expensive GPU hardware is an inefficient use of resources. The "SRE RAG FinOps Blueprint" proposes offloading embedding tasks to CPUs, leveraging optimizations like…
-
INT8 Quantization Portability Study Reveals Inconsistencies Across Hardware
A new study challenges the assumption that INT8 quantization is universally portable across different hardware platforms for AI inference. Researchers found that INT8 speedups are heavily dependent on specific CPU instr…
-
New research explores VLA model efficiency and latency trade-offs · 2 sources tracked
Two new research papers explore the efficiency and performance of Vision-Language-Action (VLA) models. The first paper analyzes SmolVLA, demonstrating how deployment optimizations like ONNX can significantly reduce late…
-
New tool streamlines AI model deployment via quantization analysis
A new paper introduces the Quantization Analysis Tool, designed to optimize AI model deployment on resource-constrained devices. This tool, built on the ONNX framework, offers layer-wise sensitivity analysis and visuali…
-
Developer replaces cloud LLM with local Ollama for cost savings
A developer has replaced the cloud-based generation component of their RAG chatbot with a local LLM, specifically Ollama running the qwen2.5-coder:32b model. This change was motivated by cost savings and privacy, tradin…
-
Dubai startups embrace On-Device AI for data sovereignty amid cloud surveillance concerns
Startups in Dubai are increasingly adopting On-Device AI to comply with the UAE's upcoming 2026 Data Sovereignty laws. This shift is driven by concerns over cloud surveillance and the need to protect corporate data priv…
-
H3DNAS framework compresses 3D point cloud models for edge hardware via ONNX
Researchers have developed H3DNAS, a novel framework designed to compress 3D point cloud models for deployment on edge hardware like the NVIDIA Jetson Orin Nano. Unlike existing methods that require original source code…
-
New H3DNAS framework compresses 3D point cloud models on ONNX binaries
Researchers have developed H3DNAS, a novel framework for compressing 3D point cloud models that operates directly on ONNX binaries without needing original source code. This method addresses the limitations of deploying…
-
MiniMax-H3 VAE optimized for ComfyUI boosts speed up to 1.7x
A new implementation of the MiniMax-H3 variational auto-encoder (VAE) is now available for ComfyUI, optimized with TensorRT. This version aims to significantly improve processing speed, with reported gains of up to 1.7 …
-
Serving YOLOv8 with NVIDIA Triton via ONNX and TensorRT
This article details how to serve the YOLOv8 object detection model using NVIDIA Triton Inference Server. It explains the process of converting the ONNX format of YOLOv8 to TensorRT, a high-performance inference optimiz…
-
Hugging Face launches $399 open-source Microduck robot trained with RL
Hugging Face's robotics team, Pollen Robotics, has launched Microduck, an open-source 25 cm bipedal robot available for pre-order at $399. This robot is designed for dynamic movement and self-recovery, with all its beha…
-
ExLlamaSharp v1.2.1-beta adds OpenAI-compatible API and EXL3 inference
Kortexio has released ExLlamaSharp v1.2.1-beta, a local LLM server for Windows that supports NVIDIA GPUs. This beta version introduces OpenAI-compatible API endpoints, a Blazor admin interface, and enhanced EXL3 inferen…
-
AWS, NVIDIA, and Heidi Health cut ASR inference costs by 75%
AWS, NVIDIA, and Heidi Health collaborated to reduce automatic speech recognition (ASR) inference costs by 75% on Amazon EC2 instances. By implementing NVIDIA MPS with the NVIDIA Triton Inference Server, they achieved a…
-
ONNX model deployment demonstrated across two apps and three processors
This article provides a practical guide to deploying a single ONNX sentiment model across multiple applications and hardware configurations. It details a hands-on walkthrough using Windows ML to demonstrate how one mode…
-
AI frameworks tackle plant disease diagnosis and fruit classification
Researchers have developed advanced AI frameworks for agricultural applications, focusing on plant disease diagnosis and fruit classification. The first study introduces H²MAF, which fuses vision models like EfficientNe…
-
New ONNX-Net system enables universal neural architecture representation
Researchers have developed ONNX-Net, a novel approach to create universal representations for neural architectures, aiming to overcome the limitations of existing methods tied to specific search spaces. This system util…
-
New SAMpLE framework integrates ML models into SystemC-AMS virtual prototypes
Researchers have developed SAMpLE, an open-source framework that integrates machine learning models into SystemC-AMS virtual prototypes. This framework allows ML models to function as first-class Timed Dataflow componen…
-
AI models now run directly in browsers using ONNX Runtime Web
ONNX Runtime Web is enabling complex AI tasks like background removal and feature extraction to be performed directly within a web browser. This client-side processing eliminates the need for powerful backend servers, r…
-
Chengtai Technology CTO unveils AI Radar, a paradigm shift for radar development
Chengtai Technology's CTO, Zhou Ke, has introduced the concept of "AI Radar," which represents a paradigm shift from traditional radar systems. Unlike conventional radars with fixed firmware, AI Radars utilize trainable…
-
Microsoft embeds invisible AI watermarks in Paint and Photos apps
Microsoft's Paint and Photos applications now embed invisible watermarks in AI-generated images, according to developer Xusheng Li. This invisible watermark, which includes a server-issued GUID, is mandatory for AI imag…