PulseAugur
EN
LIVE 13:07:57
ENTITY ONNX

ONNX

PulseAugur coverage of ONNX — every cluster mentioning ONNX across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
17
56 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
6
17 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

10 day(s) with sentiment data

RECENT · PAGE 1/5 · 85 TOTAL
  1. TOOL · CL_259691 ·

    Guide: Run AI text embeddings on CPUs, not expensive GPUs

    A guide suggests that running text embedding models on expensive GPU hardware is an inefficient use of resources. The "SRE RAG FinOps Blueprint" proposes offloading embedding tasks to CPUs, leveraging optimizations like…

  2. TOOL · CL_257113 ·

    INT8 Quantization Portability Study Reveals Inconsistencies Across Hardware

    A new study challenges the assumption that INT8 quantization is universally portable across different hardware platforms for AI inference. Researchers found that INT8 speedups are heavily dependent on specific CPU instr…

  3. RESEARCH · CL_254750 ·

    New research explores VLA model efficiency and latency trade-offs · 2 sources tracked

    Two new research papers explore the efficiency and performance of Vision-Language-Action (VLA) models. The first paper analyzes SmolVLA, demonstrating how deployment optimizations like ONNX can significantly reduce late…

  4. TOOL · CL_252036 ·

    New tool streamlines AI model deployment via quantization analysis

    A new paper introduces the Quantization Analysis Tool, designed to optimize AI model deployment on resource-constrained devices. This tool, built on the ONNX framework, offers layer-wise sensitivity analysis and visuali…

  5. TOOL · CL_235906 ·

    Developer replaces cloud LLM with local Ollama for cost savings

    A developer has replaced the cloud-based generation component of their RAG chatbot with a local LLM, specifically Ollama running the qwen2.5-coder:32b model. This change was motivated by cost savings and privacy, tradin…

  6. RESEARCH · CL_233787 ·

    Dubai startups embrace On-Device AI for data sovereignty amid cloud surveillance concerns

    Startups in Dubai are increasingly adopting On-Device AI to comply with the UAE's upcoming 2026 Data Sovereignty laws. This shift is driven by concerns over cloud surveillance and the need to protect corporate data priv…

  7. TOOL · CL_240010 ·

    H3DNAS framework compresses 3D point cloud models for edge hardware via ONNX

    Researchers have developed H3DNAS, a novel framework designed to compress 3D point cloud models for deployment on edge hardware like the NVIDIA Jetson Orin Nano. Unlike existing methods that require original source code…

  8. RESEARCH · CL_233520 ·

    New H3DNAS framework compresses 3D point cloud models on ONNX binaries

    Researchers have developed H3DNAS, a novel framework for compressing 3D point cloud models that operates directly on ONNX binaries without needing original source code. This method addresses the limitations of deploying…

  9. TOOL · CL_232071 ·

    MiniMax-H3 VAE optimized for ComfyUI boosts speed up to 1.7x

    A new implementation of the MiniMax-H3 variational auto-encoder (VAE) is now available for ComfyUI, optimized with TensorRT. This version aims to significantly improve processing speed, with reported gains of up to 1.7 …

  10. TOOL · CL_227910 ·

    Serving YOLOv8 with NVIDIA Triton via ONNX and TensorRT

    This article details how to serve the YOLOv8 object detection model using NVIDIA Triton Inference Server. It explains the process of converting the ONNX format of YOLOv8 to TensorRT, a high-performance inference optimiz…

  11. TOOL · CL_224637 ·

    Hugging Face launches $399 open-source Microduck robot trained with RL

    Hugging Face's robotics team, Pollen Robotics, has launched Microduck, an open-source 25 cm bipedal robot available for pre-order at $399. This robot is designed for dynamic movement and self-recovery, with all its beha…

  12. TOOL · CL_224551 ·

    ExLlamaSharp v1.2.1-beta adds OpenAI-compatible API and EXL3 inference

    Kortexio has released ExLlamaSharp v1.2.1-beta, a local LLM server for Windows that supports NVIDIA GPUs. This beta version introduces OpenAI-compatible API endpoints, a Blazor admin interface, and enhanced EXL3 inferen…

  13. TOOL · CL_222176 ·

    AWS, NVIDIA, and Heidi Health cut ASR inference costs by 75%

    AWS, NVIDIA, and Heidi Health collaborated to reduce automatic speech recognition (ASR) inference costs by 75% on Amazon EC2 instances. By implementing NVIDIA MPS with the NVIDIA Triton Inference Server, they achieved a…

  14. TOOL · CL_221476 ·

    ONNX model deployment demonstrated across two apps and three processors

    This article provides a practical guide to deploying a single ONNX sentiment model across multiple applications and hardware configurations. It details a hands-on walkthrough using Windows ML to demonstrate how one mode…

  15. RESEARCH · CL_221276 ·

    AI frameworks tackle plant disease diagnosis and fruit classification

    Researchers have developed advanced AI frameworks for agricultural applications, focusing on plant disease diagnosis and fruit classification. The first study introduces H²MAF, which fuses vision models like EfficientNe…

  16. TOOL · CL_221131 ·

    New ONNX-Net system enables universal neural architecture representation

    Researchers have developed ONNX-Net, a novel approach to create universal representations for neural architectures, aiming to overcome the limitations of existing methods tied to specific search spaces. This system util…

  17. TOOL · CL_221085 ·

    New SAMpLE framework integrates ML models into SystemC-AMS virtual prototypes

    Researchers have developed SAMpLE, an open-source framework that integrates machine learning models into SystemC-AMS virtual prototypes. This framework allows ML models to function as first-class Timed Dataflow componen…

  18. TOOL · CL_220638 ·

    AI models now run directly in browsers using ONNX Runtime Web

    ONNX Runtime Web is enabling complex AI tasks like background removal and feature extraction to be performed directly within a web browser. This client-side processing eliminates the need for powerful backend servers, r…

  19. RESEARCH · CL_218435 ·

    Chengtai Technology CTO unveils AI Radar, a paradigm shift for radar development

    Chengtai Technology's CTO, Zhou Ke, has introduced the concept of "AI Radar," which represents a paradigm shift from traditional radar systems. Unlike conventional radars with fixed firmware, AI Radars utilize trainable…

  20. TOOL · CL_217599 ·

    Microsoft embeds invisible AI watermarks in Paint and Photos apps

    Microsoft's Paint and Photos applications now embed invisible watermarks in AI-generated images, according to developer Xusheng Li. This invisible watermark, which includes a server-issued GUID, is mandatory for AI imag…