AWS, NVIDIA, and Heidi Health collaborated to reduce automatic speech recognition (ASR) inference costs by 75% on Amazon EC2 instances. By implementing NVIDIA MPS with the NVIDIA Triton Inference Server, they achieved a 75% reduction in required GPU instances, from 16 down to 4, while maintaining sub-second latency. This optimization addresses the inefficiency of low GPU utilization per request in ASR, where traditional CUDA time-slicing leaves significant hardware capacity idle. AI
IMPACT Optimizes inference costs for ASR, potentially lowering operational expenses for AI services relying on speech recognition.
RANK_REASON This article describes a technical optimization for using existing hardware and software to reduce costs, rather than a new product release or frontier model.
Read on AWS Machine Learning Blog →
- Amazon Elastic Compute Cloud
- AWS
- CUDA
- Heidi Health
- Nemotron
- NVIDIA
- NVIDIA L40S GPU
- NVIDIA MPS
- NVIDIA Triton Inference Server
- ONNX
- tensorrt
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →