{
  "id": 3769383,
  "title": "Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2",
  "url": "https://urgent.news/2026/08/27/reduce-asr-inference-costs-by-75-with-nvidia-mps-on-amazon-ec2",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-27T16:05:10.000Z",
  "source": {
    "name": "AWS Machine Learning",
    "slug": "aws-machine-learning",
    "url": "https://aws.amazon.com/blogs/machine-learning/reduce-asr-inference-costs-by-75-with-nvidia-mps-on-amazon-ec2/"
  },
  "original_language": "en",
  "account": "This article explores how NVIDIA Multi-Process Service (MPS) can reduce automatic speech recognition (ASR) inference costs by 75% when deployed on Amazon Elastic Compute Cloud (Amazon EC2) GPU instances. ASR inference requests typically use only 15-20% of a GPU's compute capacity, leaving 80% idle. This inefficiency requires running 16 GPU instances to sustain sub-second transcription latency at peak traffic. By leveraging NVIDIA MPS combined with NVIDIA Triton Inference Server on EC2 GPU instances, GPU infrastructure requirements can be reduced by 75% (from 16 instances to 4), while maintaining sub-second latency at 92.1 requests per second per GPU. The article also highlights model-level optimizations with ONNX and TensorRT, as well as request scheduling with Triton, to further improve performance.",
  "summary": "Serving automatic speech recognition (ASR) models at scale is costly when each request uses only a fraction of a GPU. Learn how NVIDIA CUDA Multi-Process Service (MPS) with NVIDIA Triton Inference Server on Amazon EC2 GPU instances cuts GPU infrastructure by 75% while holding sub-second latency at 92.1 requests per second per GPU.",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}