Articles liés à PRODUCTION ONNX: Model Conversion, Runtime Optimization,...

PRODUCTION ONNX: Model Conversion, Runtime Optimization, Hardware Acceleration, and AI Deployment - Couverture souple

Aurev, Zyren

 
9798171013226: PRODUCTION ONNX: Model Conversion, Runtime Optimization, Hardware Acceleration, and AI Deployment

Synopsis

MASTER ONNX AND BUILD PRODUCTION-READY AI INFERENCE SYSTEMS.

Machine learning does not end with model training. To use an AI model in the real world, you need reliable model conversion, optimized inference, hardware acceleration, scalable deployment, and production monitoring.

PRODUCTION ONNX: Model Conversion, Runtime Optimization, Hardware Acceleration, and AI Deployment is a practical technical guide to ONNX and ONNX Runtime, designed for developers, AI engineers, machine-learning engineers, and MLOps practitioners who want to move models from development into production.

Inside, you’ll learn how to:

  • Understand ONNX models, tensors, operators, and computational graphs

  • Convert PyTorch and other machine-learning models to ONNX

  • Validate and troubleshoot ONNX model conversion

  • Optimize ONNX graphs and improve AI inference performance

  • Apply quantization, INT8, FP16, and reduced-precision inference

  • Accelerate inference with CUDA, NVIDIA GPUs, TensorRT, OpenVINO, DirectML, and other execution providers

  • Build efficient ONNX Runtime inference pipelines and APIs

  • Deploy AI inference services using Docker, Kubernetes, and cloud infrastructure

  • Improve throughput with batching, concurrency, scheduling, and horizontal scaling

  • Implement monitoring, logging, tracing, metrics, and AI observability

  • Secure ONNX models, inference APIs, and production deployments

  • Perform advanced ONNX Runtime performance tuning and optimization

  • Build cross-platform ONNX deployment and optimization pipelines

  • Manage model versions, rollouts, rollbacks, testing, and production reliability

  • Balance inference speed, model accuracy, infrastructure cost, and maintainability

From ONNX model conversion and runtime optimization to GPU acceleration, AI deployment, MLOps, and production inference, this book provides a complete path for turning trained machine-learning models into reliable production systems.

If you want to understand ONNX, ONNX Runtime, model optimization, quantization, hardware acceleration, and scalable AI inference, PRODUCTION ONNX gives you the practical foundation to build and operate modern machine-learning inference systems.

Convert. Optimize. Accelerate. Deploy.

Les informations fournies dans la section « Synopsis » peuvent faire référence à une autre édition de ce titre.