AI Inference Optimization Engineering (Paperback)

Langue : anglais

Edité par Independently Published, 2026

9798199720021

Série : Livre 6 sur 21 - Production AI Engineering Series

Vendeur : CitiRetail, Stevenage, Royaume-UniCitiRetail

Vendeur avec une évaluation de 5 étoiles

Vendeur AbeBooks depuis 29 juin 2022

Afficher les articles de ce vendeur
Livre broché

Etat: Neuf

EUR 16,80

EUR 43,13 expédition 
Expédition depuis Royaume-Uni vers Etats-Unis

Quantité disponible : 1 disponible(s)

Ajouter au panier
Retours gratuits sous 30 jours

Item description from seller

Paperback. Slash LLM Deployment Costs and LatencyDeploying Large Language Models (LLMs) in production is a massive economic and engineering hurdle. AI Inference Optimization Engineering is your comprehensive, hands-on guide to mastering the full stack of modern LLM optimization techniques. From memory-bandwidth solutions to hardware-specific compilation, this book bridges the gap between research-level models and enterprise-grade execution.What you will master inside this book: Hardware-Aware Optimization: Dive deep into KV cache mechanics, autoregressive decoding, and GPU memory hierarchies to eliminate latency bottlenecks.State-of-the-Art Quantization: Apply GPTQ, AWQ, and GGUF compression algorithms to scale down massive neural networks without sacrificing model accuracy.Advanced Acceleration Methods: Implement speculative decoding with draft models (like Medusa and Eagle), PagedAttention, and FlashAttention to boost throughput by 2-3x.Production-Grade Serving: Build ultra-low-latency deployment infrastructures using vLLM, Triton Inference Server, and continuous batching.Cross-Platform Deployment: Optimize models for specific target hardware, including NVIDIA H100 (TensorRT-LLM), Apple Silicon (llama.cpp/Metal), and Qualcomm mobile/edge accelerators.Whether you are an ML infrastructure engineer, an AI platform architect, or a technical leader looking to scale LLMs cost-effectively, this book provides the production-ready code, equations, and architectural patterns you need to build hyper-efficient AI pipelines. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.

N° de réf. du vendeur 9798199720021

Titre
AI Inference Optimization Engineering (Paperback)
Auteur
Chatvariety Team
Éditeur
Independently Published
Année de publication
2026
État de l'article
new
Reliure
Paperback
Langue
anglais
ISBN à 13 chiffres
9798199720021
Série
Livre 6 sur 21: Production AI Engineering Series

CitiRetail

Stevenage, Royaume-Uni

Vendeur avec une évaluation de 5 étoiles

Vendeur AbeBooks depuis 29 juin 2022

Frais d'expédition de Royaume-Uni vers Etats-Unis

Article7 à 14 jours ouvrés7 à 60 jours ouvrés
Premier articleEUR 43,13EUR 43,13
Les délais de livraison sont fixés par les vendeurs et varient en fonction du transporteur et du lieu. Les commandes transitant par les douanes peuvent être retardées et les acheteurs sont responsables de tous les droits ou frais associés. Les vendeurs peuvent vous contacter au sujet de frais supplémentaires afin de couvrir toute augmentation des coûts d'expédition de vos articles.

Modes de paiement

  • Visa
  • Mastercard
  • American Express
  • Carte Bleue
  • Apple Pay
  • Google Pay

Description de la boutique

Online business

Profil professionnel du vendeur

ABC BOOKS LIMITED

10 John Street
London, Royaume-Uni WC1N 2EB