Isbn: 9798199720021 - ai inference optimization engineering: quantization, speculative decoding, and hardware-specific llm deployment (5 résultats)

Affiner la recherche

  • Livres (5)

  • Neuf (5)

à

Fourchette de prix personnalisée (EUR)

à

  • Langue : anglais

    Edité par Independently published, 2026

    9798199720021

    Série : Livre 6 sur 20 - Production AI Engineering Series

    • Couverture souple

    Vendeur : PBShop.store US, Wood Dale, IL, Etats-UnisPBShop.store US

    Vendeur avec une évaluation de 5 étoiles
    Contacter le vendeur

    Etat: Neuf

    EUR 14,03

     Frais de port gratuits 
    Expédition nationale : Etats-Unis

    Quantité disponible : Plus de 20 disponibles

    PAP. Etat : New. New Book. Shipped from UK. Established seller since 2000.

  • Langue : anglais

    Edité par Independently published, 2026

    9798199720021

    Série : Livre 6 sur 20 - Production AI Engineering Series

    • Couverture souple

    Vendeur : PBShop.store UK, Fairford, GLOS, Royaume-UniPBShop.store UK

    Vendeur avec une évaluation de 5 étoiles
    Contacter le vendeur

    Etat: Neuf

    EUR 13,38

    EUR 3,83 expédition 
    Expédition depuis Royaume-Uni vers Etats-Unis

    Quantité disponible : Plus de 20 disponibles

    PAP. Etat : New. New Book. Shipped from UK. Established seller since 2000.

  • Langue : anglais

    Edité par Independently Published Jun 2026, 2026

    9798199720021

    Série : Livre 6 sur 20 - Production AI Engineering Series

    • Couverture souple

    Vendeur : AHA-BUCH GmbH, Einbeck, AllemagneAHA-BUCH GmbH

    Vendeur avec une évaluation de 5 étoiles
    Contacter le vendeur

    Etat: Neuf

    EUR 15,40

    EUR 35,00 expédition 
    Expédition depuis Allemagne vers Etats-Unis

    Quantité disponible : 2 disponible(s)

    Taschenbuch. Etat : Neu. Neuware - Slash LLM Deployment Costs and LatencyDeploying Large Language Models (LLMs) in production is a massive economic and engineering hurdle. AI Inference Optimization Engineering is your comprehensive, hands-on guide to mastering the full stack of modern LLM optimization techniques. From memory-bandwidth solutions to hardware-specific compilation, this book bridges the gap between research-level models and enterprise-grade execution.What you will master inside this book: - Hardware-Aware Optimization: Dive deep into KV cache mechanics, autoregressive decoding, and GPU memory hierarchies to eliminate latency bottlenecks.- State-of-the-Art Quantization: Apply GPTQ, AWQ, and GGUF compression algorithms to scale down massive neural networks without sacrificing model accuracy.- Advanced Acceleration Methods: Implement speculative decoding with draft models (like Medusa and Eagle), PagedAttention, and FlashAttention to boost throughput by 2-3x.- Production-Grade Serving: Build ultra-low-latency deployment infrastructures using vLLM, Triton Inference Server, and continuous batching.- Cross-Platform Deployment: Optimize models for specific target hardware, including NVIDIA H100 (TensorRT-LLM), Apple Silicon (llama.cpp/Metal), and Qualcomm mobile/edge accelerators.Whether you are an ML infrastructure engineer, an AI platform architect, or a technical leader looking to scale LLMs cost-effectively, this book provides the production-ready code, equations, and architectural patterns you need to build hyper-efficient AI pipelines.

  • Langue : anglais

    Edité par Independently published, 2026

    9798199720021

    Série : Livre 6 sur 20 - Production AI Engineering Series

    • Couverture souple
    • impression à la demande

    Vendeur : California Books, Miami, FL, Etats-UnisCalifornia Books

    Vendeur avec une évaluation de 4 étoiles
    Contacter le vendeur

    Etat: Neuf

    EUR 13,50

     Frais de port gratuits 
    Expédition nationale : Etats-Unis

    Quantité disponible : Plus de 20 disponibles

    Etat : New. Print on Demand.

  • Langue : anglais

    Edité par Independently Published, 2026

    9798199720021

    Série : Livre 6 sur 20 - Production AI Engineering Series

    • Couverture souple
    • impression à la demande

    Vendeur : CitiRetail, Stevenage, Royaume-UniCitiRetail

    Vendeur avec une évaluation de 5 étoiles
    Contacter le vendeur

    Etat: Neuf

    EUR 16,80

    EUR 43,13 expédition 
    Expédition depuis Royaume-Uni vers Etats-Unis

    Quantité disponible : 1 disponible(s)

    Paperback. Etat : new. Paperback. Slash LLM Deployment Costs and LatencyDeploying Large Language Models (LLMs) in production is a massive economic and engineering hurdle. AI Inference Optimization Engineering is your comprehensive, hands-on guide to mastering the full stack of modern LLM optimization techniques. From memory-bandwidth solutions to hardware-specific compilation, this book bridges the gap between research-level models and enterprise-grade execution.What you will master inside this book: Hardware-Aware Optimization: Dive deep into KV cache mechanics, autoregressive decoding, and GPU memory hierarchies to eliminate latency bottlenecks.State-of-the-Art Quantization: Apply GPTQ, AWQ, and GGUF compression algorithms to scale down massive neural networks without sacrificing model accuracy.Advanced Acceleration Methods: Implement speculative decoding with draft models (like Medusa and Eagle), PagedAttention, and FlashAttention to boost throughput by 2-3x.Production-Grade Serving: Build ultra-low-latency deployment infrastructures using vLLM, Triton Inference Server, and continuous batching.Cross-Platform Deployment: Optimize models for specific target hardware, including NVIDIA H100 (TensorRT-LLM), Apple Silicon (llama.cpp/Metal), and Qualcomm mobile/edge accelerators.Whether you are an ML infrastructure engineer, an AI platform architect, or a technical leader looking to scale LLMs cost-effectively, this book provides the production-ready code, equations, and architectural patterns you need to build hyper-efficient AI pipelines. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.