Distributed machine learning patterns par carter jazper (8 résultats)

Auteur
Titre
Affiner les résultats avec une recherche avancée

Affiner la recherche

  • Livres (8)

  • Neuf (8)

à

Fourchette de prix personnalisée (EUR)

à

  • Langue : anglais

    Edité par Cybersoft Publishing LLc, 2026

    9798904980030

    • Couverture souple

    Vendeur : California Books, Miami, FL, Etats-UnisCalifornia Books

    Vendeur avec une évaluation de 4 étoiles
    Contacter le vendeur

    Etat: Neuf

    EUR 26,02

     Frais de port gratuits 
    Expédition nationale : Etats-Unis

    Quantité disponible : Plus de 20 disponibles

    Etat : New.

  • Langue : anglais

    Edité par Cybersoft Publishing LLC, 2026

    9798904980030

    • Couverture souple

    Vendeur : PBShop.store US, Wood Dale, IL, Etats-UnisPBShop.store US

    Vendeur avec une évaluation de 5 étoiles
    Contacter le vendeur

    Etat: Neuf

    EUR 28,91

     Frais de port gratuits 
    Expédition nationale : Etats-Unis

    Quantité disponible : Plus de 20 disponibles

    PAP. Etat : New. New Book. Shipped from UK. Established seller since 2000.

  • Langue : anglais

    Edité par Cybersoft Publishing LLC, 2026

    9798904980030

    • Couverture souple

    Vendeur : PBShop.store UK, Fairford, GLOS, Royaume-UniPBShop.store UK

    Vendeur avec une évaluation de 5 étoiles
    Contacter le vendeur

    Etat: Neuf

    EUR 25,49

    EUR 5,86 expédition 
    Expédition depuis Royaume-Uni vers Etats-Unis

    Quantité disponible : Plus de 20 disponibles

    PAP. Etat : New. New Book. Shipped from UK. Established seller since 2000.

  • Langue : anglais

    Edité par Cybersoft Publishing LLC, 2026

    9798904980030

    • Couverture souple

    Vendeur : Rarewaves.com USA, London, LONDO, Royaume-UniRarewaves.com USA

    Vendeur avec une évaluation de 5 étoiles
    Contacter le vendeur

    Etat: Neuf

    EUR 32,09

     Frais de port gratuits 
    Expédition depuis Royaume-Uni vers Etats-Unis

    Quantité disponible : Plus de 20 disponibles

    Paperback. Etat : New.

  • Langue : anglais

    Edité par Independently Published Mai 2026, 2026

    9798904980030

    • Couverture souple

    Vendeur : AHA-BUCH GmbH, Einbeck, AllemagneAHA-BUCH GmbH

    Vendeur avec une évaluation de 5 étoiles
    Contacter le vendeur

    Etat: Neuf

    EUR 54,90

    EUR 30,50 expédition 
    Expédition depuis Allemagne vers Etats-Unis

    Quantité disponible : 2 disponible(s)

    Taschenbuch. Etat : Neu. Neuware.

  • Langue : anglais

    Edité par Cybersoft Publishing LLC, 2026

    9798904980030

    • Couverture souple

    Vendeur : Rarewaves.com UK, London, Royaume-UniRarewaves.com UK

    Vendeur avec une évaluation de 5 étoiles
    Contacter le vendeur

    Etat: Neuf

    EUR 30,43

    EUR 75,82 expédition 
    Expédition depuis Royaume-Uni vers Etats-Unis

    Quantité disponible : Plus de 20 disponibles

    Paperback. Etat : New.

  • Langue : anglais

    Edité par Cybersoft Publishing LLC, 2026

    9798904980030

    • Couverture souple
    • impression à la demande

    Vendeur : Grand Eagle Retail, Bensenville, IL, Etats-UnisGrand Eagle Retail

    Vendeur avec une évaluation de 5 étoiles
    Contacter le vendeur

    Etat: Neuf

    EUR 28,90

     Frais de port gratuits 
    Expédition nationale : Etats-Unis

    Quantité disponible : 1 disponible(s)

    Paperback. Etat : new. Paperback. Distributed machine learning systems fail in ways single-node systems never do. A 1024-GPU training job stalls for four hours while every worker reports healthy; gradient synchronization deadlocks leave no stack trace and no alert. A serving cluster absorbs a traffic spike, then silently doubles inference cost because the KV cache policy was tuned for a model half the size. Checkpoint corruption surfaces only after twelve hours of resumed training. These are the predictable failure modes of distributed systems, and the teams that ship reliable distributed ML design against them with patterns that hold across frameworks, clouds, and model scales.Inside this book, readers will learn how to: Design parallelism strategies that fit workload shape and hardware, selecting among data, tensor, pipeline, and expert axes based on architecture, memory budget, and interconnect topology.Tune gradient synchronization and sharding applying ZeRO, FSDP, and pipeline schedules to keep accelerator utilization high without amplifying communication overhead as cluster size grows.Build fault-tolerant training pipelines with checkpoint strategies, elastic cluster patterns, and spot instance management that recover from mid-run hardware failures without restarting from epoch zero.Operate inference at scale using continuous batching, paged attention, and KV cache management to maximize throughput and meet latency SLOs under variable load.Instrument distributed jobs for observability tracing per-rank metrics, gradient norms, and communication timings so silent failures surface before consuming days of compute budget.Manage multi-tenant clusters securely with workload isolation, quota enforcement, and cost attribution that keep shared GPU infrastructure safe and financially accountable.Apply LLM and foundation model patterns for distributed pre-training, RLHF infrastructure, and large-scale inference that generalize across architectures as hardware generations turn over.Assess platform maturity using the book's maturity model to locate gaps in reliability, cost efficiency, and operational readiness across the distributed ML stack.Frameworks rotate; the parallelism decisions, synchronization tradeoffs, and fault-tolerance designs that determine whether a distributed ML system works at scale do not. As foundation models grow larger and serving loads grow steeper, the distance between teams that reason in patterns and teams that copy configurations will only widen.The book is organized in four parts: Foundations, covering parallelism patterns, data sharding, I/O, and orchestration; Training at Scale, addressing fault-tolerant training, checkpoint management, and spot scheduling; Serving and Operations, covering inference architecture, cost control, observability, and multi-tenant security; and Frontier Patterns, applying everything to LLMs and foundation models and closing with end-to-end case studies and a full platform synthesis.This book is for ML architects who design distributed systems others depend on, ML engineers and data engineers who build and operate them, and technical team leads who set reliability and cost standards, with platform and SRE engineers as a strong secondary audience. Every chapter opens with a production incident scenario, teaches canonical patterns by name, and closes with a checklist the team can apply immediately. Readers finish with the vocabulary, playbook, and pattern library to ship reliable distributed ML systems with confidence. This item is printed on demand. Shipping may be from multiple locations in the US or from the UK, depending on stock availability.

  • Langue : anglais

    Edité par Cybersoft Publishing LLC, 2026

    9798904980030

    • Couverture souple
    • impression à la demande

    Vendeur : CitiRetail, Stevenage, Royaume-UniCitiRetail

    Vendeur avec une évaluation de 5 étoiles
    Contacter le vendeur

    Etat: Neuf

    EUR 29,42

    EUR 43,16 expédition 
    Expédition depuis Royaume-Uni vers Etats-Unis

    Quantité disponible : 1 disponible(s)

    Paperback. Etat : new. Paperback. Distributed machine learning systems fail in ways single-node systems never do. A 1024-GPU training job stalls for four hours while every worker reports healthy; gradient synchronization deadlocks leave no stack trace and no alert. A serving cluster absorbs a traffic spike, then silently doubles inference cost because the KV cache policy was tuned for a model half the size. Checkpoint corruption surfaces only after twelve hours of resumed training. These are the predictable failure modes of distributed systems, and the teams that ship reliable distributed ML design against them with patterns that hold across frameworks, clouds, and model scales.Inside this book, readers will learn how to: Design parallelism strategies that fit workload shape and hardware, selecting among data, tensor, pipeline, and expert axes based on architecture, memory budget, and interconnect topology.Tune gradient synchronization and sharding applying ZeRO, FSDP, and pipeline schedules to keep accelerator utilization high without amplifying communication overhead as cluster size grows.Build fault-tolerant training pipelines with checkpoint strategies, elastic cluster patterns, and spot instance management that recover from mid-run hardware failures without restarting from epoch zero.Operate inference at scale using continuous batching, paged attention, and KV cache management to maximize throughput and meet latency SLOs under variable load.Instrument distributed jobs for observability tracing per-rank metrics, gradient norms, and communication timings so silent failures surface before consuming days of compute budget.Manage multi-tenant clusters securely with workload isolation, quota enforcement, and cost attribution that keep shared GPU infrastructure safe and financially accountable.Apply LLM and foundation model patterns for distributed pre-training, RLHF infrastructure, and large-scale inference that generalize across architectures as hardware generations turn over.Assess platform maturity using the book's maturity model to locate gaps in reliability, cost efficiency, and operational readiness across the distributed ML stack.Frameworks rotate; the parallelism decisions, synchronization tradeoffs, and fault-tolerance designs that determine whether a distributed ML system works at scale do not. As foundation models grow larger and serving loads grow steeper, the distance between teams that reason in patterns and teams that copy configurations will only widen.The book is organized in four parts: Foundations, covering parallelism patterns, data sharding, I/O, and orchestration; Training at Scale, addressing fault-tolerant training, checkpoint management, and spot scheduling; Serving and Operations, covering inference architecture, cost control, observability, and multi-tenant security; and Frontier Patterns, applying everything to LLMs and foundation models and closing with end-to-end case studies and a full platform synthesis.This book is for ML architects who design distributed systems others depend on, ML engineers and data engineers who build and operate them, and technical team leads who set reliability and cost standards, with platform and SRE engineers as a strong secondary audience. Every chapter opens with a production incident scenario, teaches canonical patterns by name, and closes with a checklist the team can apply immediately. Readers finish with the vocabulary, playbook, and pattern library to ship reliable distributed ML systems with confidence. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.