Articles liés à Building LLM Inference Engines and Agentic Runtimes...

Building LLM Inference Engines and Agentic Runtimes from Scratch: Qwen Dense and MoE Models with TensorSharp and TensorAgent - Couverture souple

Fu, Zhongkai

 
9798174165502: Building LLM Inference Engines and Agentic Runtimes from Scratch: Qwen Dense and MoE Models with TensorSharp and TensorAgent

Synopsis

Build the machinery behind a local AI application—from loading model weights to executing a controlled agent workflow.

Using TensorSharp and TensorAgent as a concrete engineering reference, this book follows the path from tensors and tokenization to complete inference and agentic runtimes. Qwen dense and mixture-of-experts models provide examples for attention, expert routing, quantization, cache management, and the tradeoffs between memory, latency, and throughput.

The inference chapters explain model loading, GPU execution, batching, speculative decoding, distributed execution, and multimodal inputs. The agent chapters connect model output to structured tools, skills, generated code, patching, execution, and sandbox boundaries. Cross-platform chapters examine desktop backends and mobile deployment, including the constraints that shape TensorAgent on phones.

C# examples, equations, architecture diagrams, implementation exercises, and references connect each design to the engineering decisions behind it. Repository behavior is distinguished from simplified teaching implementations and proposed extensions, so readers can understand what to reproduce, what to measure, and what still requires platform-specific work.

For developers comfortable with programming who want to understand how local AI systems work below the API—from a single generated token to an agent that can act.

Les informations fournies dans la section « Synopsis » peuvent faire référence à une autre édition de ce titre.