Build evaluation systems that reveal whether your LLM applications are improving or regressing, using calibrated judges, RAG and agent metrics, CI regression tests, online evals, and cost-aware pipelines
LLM applications can look healthy in dashboards while their answers quietly become less accurate, less useful, or less reliable. AI Evals in Practice gives developers and AI engineers a systematic way to measure quality, catch regressions, and make evidence-based improvements before and after deployment.
You will build evalkit, a complete Python evaluation harness, while learning the core components of an eval: datasets, scorers, runners, and golden test sets. You will create deterministic scorers and LLM-as-judge evaluations, then calibrate judges and mitigate common biases. The book applies these foundations to prompt regression testing and CI, RAG retrieval and generation metrics, agent trajectories and tool calls, and human annotation workflows. You will then extend evaluation into production with online sampling, guardrails, and cost- and latency-aware pipelines, while comparing tools such as DeepEval, promptfoo, Langfuse, and Braintrust. A case study and final project bring the pieces together into an eval-driven development workflow with CI and a dashboard.
By the end, you will be able to design evaluation pipelines that help you ship LLM systems with measurable, repeatable quality.
This book is for AI engineers, LLM application developers, machine learning engineers, platform engineers, and technical leads who build or operate production systems using large language models. It is especially useful for teams working with prompts, RAG pipelines, agents, or AI features that need measurable quality and regression protection. Readers should be comfortable with Python and familiar with building or integrating LLM applications.
Les informations fournies dans la section « Synopsis » peuvent faire référence à une autre édition de ce titre.
Caio Incau is an Engineering Manager with experience leading software engineering teams at scale. In his day-to-day work, he combines people leadership with deep technical expertise to deliver products that impact millions of users. He started using Claude Code out of curiosity, became an advocate after seeing the productivity gains firsthand, and wrote this book so other developers wouldn't have to figure everything out on their own.
Les informations fournies dans la section « A propos du livre » peuvent faire référence à une autre édition de ce titre.
Vendeur : AHA-BUCH GmbH, Einbeck, Allemagne
Taschenbuch. Etat : Neu. Neuware. N° de réf. du vendeur 9781808817250
Quantité disponible : 2 disponible(s)