Build reliable LLM-powered systems using practical evaluation frameworks, production metrics, and deployment-ready monitoring strategies.
Move beyond benchmarks and learn how to evaluate whether LLM-powered systems actually work in production. This book gives you practical frameworks, metrics, and operational strategies to measure reliability, safety, quality, latency, and cost across modern AI systems. Guided by experienced AI leaders and researchers, you’ll build evaluation pipelines that support real business decisions instead of isolated leaderboard scores.
The book takes a product-first approach to evaluation, treating it as a continuous operational capability rather than a one-time testing exercise. You’ll explore how evaluation changes across training, inference, and end-to-end system operation while learning how to connect metrics directly to deployment gates, rollback criteria, monitoring systems, and production reliability goals.
Using practical examples and real-world workflows, the book covers evaluation strategies for text LLMs, vision-language models, multimodal conversational systems, Mixture-of-Experts architectures, agentic systems, reasoning models, Text2SQL and Text2Cypher systems, retrieval pipelines, embedding models, OCR workflows, and guardrail SLMs.
By the end of this book, you’ll be able to design and operate reliable, safe, and cost-effective LLM-powered applications with confidence.
ML engineers, GenAI engineers, AI architects, data scientists, platform engineers, and engineering managers responsible for deploying LLM-powered systems in production will benefit from this book. Applied AI researchers and technical decision-makers looking to measure reliability, safety, and operational readiness across modern AI systems will also find it valuable. Readers should have a working understanding of machine learning, Python, and modern LLM concepts.
Les informations fournies dans la section « Synopsis » peuvent faire référence à une autre édition de ce titre.
Ammar Mohanna, PhD, is an AI and machine learning specialist based in Beirut, Lebanon. His work focuses on practical LLM systems, evaluation, MLOps/LLMOps, and applied generative AI. He teaches and consults on production AI, AI agents, and graph-based machine learning, with an emphasis on turning research ideas into reliable, usable systems for real-world teams.
Indrajit Kar comes with 18 years of various Industry experience, leading all three division, AI consulting R&D and solution engineering. He and his team build cutting edge AI and deep learning solutions to address some of the toughest problems for his customers.
He has 14 research papers and 12 patents in NLP, Timeseries, Computer Vision, and Deep learning.
In his spare time, Indrajit enjoys giving advice to small and medium-sized entrepreneurs on how to enter the AI and data science markets, attract customers, develop their products, and monetize their existing data. He's won many accolades in his career from ace innovator, services excellence awards, and 40 top data scientist under the age of 40 award.
He has enabled AI & Data science program for sectors like Smart Cities, Retail, supply chain, automotive factories, Healthcare, pharma, infrastructure & utilities. Also heading research and development in the area of Deep learning, predictive maintenance using IIoT/sensor data, edgeAi, Lidar tech, NLP and GPU powered computer vision.
In the past, he spearheaded complex Analytics projects helping industries like BFSI, Retail, CPG, FMCG, petroleum/oil & gas, to take data driven decision, predict business outcomes, allocate budget, predict customer behaviour, retention customers, acquire new customers, maximize revenue & forecasting for key areas Pricing, marketing, sale, advertisement and promotion.
Zonunfeli Ralte is an Artificial Intelligence entrepreneur, researcher, and technology leader. She founded RastrAI Private Limited, the first AI startup from India's North East region, advancing innovation in emerging technologies. Recognized as Mizoram's first woman specializing in Artificial Intelligence and Machine Learning, she has authored three books on Artificial Intelligence, Generative AI, and Computer Vision.
She is also an accomplished researcher with 16 published research papers and six Best Research Awards, reflecting her significant contributions to Artificial Intelligence, Deep Learning, and applied AI innovation.
Les informations fournies dans la section « A propos du livre » peuvent faire référence à une autre édition de ce titre.
Vendeur : ThriftBooks-Atlanta, AUSTELL, GA, Etats-Unis
Paperback. Etat : Very Good. No Jacket. May have limited writing in cover pages. Pages are unmarked. ~ ThriftBooks: Read More, Spend Less. N° de réf. du vendeur G1807423891I4N00
Quantité disponible : 1 disponible(s)
Vendeur : California Books, Miami, FL, Etats-Unis
Etat : New. N° de réf. du vendeur I-9781807423896
Quantité disponible : Plus de 20 disponibles
Vendeur : Grand Eagle Retail, Bensenville, IL, Etats-Unis
Paperback. Etat : new. Paperback. Build reliable Build reliable AI evaluation frameworks that measure quality, safety, grounding, and production readiness across modern LLM and SLM applicationsFree with your book: DRM-free PDF version + access to Packt's next-gen Reader*Key FeaturesDesign evaluation frameworks for LLMs, SLMs, multimodal, reasoning, and agentic AI systemsMeasure quality, safety, grounding, robustness, and production readiness with practical metricsApply unified evaluation methods to text, multimodal, and agentic AI systemsBook DescriptionModern AI systems are expected to do far more than generate fluent text. They should be able to retrieve information, reason through complex problems, understand images and documents, call external tools, execute workflows, and support critical business decisions. Evaluating these systems requires methods that go beyond traditional NLP benchmarks.Taking a product-first approach, this book presents evaluation as a continuous operational capability spanning training, inference, and end-to-end system operation. You'll learn how to connect evaluation metrics directly to deployment gates, rollback criteria, monitoring systems, and production reliability objectives.Using practical examples and real-world workflows, you'll explore evaluation strategies for text LLMs, vision-language models, multimodal conversational systems, mixture-of-experts architectures, reasoning models, agentic systems, retrieval pipelines, Text2SQL and Text2Cypher systems, embedding models, OCR workflows, and guardrail SLMs. You'll also learn how to manage non-determinism, design repeatable test suites, validate tool execution, and measure long-horizon agent behavior in production.By the end of the book, you'll be able to design robust evaluation systems that help teams deploy reliable, safe, and economically viable LLM-powered applications with confidence.*Email sign-up and proof of purchase requiredWhat you will learnDesign repeatable evaluation pipelines for LLM systemsAssess inference quality, latency, and operational costEvaluate multimodal, agentic, and reasoning AI systemsBuild regression gates and deployment evaluation workflowsDetect hallucinations and grounding failures in VLMsAssess routing stability in mixture-of-experts modelsEvaluate Text2SQL, OCR, and retrieval-based systemsTranslate evaluation signals into production decisionsWho this book is forML engineers, GenAI engineers, AI architects, data scientists, platform engineers, and engineering managers responsible for deploying LLM-powered systems in production will benefit from this book. Applied AI researchers and technical decision-makers looking to measure reliability, safety, and operational readiness across modern AI systems will also find it valuable. Readers should have a working understanding of machine learning, Python, and modern LLM concepts. This item is printed on demand. Shipping may be from multiple locations in the US or from the UK, depending on stock availability. N° de réf. du vendeur 9781807423896
Quantité disponible : 1 disponible(s)
Vendeur : PBShop.store US, Wood Dale, IL, Etats-Unis
PAP. Etat : New. New Book. Shipped from UK. Established seller since 2000. N° de réf. du vendeur L2-9781807423896
Quantité disponible : Plus de 20 disponibles
Vendeur : PBShop.store UK, Fairford, GLOS, Royaume-Uni
PAP. Etat : New. New Book. Shipped from UK. Established seller since 2000. N° de réf. du vendeur L2-9781807423896
Quantité disponible : Plus de 20 disponibles
Vendeur : Books Puddle, New York, NY, Etats-Unis
Etat : New. N° de réf. du vendeur 26406720745
Quantité disponible : 4 disponible(s)
Vendeur : Majestic Books, Hounslow, Royaume-Uni
Etat : New. Print on Demand. N° de réf. du vendeur 407482166
Quantité disponible : 4 disponible(s)
Vendeur : CitiRetail, Stevenage, Royaume-Uni
Paperback. Etat : new. Paperback. Build reliable Build reliable AI evaluation frameworks that measure quality, safety, grounding, and production readiness across modern LLM and SLM applicationsFree with your book: DRM-free PDF version + access to Packt's next-gen Reader*Key FeaturesDesign evaluation frameworks for LLMs, SLMs, multimodal, reasoning, and agentic AI systemsMeasure quality, safety, grounding, robustness, and production readiness with practical metricsApply unified evaluation methods to text, multimodal, and agentic AI systemsBook DescriptionModern AI systems are expected to do far more than generate fluent text. They should be able to retrieve information, reason through complex problems, understand images and documents, call external tools, execute workflows, and support critical business decisions. Evaluating these systems requires methods that go beyond traditional NLP benchmarks.Taking a product-first approach, this book presents evaluation as a continuous operational capability spanning training, inference, and end-to-end system operation. You'll learn how to connect evaluation metrics directly to deployment gates, rollback criteria, monitoring systems, and production reliability objectives.Using practical examples and real-world workflows, you'll explore evaluation strategies for text LLMs, vision-language models, multimodal conversational systems, mixture-of-experts architectures, reasoning models, agentic systems, retrieval pipelines, Text2SQL and Text2Cypher systems, embedding models, OCR workflows, and guardrail SLMs. You'll also learn how to manage non-determinism, design repeatable test suites, validate tool execution, and measure long-horizon agent behavior in production.By the end of the book, you'll be able to design robust evaluation systems that help teams deploy reliable, safe, and economically viable LLM-powered applications with confidence.*Email sign-up and proof of purchase requiredWhat you will learnDesign repeatable evaluation pipelines for LLM systemsAssess inference quality, latency, and operational costEvaluate multimodal, agentic, and reasoning AI systemsBuild regression gates and deployment evaluation workflowsDetect hallucinations and grounding failures in VLMsAssess routing stability in mixture-of-experts modelsEvaluate Text2SQL, OCR, and retrieval-based systemsTranslate evaluation signals into production decisionsWho this book is forML engineers, GenAI engineers, AI architects, data scientists, platform engineers, and engineering managers responsible for deploying LLM-powered systems in production will benefit from this book. Applied AI researchers and technical decision-makers looking to measure reliability, safety, and operational readiness across modern AI systems will also find it valuable. Readers should have a working understanding of machine learning, Python, and modern LLM concepts. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability. N° de réf. du vendeur 9781807423896
Quantité disponible : 1 disponible(s)
Vendeur : Biblios, Frankfurt am main, HESSE, Allemagne
Etat : New. PRINT ON DEMAND. N° de réf. du vendeur 18406720739
Quantité disponible : 4 disponible(s)
Vendeur : AussieBookSeller, Truganina, VIC, Australie
Paperback. Etat : new. Paperback. Build reliable Build reliable AI evaluation frameworks that measure quality, safety, grounding, and production readiness across modern LLM and SLM applicationsFree with your book: DRM-free PDF version + access to Packt's next-gen Reader*Key FeaturesDesign evaluation frameworks for LLMs, SLMs, multimodal, reasoning, and agentic AI systemsMeasure quality, safety, grounding, robustness, and production readiness with practical metricsApply unified evaluation methods to text, multimodal, and agentic AI systemsBook DescriptionModern AI systems are expected to do far more than generate fluent text. They should be able to retrieve information, reason through complex problems, understand images and documents, call external tools, execute workflows, and support critical business decisions. Evaluating these systems requires methods that go beyond traditional NLP benchmarks.Taking a product-first approach, this book presents evaluation as a continuous operational capability spanning training, inference, and end-to-end system operation. You'll learn how to connect evaluation metrics directly to deployment gates, rollback criteria, monitoring systems, and production reliability objectives.Using practical examples and real-world workflows, you'll explore evaluation strategies for text LLMs, vision-language models, multimodal conversational systems, mixture-of-experts architectures, reasoning models, agentic systems, retrieval pipelines, Text2SQL and Text2Cypher systems, embedding models, OCR workflows, and guardrail SLMs. You'll also learn how to manage non-determinism, design repeatable test suites, validate tool execution, and measure long-horizon agent behavior in production.By the end of the book, you'll be able to design robust evaluation systems that help teams deploy reliable, safe, and economically viable LLM-powered applications with confidence.*Email sign-up and proof of purchase requiredWhat you will learnDesign repeatable evaluation pipelines for LLM systemsAssess inference quality, latency, and operational costEvaluate multimodal, agentic, and reasoning AI systemsBuild regression gates and deployment evaluation workflowsDetect hallucinations and grounding failures in VLMsAssess routing stability in mixture-of-experts modelsEvaluate Text2SQL, OCR, and retrieval-based systemsTranslate evaluation signals into production decisionsWho this book is forML engineers, GenAI engineers, AI architects, data scientists, platform engineers, and engineering managers responsible for deploying LLM-powered systems in production will benefit from this book. Applied AI researchers and technical decision-makers looking to measure reliability, safety, and operational readiness across modern AI systems will also find it valuable. Readers should have a working understanding of machine learning, Python, and modern LLM concepts. This item is printed on demand. Shipping may be from our Sydney, NSW warehouse or from our UK or US warehouse, depending on stock availability. N° de réf. du vendeur 9781807423896
Quantité disponible : 1 disponible(s)