Practical LLM Evaluation for Production Systems
Language: English
Published by Packt Publishing Limited, GB, 2026
- Softcover
- New

Seller: Rarewaves.com USA, London, London, United KingdomRarewaves.com USA
AbeBooks seller since June 11, 2025
Condition: New
£ 53.69
Quantity: Over 20 available
Add to basketItem description from seller
Build reliable Build reliable AI evaluation frameworks that measure quality, safety, grounding, and production readiness across modern LLM and SLM applicationsFree with your book: DRM-free PDF version + access to Packt's next-gen Reader*Key FeaturesDesign evaluation frameworks for LLMs, SLMs, multimodal, reasoning, and agentic AI systemsMeasure quality, safety, grounding, robustness, and production readiness with practical metricsApply unified evaluation methods to text, multimodal, and agentic AI systemsBook DescriptionModern AI systems are expected to do far more than generate fluent text. They should be able to retrieve information, reason through complex problems, understand images and documents, call external tools, execute workflows, and support critical business decisions. Evaluating these systems requires methods that go beyond traditional NLP benchmarks.Taking a product-first approach, this book presents evaluation as a continuous operational capability spanning training, inference, and end-to-end system operation. You'll learn how to connect evaluation metrics directly to deployment gates, rollback criteria, monitoring systems, and production reliability objectives.Using practical examples and real-world workflows, you'll explore evaluation strategies for text LLMs, vision-language models, multimodal conversational systems, mixture-of-experts architectures, reasoning models, agentic systems, retrieval pipelines, Text2SQL and Text2Cypher systems, embedding models, OCR workflows, and guardrail SLMs. You'll also learn how to manage non-determinism, design repeatable test suites, validate tool execution, and measure long-horizon agent behavior in production.By the end of the book, you'll be able to design robust evaluation systems that help teams deploy reliable, safe, and economically viable LLM-powered applications with confidence.*Email sign-up and proof of purchase requiredWhat you will learnDesign repeatable evaluation pipelines for LLM systemsAssess inference quality, latency, and operational costEvaluate multimodal, agentic, and reasoning AI systemsBuild regression gates and deployment evaluation workflowsDetect hallucinations and grounding failures in VLMsAssess routing stability in mixture-of-experts modelsEvaluate Text2SQL, OCR, and retrieval-based systemsTranslate evaluation signals into production decisionsWho this book is forML engineers, GenAI engineers, AI architects, data scientists, platform engineers, and engineering managers responsible for deploying LLM-powered systems in production will benefit from this book. Applied AI researchers and technical decision-makers looking to measure reliability, safety, and operational readiness across modern AI systems will also find it valuable. Readers should have a working understanding of machine learning, Python, and modern LLM concepts. …
Seller Inventory # LU-9781807423896
- Title
- Practical LLM Evaluation for Production Systems
- Author
- Ammar Mohanna, Indrajit Kar, Zonunfeli Ralte
- Publisher
- Packt Publishing Limited, GB
- Publication year
- 2026
- Condition
- New
- Binding
- Paperback
- Language
- English
- ISBN 10
- 1807423891
- ISBN 13
- 9781807423896
Build reliable Build reliable AI evaluation frameworks that measure quality, safety, grounding, and production readiness across modern LLM and SLM applications
Free with your book: DRM-free PDF version + access to Packt's next-gen Reader*
Key Features
- Design evaluation frameworks for LLMs, SLMs, multimodal, reasoning, and agentic AI systems
- Measure quality, safety, grounding, robustness, and production readiness with practical metrics
- Apply unified evaluation methods to text, multimodal, and agentic AI systems
Book Description
Modern AI systems are expected to do far more than generate fluent text. They should be able to retrieve information, reason through complex problems, understand images and documents, call external tools, execute workflows, and support critical business decisions. Evaluating these systems requires methods that go beyond traditional NLP benchmarks.
Taking a product-first approach, this book presents evaluation as a continuous operational capability spanning training, inference, and end-to-end system operation. You'll learn how to connect evaluation metrics directly to deployment gates, rollback criteria, monitoring systems, and production reliability objectives.
Using practical examples and real-world workflows, you'll explore evaluation strategies for text LLMs, vision-language models, multimodal conversational systems, mixture-of-experts architectures, reasoning models, agentic systems, retrieval pipelines, Text2SQL and Text2Cypher systems, embedding models, OCR workflows, and guardrail SLMs. You'll also learn how to manage non-determinism, design repeatable test suites, validate tool execution, and measure long-horizon agent behavior in production.
By the end of the book, you'll be able to design robust evaluation systems that help teams deploy reliable, safe, and economically viable LLM-powered applications with confidence.
*Email sign-up and proof of purchase required
What you will learn
- Design repeatable evaluation pipelines for LLM systems
- Assess inference quality, latency, and operational cost
- Evaluate multimodal, agentic, and reasoning AI systems
- Build regression gates and deployment evaluation workflows
- Detect hallucinations and grounding failures in VLMs
- Assess routing stability in mixture-of-experts models
- Evaluate Text2SQL, OCR, and retrieval-based systems
- Translate evaluation signals into production decisions
Who this book is for
ML engineers, GenAI engineers, AI architects, data scientists, platform engineers, and engineering managers responsible for deploying LLM-powered systems in production will benefit from this book. Applied AI researchers and technical decision-makers looking to measure reliability, safety, and operational readiness across modern AI systems will also find it valuable. Readers should have a working understanding of machine learning, Python, and modern LLM concepts.
Table of Contents
- Foundations of LLM Evaluation: Core Concepts and Primitives
- Building Reliable Text-Only LLMs Through Training-Time Evaluation
- Controlling Text-Only LLM Behavior at Inference Time
- Grounding and Reliability in Vision Language Models During Training
- Evaluating Visual Grounding and Reliability at Inference Time
- Evaluating Multimodal Conversational LLMs Across Training and Inference
- Evaluating Routing and Reliability in Mixture of Experts LLMs
- Evaluating Reliability and Control in Computer-Using Agent Systems
- Evaluating Information Extraction and Document-Understanding LLMs
- Evaluating Reasoning LLMs in Depth
- Evaluating Specialized LLM Systems
"Synopsis" may belong to another edition of this title.
About the Author
Ammar Mohanna, PhD, is an AI and machine learning specialist based in Beirut, Lebanon. His work focuses on practical LLM systems, evaluation, MLOps/LLMOps, and applied generative AI. He teaches and consults on production AI, AI agents, and graph-based machine learning, with an emphasis on turning research ideas into reliable, usable systems for real-world teams.
Indrajit Kar comes with 18 years of various Industry experience, leading all three division, AI consulting R&D and solution engineering. He and his team build cutting edge AI and deep learning solutions to address some of the toughest problems for his customers. He has 14 research papers and 12 patents in NLP, Timeseries, Computer Vision, and Deep learning. In his spare time, Indrajit enjoys giving advice to small and medium-sized entrepreneurs on how to enter the AI and data science markets, attract customers, develop their products, and monetize their existing data. He's won many accolades in his career from ace innovator, services excellence awards, and 40 top data scientist under the age of 40 award. He has enabled AI & Data science program for sectors like Smart Cities, Retail, supply chain, automotive factories, Healthcare, pharma, infrastructure & utilities. Also heading research and development in the area of Deep learning, predictive maintenance using IIoT/sensor data, edgeAi, Lidar tech, NLP and GPU powered computer vision. In the past, he spearheaded complex Analytics projects helping industries like BFSI, Retail, CPG, FMCG, petroleum/oil & gas, to take data driven decision, predict business outcomes, allocate budget, predict customer behaviour, retention customers, acquire new customers, maximize revenue & forecasting for key areas Pricing, marketing, sale, advertisement and promotion.
Zonunfeli Ralte is an Artificial Intelligence entrepreneur, researcher, and technology leader. She founded RastrAI Private Limited, the first AI startup from India's North East region, advancing innovation in emerging technologies. Recognized as Mizoram's first woman specializing in Artificial Intelligence and Machine Learning, she has authored three books on Artificial Intelligence, Generative AI, and Computer Vision. She is also an accomplished researcher with 16 published research papers and six Best Research Awards, reflecting her significant contributions to Artificial Intelligence, Deep Learning, and applied AI innovation.
"About the title" may belong to another edition of this title.
Rarewaves.com USA
London, London, United Kingdom
AbeBooks seller since June 11, 2025
Shipping rates from United Kingdom to U.S.A.
| Item | 9 to 14 business days | 9 to 14 business days |
|---|---|---|
| First item | £ 0.00 | £ 0.00 |
Payment methods
Seller's business information
RAREWAVES.COM LIMITED
Elsley Court, 20-22 Great Titchfield Street
London, United Kingdom W1W 8BE
Right of withdrawal
If you are a consumer you can withdraw from the contract in accordance with the following. Consumer means any natural person who is acting for purposes which are outside his trade, business, craft or profession.
Information regarding the right of withdrawal
Statutory right to withdraw
You have the right to withdraw from this contract within 14 days without giving any reason.
The withdrawal period will expire after 14 days from the day on which you acquire, or a third party other than the carrier and indicated by you acquires, physical possession of the last good or the last lot or piece.
To exercise the right of withdrawal, electronically fill in and submit a clear statement on our website, under "My Purchases" in "My Account". We will communicate to you an acknowledgement of receipt of such a withdrawal on a durable medium (e.g. by e-mail) without delay.
To meet the withdrawal deadline, it is sufficient for you to send your communication concerning your exercise of the right of withdrawal before the withdrawal period has expired.
Effects of withdrawal
If you withdraw from this contract, we will reimburse to you all payments received from you, including the costs of delivery (except for the supplementary costs arising if you chose a type of delivery other than the least expensive type of standard delivery offered by us).
We may make a deduction from the reimbursement for loss in value of any goods supplied, if the loss is the result of unnecessary handling by you.
We will make the reimbursement without undue delay, and not later than 14 days after the day on which we are informed about your decision to withdraw from this contract.
We will make the reimbursement using the same means of payment as you used for the initial transaction, unless you have expressly agreed otherwise; in any event, you will not incur any fees as a result of such reimbursement.
We may withhold reimbursement until we have received the goods back, or you have supplied evidence of having sent back the goods, whichever is the earliest.
You shall send back the goods or hand them over to Rarewaves.com USA, Unit 144 The Lightbox, 111 Power Road, W4 5PY, London, London, United Kingdom, without undue delay and in any event not later than 14 days from the day on which you communicate your withdrawal from this contract to us. The deadline is met if you send back the goods before the period of 14 days has expired. You will have to bear the direct cost of returning the goods. You are only liable for any diminished value of the goods resulting from the handling other than what is necessary to establish the nature, characteristics and functioning of the goods.
Exceptions to the right of withdrawal
The right of withdrawal does not apply to:
- The delivery of newspapers, journals or magazines with the exception of subscription contracts; and
- The supply of digital content which is not supplied on a tangible medium (e.g. on a CD or DVD) if you accepted when you placed your order that we could start to deliver it, and that you could not withdraw once delivery had started.
Shipping terms
Please note that we do not offer Priority shipping to any country.
We currently do not ship to the below countries:
Russia
Belarus
Ukraine
Please do not attempt to place orders with any of these countries as a ship to address - they will be cancelled.