AI Inference Optimization Engineering (Paperback)

Language: English

Published by Independently Published, 2026

9798199720021

Series: Book 6 of 21 - Production AI Engineering Series

  • Softcover
  • New
See all details

Seller: CitiRetail, Stevenage, United KingdomCitiRetail

5-star seller

AbeBooks seller since June 29, 2022

View this seller's items
Softcover

Condition: New

£ 13.99

£ 37.00 shipping 
Ships from United Kingdom to U.S.A.

Quantity: 1 available

Add to basket
Free 30-day returns

Item description from seller

Paperback. Slash LLM Deployment Costs and LatencyDeploying Large Language Models (LLMs) in production is a massive economic and engineering hurdle. AI Inference Optimization Engineering is your comprehensive, hands-on guide to mastering the full stack of modern LLM optimization techniques. From memory-bandwidth solutions to hardware-specific compilation, this book bridges the gap between research-level models and enterprise-grade execution.What you will master inside this book: Hardware-Aware Optimization: Dive deep into KV cache mechanics, autoregressive decoding, and GPU memory hierarchies to eliminate latency bottlenecks.State-of-the-Art Quantization: Apply GPTQ, AWQ, and GGUF compression algorithms to scale down massive neural networks without sacrificing model accuracy.Advanced Acceleration Methods: Implement speculative decoding with draft models (like Medusa and Eagle), PagedAttention, and FlashAttention to boost throughput by 2-3x.Production-Grade Serving: Build ultra-low-latency deployment infrastructures using vLLM, Triton Inference Server, and continuous batching.Cross-Platform Deployment: Optimize models for specific target hardware, including NVIDIA H100 (TensorRT-LLM), Apple Silicon (llama.cpp/Metal), and Qualcomm mobile/edge accelerators.Whether you are an ML infrastructure engineer, an AI platform architect, or a technical leader looking to scale LLMs cost-effectively, this book provides the production-ready code, equations, and architectural patterns you need to build hyper-efficient AI pipelines. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.

Seller Inventory # 9798199720021

Title
AI Inference Optimization Engineering (Paperback)
Author
Chatvariety Team
Publisher
Independently Published
Publication year
2026
Condition
new
Binding
Paperback
Language
English
ISBN 13
9798199720021
Series
Book 6 of 21: Production AI Engineering Series

CitiRetail

Stevenage, United Kingdom

5-star seller

AbeBooks seller since June 29, 2022

Shipping rates from United Kingdom to U.S.A.

Item7 to 14 business days7 to 60 business days
First item£ 37.00£ 37.00
Delivery times are set by sellers and vary by carrier and location. Orders passing through Customs may face delays and buyers are responsible for any associated duties or fees. Sellers may contact you regarding additional charges to cover any increased costs to ship your items.

Payment methods

  • Visa
  • Mastercard
  • American Express
  • Apple Pay
  • Google Pay

Store description

Online business

Seller's business information

ABC BOOKS LIMITED

10 John Street
London, United Kingdom WC1N 2EB