AI Inference Optimization Engineering : Quantization, Speculative Decoding, and Hardware-Specific LLM Deployment

Language: English

Published by Independently Published Jun 2026, 2026

9798199720021

Series: Book 6 of 21 - Production AI Engineering Series

  • Softcover
  • New
See all details

Seller: AHA-BUCH GmbH, Einbeck, GermanyAHA-BUCH GmbH

5-star seller

AbeBooks seller since August 14, 2006

View this seller's items
Softcover

Condition: New

£ 13.64

£ 30.09 shipping 
Ships from Germany to U.S.A.

Quantity: 2 available

Add to basket
Free 30-day returns

Item description from seller

Neuware - Slash LLM Deployment Costs and LatencyDeploying Large Language Models (LLMs) in production is a massive economic and engineering hurdle. AI Inference Optimization Engineering is your comprehensive, hands-on guide to mastering the full stack of modern LLM optimization techniques. From memory-bandwidth solutions to hardware-specific compilation, this book bridges the gap between research-level models and enterprise-grade execution.What you will master inside this book: - Hardware-Aware Optimization: Dive deep into KV cache mechanics, autoregressive decoding, and GPU memory hierarchies to eliminate latency bottlenecks.- State-of-the-Art Quantization: Apply GPTQ, AWQ, and GGUF compression algorithms to scale down massive neural networks without sacrificing model accuracy.- Advanced Acceleration Methods: Implement speculative decoding with draft models (like Medusa and Eagle), PagedAttention, and FlashAttention to boost throughput by 2-3x.- Production-Grade Serving: Build ultra-low-latency deployment infrastructures using vLLM, Triton Inference Server, and continuous batching.- Cross-Platform Deployment: Optimize models for specific target hardware, including NVIDIA H100 (TensorRT-LLM), Apple Silicon (llama.cpp/Metal), and Qualcomm mobile/edge accelerators.Whether you are an ML infrastructure engineer, an AI platform architect, or a technical leader looking to scale LLMs cost-effectively, this book provides the production-ready code, equations, and architectural patterns you need to build hyper-efficient AI pipelines.

Seller Inventory # 9798199720021

Title
AI Inference Optimization Engineering : Quantization, Speculative Decoding, and Hardware-Specific LLM Deployment
Author
Chatvariety Team
Publisher
Independently Published Jun 2026
Publication year
2026
Condition
Neu
Binding
Taschenbuch
Language
English
ISBN 13
9798199720021
Item weight
142 grams
Dimensions
229x152x5 mm
Series
Book 6 of 21: Production AI Engineering Series

AHA-BUCH GmbH

Einbeck, Germany

5-star seller

AbeBooks seller since August 14, 2006

Shipping rates from Germany to U.S.A.

Item7 to 10 business days5 to 7 business days
First item£ 30.09£ 38.69
Delivery times are set by sellers and vary by carrier and location. Orders passing through Customs may face delays and buyers are responsible for any associated duties or fees. Sellers may contact you regarding additional charges to cover any increased costs to ship your items.

Payment methods

  • Visa
  • Mastercard
  • American Express
  • Apple Pay
  • Google Pay
  • Bank Wire Transfer
  • Check
  • Paypal

Store description

Das Unternehmen AHA-BUCH GmbH: Seit der Gründung von AHA-BUCH im Juli 2005 ist unser Hauptziel, zufriedenen Kunden so schnell und so preisgünstig wie möglich ihren Bücherwunsch zu erfüllen. Unsere Firma beschäftigt 16 Mitarbeiter, die nur ein Ziel kennen: den Kunden und seine Wünsche! Auf über 3700 m2 Fläche haben wir über 100.000 Bücher, Modernes Antiquariat und Spiele auf Lager.

Specialty

Kinderbücher & Kinderhör Casetten, German Books, Software, Natur & Tiere, Ratgeber, Sachbücher, Englische Bücher, Medizin & Gesundheit, Universität & Studium

Seller's business information

AHA-BUCH GmbH

Garlebsen 48
Einbeck, Germany 37574