Isbn: 9798199720021 - Ai Inference Optimization Engineering: Quantization, Speculative Decoding, and Hardware-specific Llm Deployment (production Ai Engineering Series) (5 results)

ISBN
Refine with Advanced Search

Refine your search

  • Books (5)

  • New (5)

to

Custom price range (£)

to

  • Language: English

    Published by Independently published, 2026

    9798199720021

    Series: Book 6 of 20 - Production AI Engineering Series

    • Softcover

    Seller: PBShop.store US, Wood Dale, IL, U.S.A.PBShop.store US

    5-star seller
    Contact seller

    Condition: New

    £ 12.04

     Free Shipping 
    Ships within U.S.A.

    Quantity: Over 20 available

    PAP. Condition: New. New Book. Shipped from UK. Established seller since 2000.

  • Language: English

    Published by Independently published, 2026

    9798199720021

    Series: Book 6 of 20 - Production AI Engineering Series

    • Softcover

    Seller: PBShop.store UK, Fairford, GLOS, United KingdomPBShop.store UK

    5-star seller
    Contact seller

    Condition: New

    £ 11.15

    £ 3.29 shipping 
    Ships from United Kingdom to U.S.A.

    Quantity: Over 20 available

    PAP. Condition: New. New Book. Shipped from UK. Established seller since 2000.

  • Language: English

    Published by Independently Published Jun 2026, 2026

    9798199720021

    Series: Book 6 of 20 - Production AI Engineering Series

    • Softcover

    Seller: AHA-BUCH GmbH, Einbeck, GermanyAHA-BUCH GmbH

    5-star seller
    Contact seller

    Condition: New

    £ 13.61

    £ 30.03 shipping 
    Ships from Germany to U.S.A.

    Quantity: 2 available

    Taschenbuch. Condition: Neu. Neuware - Slash LLM Deployment Costs and LatencyDeploying Large Language Models (LLMs) in production is a massive economic and engineering hurdle. AI Inference Optimization Engineering is your comprehensive, hands-on guide to mastering the full stack of modern LLM optimization techniques. From memory-bandwidth solutions to hardware-specific compilation, this book bridges the gap between research-level models and enterprise-grade execution.What you will master inside this book: - Hardware-Aware Optimization: Dive deep into KV cache mechanics, autoregressive decoding, and GPU memory hierarchies to eliminate latency bottlenecks.- State-of-the-Art Quantization: Apply GPTQ, AWQ, and GGUF compression algorithms to scale down massive neural networks without sacrificing model accuracy.- Advanced Acceleration Methods: Implement speculative decoding with draft models (like Medusa and Eagle), PagedAttention, and FlashAttention to boost throughput by 2-3x.- Production-Grade Serving: Build ultra-low-latency deployment infrastructures using vLLM, Triton Inference Server, and continuous batching.- Cross-Platform Deployment: Optimize models for specific target hardware, including NVIDIA H100 (TensorRT-LLM), Apple Silicon (llama.cpp/Metal), and Qualcomm mobile/edge accelerators.Whether you are an ML infrastructure engineer, an AI platform architect, or a technical leader looking to scale LLMs cost-effectively, this book provides the production-ready code, equations, and architectural patterns you need to build hyper-efficient AI pipelines.

  • Language: English

    Published by Independently published, 2026

    9798199720021

    Series: Book 6 of 20 - Production AI Engineering Series

    • Softcover
    • Print on Demand

    Seller: California Books, Miami, FL, U.S.A.California Books

    4-star seller
    Contact seller

    Condition: New

    £ 11.58

     Free Shipping 
    Ships within U.S.A.

    Quantity: Over 20 available

    Condition: New. Print on Demand.

  • Language: English

    Published by Independently Published, 2026

    9798199720021

    Series: Book 6 of 20 - Production AI Engineering Series

    • Softcover
    • Print on Demand

    Seller: CitiRetail, Stevenage, United KingdomCitiRetail

    5-star seller
    Contact seller

    Condition: New

    £ 13.99

    £ 37.00 shipping 
    Ships from United Kingdom to U.S.A.

    Quantity: 1 available

    Paperback. Condition: new. Paperback. Slash LLM Deployment Costs and LatencyDeploying Large Language Models (LLMs) in production is a massive economic and engineering hurdle. AI Inference Optimization Engineering is your comprehensive, hands-on guide to mastering the full stack of modern LLM optimization techniques. From memory-bandwidth solutions to hardware-specific compilation, this book bridges the gap between research-level models and enterprise-grade execution.What you will master inside this book: Hardware-Aware Optimization: Dive deep into KV cache mechanics, autoregressive decoding, and GPU memory hierarchies to eliminate latency bottlenecks.State-of-the-Art Quantization: Apply GPTQ, AWQ, and GGUF compression algorithms to scale down massive neural networks without sacrificing model accuracy.Advanced Acceleration Methods: Implement speculative decoding with draft models (like Medusa and Eagle), PagedAttention, and FlashAttention to boost throughput by 2-3x.Production-Grade Serving: Build ultra-low-latency deployment infrastructures using vLLM, Triton Inference Server, and continuous batching.Cross-Platform Deployment: Optimize models for specific target hardware, including NVIDIA H100 (TensorRT-LLM), Apple Silicon (llama.cpp/Metal), and Qualcomm mobile/edge accelerators.Whether you are an ML infrastructure engineer, an AI platform architect, or a technical leader looking to scale LLMs cost-effectively, this book provides the production-ready code, equations, and architectural patterns you need to build hyper-efficient AI pipelines. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.