Inference Ollama Llama Cpp Vllm by Marballi (10 results)

- Softcover
Seller: Rarewaves.com USA, London, LONDO, United KingdomRarewaves.com USA
Contact seller5-star sellerCondition: New
£ 24.27
Free ShippingShips from United Kingdom to U.S.A.Quantity: Over 20 available
Paperback. Condition: New.

- Softcover
Seller: California Books, Miami, FL, U.S.A.California Books
Contact seller5-star sellerCondition: New
£ 24.97
Free ShippingShips within U.S.A.Quantity: Over 20 available
Condition: New.

- Softcover
Seller: PBShop.store US, Wood Dale, IL, U.S.A.PBShop.store US
Contact seller5-star sellerCondition: New
£ 25.42
Free ShippingShips within U.S.A.Quantity: Over 20 available
PAP. Condition: New. New Book. Shipped from UK. Established seller since 2000.

- Softcover
Seller: PBShop.store UK, Fairford, GLOS, United KingdomPBShop.store UK
Contact seller5-star sellerCondition: New
£ 22.92
£ 4.16 shippingShips from United Kingdom to U.S.A.Quantity: Over 20 available
PAP. Condition: New. New Book. Shipped from UK. Established seller since 2000.

- Softcover
Seller: Rarewaves.com UK, London, United KingdomRarewaves.com UK
Contact seller5-star sellerCondition: New
£ 21.80
£ 65.00 shippingShips from United Kingdom to U.S.A.Quantity: Over 20 available
Paperback. Condition: New.

- Softcover
- Print on Demand
Seller: Grand Eagle Retail, Bensenville, IL, U.S.A.Grand Eagle Retail
Contact seller5-star sellerCondition: New
£ 27.42
Free ShippingShips within U.S.A.Quantity: 1 available
Paperback. Condition: new. Paperback. The era of cloud-dependent AI is over. Today's developers can run state-of-the-art language models on their own hardware-from laptops to GPU clusters-without ever sending data to a third party. But the gap between downloading a model and deploying it efficiently is filled with questions about quantization, memory bandwidth, batching strategies, and tool selection. This book is your guide through that gap, showing you how to build scalable, cost-effective inference systems using the three pillars of open-source AI: Ollama, llama.cpp, and vLLM. AI Inference with Ollama, llama.cpp, and vLLM takes you from running your first local model in minutes to optimizing production deployments serving thousands of requests per second. You'll learn when to use each tool, how to navigate the memory wall that bottlenecks LLM performance, and how to choose the right hardware and quantization strategy for your use case. Whether you're building RAG systems, deploying chatbots, or scaling inference across GPU clusters, this book gives you the practical knowledge to move from experimentation to production with confidence. About the Author GK Marballi has spent 20+ years turning data into competitive advantage for global brands from Priceline to S&P Global and Barnes & Noble. He has led high-impact product and analytics teams, and navigated the front lines of the AI revolution. He is based in New York City and holds an MBA from Harvard Business School. This item is printed on demand. Shipping may be from multiple locations in the US or from the UK, depending on stock availability.…

- Softcover
- Print on Demand
Seller: AussieBookSeller, Truganina, VIC, AustraliaAussieBookSeller
Contact seller5-star sellerCondition: New
£ 28.52
£ 28.04 shippingShips from Australia to U.S.A.Quantity: 1 available
Paperback. Condition: new. Paperback. The era of cloud-dependent AI is over. Today's developers can run state-of-the-art language models on their own hardware-from laptops to GPU clusters-without ever sending data to a third party. But the gap between downloading a model and deploying it efficiently is filled with questions about quantization, memory bandwidth, batching strategies, and tool selection. This book is your guide through that gap, showing you how to build scalable, cost-effective inference systems using the three pillars of open-source AI: Ollama, llama.cpp, and vLLM. AI Inference with Ollama, llama.cpp, and vLLM takes you from running your first local model in minutes to optimizing production deployments serving thousands of requests per second. You'll learn when to use each tool, how to navigate the memory wall that bottlenecks LLM performance, and how to choose the right hardware and quantization strategy for your use case. Whether you're building RAG systems, deploying chatbots, or scaling inference across GPU clusters, this book gives you the practical knowledge to move from experimentation to production with confidence. About the Author GK Marballi has spent 20+ years turning data into competitive advantage for global brands from Priceline to S&P Global and Barnes & Noble. He has led high-impact product and analytics teams, and navigated the front lines of the AI revolution. He is based in New York City and holds an MBA from Harvard Business School. This item is printed on demand. Shipping may be from our Sydney, NSW warehouse or from our UK or US warehouse, depending on stock availability.…

- Softcover
- Print on Demand
Seller: AHA-BUCH GmbH, Einbeck, GermanyAHA-BUCH GmbH
Contact seller5-star sellerCondition: New
£ 32.86
£ 29.83 shippingShips from Germany to U.S.A.Quantity: 2 available
Taschenbuch. Condition: Neu. nach der Bestellung gedruckt Neuware - Printed after ordering - The era of cloud-dependent AI is over. Today's developers can run state-of-the-art language models on their own hardware-from laptops to GPU clusters-without ever sending data to a third party. But the gap between downloading a model and deploying it efficiently is filled with questions about quantization, memory bandwidth, batching strategies, and tool selection. This book is your guide through that gap, showing you how to build scalable, cost-effective inference systems using the three pillars of open-source AI: Ollama, llama.cpp, and vLLM.AI Inference with Ollama, llama.cpp, and vLLM takes you from running your first local model in minutes to optimizing production deployments serving thousands of requests per second. You'll learn when to use each tool, how to navigate the memory wall that bottlenecks LLM performance, and how to choose the right hardware and quantization strategy for your use case. Whether you're building RAG systems, deploying chatbots, or scaling inference across GPU clusters, this book gives you the practical knowledge to move from experimentation to production with confidence.About the AuthorGK Marballi has spent 20+ years turning data into competitive advantage for global brands from Priceline to S&P Global and Barnes & Noble. He has led high-impact product and analytics teams, and navigated the front lines of the AI revolution. He is based in New York City and holds an MBA from Harvard Business School.…

- Softcover
- Print on Demand
Seller: CitiRetail, Stevenage, United KingdomCitiRetail
Contact seller5-star sellerCondition: New
£ 25.99
£ 37.00 shippingShips from United Kingdom to U.S.A.Quantity: 1 available
Paperback. Condition: new. Paperback. The era of cloud-dependent AI is over. Today's developers can run state-of-the-art language models on their own hardware-from laptops to GPU clusters-without ever sending data to a third party. But the gap between downloading a model and deploying it efficiently is filled with questions about quantization, memory bandwidth, batching strategies, and tool selection. This book is your guide through that gap, showing you how to build scalable, cost-effective inference systems using the three pillars of open-source AI: Ollama, llama.cpp, and vLLM. AI Inference with Ollama, llama.cpp, and vLLM takes you from running your first local model in minutes to optimizing production deployments serving thousands of requests per second. You'll learn when to use each tool, how to navigate the memory wall that bottlenecks LLM performance, and how to choose the right hardware and quantization strategy for your use case. Whether you're building RAG systems, deploying chatbots, or scaling inference across GPU clusters, this book gives you the practical knowledge to move from experimentation to production with confidence. About the Author GK Marballi has spent 20+ years turning data into competitive advantage for global brands from Priceline to S&P Global and Barnes & Noble. He has led high-impact product and analytics teams, and navigated the front lines of the AI revolution. He is based in New York City and holds an MBA from Harvard Business School. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability.…

- Softcover
- Print on Demand
Seller: preigu, Osnabrück, Germanypreigu
Contact seller5-star sellerCondition: New
£ 30.85
£ 59.66 shippingShips from Germany to U.S.A.Quantity: 5 available
Taschenbuch. Condition: Neu. AI Inference with Ollama, [.], and vLLM | Gk Marballi | Taschenbuch | Englisch | 2026 | [.] | EAN 9781105842733 | Verantwortliche Person für die EU: Libri GmbH, Europaallee 1, 36244 Bad Hersfeld, gpsr[at]libri[dot]de | Anbieter: preigu Print on Demand.