Items related to Vision Language Models: Building Vlms with Hugging...

Vision Language Models: Building Vlms with Hugging Face - Softcover

Noyan, Merve; Marafioti, Andrés; Farré, Miquel; Zohar, Orr

 
9798341624047: Vision Language Models: Building Vlms with Hugging Face

Synopsis

Vision language models (VLMs) combine computer vision and natural language processing to create powerful systems that can interpret, generate, and respond in multimodal contexts. Vision Language Models is a hands-on guide to building real-world VLMs using the most up-to-date stack of machine learning tools from Hugging Face, Meta (PyTorch), NVIDIA (Cuda), and others, written by leading researchers and practitioners Merve Noyan, Miquel Farré, Andrés Marafioti, and Orr Zohar. From image captioning and document understanding to advanced zero-shot inference and retrieval-augmented generation (RAG), this book covers the full VLM application and development lifecycle.

"synopsis" may belong to another edition of this title.

About the Author

Merve Noyan is a machine learning engineer working in the ML advocacy engineering team at Hugging Face. She builds tools to enable people to build with vision language models across the Hugging Face ecosystem (transformers, TRL, smolagents). Previously she worked for different companies building natural language understanding based solutions on information retrieval and conversational agents.

"About this title" may belong to another edition of this title.