Vision Language Models and PDFs: What You See Is What You Search - Haystack EU 2024
Offered By: OpenSource Connections via YouTube
Course Description
Overview
Explore a groundbreaking approach to information extraction from complex PDF documents in this conference talk from Haystack EU 2024. Discover how Vision Language Models (VLMs) are revolutionizing the traditional multi-step process of text extraction, OCR, layout analysis, chunking, and embedding. Learn about ColPali, a new retrieval model that efficiently embeds entire PDF pages, including text, figures, and charts, resulting in improved retrieval quality and a simplified extraction and indexing process. Gain insights into representing ColPali in Vespa and its superior performance on the Visual Document Retrieval (ViDoRe) Benchmark. Benefit from the expertise of Jo Kristian, Chief Scientist at Vespa.ai, as he shares his two decades of experience in building and deploying search and recommender systems.
Syllabus
Haystack EU 2024 - Jo Kristian Bergum:What You See Is What You Search: Vision Language Models & PDFs
Taught by
OpenSource Connections
Related Courses
Mastering Google's PaliGemma VLM: Tips and Tricks for Success and Fine-TuningSam Witteveen via YouTube Fine-tuning PaliGemma for Custom Object Detection
Roboflow via YouTube Florence-2: The Best Small Vision Language Model - Capabilities and Demo
Sam Witteveen via YouTube Fine-tuning Florence-2: Microsoft's Multimodal Model for Custom Object Detection
Roboflow via YouTube OpenVLA: An Open-Source Vision-Language-Action Model - Research Presentation
HuggingFace via YouTube