Deconstructing Text Embedding Models - Understanding Tokenizers and Model Selection
Offered By: EuroPython Conference via YouTube
Course Description
Overview
Explore the intricacies of text embedding models in this 44-minute EuroPython Conference talk. Delve into the critical role of tokenizers in model selection, moving beyond reliance on benchmarks like the Massive Text Embedding Benchmark (MTEB). Learn to assess model suitability for specific datasets based on tokenizer performance, and discover strategies for optimizing tokenizers during the fine-tuning process of embedding models. Gain insights into making informed decisions when choosing text embedding models for unique data characteristics.
Syllabus
Deconstructing the text embedding models — Kacper Łukawski
Taught by
EuroPython Conference
Related Courses
Social Network AnalysisUniversity of Michigan via Coursera Intro to Algorithms
Udacity Data Analysis
Johns Hopkins University via Coursera Computing for Data Analysis
Johns Hopkins University via Coursera Health in Numbers: Quantitative Methods in Clinical & Public Health Research
Harvard University via edX