LLM Efficient Inference in CPUs and Intel GPUs - Intel Neural Speed
Offered By: The Machine Learning Engineer via YouTube
Course Description
Overview
Explore efficient inference techniques for Large Language Models (LLMs) on CPUs and Intel GPUs using Intel Neural Speed in this 30-minute video. Dive into the performance capabilities of Intel Extension for Transformers and gain practical insights through provided Jupyter notebooks. Learn how to optimize LLM inference for data science and machine learning applications, leveraging Intel's hardware-specific solutions. Access accompanying resources, including a Medium article and GitHub repositories, to deepen your understanding and implement the techniques discussed.
Syllabus
LLM Efficient Inference In CPUs and Intel GPUs. Intel Neural Speed #datascience #machinelearning
Taught by
The Machine Learning Engineer
Related Courses
Introduction to Artificial IntelligenceStanford University via Udacity Natural Language Processing
Columbia University via Coursera Probabilistic Graphical Models 1: Representation
Stanford University via Coursera Computer Vision: The Fundamentals
University of California, Berkeley via Coursera Learning from Data (Introductory Machine Learning course)
California Institute of Technology via Independent