|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/llm-fine-tuning-grpo-sft-dpo-with-reinforcement-learning/
课程评论:没有评论
课程名称:LLM 强化学习微调深度寻求方法 GRPO 概述:在本课程中,您将步入大型语言模型(LLMs)的世界,学习基础和高级的端到端优化方法。课程开始于 SFT(监督微调)方法,您将了解如何有效地准备数据,并通过实践示例使用标记器和数据整理器创建定制数据集。在 SFT 过程中,您将学习使大型模型变得更轻巧和高效的关键技术,如 LoRA(低秩适应)和量化,并逐步探索如何将它们集成到您的项目中。 在巩固 SFT 基础知识后,我们将转向 DPO(直接偏好优化)。DPO 允许您通过直接反映用户反馈在模型中获得以用户为中心的结果。您将学习如何为该方法格式化数据,如何设计奖励机制,以及如何将训练好的模型共享到 Hugging Face 等热门平台。此外,您将深入了解数据整理器在 DPO 过程中的工作原理,学习在各种场景中准备和转换数据集的实用技巧。 课程的最重要阶段是 GRPO(群体相对策略优化),该方法因获得显著成果而受到广泛关注。通过 GRPO,您将学习优化模型行为的方法,既可以在个人层面进行,也可以在社区或不同用户群体中进行。这使得大型语言模型更系统、有效地服务于多样化的受众或目的。在本课程中,您将学习 GRPO 的基本原理,然后通过实际数据集应用这一技术来巩固您的知识。 在整个培训期间,我们将涵盖 LoRA、量化、SFT、DPO,特别是 GRPO 等关键主题,并通过项目导向的应用来支持每个主题。到本课程结束时,您将完全有能力从端到端数据准备到微调及基于群体的政策优化的每个阶段进行管理。开发现代且具有竞争力的 LLM 解决方案,专注于性能和用户满意度将变得更加简单。
In this course, you will step into the world of Large Language Models (LLMs) and learn both fundamental and advanced end-to-end optimization methods. You'll begin with the SFT (Supervised Fine-Tuning) approach, where you'll discover how to properly prepare your data and create customized datasets using tokenizers and data collators through practical examples. During the SFT process, you'll learn the key techniques for making large models lighter and more efficient with LoRA (Low-Rank Adaptation) and quantization, and explore step by step how to integrate them into your projects.After solidifying the basics of SFT, we will move on to DPO (Direct Preference Optimization). DPO allows you to obtain user-focused results by directly reflecting user feedback in the model. You'll learn how to format your data for this method, how to design a reward mechanism, and how to share models trained on popular platforms such as Hugging Face. Additionally, you'll gain a deeper understanding of how data collators work in DPO processes, learning practical techniques for preparing and transforming datasets in various scenarios.The most significant phase of the course is GRPO (Group Relative Policy Optimization), which has been gaining popularity for producing strong results. With GRPO, you will learn methods to optimize model behavior not only at the individual level but also within communities or across different user groups. This makes it more systematic and effective for large language models to serve diverse audiences or purposes. In this course, you'll learn the fundamental principles of GRPO, and then solidify your knowledge by applying this technique with real-world datasets.Throughout the training, we will cover key topics-LoRA, quantization, SFT, DPO, and especially GRPO-together, supporting each topic with project-oriented applications. By the end of this course, you will be fully equipped to manage every stage with confidence, from end-to-end data preparation to fine-tuning and group-based policy optimization. Developing modern and competitive LLM solutions that focus on both performance and user satisfaction in your own projects will become much easier.