Deep Learning for NLP - Part 7

所在平台: Udemy

课程主页: https://www.udemy.com/course/ahol-dl4nlp7/

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:深度学习与自然语言处理(NLP) - 第七部分 概述:近年来,自然语言处理(NLP)和信息检索(IR)领域取得了巨大进展,这得益于深度学习模型的发展,例如循环神经网络(RNN)、门控循环单元(GRU)、长短期记忆网络(LSTM)以及基于Transformer的模型,如双向编码器表示(BERT)、生成预训练变换器(GPT-2)、多任务深度神经网络(MT-DNN)、超长网络(XLNet)、文本到文本转换器(T5)、T-NLG和GShard等。这些模型的参数规模庞大,例如BERT(3.4亿参数)、GPT-2(15亿参数)、T5(110亿参数,21.7GB)等。然而,现实世界的应用却需要模型体积小、响应时间快速且计算功耗低。 本课程讨论了五种模型压缩方法(剪枝、量化、知识蒸馏、参数共享、张量分解),以便在实际的行业NLP项目中应用这些模型。考虑到构建高效小模型应用程序的迫切需求,以及近期在该领域发表的大量研究工作,我们认为这门课程将深度学习与NLP领域的丰富研究整理成一个连贯的故事。 近年来,针对深度学习文本模型的压缩引起了研究界和工业界的广泛关注。许多企业主因模型大小和基础设施要求而犹豫不决,移动应用程序需要低内存占用,并显然需要小的功耗。物联网(IoT)和嵌入式系统相关组织在为资源受限环境(如传感器)设计机器学习解决方案方面投入了相当大的精力。参与应用深度学习研究的人员将从本课程中获益匪浅,因为该教程将为他们提供有关实际深度学习的全面概述。我们相信该教程将为新手提供当前工作的完整视图,介绍该领域的重要研究主题,并激励他们深入学习。业界从业者也将从方法和应用角度的讨论中受益,了解这些机制的应用现状。该教程可视为中级教程,假设听众具有一些基本的深度学习架构知识。先修知识包括对深度学习的入门知识的了解,特别是循环神经网络模型和变换器模型。此外,期待对自然语言处理和机器学习概念有基本的理解。

课程评论(0条)

课程详情

In recent years, the fields of natural language processing (NLP) and information retrieval (IR) have made tremendous progress thanks to deep learning models like Recurrent Neural Networks (RNNs), Gated Recurrent Units (GRUs) and Long Short-Term Memory (LSTMs) networks, and Transformer based models like Bidirectional Encoder Representations from Transformers (BERT), Generative Pre-training Transformer (GPT-2), Multi-task Deep Neural Network (MT-DNN), Extra-Long Network (XLNet), Text-to-text transfer transformer (T5), T-NLG and GShard. These models are humongous in size: BERT (340M parameters), GPT-2 (1.5B parameters), T5 (11B parameters, 21.7GB), etc. On the other hand, real world applications demand small model size, low response times and low computational power wattage. In this course, we discuss five different types of methods (Pruning, Quantization, Knowledge Distillation, Parameter Sharing, Tensor Decomposition) for compression of such models to enable their deployment in real industry NLP projects. Given the critical need of building applications with efficient and small models, and the large amount of recently published work in this area, we believe that this course organizes the plethora of work done by the "deep learning for NLP" community in the past few years and presents it as a coherent story.Compression for deep learning text models has gained a lot of interest in recent years both from the research community and the industry. Many business owners shy away from using deep learning models fearing the model sizes and infrastructure requirements. Mobile apps need to have a low RAM footprint and clearly a small power envelope. IoT (Internet of Things) and embedded systems related organizations have been investing significantly in designing machine learning solutions for resource constrained environments like sensors. Researchers in the field of applied deep learning for text will benefit the most, as this tutorial will give them an exhaustive overview of the research in the direction of practical deep learning. We believe that the tutorial will give the newcomers a complete picture of the current work, introduce important research topics in this field, and inspire them to learn more. Practitioners and people from the industry will clearly benefit from the discussions both from the methods perspective, as well from the point of view of applications where such mechanisms are starting to be deployed. This tutorial can be considered an intermediate level tutorial where we assume the folks in audience to know some basic deep learning architectures. Prerequisite knowledge includes introductory level knowledge in deep learning, specifically recurrent neural networks models, and transformers. Also, basic understanding of natural language processing and machine learning concepts is expected.

课程标签

0人关注该课程

主题相关的课程