Unsupervised Text Classification for Marketing Analytics

所在平台: Coursera

课程主页: https://www.coursera.org/learn/unsupervised-text-classification-for-marketing-analytics

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:无监督文本分类与市场分析 课程概述:市场数据通常庞大到人类无法阅读或分析其代表性样本来理解其中潜在的见解。在本课程中,学习者将使用无监督深度学习来训练算法,从文本数据中提取主题和洞察。学习者将通过教师引导的Python教程,了解无监督机器学习的概念概述,并深入真实世界的数据集。课程最后将以一个主要项目作为结束。 本课程使用Jupyter Notebooks及基于浏览器的Jupyter Notebook环境Google Colab,文件存储在Google Drive。 此课程可作为CU Boulder数据科学硕士(MS-DS)学位的一部分获得学分,该学位在Coursera平台上提供。MS-DS是一个跨学科的学位,汇聚了CU Boulder应用数学、计算机科学、信息科学等多个部门的教师。该项目采用基于表现的录取方式,无需申请流程,适合具有广泛本科教育背景和/或计算机科学、信息科学、数学和统计学专业经验的个人。了解更多关于MS-DS项目的信息,请访问:https://www.coursera.org/degrees/master-of-science-data-science-boulder。 课程大纲: 1. 模块标题:主题建模是什么? 描述:本模块将介绍主题建模的基本概念,也称为对非结构化文本文档的无监督机器学习。将无监督方法与监督方法进行对比,并调查主题建模的常见应用。 2. 模块标题:主题模型的假设、词袋模型与自然语言处理 描述:本模块深入探讨主题建模方法,理解驱动主题模型拟合的假设。还将揭示词袋模型在主题建模中的工作原理,以及进行 meaningful 主题建模特征所需的自然语言处理。 3. 模块标题:准备亚马逊评论数据 描述:本模块将介绍如何解析 JSON 格式的数据并进行分割,以创建一个准备进行主题建模过程的语料库。将涵盖您的项目数据结构及其分类法。 4. 模块标题:文本预处理与主题模型训练 描述:本模块将使用亚马逊评论数据加载到语料库中进行预处理。将涵盖如何从数据中构建主题模型并保存这些主题模型。 5. 模块标题:主题建模的评估、分类和神经网络方法 描述:本模块将学习如何评估主题模型的拟合情况,并使用最佳主题模型对文档进行分类。同时将介绍如何利用预训练的神经网络构建主题模型。

课程大纲

Part: 1

Title:What is topic modeling?

Description:In this module, we will cover the fundamental concepts of topic modeling, also known as unsupervised machine learning on unstructured text documents. We will contrast unsupervised methods to supervised ones and survey common applications of topic modeling.

Part: 2

Title:The Assumptions of a Topic Model, Bag of Words, and Natural Language Processing

Description:In this module, we will go under the hood inside a topic modeling approach and understand what assumptions drive topic model fit. We will also uncover how bag-of-words approaches to topic modeling work, and the natural language processing required to produce meaningful topic modeling features.

Part: 3

Title:Prepping Amazon Review Data

Description:In this module, we will cover how to parse through JSON-like data and segment it to create a corpus that is ready for the topic modeling process. We will cover how the data for your project is structured and its taxonomy.

Part: 4

Title:Pre-Processing Text and Training a Topic Model

Description:In this module, we will take Amazon review data and load it into a corpus to preprocess it. We will cover how to build topic models from the data and also save those topic models.

Part: 5

Title:Topic Modeling Evaluation, Classification, and Neural Network Approaches

Description:In this module, we will learn how to evaluate the fit of topic models and use the best topic model to classify documents. We will also cover how to build topic models with pre-trained neural networks.

课程评论(0条)

课程详情

Marketing data is often so big that humans cannot read or analyze a representative sample of it to understand what insights might lie within. In this course, learners use unsupervised deep learning to train algorithms to extract topics and insights from text data. Learners walk through a conceptual overview of unsupervised machine learning and dive into real-world datasets through instructor-led tutorials in Python. The course concludes with a major project. This course uses Jupyter Notebooks and the coding environment Google Colab, a browser-based Jupyter notebook environment. Files are stored in Google Drive. This course can be taken for academic credit as part of CU Boulder’s Master of Science in Data Science (MS-DS) degree offered on the Coursera platform. The MS-DS is an interdisciplinary degree that brings together faculty from CU Boulder’s departments of Applied Mathematics, Computer Science, Information Science, and others. With performance-based admissions and no application process, the MS-DS is ideal for individuals with a broad range of undergraduate education and/or professional experience in computer science, information science, mathematics, and statistics. Learn more about the MS-DS program at https://www.coursera.org/degrees/master-of-science-data-science-boulder.

课程标签

0人关注该课程

主题相关的课程