|
所在平台: Coursera |
课程主页: https://www.coursera.org/learn/text-mining
课程评论:没有评论
课程名称:文本挖掘与分析 课程概述:本课程将涵盖挖掘和分析文本数据的主要技术,以发现有趣的模式、提取有用的知识以及支持决策,重点强调统计方法,这些方法可以在任何自然语言的文本数据中普遍应用,且需最低限度的人为干预。深入分析文本数据需要对自然语言文本的理解,这对计算机来说是一个具有挑战性的任务。然而,许多统计方法已被证明在模式发现和知识发现的“浅层”但稳健的文本数据分析中表现良好。学员将学习文本挖掘的基本概念、原理及主要算法及其潜在应用。 课程大纲: 1. 课程介绍:熟悉课程内容、同学及学习环境,培养完成课程所需的技术技能。 2. 第1周:学习自然语言处理技术和文本表示的概述,这些是所有文本挖掘应用的基础,重点分析词汇关联挖掘中的典型关系。 3. 第2周:深入了解词汇关联挖掘的另一种形式(序列关系),并开始学习主题分析,重点在于从文本中挖掘单一主题的技术。 4. 第3周:深入学习主题分析,包括混合模型及其工作原理,期望最大化(EM)算法及其在混合模型参数估计中的应用,基本主题模型,概率潜在语义分析(PLSA)及其扩展——潜在狄利克雷分配(LDA)。 5. 第4周:学习文本聚类的基本概念、主要聚类技术(包括概率方法和基于相似度的方法)以及如何评估文本聚类。同时,开始学习文本分类,文本分类与文本聚类相关,但有预定义的类别,相当于预定义的聚类。 6. 第5周:继续学习各种文本分类方法,包括多个判别分类器下的分类技术,并学习情感分析和意见挖掘,包括对情感分类的特定技术(即有序回归)的详细介绍。 7. 第6周:继续情感分析和意见挖掘,重点关注潜在方面评分分析(LARA),学习文本和非文本数据的联合挖掘技术,包括结合时间、地点、作者和数据来源等各种上下文信息分析文本主题的上下文文本挖掘技术。最后,将总结整个课程。 通过本课程,学员将掌握分析文本数据的关键技术和方法,为处理和理解自然语言中的信息提供坚实的基础。
Name:Orientation
Description:You will become familiar with the course, your classmates, and our learning environment. The orientation will also help you obtain the technical skills required for the course.
Name:Week 1
Description:During this module, you will learn the overall course design, an overview of natural language processing techniques and text representation, which are the foundation for all kinds of text-mining applications, and word association mining with a particular focus on mining one of the two basic forms of word associations (i.e., paradigmatic relations).
Name:Week 2
Description:During this module, you will learn more about word association mining with a particular focus on mining the other basic form of word association (i.e., syntagmatic relations), and start learning topic analysis with a focus on techniques for mining one topic from text.
Name:Week 3
Description:During this module, you will learn topic analysis in depth, including mixture models and how they work, Expectation-Maximization (EM) algorithm and how it can be used to estimate parameters of a mixture model, the basic topic model, Probabilistic Latent Semantic Analysis (PLSA), and how Latent Dirichlet Allocation (LDA) extends PLSA.
Name:Week 4
Description:During this module, you will learn text clustering, including the basic concepts, main clustering techniques, including probabilistic approaches and similarity-based approaches, and how to evaluate text clustering. You will also start learning text categorization, which is related to text clustering, but with pre-defined categories that can be viewed as pre-defining clusters.
Name:Week 5
Description:During this module, you will continue learning about various methods for text categorization, including multiple methods classified under discriminative classifiers, and you will also learn sentiment analysis and opinion mining, including a detailed introduction to a particular technique for sentiment classification (i.e., ordinal regression).
Name:Week 6
Description:During this module, you will continue learning about sentiment analysis and opinion mining with a focus on Latent Aspect Rating Analysis (LARA), and you will learn about techniques for joint mining of text and non-text data, including contextual text mining techniques for analyzing topics in text in association with various context information such as time, location, authors, and sources of data. You will also see a summary of the entire course.
This course will cover the major techniques for mining and analyzing text data to discover interesting patterns, extract useful knowledge, and support decision making, with an emphasis on statistical approaches that can be generally applied to arbitrary text data in any natural language with no or minimum human effort. Detailed analysis of text data requires understanding of natural language text, which is known to be a difficult task for computers. However, a number of statistical approaches have been shown to work well for the "shallow" but robust analysis of text data for pattern finding and knowledge discovery. You will learn the basic concepts, principles, and major algorithms in text mining and their potential applications.