|
所在平台: Coursera |
课程主页: https://www.coursera.org/learn/big-data-integration-processing
课程评论:没有评论
课程名称:大数据集成与处理 概述:完成本课程后,您将能够: * 从示例数据库和大数据管理系统中检索数据 * 描述数据管理操作与大数据处理模式之间的关联,以及如何在大规模分析应用程序中利用它们 * 识别何时需要数据集成以解决大数据问题 * 在Hadoop和Spark平台上执行简单的大数据集成与处理 该课程面向数据科学新手,建议完成《大数据入门》课程。无需具备编程经验,但需具备安装应用程序和使用虚拟机的能力以完成实践作业。有关完整的硬件和软件规范,请参阅专业技术要求。 硬件要求: (A) 四核处理器(建议支持VT-x或AMD-V),64位; (B) 8GB RAM; (C) 20GB可用磁盘空间。查找硬件信息的方法为:在Windows中,点击开始按钮,右键单击计算机,然后点击属性;在Mac中,点击苹果菜单,选择“关于这台Mac”。大多数在过去3年内购买的8GB RAM电脑均符合最低要求。您需要高速互联网连接,因为将下载高达4GB大小的文件。 软件要求: 本课程依赖于多个开源软件工具,包括Apache Hadoop。所有所需软件均可免费下载安装(不包括互联网提供商的流量费用)。软件要求包括:Windows 7+、Mac OS X 10.10+、Ubuntu 14.04+ 或 CentOS 6+,以及VirtualBox 5+。 课程大纲: - 名称:欢迎加入大数据集成与处理 描述:欢迎来到大数据专业的第三门课程。本周将介绍大数据集成和处理的基本概念。您将学习如何安装Docker、下载本课程所需的数据集,并学习如何使用Jupyter Notebook。 - 名称:大数据检索(第一部分) 描述:本模块涵盖数据检索和关系查询的各个方面,同时您将接触Postgres数据库。 - 名称:大数据检索(第二部分) 描述:本模块涵盖NoSQL数据的检索、数据聚合及数据帧的使用。您将接触MongoDB和Aerospike,并学习如何使用Pandas从中检索数据。 - 名称:大数据集成 描述:在本模块中,您将学习数据集成工具,包括Splunk和Datameer,并获得信息集成过程的实用见解。 - 名称:处理大数据 描述:本模块介绍大数据管道和工作流,以及使用Apache Spark处理与分析大数据的方法。 - 名称:使用Spark进行大数据分析 描述:在本模块中,您将深入了解大数据处理,学习Spark Core的内部工作原理,并接触Spark工具包中的两个关键工具:Spark MLlib和GraphX。 - 名称:通过实践学习:将MongoDB和Spark应用于工作 描述:在本模块中,您将获得实践经验,将所学内容应用于分析Twitter数据,利用Spark和MongoDB进行研究。
Name:Welcome to Big Data Integration and Processing
Description:Welcome to the third course in the Big Data Specialization. This week you will be introduced to basic concepts in big data integration and processing. You will be guided through installing Docker, downloading the data sets to be used for this course, and learning how to work with Jupyter notebooks.
Name:Retrieving Big Data (Part 1)
Description:This module covers the various aspects of data retrieval and relational querying. You will also be introduced to the Postgres database.
Name:Retrieving Big Data (Part 2)
Description:This module covers the various aspects of data retrieval for NoSQL data, as well as data aggregation and working with data frames. You will be introduced to MongoDB and Aerospike, and you will learn how to use Pandas to retrieve data from them.
Name:Big Data Integration
Description:In this module you will be introduced to data integration tools including Splunk and Datameer, and you will gain some practical insight into how information integration processes are carried out.
Name:Processing Big Data
Description:This module introduces Learners to big data pipelines and workflows as well as processing and analysis of big data using Apache Spark.
Name:Big Data Analytics using Spark
Description:In this module, you will go deeper into big data processing by learning the inner workings of the Spark Core. You will be introduced to two key tools in the Spark toolkit: Spark MLlib and GraphX.
Name:Learn By Doing: Putting MongoDB and Spark to Work
Description:In this module you will get some practical hands-on experience applying what you learned about Spark and MongoDB to analyze Twitter data.
At the end of the course, you will be able to: *Retrieve data from example database and big data management systems *Describe the connections between data management operations and the big data processing patterns needed to utilize them in large-scale analytical applications *Identify when a big data problem needs data integration *Execute simple big data integration and processing on Hadoop and Spark platforms This course is for those new to data science. Completion of Intro to Big Data is recommended. No prior programming experience is needed, although the ability to install applications and utilize a virtual machine is necessary to complete the hands-on assignments. Refer to the specialization technical requirements for complete hardware and software specifications. Hardware Requirements: (A) Quad Core Processor (VT-x or AMD-V support recommended), 64-bit; (B) 8 GB RAM; (C) 20 GB disk free. How to find your hardware information: (Windows): Open System by clicking the Start button, right-clicking Computer, and then clicking Properties; (Mac): Open Overview by clicking on the Apple menu and clicking “About This Mac.” Most computers with 8 GB RAM purchased in the last 3 years will meet the minimum requirements.You will need a high speed internet connection because you will be downloading files up to 4 Gb in size. Software Requirements: This course relies on several open-source software tools, including Apache Hadoop. All required software can be downloaded and installed free of charge (except for data charges from your internet provider). Software requirements include: Windows 7+, Mac OS X 10.10+, Ubuntu 14.04+ or CentOS 6+ VirtualBox 5+.