|
所在平台: Coursera |
课程主页: https://www.coursera.org/learn/big-data-emerging-technologies
课程评论:没有评论
课程名称:大数据新兴技术 课程概述:在每天使用谷歌搜索、社交网络服务(如Facebook和Twitter)以及从亚马逊等平台购买推荐商品时,您无时无刻不在使用大数据系统。大数据技术不仅支持智能手机、智能手表、语音助手如Alexa和Siri,也在现代汽车中扮演着重要角色。如今,全球顶尖公司均在应用大数据技术,以应对生存与发展的需求。因此,了解大数据及其应用对公司发展至关重要。本课程共分六个模块,首先介绍大数据硬件、软件及专业服务的市场份额排名,接着阐述主要大数据公司的产品线和服务类型,随后深入讲解市场上最流行的三种大数据技术:Hadoop、Spark和Storm,最后提供IBM SPSS Statistics的实际操作体验,帮助学员在即将到来的大数据时代中更好地进行商业战略规划。欢迎走进令人惊叹的大数据世界! 课程大纲: 1. 大数据排名与产品:探讨大数据硬件、软件和专业服务的市场关系及其未来对行业、产品、服务与政府组织的影响。 2. 大数据与Hadoop:介绍Hadoop的特性及操作,解读MapReduce和HDFS的功能,并比较Hadoop与SQL的差异。 3. Spark:讲解Spark的特性及操作,分析Spark与Hadoop的数据分析特点,以及分布式数据集(RDD)等核心组件。 4. Spark ML与流处理:分析Spark的机器学习算法及实时流处理的工作机制,介绍Spark ML基本统计算法和流处理特点。 5. Storm:讲解Storm技术的操作与特性,比较Storm与其他技术的数据分析特色,以及Storm在实时应用中的优势。 6. IBM SPSS Statistics项目:提供IBM SPSS Statistics的使用经验,进行基于该系统的数据处理和统计分析项目,探索数据集之间的关系及其统计结果。 通过本课程的学习,您将能够掌握大数据技术并有效应用于商业战略规划中。
Name:Big Data Rankings & Products
Description:The first module “Big Data Rankings & Products” focuses on the relation and market shares of big data hardware, software, and professional services. This information provides an insight to how future industry, products, services, schools, and government organizations will be influenced by big data technology. To have a deeper view into the world’s top big data products line and service types, the lecture provides an overview on the major big data company, which include IBM, SAP, Oracle, HPE, Splunk, Dell, Teradata, Microsoft, Cisco, and AWS. In order to understand the power of big data technology, the difference of big data analysis compared to traditional data analysis is explained. This is followed by a lecture on the 4 V big challenges of big data technology, which deal with issues in the volume, variety, velocity, and veracity of the massive data. Based on this introduction information, big data technology used in adding global insights on investments, help locate new stores and factories, and run real-time recommendation systems by Wal-Mart, Amazon, and Citibank is introduced.
Name:Big Data & Hadoop
Description:The second module “Big Data & Hadoop” focuses on the characteristics and operations of Hadoop, which is the original big data system that was used by Google. The lectures explain the functionality of MapReduce, HDFS (Hadoop Distributed FileSystem), and the processing of data blocks. These functions are executed on a cluster of nodes that are assigned the role of NameNode or DataNodes, where the data processing is conducted by the JobTracker and TaskTrackers, which are explained in the lectures. In addition, the characteristics of metadata types and the differences in the data analysis processes of Hadoop and SQL (Structured Query Language) are explained. Then the Hadoop Release Series is introduced which include the descriptions of Hadoop YARN (Yet Another Resource Negotiator), HDFS Federation, and HDFS HA (High Availability) big data technology.
Name:Spark
Description:The third module “Spark” focuses on the operations and characteristics of Spark, which is currently the most popular big data technology in the world. The lecture first covers the differences in data analysis characteristics of Spark and Hadoop, then goes into the features of Spark big data processing based on the RDD (Resilient Distributed Datasets), Spark Core, Spark SQL, Spark Streaming, MLlib (Machine Learning Library), and GraphX core units. Details of the features of Spark DAG (Directed Acyclic Graph) stages and pipeline processes that are formed based on Spark transformations and actions are explained. Especially, the definition and advantages of lazy transformations and DAG operations are described along with the characteristics of Spark variables and serialization. In addition, the process of Spark cluster operations based on Mesos, Standalone, and YARN are introduced.
Name:Spark ML & Streaming
Description:The fourth module “Spark ML & Streaming” focuses on how Spark ML (Machine Learning) works and how Spark streaming operations are conducted. The Spark ML algorithms include featurization, pipelines, persistence, and utilities which operate on the RDDs (Resilient Distributed Datasets) to extract information form the massive datasets. The lectures explain the characteristics of the DataFrame-based API, which is the primary ML API in the spark.ml package. Spark ML basic statistics algorithms based on correlation and hypothesis testing (P-value) are first introduced followed by the Spark ML classification and regression algorithms based on linear models, naive Bayes, and decision tree techniques. Then the characteristics of Spark streaming, streaming input and output, as well as streaming receiver types (which include basic, custom, and advanced) are explained, followed by how the Spark Streaming process and DStream (Discretized Stream) enable big data streaming operations for real-time and near-real-time applications.
Name:Storm
Description:The fifth module “Storm” focuses on the characteristics and operations of Storm big data systems. The lecture first covers the differences in data analysis characteristics of Storm, Spark, and Hadoop technology. Then the features of Storm big data processing based on the nimbus, spouts, and bolts are described followed by the Storm streams, supervisor, and ZooKeeper details. Further details on Storm reliable and unreliable spouts and bolts are provided followed by the advantages of Storm DAG (Directed Acyclic Graph) and data stream queue management. In addition, the advantages of using Storm based fast real-time applications, which include real-time analytics, online ML (Machine Learning), continuous computation, DRPC (Distributed Remote Procedure Call), and ETL (Extract, Transform, Load) are introduced.
Name:IBM SPSS Statistics Project
Description:The sixth and last module “IBM SPSS Statistics Project” focuses on providing experience on one of the most famous and widely used big data statistical analysis systems in the world. First, the lecture starts with how to setup and use IBM SPSS Statistics, and continues on to describe how IBM SPSS Statistics can be used to gain corporate data analysis experience. Then the data processing statistical results of two projects based on using the IBM SPSS Statistics big data system is conducted. The projects are conducted so the student can discover new ways to use, analyze, and draw charts of the relationship between datasets, and also compare the statistical results using IBM SPSS Statistics.
Every time you use Google to search something, every time you use Facebook, Twitter, Instagram or any other SNS (Social Network Service), and every time you buy from a recommended list of products on Amazon.com you are using a big data system. In addition, big data technology supports your smartphone, smartwatch, Alexa, Siri, and automobile (if it is a newer model) every day. The top companies in the world are currently using big data technology, and every company is in need of advanced big data technology support. Simply put, big data technology is not an option for your company, it is a necessity for survival and growth. So now is the right time to learn what big data is and how to use it in advantage of your company. This 6 module course first focuses on the world’s industry market share rankings of big data hardware, software, and professional services, and then covers the world’s top big data product line and service types of the major big data companies. Then the lectures focused on how big data analysis is possible based on the world’s most popular three big data technologies Hadoop, Spark, and Storm. The last part focuses on providing experience on one of the most famous and widely used big data statistical analysis systems in the world, the IBM SPSS Statistics. This course was designed to prepare you to be more successful in businesses strategic planning in the upcoming big data era. Welcome to the amazing Big Data world!