|
所在平台: Coursera |
课程主页: https://www.coursera.org/learn/hadoop
课程评论:没有评论
课程名称:Hadoop平台与应用框架 概述:本课程针对初学者程序员或希望了解大数据分析核心工具的商业人士。即使没有相关经验,您也可以通过实际示例学习Hadoop和Spark框架,这两个是业界最常用的工具。课程将使您能够清楚地解释Hadoop架构、软件栈和执行环境的具体组件和基本流程。在作业中,您将学习数据科学家如何应用重要概念和技术,如Map-Reduce,以解决大数据中的基础问题。通过本课程,您将能够自信地与他人讨论大数据及数据分析流程。 课程大纲: 1. 模块名称:Hadoop基础 描述:本模块将提供关于大数据热潮、其技术、机遇和挑战的洞察,深入研究Hadoop栈以及与大数据解决方案相关的工具和技术。 2. 模块名称:Hadoop栈介绍 描述:本模块将详细介绍Hadoop栈,从基本的HDFS组件到应用执行框架、语言和服务。 3. 模块名称:Hadoop分布式文件系统(HDFS)介绍 描述:本模块将深度探讨HDFS,涵盖其主要设计目标、读写过程、HDFS性能与鲁棒性的配置参数,并概述访问HDFS数据的不同方法。 4. 模块名称:Map/Reduce介绍 描述:本模块将介绍Map/Reduce的概念与实践,学习Map/Reduce的核心理念,如何设计、实现及执行map/reduce框架中的任务,以及在map/reduce中的权衡与其他工具的关系。 5. 模块名称:Spark 描述:本模块将介绍Apache Spark集群计算框架,这是Hadoop MapReduce在大数据领域的重要竞争者。Spark相较于Hadoop MapReduce在性能上具有显著优势,特别是对于迭代算法,得益于其内存缓存特性。此外,Spark使数据科学家更容易以Python和Scala编写分析管道,甚至提供交互式环境即时处理数据。
Name:Hadoop Basics
Description:Welcome to the first module of the Big Data Platform course. This first module will provide insight into Big Data Hype, its technologies opportunities and challenges. We will take a deeper look into the Hadoop stack and tool and technologies associated with Big Data solutions.
Name:Introduction to the Hadoop Stack
Description:In this module we will take a detailed look at the Hadoop stack ranging from the basic HDFS components, to application execution frameworks, and languages, services.
Name:Introduction to Hadoop Distributed File System (HDFS)
Description:In this module we will take a detailed look at the Hadoop Distributed File System (HDFS). We will cover the main design goals of HDFS, understand the read/write process to HDFS, the main configuration parameters that can be tuned to control HDFS performance and robustness, and get an overview of the different ways you can access data on HDFS.
Name:Introduction to Map/Reduce
Description:This module will introduce Map/Reduce concepts and practice. You will learn about the big idea of Map/Reduce and you will learn how to design, implement, and execute tasks in the map/reduce framework. You will also learn the trade-offs in map/reduce and how that motivates other tools.
Name:Spark
Description:Welcome to module 5, Introduction to Spark, this week we will focus on the Apache Spark cluster computing framework, an important contender of Hadoop MapReduce in the Big Data Arena. Spark provides great performance advantages over Hadoop MapReduce,especially for iterative algorithms, thanks to in-memory caching. Also, gives Data Scientists an easier way to write their analysis pipeline in Python and Scala,even providing interactive shells to play live with data.
This course is for novice programmers or business people who would like to understand the core tools used to wrangle and analyze big data. With no prior experience, you will have the opportunity to walk through hands-on examples with Hadoop and Spark frameworks, two of the most common in the industry. You will be comfortable explaining the specific components and basic processes of the Hadoop architecture, software stack, and execution environment. In the assignments you will be guided in how data scientists apply the important concepts and techniques such as Map-Reduce that are used to solve fundamental problems in big data. You'll feel empowered to have conversations about big data and the data analysis process.