|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/hands-on-with-hadoop-2-3-in-1/
课程评论:没有评论
课程名称:动手实践Hadoop 2:3合1 概述:Hadoop是目前最受欢迎、可靠和可扩展的大数据解决方案的分布式计算和存储系统。它由多个专为在多台服务器和数千台机器之间进行分布式任务而设计的组件组成。本课程提供全面的3合1培训,借助真实世界的实例为您打下坚实基础。您将学习如何设置HDFS集群,以及如何在本地存储和Hadoop文件系统之间进行格式化和数据传输。此外,课程中还将提供10个使用Hadoop的真实案例的实践解决方案。 课程内容概述: 本培训项目包含三个完整的课程,旨在提供尽可能全面的培训。第一个课程《初识Hadoop 2.x》介绍Hadoop的基本概念,包括节点、数据集和操作(如map和reduce)。第二个模块专注于HDFS,Hadoop用于存储数据的文件系统。您将学习任务与作业之间的区别,并熟悉Hadoop用户界面。随后,我们将学习如何在HDFS中存储数据和进行数据转换,并最终通过Hadoop的map-reduce方式实现算法并分析整体性能。 第二个课程《Hadoop管理与集群管理》从Apache Hadoop的集群安装和所需服务的配置开始。您将学习各种集群操作,如验证、扩展和缩减Hadoop服务。该课程还涵盖了计划集群、监控、日志记录、安全性、故障排除和最佳实践等管理任务,确保您的Hadoop集群高度可用和可靠。 第三个课程《解决10个Hadoop问题》介绍Hadoop生态系统的核心部分,帮助您快速上手。接下来,课程将描述多个共同问题,作为Hadoop能够解决的案例研究项目。每个项目都被划分为特定的用例,用于解决大数据问题。在学习路径结束时,您将能够计划、部署、管理和监控您的Hadoop集群,并对性能进行调整。 关于作者: A K M Zahiduzzaman是一名在NewsCred Dhaka工作的软件工程师,曾专注于Ruby on Rails开发,现在致力于NodeJS、AngularJS和Python。他对Spark有丰富的经验,热衷于分享知识。Gurmukh Singh是一名拥有14年以上行业经验的技术专业人士,在基础设施设计、分布式系统、性能优化和网络领域有丰富的背景,并在大数据领域工作了5年。Tomasz Lelek是InitLearn的联合创始人,主要从事Java和Scala编程,对软件开发充满热情,也曾在多个会议上担任演讲者。 通过本课程,您将全面掌握Hadoop的各个方面,为大数据解决方案做好准备。
Hadoop is the most popular, reliable and scalable distributed computing and storage for Big Data solutions. It comprises of components designed to enable tasks on a distributed scale, across multiple servers and thousands of machines. This comprehensive 3-in-1 training course gives you a strong foundation by exploring Hadoop ecosystem with real-world examples. You'll discover the process to set up an HDFS cluster along with formatting and data transfer in between your local storage and the Hadoop filesystem. Also get a hands-on solution to 10 real-world use-cases using Hadoop. Contents and Overview This training program includes 3 complete courses, carefully chosen to give you the most comprehensive training possible. The first course, Getting Started with Hadoop 2.x, opens with an introduction to the world of Hadoop, where you will learn Nodes, Data Sets, and operations such as map and reduce. The second section deals HDFS, Hadoop's file-system used to store data. Further on, you'll discover the differences between jobs and tasks, and get to know about the Hadoop UI. After this, we turn our attention to storing data in HDFS and Data Transformations. Lastly, we will learn how to implement an algorithm in Hadoop map-reduce way and analyze the overall performance. The second course, Hadoop Administration and Cluster Management, starts by installing the Apache Hadoop for cluster installation and configuring the required services. Learn various cluster operations like validations, and expanding and shrinking Hadoop services. You will then move onto gain a better understanding of administrative tasks like planning your cluster, monitoring, logging, security, troubleshooting and best practices. Techniques to keep your Hadoop clusters highly available and reliant are also covered in this course. The third course, Solving 10 Hadoop'able Problems, covers the core parts of the Hadoop ecosystem, helping to give a broad understanding and get you up-and-running fast. Next, it describes a number of common problems as case-study projects Hadoop is able to solve. These sections are broken down into sections by different projects, each serving as a specific use case for solving big data problems. By the end of this Learning Path, you'll be able to plan, deploy, manage and monitor and performance-tune your Hadoop Cluster with Apache Hadoop. About the Author A K M Zahiduzzaman is a software engineer with NewsCred Dhaka. He is a software developer and technology enthusiast. He was a Ruby on Rails developer, but now working on NodeJS and angularJS and python. He is also working with a much wider vision as a technology company. The next goal is introducing SOA within the current applications to scale development via microservices. Zahiduzzaman has a lot of experience with Spark and is passionate about it. He is also a guitarist and has a band too. He was also a speaker for an international event in Dhaka. He is very enthusiastic and love to share his knowledge. Gurmukh Singh is a technology professional with 14+ years of industry experience in infrastructure design, distributed systems, performance optimization, and networks. He has worked in big data domain for the last 5 years and provides consultancy and training on various technologies. He has worked with companies such as HP, JP Morgan, and Yahoo and has authored the book Monitoring Hadoop. Tomasz Lelek is a Software Engineer and Co-Founder of InitLearn. He mostly does programming in Java and Scala. He dedicates his time and efforts to get better at everything. He is currently delving into big data technologies. Tomasz is very passionate about everything associated with software development. He has been a speaker at a few conferences in Poland-Confitura and JDD, and at the Krakow Scala User Group. He has also conducted a live coding session at Geecon Conference. He was also a speaker at an international event in Dhaka. He is very enthusiastic and loves to share his knowledge.