Projects in Hadoop and Big Data - Learn by Building Apps

所在平台: Udemy

课程主页: https://www.udemy.com/course/projects-in-hadoop-and-big-data-learn-by-building-apps/

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:基于Hadoop和大数据的项目 - 通过构建应用程序学习 课程概述:这是一门备受期待的大数据课程,涵盖Hadoop生态系统中的所有主要大数据技术,并通过实际项目将它们结合在一起。在学习过程中,您不仅会掌握Hadoop及其相关技术的细节,还能看到这些技术如何解决实际问题,并被全球公司使用。本课程将帮助您实现飞跃,构建能够解决现实世界问题的Hadoop解决方案。课程内容丰富,挑战性十足,适合希望在前沿技术领域深入学习的学生。 课程重点包括: 1. **为现有数据增值**:学习如何运用Mapreduce技术解决聚类问题,通过Mapreduce项目剔除大型数据集中的重复值。 2. **Hadoop分析与NoSQL**:使用Python解析Twitter流,利用Apache Pig提取关键词并映射到HDFS,然后推送至MongoDB,最后通过Node.js进行数据可视化。 3. **Kafka流与Yarn及Zookeeper**:建立Twitter流与Kafka流,使用Java代码为生产者和消费者编写代码,并借助Apache Samza进行打包和部署。 4. **实时流处理**:使用Apache Kafka和Apache Storm处理Twitter流,掌握两者的有效使用。 5. **医疗行业的大数据应用**:搭建美国退伍军人事务部的健康数据字典关系模式,展示技术框架,解决MySQL中的SQL连接查询问题,通过Scoop和HCatalog映射到Hadoop/Hive堆栈成功执行查询。 6. **日志收集与分析**:利用Apache Flume和HCatalog将实时日志流映射到HDFS。 7. **数据科学与Hadoop预测分析**:使用Mapreduce创建结构化数据,并在Python中执行机器学习逻辑回归。 8. **使用Apache Spark进行可视化分析**:将HDFS中的数据映射到Python,并进行数据可视化。 9. **客户360度视图与电子商务的大数据分析**:通过电子商务工具“Datameer”执行多种分析查询,支持情感分析与Twitter流。 10. **全面整合大数据与Amazon Elastic Map Reduce**:在AWS Mapreduce集群上运行聚类代码。 完成本课程后,您将能够自信地在Hadoop技术体系内构建几乎任何系统。课程提供完整的源代码和预配置的虚拟机,使您能够快速构建项目,而无需浪费过多时间在系统设置上,同时提供英语字幕。快来加入我们,开启大数据的旅程吧!

课程评论(0条)

课程详情

The most awaited Big Data course on the planet is here. The course covers all the major big data technologies within the Hadoop ecosystem and weave them together in real life projects. So while doing the course you not only learn the nuances of the hadoop and its associated technologies but see how they solve real world problems and how they are being used by companies worldwide. This course will help you take a quantum jump and will help you build Hadoop solutions that will solve real world problems. However we must warn you that this course is not for the faint hearted and will test your abilities and knowledge while help you build a cutting edge knowhow in the most happening technology space. The course focuses on the following topics Add Value to Existing Data - Learn how technologies such as Mapreduce applies to Clustering problems. The project focus on removing duplicate or equivalent values from a very large data set with Mapreduce. Hadoop Analytics and NoSQL - Parse a twitter stream with Python, extract keyword with apache pig and map to hdfs, pull from hdfs and push to mongodb with pig, visualise data with node js. Learn all this in this cool project. Kafka Streaming with Yarn and Zookeeper - Set up a twitter stream with Python, set up a Kafka stream with java code for producers and consumers, package and deploy java code with apache samza. Real-Time Stream Processing with Apache Kafka and Apache Storm - This project focus on twitter streaming but uses Kafka and apache storm and you will learn to use each of them effectively. Big Data Applications for the Healthcare Industry with Apache Sqoop and Apache Solr - Set up the relational schema for a Health Care Data dictionary used by the US Dept of Veterans Affairs, demonstrate underlying technology and conceptual framework. Demonstrate issues with certain join queries that fail on MySQL, map technology to a Hadoop/Hive stack with Scoop and HCatalog, show how this stack can perform the query successfully. Log collection and analytics with the Hadoop Distributed File System using Apache Flume and Apache HCatalog - Use Apache Flume and Apache HCatalog to map real time log stream to hdfs and tail this file as Flume event stream. , Map data from hdfs to Python with Pig, use Python modules for analytic queries Data Science with Hadoop Predictive Analytics - Create structured data with Mapreduce, Map data from hdfs to Python with Pig, run Python Machine Learning logistic regression, use Python modules for regression matrices and supervise training Visual Analytics with Apache Spark on Yarn - Create structured data with Mapreduce, Map data from hdfs to Python with Spark, convert Spark dataframes and RDD's to Python datastructures, Perform Python visualisations Customer 360 degree view, Big Data Analytics for e-commerce - Demonstrate use of EComerce tool ‘Datameer' to perform many fof the analytic queries from part 6,7 and 8. Perform queries in the context of Senitment analysis and Twiteer stream. Putting it all together Big Data with Amazon Elastic Map Reduce - Rub clustering code on AWS Mapreduce cluster. Using AWS Java sdk spin up a Dedicated task cluster with the same attributes. So after this course you can confidently built almost any system within the Hadoop family of technologies. This course comes with complete source code and fully operational Virtual machines which will help you build the projects quickly without wasting too much time on system setup. The course also comes with English captions. So buckle up and join us on our journey into the Big Data.

课程标签

0人关注该课程

主题相关的课程