|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/build-real-world-big-data-projects/
课程评论:没有评论
课程名称:大数据项目 课程概述: 《大数据项目》课程旨在为学生提供处理和分析大规模数据所需的各种工具和技术的深入理解。课程内容涵盖数据预处理、数据可视化、统计分析以及用于数据分析的机器学习和深度学习技术。在整个课程中,学生将学习Hadoop生态系统,包括Hadoop分布式文件系统(HDFS)、MapReduce和Apache Spark等技术。学员们还将获得使用Apache Hive、Pig和Impala等大数据工具的实践经验。 课程结束时,学生将具备有效处理和分析大规模数据的必要技能和知识,同时对Hadoop生态系统及行业中常用的各种大数据工具有全面的了解。真实的数据工程项目通常涉及多个组件。建立一个符合最佳实践的数据工程项目可能会非常耗时。如果你是一名数据分析师、学生、科学家或工程师,寻找获得数据工程经验的机会,但又无法找到合适的入门项目,这门课程将非常适合你: 1. 想要参与一个模拟真实项目的数据工程项目。 2. 寻找一个端到端的数据工程项目。 3. 希望获得数据工程经验以备工作面试。 在本课程中,你将学习如何设置数据基础设施,如Airflow、Redshift、Snowflake等,掌握数据管道的最佳实践,识别数据管道中的故障点并构建具有抗故障能力的系统,从商业需求出发设计和构建数据管道,构建端到端的ETL管道,设置Apache Airflow、AWS EMR、AWS Redshift、AWS Spectrum和AWS S3。 技术栈: - 编程语言:Python - 包:PySpark - 服务:Docker、Kafka、Amazon Redshift、S3、IICS、DBT等 课程要求: 本课程假设学生具备AWS或其大数据服务的基础知识,了解Python和SQL虽有帮助,但并非强制要求。课程每月会新增项目,以保证学习内容的丰富性和前沿性。
The Big Data Projects course is designed to provide students with an in-depth understanding of the various tools and techniques used to handle and analyze large-scale data. The course will cover topics such as data preprocessing, data visualization, and statistical analysis, as well as machine learning and deep learning techniques for data analysis.Throughout the course, students will be introduced to the Hadoop ecosystem, including technologies such as Hadoop Distributed File System (HDFS), MapReduce, and Apache Spark. Students will also gain hands-on experience working with big data tools such as Apache Hive, Pig, and Impala.At the end of the course, students will have the necessary skills and knowledge to handle large-scale data and analyze it effectively. Students will also have a solid understanding of the Hadoop ecosystem and various big data tools that are commonly used in the industry.A real data engineering project usually involves multiple components. Setting up a data engineering project, while conforming to best practices can be extremely time-consuming. If you areA data analyst, student, scientist, or engineer looking to gain data engineering experience, but are unable to find a good starter project.1. Wanting to work on a data engineering project that simulates a real-life project.2. Looking for an end-to-end data engineering project.3. Looking for a good project to get data engineering experience for job interviews.Then this Course is for you. In this Course, you willLearn How to Set up data infrastructure such as Airflow, Redshift, Snowflake, etcLearn data pipeline best practices.Learn how to spot failure points in data pipelines and build systems resistant to failures.Learn how to design and build a data pipeline from business requirements.Learn How to Build End to End ETL PipelineSet up Apache Airflow, AWS EMR, AWS Redshift, AWS Spectrum, and AWS S3.Tech stack: ➔Language: Python➔Package: PySpark➔Services: Docker, Kafka, Amazon Redshift,S3, IICS, DBT Many MoreRequirementsThis course presume that students have prior knowledge of AWS or its Big Data services.Having a fair understanding of Python and SQL would help but it is not mandatory.Every Month New Projects will be added