Managing Big Data on Google's Cloud Platform

所在平台: Udemy

课程主页: https://www.udemy.com/course/managing-big-data-on-googles-cloud-platform/

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程简介:本课程名为《在谷歌云平台上管理大数据》,是系列课程中的第二门,旨在帮助您获得谷歌认证数据工程师资格。本课程将展示数据工程师在谷歌云平台上的角色,强调谷歌认证数据工程师是数据和机器学习工程师中唯一的现实世界认证。需要注意的是,这不是关于大数据的课程,而是专注于谷歌云Dataproc这一特定云服务的课程。 课程内容主要围绕将本地Hadoop作业迁移到谷歌云平台,探讨结构化数据与非结构化数据的区别。现实中,约90%的企业数据为非结构化数据,课程的目标是帮助您为这些数据构建结构,以便进行分析。同时,课程将重点介绍使用Cloud Dataproc这一谷歌管理的Hadoop和Spark平台进行大规模数据分析与机器学习项目。 选择本课程的五大理由: 1) 数据工程师是全球最需的职业。 2) 谷歌在数据领域的领导地位无可争议,尤其是在开源人工智能方面。 3) 90%的组织数据是非结构化的,随着数据的增长,企业需要云平台进行扩展。 4) 数据革命正在进行,数据被视为真相的来源,数据工程师在数据管理中扮演着桥梁角色。 5) 数据工程师为数据的流动打基础,确保数据的清晰和可用,数据科学家则通过分析与可视化来赋予数据意义。 感谢您对《在谷歌云平台上管理大数据》的关注,我们期待在课程中见到您!

课程评论(0条)

课程详情

Welcome to Managing Big Data on Google's Cloud Platform. This is the second course in a series of courses designed to help you attain the coveted Google Certified Data Engineer. Additionally, the series of courses is going to show you the role of the data engineer on the Google Cloud Platform. At this juncture the Google Certified Data Engineer is the only real world certification for data and machine learning engineers. NOTE: This is NOT a course on Big Data. This is a course on a specific cloud service called Google Cloud Dataproc. The course was designed to be part of a series for those who want to become data engineers on Google's Cloud Platform. This course is all about Google's Cloud and migrating on-premise Hadoop jobs to GCP. In reality, Big Data is simply about unstructured data. There are two core types of data in the real world. The first is structured data, this is the kind of data found in a relational database. The second is unstructured, this is a file sitting on a file system. Approximately 90% of all data in the enterprise is unstructured and our job is to give it structure. Why do we want to give it structure? We want to give is structure so we can analyze it. Recall that 99% of all applied machine learning is supervised learning. That simply means we have a data set and we point our machine learning models at that data set in order to gain insight into that data. In the course we will spend much of the time working in Cloud Dataproc. This is Google's managed Hadoop and Spark platform. Recall the end goal of big data is to get that data into a state where it can be analyzed and modeled. Therefore, we are also going to cover how to work on machine learning projects with big data at scale. Please keep in mind this course alone will not give you the knowledge and skills to pass the exam. The course will provide you with the big data knowledge you need for working with Cloud Dataproc and for moving existing projects to the Google Cloud Platform. *Five Reasons to take this Course.* 1) The Top Job in the World The data engineer role is the single most needed role in the world. Many believe that it's the data scientist but several studies have broken down the job descriptions and the most needed position is that of the data engineer. 2) Google's the World Leader in Data Amazon's AWS is the most used cloud and Azure has the best UI but no cloud vendor in the world understands data like Google. They are the world leader in open sources artificial intelligence. You can't be the leader in AI without being the leader in data. 3) 90% of all Organizational Data is Unstructured The study of big data is the study of unstructured data. As the data in companies grows most will need to scale to unprecedented level. Without a significant investment in infrastructure and talent this won't be possible without the cloud. 4) The Data Revolution is Now We are in a data revolution. Data used to be viewed as a simple necessity and lower on the totem pole. Now it is more widely recognized as the source of truth. As we move into more complex systems of data management, the role of the data engineer becomes extremely important as a bridge between the DBA and the data consumer. Beyond the ubiquitous spreadsheet, graduating from RDBMS (which will always have a place in the data stack), we now work with NoSQL and Big Data technologies. 5) Data is Foundation Data engineers are the plumbers building a data pipeline, while data scientists are the painters and storytellers giving meaning to an otherwise static entity. Simply put, data engineers clean, prepare and optimize data for consumption. Once the data becomes useful, data scientists can perform a variety of analyses and visualization techniques to truly understand the data, and eventually, tell a story from the data. Thank you for your interest in Managing Big Data on Google's Cloud Platform and we will see you in the course!!

课程标签

0人关注该课程

主题相关的课程