Google GCP-DE Exam Overview:
| Certification Vendor: | Google Cloud |
| Exam Name: | Professional Data Engineer |
| Exam Number: | GCP-DE |
| Passing Score: | Not publicly disclosed |
| Exam Duration: | 120 minutes |
| Certificate Validity Period: | 2 years |
| Exam Format: | Multiple choice, Multiple select |
| Available Languages: | English, Japanese |
| Exam Price: | $200 USD |
| Real Exam Qty: | 50-60 |
| Related Certifications: | Google Cloud Certified Professional Data Engineer |
| Sample Questions: | Google GCP-DE Sample Questions |
| Exam Way: | Online proctored or onsite testing center |
| Pre Condition: | No formal prerequisites. Google recommends 3+ years of industry experience including 1+ years designing and managing solutions using Google Cloud. |
| Official Syllabus URL: | https://cloud.google.com/learn/certification/data-engineer |
Google GCP-DE Exam Syllabus Topics:
| Section | Weight | Objectives |
|---|---|---|
| Building and operationalizing data processing systems | 28%-33% | - Operationalizing systems
|
| Designing data processing systems | 22%-27% | - Designing data pipelines
|
| Ensuring solution quality | 20%-25% | - Data quality management
|
| Managing and optimizing solutions | 20%-25% | - Reliability and scalability
|
Google Data Engineer Sample Questions:
1. You are planning to use Google's Dataflow SDK to analyze customer data such as displayed below. Your project requirement is to extract only the customer name Passing Certification Exams Made Easy visit - https://www.2PassEasy.com from the data source and then write to an output PCollection.
Tom,555 X street Tim,553 Y street Sam, 111 Z street
Which operation is best suited for the above data processing requirement?
A) Data extraction
B) Source API
C) Sink API
D) ParDo
2. The YARN ResourceManager and the HDFS NameNode interfaces are available on a Cloud Dataproc cluster .
A) application node
B) master node
C) worker node
D) conditional node
3. What are two methods that can be used to denormalize tables in BigQuery?
A) 1) Use a partitioned table; 2) Join tables into one table
B) 1) Split table into multiple tables; 2) Use a partitioned table
C) 1) Use nested repeated fields; 2) Use a partitioned table
D) 1) Join tables into one table; 2) Use nested repeated fields
4. You want to migrate an on-premises Hadoop system to Cloud Dataproc. Hive is the primary tool in use, and the data format is Optimized Row Columnar (ORC). All ORC files have been successfully copied to a Cloud Storage bucket. You need to replicate some data to the cluster's local Hadoop Distributed File System (HDFS) to maximize performance. What are two ways to start using Hive in Cloud Dataproc? (Choose two.)
A) Mount the Hive tables from HDFS.
B) Run the gsutil utility to transfer all ORC files from the Cloud Storage bucket to the master node of the Dataproc cluste
C) Leverage Cloud Storage connector for Hadoop to mount the ORC files as external Hive table
D) Replicate external Hive tables to the native ones.
E) Leverage BigQuery connector for Hadoop to mount the BigQuery tables as external Hive table
F) Load the ORC files into BigQuer
G) Run the gsutil utility to transfer all ORC files from the Cloud Storage bucket to any node of the Dataproc cluste
H) Replicate external Hive tables to the native ones.
I) Then run the Hadoop utility to copy them do HDF
J) Run the gsutil utility to transfer all ORC files from the Cloud Storage bucket to HDF
K) Mount the Hive tables locally.
L) Mount the Hive tables locally.
5. Your company produces 20,000 files every hour. Each data file is formatted as a comma separated values (CSV) file that is less than 4 KB. All files must be ingested on Google Cloud Platform before they can be processed. Your company site has a 200 ms latency to Google Cloud, and your Internet connection bandwidth is limited as 50 Mbps. You currently deploy a secure FTP (SFTP) server on a virtual machine in Google Compute Engine as the data ingestion point. A local SFTP client runs on a dedicated machine to transmit the CSV files as is. The goal is to make reports with data from the previous day available to the executives by 10:00 a.m. each day. This design is barely able to keep up with the current volume, even though the bandwidth utilization is rather low.
You are told that due to seasonality, your company expects the number of files to double for the next three months. Which two actions should you take? (choose two.)
A) Create an S3-compatible storage endpoint in your network, and use Google Cloud Storage Transfer Service to transfer on-premices data to the designated storage bucket.
B) Assemble 1,000 files into a tape archive (TAR) fil
C) Redesign the data ingestion process to use gsutil tool to send the CSV files to a storage bucket in parallel.
D) Introduce data compression for each file to increase the rate file of file transfer.
E) Transmit the TAR files instead, and disassemble the CSV files in the cloud upon receiving them.
F) Contact your internet service provider (ISP) to increase your maximum bandwidth to at least 100 Mbps.
Solutions:
| Question # 1 Answer: D | Question # 2 Answer: B | Question # 3 Answer: D | Question # 4 Answer: G,K | Question # 5 Answer: C,E |

We're so confident of our products that we provide no hassle product exchange.


By Yves


