Google Professional Data Engineer exam dumps

Google Professional Data Engineer practice question 90 of 279

Professional Data Engineer. Professional level, Google Cloud. Free question with the correct answer and a full explanation.

Google Professional Data Engineer Question 90

Select 2Google Cloud Platform

You are designing a data processing pipeline for a retail company. The pipeline must handle both real-time streaming data from their online store and batch data from their offline point-of-sale systems. The real-time data requires low-latency processing, while the batch data involves large-scale ETL transformations. Which combination of services should you use to meet these requirements?

  1. A

    Cloud Dataflow for real-time streaming data and Cloud Data Fusion for batch ETL transformations

  2. B

    Pub/Sub for real-time streaming data ingestion and BigQuery for batch ETL transformations

  3. C

    Apache Kafka for real-time streaming data and Dataproc for batch ETL transformations

  4. D

    Cloud Dataflow for both real-time streaming data and batch ETL transformations

  5. E

    Apache Spark on Dataproc for both real-time streaming data and batch ETL transformations

Show answer and explanation

Correct answers: A, D

Explanation

To handle real-time streaming data with low latency and perform large-scale batch ETL transformations, a combination of Cloud Dataflow and Cloud Data Fusion is an excellent choice. Cloud Dataflow provides a unified, serverless solution for both streaming and batch processing, while Cloud Data Fusion simplifies complex batch ETL workflows. Alternatively, Cloud Dataflow alone can be used for both streaming and batch processing due to its flexibility and support for Apache Beam's unified model.

  • A. Correct.

    Correct: Cloud Dataflow is a managed service ideal for both real-time stream processing and batch ETL jobs. Cloud Data Fusion is a fully managed ETL service that simplifies batch data transformations with its visual interface.

  • B. Incorrect.

    Incorrect: While Pub/Sub is excellent for real-time streaming ingestion, BigQuery is primarily used for data warehousing and analytics, not for ETL transformations.

  • C. Incorrect.

    Incorrect: Apache Kafka is suitable for streaming data ingestion, but Dataproc, while good for batch ETL, does not suit low-latency real-time processing as well as Dataflow.

  • D. Correct.

    Correct: Cloud Dataflow supports unified processing for both real-time streaming and batch data using Apache Beam, making it a versatile choice.

  • E. Incorrect.

    Incorrect: While Apache Spark on Dataproc is powerful for both streaming and batch workloads, it requires more effort for setup and maintenance compared to Cloud Dataflow, which is serverless and fully managed.

Timed practice exam

Take a Google Professional Data Engineer practice test under exam conditions

60 questions in 120 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam