Google Professional Machine Learning Engineer exam dumps

Google Professional Machine Learning Engineer practice question 94 of 522

Professional Machine Learning Engineer. Professional level, Google Cloud. Free question with the correct answer and a full explanation.

Google Professional Machine Learning Engineer Question 94

Select 2Google Cloud Platform

You are building a recommendation system for an e-commerce platform and need to process large-scale transactional data for training a machine learning model. The data is stored in a globally distributed, strongly consistent database. Additionally, you need to preprocess this data using distributed data processing tools. Which combination of Google Cloud tools should you use to meet these requirements?

  1. A

    Google Spanner for data storage and Apache Spark for distributed data processing

  2. B

    Cloud SQL for data storage and Apache Hadoop for distributed data processing

  3. C

    Google Spanner for data storage and Google Cloud Dataflow for distributed data processing

  4. D

    Cloud SQL for data storage and Apache Spark for distributed data processing

  5. E

    Google BigQuery for data storage and Apache Spark for distributed data processing

Show answer and explanation

Correct answers: A, C

Explanation

For globally distributed, strongly consistent transactional data, Google Spanner is the best choice among Google Cloud's storage solutions. For distributed data processing, both Apache Spark and Google Cloud Dataflow are suitable options. Apache Spark is widely used for its flexibility and ecosystem, while Cloud Dataflow offers a managed solution built for Google Cloud. Cloud SQL and BigQuery do not meet the requirements for globally distributed transactional data storage.

  • A. Correct.

    Correct: Google Spanner is designed for globally distributed, strongly consistent data storage, and Apache Spark is a distributed data processing framework suitable for large-scale data preprocessing.

  • B. Incorrect.

    Incorrect: Cloud SQL is not a globally distributed database and is not designed for handling large-scale transactional data efficiently. Apache Hadoop is a valid data processing tool but not the best fit compared to Apache Spark or Dataflow for this scenario.

  • C. Correct.

    Correct: Google Spanner provides globally distributed, strongly consistent storage, and Google Cloud Dataflow is a managed service suitable for distributed data processing and machine learning preprocessing tasks.

  • D. Incorrect.

    Incorrect: Cloud SQL is not suitable for globally distributed, large-scale transactional data. While Apache Spark is a valid processing tool, the storage solution is inadequate for this scenario.

  • E. Incorrect.

    Incorrect: Google BigQuery is optimized for analytical queries rather than transactional data storage. While Apache Spark is a valid data processing tool, BigQuery does not meet the requirements for this scenario.

Timed practice exam

Take a Google Professional Machine Learning Engineer practice test under exam conditions

60 questions in 120 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam