Google Professional Machine Learning Engineer exam dumps

Google Professional Machine Learning Engineer practice question 95 of 522

Professional Machine Learning Engineer. Professional level, Google Cloud. Free question with the correct answer and a full explanation.

Google Professional Machine Learning Engineer Question 95

Select 2Google Cloud Platform

You are tasked with building a machine learning pipeline that processes terabytes of semi-structured data daily. The pipeline must provide high availability, ensure scalability for future data growth, and integrate with an existing Apache Spark ecosystem. Additionally, the processed data must be stored in a database that supports SQL queries and global consistency. Which combination of services should you use?

  1. A

    Google Cloud Spanner for storage and Apache Spark for data processing

  2. B

    Google Cloud SQL for storage and Apache Hadoop for data processing

  3. C

    Google Cloud Spanner for storage and Apache Hadoop for data processing

  4. D

    Google Cloud SQL for storage and Apache Spark for data processing

  5. E

    Google Cloud BigQuery for storage and Apache Spark for data processing

Show answer and explanation

Correct answers: A, E

Explanation

The correct combination of services must address both the storage and processing requirements. Google Cloud Spanner provides high availability, global consistency, and SQL query support, making it suitable for storing the processed data. Apache Spark is an excellent choice for large-scale data processing and integrates well with Spanner. Alternatively, Google Cloud BigQuery is another excellent storage option for large data sets due to its scalability and support for SQL queries. Combining it with Apache Spark satisfies the processing needs, making both combinations (Spanner + Spark and BigQuery + Spark) valid solutions.

  • A. Correct.

    Google Cloud Spanner offers high availability, global consistency, and SQL query support. Apache Spark is well-suited for large-scale data processing and integrates with Spanner effectively, making this a suitable choice.

  • B. Incorrect.

    Google Cloud SQL is a fully managed relational database, but it is not optimized for handling terabytes of data daily or ensuring global scalability. Apache Hadoop is also less efficient compared to Apache Spark for such use cases.

  • C. Incorrect.

    While Cloud Spanner is suitable for storage due to its scalability and consistency, Apache Hadoop is not the best choice for large-scale, high-speed data processing compared to Apache Spark.

  • D. Incorrect.

    Google Cloud SQL does not provide the scalability required for processing terabytes of data daily. While Apache Spark is efficient for data processing, the storage component here is not suitable.

  • E. Correct.

    Google Cloud BigQuery is highly scalable and supports SQL queries. Combined with Apache Spark for data processing, this setup is ideal for processing large-scale data efficiently.

Timed practice exam

Take a Google Professional Machine Learning Engineer practice test under exam conditions

60 questions in 120 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam