MLS-C01 exam dumps

MLS-C01 practice question 8 of 389

AWS Certified Machine Learning - Specialty. Expert level, Amazon Web Services. Free question with the correct answer and a full explanation.

MLS-C01 Question 8

Select 2

You are a data engineer tasked with preparing a centralized data repository for a machine learning team. The data consists of structured transaction records, semi-structured user behavior logs, and unstructured images. The team plans to use this data for training models and performing exploratory data analysis. Which combination of AWS services would best meet the requirements of storing and retrieving this diverse dataset effectively?

  1. A

    Amazon S3 for storing all types of data and Amazon Athena for querying structured and semi-structured data

  2. B

    Amazon RDS for storing structured data, Amazon OpenSearch Service for semi-structured data, and Amazon S3 for unstructured images

  3. C

    Amazon DynamoDB for storing all types of data and Amazon SageMaker for querying the data

  4. D

    Amazon Redshift for storing structured and semi-structured data, and Amazon S3 for unstructured images

  5. E

    Amazon S3 for storing all data types and AWS Glue for cataloging and transforming data

Show answer and explanation

Correct answers: A, E

Explanation

To create a centralized data repository for machine learning, you need a solution that can handle structured, semi-structured, and unstructured data effectively. Amazon S3 is highly scalable and versatile for storing diverse data types, while Amazon Athena provides an SQL-like querying capability for structured and semi-structured data stored in S3. Additionally, AWS Glue can catalog and transform data, making it easier to prepare and analyze for machine learning workflows. Together, these services provide a robust and flexible solution for diverse ML data requirements.

  • A. Correct.

    Correct. Amazon S3 is a versatile storage service that can handle structured, semi-structured, and unstructured data. Amazon Athena can query data stored in S3 directly, making it suitable for exploring structured and semi-structured data.

  • B. Incorrect.

    Partially correct but not the best option. While Amazon RDS is suitable for structured data and OpenSearch Service can handle semi-structured data, this approach is less optimal because it doesn’t leverage a unified repository like S3 for diverse data types, and querying across these services can become complex.

  • C. Incorrect.

    Incorrect. Amazon DynamoDB is a NoSQL database optimized for key-value and document use cases, but it is not well-suited for handling unstructured data or querying diverse datasets. SageMaker is not designed for querying data directly.

  • D. Incorrect.

    Partially correct but not ideal. Amazon Redshift is optimized for analytics on structured and semi-structured data, but it is not designed for unstructured data like images. S3 is appropriate for unstructured data, but the combination lacks flexibility for diverse data types.

  • E. Correct.

    Correct. Amazon S3 can store structured, semi-structured, and unstructured data effectively. AWS Glue can catalog the data and perform ETL (Extract, Transform, Load) operations, making it easier to prepare and query the data for machine learning workflows.

Timed practice exam

Take a MLS-C01 practice test under exam conditions

65 questions in 180 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam