MLA-C01 exam dumps

MLA-C01 practice question 334 of 458

AWS Certified Machine Learning Engineer - Associate. Associate level, Amazon Web Services. Free question with the correct answer and a full explanation.

MLA-C01 Question 334

Select 4

A company is building a machine learning pipeline that processes large amounts of streaming data from IoT devices. The data must be ingested, transformed, and stored in near real-time for model training and analytics. They want to automate and orchestrate the entire process using AWS services. Which combination of services should the company use to achieve this?

  1. A

    AWS Kinesis Data Streams for data ingestion

  2. B

    AWS Glue for data transformation and ETL

  3. C

    AWS Step Functions for orchestration

  4. D

    Amazon Comprehend for data ingestion

  5. E

    Amazon S3 for data storage

Show answer and explanation

Correct answers: A, B, C, E

Explanation

To automate and integrate data ingestion with orchestration services in a real-time machine learning pipeline, a combination of services such as AWS Kinesis Data Streams, AWS Glue, AWS Step Functions, and Amazon S3 is ideal. Kinesis handles real-time data ingestion, Glue performs data transformation, Step Functions orchestrates the entire workflow, and S3 stores the data. Amazon Comprehend, while useful for NLP tasks, is not relevant in this scenario.

  • A. Correct.

    AWS Kinesis Data Streams is a fully managed service that allows for the real-time ingestion of streaming data, making it well-suited for IoT use cases.

  • B. Correct.

    AWS Glue is a fully managed ETL (Extract, Transform, Load) service that can be used to automate data transformation tasks, which are necessary for preparing the data for model training.

  • C. Correct.

    AWS Step Functions is an orchestration service that helps automate workflows by coordinating the execution of multiple AWS services, making it suitable for orchestrating the machine learning pipeline.

  • D. Incorrect.

    Amazon Comprehend is a natural language processing service, not designed for data ingestion or orchestration in this context.

  • E. Correct.

    Amazon S3 is a scalable storage service that can be used to store both raw and transformed data, making it a key component for storing data in this pipeline.

Timed practice exam

Take a MLA-C01 practice test under exam conditions

65 questions in 130 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam