DEA-C01 exam dumps

DEA-C01 practice question 46 of 550

AWS Certified Data Engineer - Associate. Associate level, Amazon Web Services. Free question with the correct answer and a full explanation.

DEA-C01 Question 46

Select 2

You are designing a data pipeline for processing large volumes of real-time clickstream data. The data needs to be ingested, transformed, and stored for analytics. Given the requirements below, which combination of AWS services would best meet your needs?

  • The pipeline should handle scalable real-time ingestion.
  • Minimal latency is required for transformation.
  • Data must be stored in a format optimized for analytical queries.

Which services would you choose?

  1. A

    Amazon Kinesis Data Streams for ingestion, AWS Lambda for transformation, and Amazon S3 with Parquet format for storage

  2. B

    Amazon SQS for ingestion, AWS Glue for transformation, and Amazon RDS for storage

  3. C

    Amazon Kinesis Data Firehose for ingestion and transformation, and Amazon Redshift for storage

  4. D

    Amazon MQ for ingestion, Amazon EMR for transformation, and Amazon DynamoDB for storage

  5. E

    Amazon MSK (Managed Streaming for Apache Kafka) for ingestion, AWS Lambda for transformation, and Amazon Aurora for storage

Show answer and explanation

Correct answers: A, C

Explanation

The combination of Amazon Kinesis Data Streams (or Kinesis Data Firehose) for ingestion, AWS Lambda (or Firehose's native transformation capabilities) for transformation, and Amazon S3 (using Parquet) or Amazon Redshift for storage aligns perfectly with the requirements of scalable ingestion, minimal latency transformation, and storage optimized for analytics. SQS, MQ, and transactional databases like Aurora or DynamoDB are not suitable for this scenario because they are not designed for high-throughput real-time data ingestion or analytical querying.

  • A. Correct.

    This is a valid option. Amazon Kinesis Data Streams is well-suited for real-time ingestion, AWS Lambda can handle low-latency transformations, and storing the data in Amazon S3 using the Parquet format optimizes it for analytical queries.

  • B. Incorrect.

    This is not an ideal choice. Amazon SQS is not designed for high-throughput real-time streaming, AWS Glue may introduce latency during transformations, and Amazon RDS is not optimized for large-scale analytical querying.

  • C. Correct.

    This is a valid option. Amazon Kinesis Data Firehose can handle both real-time ingestion and transformations with minimal latency, and Amazon Redshift is highly optimized for storage and analytics.

  • D. Incorrect.

    This is not an appropriate solution. Amazon MQ is better suited for messaging, not high-throughput ingestion, and while Amazon EMR is powerful for transformations, Amazon DynamoDB is not designed for analytical workloads.

  • E. Incorrect.

    This option is partially correct but not optimal. Amazon MSK is a good choice for ingestion, but Amazon Aurora is primarily a transactional database and not optimized for large-scale analytics.

Timed practice exam

Take a DEA-C01 practice test under exam conditions

65 questions in 130 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam