DEA-C01 exam dumps

DEA-C01 practice question 58 of 550

AWS Certified Data Engineer - Associate. Associate level, Amazon Web Services. Free question with the correct answer and a full explanation.

DEA-C01 Question 58

Select 2

A company is building a data pipeline to process a large volume of clickstream data that is generated in real-time from their website and mobile applications. The data includes structured fields like timestamps and user IDs, as well as unstructured data such as user comments. Which combination of AWS services is best suited to handle the volume, velocity, and variety of this data for ingestion and processing?

  1. A

    Amazon Kinesis Data Streams for real-time ingestion and Amazon S3 for scalable storage

  2. B

    Amazon RDS for storing structured data and Amazon Redshift for batch analytics

  3. C

    Amazon DynamoDB for high-throughput ingestion and AWS Lambda for processing

  4. D

    Amazon S3 for storing raw data and AWS Glue for data transformation

  5. E

    Amazon Kinesis Firehose for real-time streaming and Amazon Elasticsearch Service (Amazon OpenSearch Service) for analyzing unstructured data

Show answer and explanation

Correct answers: A, E

Explanation

The scenario involves high volume, high velocity, and a mix of structured and unstructured data. Amazon Kinesis Data Streams or Firehose are ideal for real-time ingestion, while Amazon S3 and Amazon Elasticsearch Service (Amazon OpenSearch Service) can handle scalable storage and analytics for structured and unstructured data. The correct combination of services ensures the pipeline can effectively process the data as it flows through the system.

  • A. Correct.

    Amazon Kinesis Data Streams is optimized for handling high-velocity, real-time data ingestion, and Amazon S3 provides cost-effective and scalable storage for both structured and unstructured data. This combination addresses the volume, velocity, and variety of the data.

  • B. Incorrect.

    Amazon RDS and Amazon Redshift are not suitable for real-time ingestion or handling unstructured data. They work well for structured data and batch analytics but fail to address the velocity and variety requirements in this scenario.

  • C. Incorrect.

    Amazon DynamoDB is highly performant for specific structured data use cases, but it is not designed for real-time ingestion of large-scale streaming data. AWS Lambda is event-driven and not optimal for continuous high-throughput data processing.

  • D. Incorrect.

    While Amazon S3 is suitable for raw data storage, AWS Glue is used for ETL (Extract, Transform, Load) processes, which are typically batch-oriented and may not handle the velocity of real-time streaming data effectively.

  • E. Correct.

    Amazon Kinesis Firehose can handle real-time streaming data and is capable of delivering it to destinations like Amazon Elasticsearch Service (Amazon OpenSearch Service), which is optimized for analyzing unstructured data. This makes it a good fit for the scenario.

Timed practice exam

Take a DEA-C01 practice test under exam conditions

65 questions in 130 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam