MLS-C01 exam dumps

MLS-C01 practice question 81 of 389

AWS Certified Machine Learning - Specialty. Expert level, Amazon Web Services. Free question with the correct answer and a full explanation.

MLS-C01 Question 81

Select 3

A data science team at a retail company is tasked with building a machine learning model to predict customer churn. They have collected a dataset with 10,000 samples, each containing 20 features. However, only 500 samples are labeled with the churn status (churned or not churned). After an initial review, the team is concerned about the sufficiency of labeled data for training the model. How should the team determine if the labeled data is sufficient for their problem?

  1. A

    Perform a baseline model evaluation using the labeled data to assess model performance and generalization.

  2. B

    Use techniques such as k-fold cross-validation to estimate how well the model performs with the available labeled data.

  3. C

    Assess if the labeled data has sufficient representation across all classes (churned and not churned).

  4. D

    Assume the labeled data is sufficient if the dataset contains more than 10,000 samples in total, regardless of labels.

  5. E

    Leverage data augmentation techniques to synthetically increase the size of the labeled dataset.

Show answer and explanation

Correct answers: A, B, C

Explanation

To determine if labeled data is sufficient, the team needs to evaluate the labeled data's ability to produce a model with acceptable performance and generalization. Techniques such as baseline evaluations, cross-validation, and assessing class representation can provide valuable insights. However, assumptions based on total dataset size or inappropriate augmentation strategies should be avoided.

  • A. Correct.

    Performing a baseline model evaluation can help assess if the labeled data is sufficient to produce a model with acceptable performance. If the model performs poorly, it may indicate insufficient labeled data.

  • B. Correct.

    Using techniques like k-fold cross-validation provides insights into how well the model generalizes with the given labeled data. Poor generalization could point to insufficient labeled data.

  • C. Correct.

    It's important to ensure that the labeled data has sufficient representation for each class. If one class is underrepresented, the model may fail to learn patterns for that class effectively.

  • D. Incorrect.

    The total dataset size is not relevant when assessing sufficiency of labeled data. What matters is the quantity and quality of labeled data available for the specific problem.

  • E. Incorrect.

    Data augmentation is typically used for image, text, or time-series data, but it does not apply to this scenario since customer churn data is likely tabular and cannot be augmented synthetically in a meaningful way.

Timed practice exam

Take a MLS-C01 practice test under exam conditions

65 questions in 180 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam