Databricks Machine Learning Associate exam dumps

Databricks Machine Learning Associate practice question 178 of 656

Databricks Certified Machine Learning Associate. Associate level, Databricks. Free question with the correct answer and a full explanation.

Databricks Machine Learning Associate Question 178

Select 2

In which of the following scenarios is replacing missing values with the mode an appropriate approach?

  1. A

    When the feature is categorical and exhibits a strong class imbalance

  2. B

    When the feature is numerical and normally distributed

  3. C

    When the feature is categorical and missing values are frequent

  4. D

    When the feature is numerical and contains outliers

  5. E

    When the missing values are randomly distributed across the dataset

Show answer and explanation

Correct answers: A, C

Explanation

Replacing missing values with the mode is most appropriate for categorical features because the mode represents the most frequent category and aligns with the data's distribution. It is especially useful when there is a strong class imbalance or frequent missing values, ensuring the imputation does not distort the feature's distribution. However, it is not suitable for numerical features or scenarios involving outliers, as the mode does not accurately capture numerical central tendencies.

  • A. Correct.

    Replacing missing values with the mode is appropriate for categorical features with a strong class imbalance because the mode represents the most frequent category, which aligns with the data's overall distribution.

  • B. Incorrect.

    Replacing missing values with the mode is generally not suitable for numerical features, as the mode may not effectively represent the central tendency of numerical data, especially if it is normally distributed.

  • C. Correct.

    For categorical features with frequent missing values, replacing them with the mode can be effective because it ensures the most common category is preserved and maintains the majority distribution.

  • D. Incorrect.

    Replacing missing values with the mode is not recommended for numerical features with outliers, as the presence of outliers does not align with the concept of choosing the most frequent value.

  • E. Incorrect.

    While random distribution of missing values across the dataset is a favorable scenario for imputation, using the mode is not necessarily the best approach unless the feature is categorical.

Timed practice exam

Take a Databricks Machine Learning Associate practice test under exam conditions

48 questions in 90 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam