Databricks Generative AI Engineer Associate exam dumps

Databricks Generative AI Engineer Associate practice question 247 of 306

Databricks Certified Generative AI Engineer Associate. Free level, Databricks. Free question with the correct answer and a full explanation.

Databricks Generative AI Engineer Associate Question 247

Select 2

You are training a generative AI model on a large dataset containing sensitive customer information such as credit card numbers and social security numbers. To ensure the model meets both privacy and performance objectives, which of the following strategies would be most appropriate for using masking techniques?

  1. A

    Replace sensitive fields with random noise before training the model.

  2. B

    Mask sensitive fields with consistent placeholders, such as 'XXXX', to preserve data structure.

  3. C

    Remove all rows containing any sensitive information from the dataset to prevent bias.

  4. D

    Use tokenization to replace sensitive data with unique, reversible identifiers.

  5. E

    Mask sensitive fields with domain-specific patterns to maintain model understanding of data context.

Show answer and explanation

Correct answers: B, E

Explanation

Masking techniques are essential for protecting sensitive data while ensuring that the model maintains its performance. Using consistent placeholders or domain-specific patterns preserves the structure and context of the data, which is critical for generative AI tasks. These strategies strike a balance between privacy and model performance, avoiding the pitfalls of noisy replacements or outright data removal.

  • A. Incorrect.

    Replacing sensitive fields with random noise can result in a loss of data structure, leading to reduced model performance and generalization ability.

  • B. Correct.

    Masking sensitive fields with consistent placeholders like 'XXXX' helps preserve the structure of the data while protecting sensitive information. This allows the model to focus on the task-relevant patterns without being exposed to sensitive data.

  • C. Incorrect.

    Removing all rows containing sensitive information may significantly reduce the size of the dataset, introducing bias and negatively impacting model performance.

  • D. Incorrect.

    Using tokenization to replace sensitive data with unique, reversible identifiers does not ensure privacy because the sensitive data can potentially be reconstructed, which is not aligned with privacy requirements.

  • E. Correct.

    Masking sensitive fields with domain-specific patterns (e.g., replacing a credit card number with '####-####-####-####') maintains the contextual understanding of the data, allowing the model to learn effectively while protecting sensitive information.

Timed practice exam

Take a Databricks Generative AI Engineer Associate practice test under exam conditions

45 questions in 90 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam