Databricks Generative AI Engineer Associate exam dumps

Databricks Generative AI Engineer Associate practice question 249 of 306

Databricks Certified Generative AI Engineer Associate. Free level, Databricks. Free question with the correct answer and a full explanation.

Databricks Generative AI Engineer Associate Question 249

Select 2

You are tasked with fine-tuning a large language model (LLM) on a customer feedback dataset to generate summaries of customer reviews. To ensure sensitive customer data (e.g., personal identifiers) is not leaked and the model focuses on general insights, you decide to use masking techniques. Which of the following approaches would help you meet your objective while maintaining model performance?

  1. A

    Replace sensitive customer identifiers (e.g., names and email addresses) with generic placeholders like '' or '' before fine-tuning.

  2. B

    Mask all numerical values in the dataset, including ratings and quantities, with '' to ensure no numerical data is included.

  3. C

    Use token masking to replace sensitive data but ensure that the masked tokens maintain semantic relevance to the task.

  4. D

    Remove all sensitive data entirely from the dataset without using any replacement tokens.

  5. E

    Perform masking only during inference and retain sensitive data during fine-tuning for better performance.

Show answer and explanation

Correct answers: A, C

Explanation

Using masking techniques such as replacing sensitive data with generic placeholders or semantically relevant tokens ensures that the training data remains anonymized while preserving important context for the fine-tuning task. These approaches strike a balance between privacy and performance, making them suitable for scenarios involving sensitive information.

  • A. Correct.

    Replacing sensitive customer identifiers with generic placeholders helps anonymize the data while preserving the context necessary for the model to learn patterns effectively.

  • B. Incorrect.

    Masking all numerical values indiscriminately, including relevant information like ratings or quantities, could degrade the model's performance as these numbers may be essential for understanding the task.

  • C. Correct.

    Using token masking that maintains semantic relevance ensures that the model can still understand the context of the task while protecting sensitive data.

  • D. Incorrect.

    Completely removing sensitive data without any replacement can result in a loss of valuable context, which may negatively impact model performance.

  • E. Incorrect.

    Masking only during inference does not protect sensitive data during fine-tuning, which could lead to data leakage and compromise compliance or security requirements.

Timed practice exam

Take a Databricks Generative AI Engineer Associate practice test under exam conditions

45 questions in 90 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam