AIF-C01 exam dumps

AIF-C01 practice question 154 of 231

AWS Certified AI Practitioner. Associate level, Amazon Web Services. Free question with the correct answer and a full explanation.

AIF-C01 Question 154

Select 3

A data scientist is building a machine translation application using a foundation model. To evaluate the performance of the model's translations, they need to compare the generated text against reference translations. Which of the following metrics would be most relevant to assess the model's performance?

  1. A

    BLEU (Bilingual Evaluation Understudy)

  2. B

    ROUGE (Recall-Oriented Understudy for Gisting Evaluation)

  3. C

    F1 Score

  4. D

    BERTScore

  5. E

    Mean Squared Error (MSE)

Show answer and explanation

Correct answers: A, B, D

Explanation

To assess the performance of a foundation model on a text generation task like machine translation, metrics such as BLEU, ROUGE, and BERTScore are relevant. BLEU and ROUGE focus on n-gram and overlap-based comparisons, while BERTScore leverages semantic similarity. F1 Score and Mean Squared Error are not applicable as they are used for classification and regression tasks, respectively.

  • A. Correct.

    BLEU is widely used for evaluating machine translation models by comparing generated text to reference text using n-gram overlap.

  • B. Correct.

    ROUGE is commonly used to evaluate text generation tasks, such as summarization, by measuring overlap between the generated and reference text (e.g., recall).

  • C. Incorrect.

    F1 Score is generally used for classification problems and is not designed for evaluating text generation or translation tasks.

  • D. Correct.

    BERTScore evaluates semantic similarity by using contextual embeddings, making it suitable for assessing text generation or translation tasks.

  • E. Incorrect.

    Mean Squared Error (MSE) is a regression evaluation metric and is not relevant for text generation or translation tasks.

Timed practice exam

Take a AIF-C01 practice test under exam conditions

65 questions in 90 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam