NCA-GENM exam dumps

NCA-GENM practice question 14 of 228

NVIDIA-Certified Associate - Generative AI Multimodal. Associate level, NVIDIA. Free question with the correct answer and a full explanation.

NCA-GENM Question 14

Single answer

You are designing a multimodal AI model that processes both image and text data. During training, you notice that the model struggles to align features between the two modalities, leading to suboptimal performance. Which type of loss function can help improve the alignment between image and text features?

  1. A

    Cross-entropy loss

  2. B

    Contrastive loss

  3. C

    Reconstruction loss

  4. D

    Mean squared error (MSE)

Show answer and explanation

Correct answer: B

Explanation

Contrastive loss is effective in multimodal learning scenarios because it minimizes the distance between aligned features (e.g., image-text pairs) while maximizing the distance between unaligned features. This makes it ideal for tasks where feature alignment across different modalities, such as images and text, is critical.

  • A. Incorrect.

    Cross-entropy loss is commonly used for classification tasks but does not directly address the alignment of features across modalities like image and text.

  • B. Correct.

    Contrastive loss is specifically designed to encourage similar features from different modalities to be closer in the feature space, making it suitable for improving alignment between image and text features.

  • C. Incorrect.

    Reconstruction loss focuses on reconstructing input data from outputs (e.g., in autoencoders), but it does not explicitly align features across modalities.

  • D. Incorrect.

    Mean squared error (MSE) measures the average squared difference between predicted and actual values. While it is useful in regression tasks, it does not enforce alignment between features from different modalities.

Timed practice exam

Take a NCA-GENM practice test under exam conditions

50 questions in 60 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam