NCA-GENM exam dumps

NCA-GENM practice question 87 of 228

NVIDIA-Certified Associate - Generative AI Multimodal. Associate level, NVIDIA. Free question with the correct answer and a full explanation.

NCA-GENM Question 87

Single answer

You are assisting in the development of a multimodal AI model that processes both text and image data. During testing, you notice that the model performs well on text inputs but struggles to generate accurate captions for image inputs. What is the most appropriate step to address this issue?

  1. A

    Increase the size of the text dataset to improve overall model generalization.

  2. B

    Use a pre-trained vision encoder to enhance the model's ability to process image inputs.

  3. C

    Focus on optimizing the loss function specifically for text data.

  4. D

    Apply data augmentation techniques on the image dataset to diversify training samples.

Show answer and explanation

Correct answer: B

Explanation

The problem lies in the model's ability to process image inputs effectively. By using a pre-trained vision encoder, you leverage a network that has already learned robust image features, which can improve the model's performance on image-based tasks such as caption generation. This is a more efficient and targeted approach compared to increasing text data or focusing solely on loss function optimization.

  • A. Incorrect.

    Increasing the size of the text dataset would not directly address the issue with image captioning, as the problem lies with image inputs rather than text.

  • B. Correct.

    Using a pre-trained vision encoder can significantly improve the model's ability to process and understand image data, which is the core of the issue in this scenario.

  • C. Incorrect.

    Optimizing the loss function for text data would further enhance text performance but would not resolve the specific challenge of generating accurate captions for images.

  • D. Incorrect.

    While data augmentation can help diversify the image dataset, it may not fully address the model's fundamental weakness in processing image inputs. A pre-trained vision encoder offers a more targeted solution.

Timed practice exam

Take a NCA-GENM practice test under exam conditions

50 questions in 60 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam