NCA-GENM exam dumps

NCA-GENM practice question 94 of 228

NVIDIA-Certified Associate - Generative AI Multimodal. Associate level, NVIDIA. Free question with the correct answer and a full explanation.

NCA-GENM Question 94

Select 3

A company is deploying a multimodal AI model that integrates text and image inputs to detect fraudulent financial documents. To improve explainability, the team wants to identify which parts of the image and which textual phrases contributed most to the model's decision. Which of the following approaches would best leverage multimodal models to achieve this goal?

  1. A

    Use attention heatmaps on the image input to highlight regions of interest and overlay attention scores from the text input.

  2. B

    Train separate models for text and image inputs, and compare their outputs manually to identify the key contributors.

  3. C

    Leverage multimodal saliency maps to visualize how both text and image inputs influence the model's prediction.

  4. D

    Use a gradient-based explainability technique on the combined outputs of the multimodal model to trace influential inputs.

  5. E

    Ignore the need for explainability in the multimodal model and rely on final prediction accuracy to validate decisions.

Show answer and explanation

Correct answers: A, C, D

Explanation

Using multimodal models to improve explainability requires techniques that can analyze and visualize the contributions of both input modalities (text and image) to the model's decision. Attention heatmaps, multimodal saliency maps, and gradient-based methods all provide systematic ways to achieve this, fostering trust and transparency in AI-driven systems. Relying solely on prediction accuracy or separate models fails to leverage the full potential of multimodal explainability.

  • A. Correct.

    Attention heatmaps for images and overlaying attention scores for text provide interpretable visualizations for understanding how both modalities contribute to a decision, making this an effective approach.

  • B. Incorrect.

    Training separate models for text and image inputs and manually comparing their outputs is inefficient and does not fully utilize the benefits of a multimodal model.

  • C. Correct.

    Multimodal saliency maps are specifically designed to reveal the influence of both text and image inputs in a unified manner, making this a strong choice for improving explainability.

  • D. Correct.

    Gradient-based explainability techniques, such as integrated gradients, can trace the contributions of inputs to the model's combined outputs, improving transparency of the decision-making process.

  • E. Incorrect.

    Ignoring explainability in favor of accuracy undermines trust in AI systems, especially in critical applications like fraud detection.

Timed practice exam

Take a NCA-GENM practice test under exam conditions

50 questions in 60 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam