SnowPro Specialty: Gen AI exam dumps

SnowPro Specialty: Gen AI practice question 139 of 287

SnowPro® Specialty: Gen AI. Expert level, Snowflake. Free question with the correct answer and a full explanation.

SnowPro Specialty: Gen AI Question 139

Single answerConsiderations (e.g. capability, latency, and cost)

A retail company is building a customer support assistant in Snowflake that summarizes order history and answers policy questions. The team is evaluating Cortex AI COMPLETE models for production. Business requirements are: (1) response time should stay low for a chat experience, (2) monthly inference cost must be controlled, and (3) answer quality only needs to be strong enough for routine support tasks, not complex reasoning. Which approach is the MOST appropriate?

  1. A

    Use the largest available model for all requests because larger models consistently provide the best balance of latency and cost in production.

  2. B

    Select a smaller, lower-cost model first, benchmark it on the support use case, and only move to a larger model if quality is insufficient.

  3. C

    Use multiple large models in parallel for every prompt so the application can choose the best answer, which reduces both latency and cost.

  4. D

    Prioritize the model with the longest context window, because context window size is the primary driver of low latency for chat workloads.

Show answer and explanation

Correct answer: B

Explanation

For Snowflake Cortex AI use cases, model selection should be guided by workload-specific trade-offs among capability, latency, and cost. In a customer support scenario with routine questions and summarization, the best approach is usually to evaluate whether a smaller model is sufficient before moving to a larger one. This minimizes inference cost and often improves responsiveness while still meeting business needs. A larger or more expensive model should be justified by measurable gains in answer quality for the actual task. This aligns with general Snowflake Cortex guidance to choose models based on use-case requirements and to test with realistic prompts, expected output quality, and operational constraints rather than assuming the most capable model is the best fit.

  • A. Incorrect.

    Incorrect. Larger models can improve capability, but they typically come with higher latency and higher cost. For routine support tasks, starting with the largest model is usually not the best trade-off. This option reflects a common misconception that maximum model capability is automatically optimal for production workloads.

  • B. Correct.

    Correct. When balancing capability, latency, and cost, a practical best practice is to start with the smallest model that can meet quality requirements, validate it with representative prompts and evaluation criteria, and scale up only if needed. This aligns with production decision-making for GenAI workloads, where model selection should be driven by workload requirements rather than defaulting to the most capable model.

  • C. Incorrect.

    Incorrect. Running multiple large models in parallel generally increases cost and can also increase system complexity. It does not inherently reduce latency, because the application may still wait for multiple responses or incur orchestration overhead. This is an example of overengineering when the stated use case only requires routine support quality.

  • D. Incorrect.

    Incorrect. A larger context window is useful when prompts require more input data, but it is not the primary determinant of low latency. In many cases, larger-context models can be more expensive and may not improve performance for short, routine support prompts. This option confuses a specific capability consideration with overall production efficiency.

Timed practice exam

Take a SnowPro Specialty: Gen AI practice test under exam conditions

55 questions in 85 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam