SnowPro Specialty: Gen AI exam dumps

SnowPro Specialty: Gen AI practice question 248 of 287

SnowPro® Specialty: Gen AI. Expert level, Snowflake. Free question with the correct answer and a full explanation.

SnowPro Specialty: Gen AI Question 248

Select 23.3 Monitor and optimize Snowflake Cortex costs.

A retail company has deployed a customer-support assistant built on Snowflake Cortex. Over the last month, usage of LLM-powered functions increased significantly, and finance has asked the data platform team to identify which Cortex workloads are driving spend and reduce unnecessary cost without degrading answer quality. The team wants an approach that provides visibility into consumption trends and enables practical optimization of prompt and model usage. Which TWO actions should the team take?

  1. A

    Query Snowflake account usage views related to Cortex/AI function consumption to identify which users, models, or workloads are generating the highest token-based usage, then use that information to target optimization efforts.

  2. B

    Reduce cost by moving Cortex LLM calls to a larger virtual warehouse, because warehouse size directly lowers the token charges billed for Cortex model inference.

  3. C

    Review prompts and application logic to reduce unnecessary input and output tokens, such as removing redundant context, tightening instructions, and limiting response length where appropriate.

  4. D

    Disable query tagging for Cortex workloads, because query tags add overhead and prevent accurate cost attribution in Snowflake monitoring.

  5. E

    Standardize every use case on the most capable frontier model available, because using one model for all tasks minimizes overall Cortex cost management complexity.

Show answer and explanation

Correct answers: A, C

Explanation

The best answers are 1 and 3 because effective Cortex cost optimization starts with measurement and then focuses on reducing token consumption where it does not harm business outcomes. In Snowflake, teams should use available account usage/monitoring capabilities for AI and query activity to understand who is consuming Cortex services and how usage trends change over time. Query history and query tagging are also important for attributing spend to applications and teams. After identifying high-cost workloads, prompt engineering and application-level controls are key optimization levers: minimize irrelevant context, avoid oversized prompts, cap response length when suitable, and choose models appropriate to the task rather than defaulting to the most expensive option. These approaches align with Snowflake best practices for observability, governance, and efficient use of Cortex AI functions.

  • A. Correct.

    Correct. A practical first step in managing Snowflake Cortex cost is to monitor actual usage through Snowflake's account usage and usage-monitoring surfaces for AI/Cortex-related consumption. This helps identify which workloads, users, applications, or model calls are generating the most token consumption and cost. Without this visibility, optimization is guesswork. In real environments, teams use these views along with query history and tags to attribute spending and spot trends.

  • B. Incorrect.

    Incorrect. Cortex model inference pricing is not reduced by increasing virtual warehouse size. Cortex LLM functions are billed based on the service's usage model, such as token-based consumption for many AI functions, rather than becoming cheaper because a larger warehouse is used. A warehouse may still be involved for SQL execution around the workload, but scaling the warehouse up is not the right lever for reducing Cortex inference charges.

  • C. Correct.

    Correct. Token usage is a major driver of LLM cost, so prompt optimization is one of the most effective cost controls. Removing duplicated context, avoiding overly verbose instructions, constraining output length, and only sending the minimum relevant grounding data can materially reduce spend while often preserving quality. This is a standard best practice for both cost and latency optimization in generative AI workloads.

  • D. Incorrect.

    Incorrect. Query tags are useful for governance, attribution, and cost analysis. They help teams separate Cortex usage by application, environment, or business unit when reviewing history and usage data. Disabling them would make monitoring harder, not easier. The misconception here is treating metadata used for observability as a significant cost driver; in practice, tagging improves accountability.

  • E. Incorrect.

    Incorrect. Using the most capable model for every task is usually not cost-efficient. Different tasks have different quality requirements, and many can be handled acceptably by smaller or less expensive models. A common optimization practice is model selection by use case: reserve premium models for complex reasoning and use lower-cost options for summarization, classification, or simpler generation tasks.

Timed practice exam

Take a SnowPro Specialty: Gen AI practice test under exam conditions

55 questions in 85 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam