MLA-C01 exam dumps

MLA-C01 practice question 318 of 458

AWS Certified Machine Learning Engineer - Associate. Associate level, Amazon Web Services. Free question with the correct answer and a full explanation.

MLA-C01 Question 318

Select 2

You are a Machine Learning Engineer managing a SageMaker endpoint for a real-time inference application that experiences significant traffic spikes during business hours but remains idle during off-hours. To optimize costs while ensuring low latency during peak periods, which actions should you take to configure SageMaker endpoint auto scaling?

  1. A

    Set up a target tracking scaling policy based on the SageMakerVariantInvocationsPerInstance metric.

  2. B

    Configure a scheduled scaling policy to scale up during business hours and scale down during off-hours.

  3. C

    Enable multi-model endpoints to automatically adjust instance counts based on traffic.

  4. D

    Set a fixed number of instances to handle peak traffic to avoid scaling delays.

  5. E

    Use a step scaling policy to add instances incrementally as traffic increases.

Show answer and explanation

Correct answers: A, B

Explanation

To meet scalability requirements effectively, you should combine dynamic scaling (via target tracking policies) with predictable scaling (via scheduled scaling policies). This ensures low latency during peak periods and minimizes costs during idle times. Target tracking responds to real-time demand changes, while scheduled scaling prepares for known traffic patterns like business hours.

  • A. Correct.

    Correct: A target tracking scaling policy using the SageMakerVariantInvocationsPerInstance metric dynamically adjusts the number of instances based on real-time traffic, helping to manage unpredictable demand efficiently.

  • B. Correct.

    Correct: Scheduled scaling policies allow you to preemptively scale up or down based on known traffic patterns, such as business hours, ensuring resource availability and cost savings.

  • C. Incorrect.

    Incorrect: Multi-model endpoints optimize resource usage by hosting multiple models on a single instance but do not directly adjust instance counts based on traffic. This is not relevant for auto scaling policies.

  • D. Incorrect.

    Incorrect: Setting a fixed number of instances does not leverage auto scaling and could lead to unnecessary costs during off-peak hours or insufficient capacity during peak demand.

  • E. Incorrect.

    Incorrect: Step scaling policies are generally less suited for real-time inference workloads with unpredictable traffic patterns compared to target tracking policies.

Timed practice exam

Take a MLA-C01 practice test under exam conditions

65 questions in 130 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam