AI-102 exam dumps

AI-102 practice question 209 of 493

Designing and Implementing a Microsoft Azure AI Solution. Professional level, Microsoft. Free question with the correct answer and a full explanation.

AI-102 Question 209

Select 3

You are developing an Azure AI solution that uses the Speech service to generate synthesized speech for an audiobook application. You need to make the speech sound more natural by adding pauses, emphasis, and altering the pitch and rate of the voice during synthesis. Which of the following actions should you take to achieve this?

  1. A

    Use SSML to add tags to introduce pauses at specific points in the text.

  2. B

    Use SSML to add tags to adjust pitch, rate, and volume of the speech.

  3. C

    Use the built-in neural voice feature of Azure Speech without any customization.

  4. D

    Use SSML to add tags to highlight specific words or phrases in the text.

  5. E

    Use a custom voice model trained with your own dataset to achieve a natural-sounding voice.

Show answer and explanation

Correct answers: A, B, D

Explanation

To improve the naturalness of text-to-speech synthesis, you can use SSML (Speech Synthesis Markup Language), which allows fine-grained control over speech characteristics such as pauses, emphasis, pitch, and rate. The , , and tags are specifically designed to enhance the expressiveness and clarity of synthesized speech. While neural voices and custom-trained voices can improve the overall quality of the voice, SSML provides the customization required to control how the speech is delivered.

  • A. Correct.

    Correct. The tag in SSML allows you to introduce pauses at specific points, which can make the speech more natural and better paced.

  • B. Correct.

    Correct. The tag in SSML is used to adjust pitch, rate, and volume, which can make the speech sound more expressive and natural.

  • C. Incorrect.

    Incorrect. While the neural voice feature enhances the quality of synthesis, it does not allow fine-grained control over pauses, emphasis, or pitch adjustments without SSML.

  • D. Correct.

    Correct. The tag in SSML is used to stress specific words or phrases, improving clarity and naturalness in synthesized speech.

  • E. Incorrect.

    Incorrect. Custom voice models are useful for creating unique voices, but they do not inherently manage pauses, emphasis, or other speech dynamics without SSML.

Timed practice exam

Take a AI-102 practice test under exam conditions

40 questions in 120 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam