Databricks Data Engineer Professional exam dumps

Databricks Data Engineer Professional practice question 272 of 313

Databricks Certified Data Engineer Professional. Professional level, Databricks. Free question with the correct answer and a full explanation.

Databricks Data Engineer Professional Question 272

Select 3

During a nightly ETL process in Databricks, a job failed due to an intermittent network issue while processing a critical task. You have been tasked with repairing and rerunning the failed job to ensure minimal reprocessing and consistent data in the target system. Which steps should you take to repair and rerun the failed job?

  1. A

    Use the Databricks UI to select 'Repair and Rerun' and ensure the task retry policy is set appropriately for the failed task.

  2. B

    Manually delete all intermediate data generated by the job before rerunning it to avoid duplicate processing.

  3. C

    Review the job's run history to identify the failed task and determine whether retries are necessary.

  4. D

    Enable the 'Repair All Tasks' option to rerun all tasks in the job, including successful ones, to ensure data consistency.

  5. E

    Inspect the failed task's log to identify the root cause of the failure and take corrective action before rerunning the job.

Show answer and explanation

Correct answers: A, C, E

Explanation

To repair and rerun failed jobs in Databricks, it is essential to use the 'Repair and Rerun' functionality to minimize reprocessing, review the job's run history to identify failed tasks, and inspect the task logs to understand and resolve the root cause of the failure. These steps ensure efficient debugging and execution of the job while avoiding unnecessary computations or data loss.

  • A. Correct.

    This is correct. The Databricks UI provides a 'Repair and Rerun' option that allows you to rerun only failed tasks, ensuring minimal reprocessing. Configuring the retry policy can help prevent future failures.

  • B. Incorrect.

    This is incorrect. Deleting all intermediate data is unnecessary and can lead to additional processing overhead. Databricks handles intermediate data management during job reruns.

  • C. Correct.

    This is correct. Reviewing the job's run history helps identify the failed task and decide whether retries are needed, ensuring efficient debugging and rerunning.

  • D. Incorrect.

    This is incorrect. The 'Repair All Tasks' option unnecessarily reruns successful tasks, leading to redundant computation and resource usage.

  • E. Correct.

    This is correct. Inspecting the task's log helps you identify the root cause of the failure, such as network issues, and allows you to take corrective action before rerunning the job.

Timed practice exam

Take a Databricks Data Engineer Professional practice test under exam conditions

60 questions in 120 minutes, drawn from this bank, with a score report and a per-question review when you finish.

Start timed exam