Certified Data Engineer Professional — Free Practice Questions
10 free sample questions from a bank of 141, with the correct answers and explanations. No signup required — start practising right now.
1An upstream system has been configured to pass the date for a given batch of data to the Databricks Jobs API as a parameter. The notebook to be scheduled will use this parameter to load data with the following code: df = spark.read.format("parquet").load(f"/mnt/source/(date)")
Which code block should be used to create the date Python variable used in the above code block?
date = spark.conf.get("date")
input_dict = input()
date= input_dict["date"]
import sys
date = sys.argv[1]
date = dbutils.notebooks.getParam("date")
dbutils.widgets.text("date", "null")
date = dbutils.widgets.get("date")
Answer: E
2A Delta table of weather records is partitioned by date and has the below schema: date DATE, device_id INT, temp FLOAT, latitude FLOAT, longitude FLOAT
To find all the records from within the Arctic Circle, you execute a query with the below filter: latitude > 66.3
Which statement describes how the Delta engine identifies which files to load?
All records are cached to an operational database and then the filter is applied
The Parquet file footers are scanned for min and max statistics for the latitude column
All records are cached to attached storage and then the filter is applied
The Delta log is scanned for min and max statistics for the latitude column
The Hive metastore is scanned for min and max statistics for the latitude column
Answer: D
3The data engineering team has been tasked with configuring connections to an external database that does not have a supported native connector with Databricks. The external database already has data security configured by group membership. These groups map directly to user groups already created in Databricks that represent various teams within the company.
A new login credential has been created for each group in the external database. The Databricks Utilities Secrets module will be used to make these credentials available to Databricks users.
Assuming that all the credentials are configured correctly on the external database and group membership is properly configured on Databricks, which statement describes how teams can be granted the minimum necessary access to using these credentials?
"Manage" permissions should be set on a secret key mapped to those credentials that will be used by a given team.
"Read" permissions should be set on a secret key mapped to those credentials that will be used by a given team.
"Read" permissions should be set on a secret scope containing only those credentials that will be used by a given team.
"Manage" permissions should be set on a secret scope containing only those credentials that will be used by a given team.
No additional configuration is necessary as long as all users are configured as administrators in the workspace where secrets have been added.
Answer: C
4Which indicators would you look for in the Spark UI’s Storage tab to signal that a cached table is not performing optimally? Assume you are using Spark’s MEMORY_ONLY storage level.
Size on Disk is < Size in Memory
The RDD Block Name includes the “*” annotation signaling a failure to cache
Size on Disk is > 0
The number of Cached Partitions > the number of Spark Partitions
On Heap Memory Usage is within 75% of Off Heap Memory Usage
Answer: C
5What is the first line of a Databricks Python notebook when viewed in a text editor?
%python
// Databricks notebook source
# Databricks notebook source
-- Databricks notebook source
# MAGIC %python
Answer: C
6Which statement describes a key benefit of an end-to-end test?
Makes it easier to automate your test suite
Pinpoints errors in the building blocks of your application
Provides testing coverage for all code paths and branches
Closely simulates real world usage of your application
Ensures code is optimized for a real-life workflow
Answer: D
7The Databricks CLI is used to trigger a run of an existing job by passing the job_id parameter. The response that the job run request has been submitted successfully includes a field run_id.
Which statement describes what the number alongside this field represents?
The job_id and number of times the job has been run are concatenated and returned.
The total number of jobs that have been run in the workspace.
The number of times the job definition has been run in this workspace.
The job_id is returned in this field.
The globally unique ID of the newly triggered run.
Answer: E
8The data science team has created and logged a production model using MLflow. The model accepts a list of column names and returns a new column of type DOUBLE.
The following code correctly imports the production model, loads the customers table containing the customer_id key column into a DataFrame, and defines the feature columns needed for the model. Which code block will output a DataFrame with the schema "customer_id LONG, predictions DOUBLE"?
9A nightly batch job is configured to ingest all data files from a cloud object storage container where records are stored in a nested directory structure YYYY/MM/DD. The data for each date represents all records that were processed by the source system on that date, noting that some records may be delayed as they await moderator approval. Each entry represents a user review of a product and has the following schema:
user_id STRING, review_id BIGINT, product_id BIGINT, review_timestamp TIMESTAMP, review_text STRING
The ingestion job is configured to append all data for the previous date to a target table reviews_raw with an identical schema to the source system. The next step in the pipeline is a batch write to propagate all new records inserted into reviews_raw to a table where data is fully deduplicated, validated, and enriched.
Which solution minimizes the compute costs to propagate this batch of data?
Perform a batch read on the reviews_raw table and perform an insert-only merge using the natural composite key user_id, review_id, product_id, review_timestamp.
Configure a Structured Streaming read against the reviews_raw table using the trigger once execution mode to process new records as a batch job.
Use Delta Lake version history to get the difference between the latest version of reviews_raw and one version prior, then write these records to the next table.
Filter all records in the reviews_raw table based on the review_timestamp; batch append those records produced in the last 48 hours.
Reprocess all records in reviews_raw and overwrite the next table in the pipeline.
Answer: A
10Which statement describes Delta Lake optimized writes?
Before a Jobs cluster terminates, OPTIMIZE is executed on all tables modified during the most recent job.
An asynchronous job runs after the write completes to detect if files could be further compacted; if yes, an OPTIMIZE job is executed toward a default of 1 GB.
Data is queued in a messaging bus instead of committing data directly to memory; all data is committed from the messaging bus in one batch once the job is complete.
Optimized writes use logical partitions instead of directory partitions; because partition boundaries are only represented in metadata, fewer small files are written.
A shuffle occurs prior to writing to try to group similar data together resulting in fewer files instead of each executor writing multiple files based on directory partitions.