Sign In
Home/Databricks/Certified Data Engineer Associate/Free questions

Certified Data Engineer Associate — Free Practice Questions

10 free sample questions from a bank of 146, with the correct answers and explanations. No signup required — start practising right now.

1A data organization leader is upset about the data analysis team’s reports being different from the data engineering team’s reports. The leader believes the siloed nature of their organization’s data engineering and data analysis architectures is to blame. Which of the following describes how a data lakehouse could alleviate this issue?
  • Both teams would autoscale their work as data size evolves
  • Both teams would use the same source of truth for their work
  • Both teams would reorganize to report to the same department
  • Both teams would be able to collaborate on projects in real-time
  • Both teams would respond more quickly to ad-hoc requests
Answer: B
2A data engineer needs to determine whether to use the built-in Databricks Notebooks versioning or version their project using Databricks Repos. Which of the following is an advantage of using Databricks Repos over the Databricks Notebooks versioning?
  • Databricks Repos automatically saves development progress
  • Databricks Repos supports the use of multiple branches
  • Databricks Repos allows users to revert to previous versions of a notebook
  • Databricks Repos provides the ability to comment on specific changes
  • Databricks Repos is wholly housed within the Databricks Lakehouse Platform
Answer: B
3A data engineer has been given a new record of data: id STRING = 'a1' rank INTEGER = 6 rating FLOAT = 9.4 Which SQL commands can be used to append the new record to an existing Delta table my_table?
  • INSERT INTO my_table VALUES ('a1', 6, 9.4)
  • INSERT VALUES ('a1', 6, 9.4) INTO my_table
  • UPDATE my_table VALUES ('a1', 6, 9.4)
  • UPDATE VALUES ('a1', 6, 9.4) my_table
Answer: A
4A data engineer has realized that the data files associated with a Delta table are incredibly small. They want to compact the small files to form larger files to improve performance. Which keyword can be used to compact the small files?
  • OPTIMIZE
  • VACUUM
  • COMPACTION
  • REPARTITION
Answer: A
5A data engineer wants to create a data entity from a couple of tables. The data entity must be used by other data engineers in other sessions. It also must be saved to a physical location. Which of the following data entities should the data engineer create?
  • Table
  • Function
  • View
  • Temporary view
Answer: A
6A data engineer runs a statement every day to copy the previous day’s sales into the table transactions. Each day’s sales are in their own file in the location "/transactions/raw". Today, the data engineer runs the following command to complete this task: After running the command today, the data engineer notices that the number of records in table transactions has not changed. What explains why the statement might not have copied any new records into the table?
Certified Data Engineer Associate question 6
  • The format of the files to be copied were not included with the FORMAT_OPTIONS keyword.
  • The COPY INTO statement requires the table to be refreshed to view the copied rows.
  • The previous day’s file has already been copied into the table.
  • The PARQUET file format does not support COPY INTO.
Answer: C
7Which command can be used to write data into a Delta table while avoiding the writing of duplicate records?
  • DROP
  • INSERT
  • MERGE
  • APPEND
Answer: C
8A data analyst has created a Delta table sales that is used by the entire data analysis team. They want help from the data engineering team to implement a series of tests to ensure the data is clean. However, the data engineering team uses Python for its tests rather than SQL. Which command could the data engineering team use to access sales in PySpark?
  • SELECT * FROM sales
  • spark.table("sales")
  • spark.sql("sales")
  • spark.delta.table("sales")
Answer: B
9A data engineer has created a new database using the following command: CREATE DATABASE IF NOT EXISTS customer360; In which location will the customer360 database be located?
  • dbfs:/user/hive/database/customer360
  • dbfs:/user/hive/warehouse
  • dbfs:/user/hive/customer360
  • dbfs:/user/hive/database
Answer: B
10A data engineer is attempting to drop a Spark SQL table my_table and runs the following command: DROP TABLE IF EXISTS my_table; After running this command, the engineer notices that the data files and metadata files have been deleted from the file system. What is the reason behind the deletion of all these files?
  • The table was managed
  • The table's data was smaller than 10 GB
  • The table did not have a location
  • The table was external
Answer: A

Want the full bank of 146 questions for Certified Data Engineer Associate? See all practice exams.