Certified Associate Developer for Apache Spark — Free Practice Questions
10 free sample questions from a bank of 180, with the correct answers and explanations. No signup required — start practising right now.
1Which of the following describes the Spark driver?
The Spark driver is responsible for performing all execution in all execution modes – it is the entire Spark application.
The Spare driver is fault tolerant – if it fails, it will recover the entire Spark application.
The Spark driver is the coarsest level of the Spark execution hierarchy – it is synonymous with the Spark application.
The Spark driver is the program space in which the Spark application’s main method runs coordinating the Spark entire application.
The Spark driver is horizontally scaled to increase overall processing throughput of a Spark application.
Answer: D
2Which of the following DataFrame operations is classified as a wide transformation?
DataFrame.filter()
DataFrame.join()
DataFrame.select()
DataFrame.drop()
DataFrame.union()
Answer: B
3The code block shown below contains an error. The code block is intended to return the exact number of distinct values in column division in DataFrame storesDF. Identify the error.
Code block:
storesDF.agg(approx_count_distinct(col(“division”)).alias(“divisionDistinct”))
The approx_count_distinct() operation needs a second argument to set the rsd parameter to ensure it returns the exact number of distinct values.
There is no alias() operation for the approx_count_distinct() operation's output.
There is no way to return an exact distinct number in Spark because the data Is distributed across partitions.
The approx_count_distinct()operation is not a standalone function - it should be used as a method from a Column object.
The approx_count_distinct() operation cannot determine an exact number of distinct values in a column.
Answer: E
4Which of the following code blocks returns the number of rows in DataFrame storesDF for each distinct combination of values in column division and column storeCategory?
5The code block shown below contains an error. The code block is intended to return a collection of summary statistics for column sqft in Data Frame storesDF. Identify the error.
Code block:
storesDF.describes(col(“sgft”))
The column sqft should be subsetted from DataFrame storesDF prior to computing summary statistics on it alone.
The describe() operation does not accept a Column object as an argument outside of a sequence — the sequence Seq(col(“sqft”)) should be specified instead.
The describe()operation doesn’t compute summary statistics for a single column — the summary() operation should be used instead.
The describe()operation doesn't compute summary statistics for numeric columns — the summary() operation should be used instead.
The describe()operation does not accept a Column object as an argument — the column name string “sqft” should be specified instead.
Answer: E
6The code block shown below should extract the integer value for column sqft from the first row of DataFrame storesDF. Choose the response that correctly fills in the numbered blanks within the code block to complete this task.
Code block:
__1__.__2__.__3__[Int](__4__)
1. storesDF
2. first()
3. getAs()
4. “sqft”
1. storesDF
2. first
3. getAs
4. sqft
1. storesDF
2. first()
3. getAs
4. col(“sqft”)
1. storesDF
2. first
3. getAs
4. “sqft”
Answer: D
7The code block shown below should print the schema of DataFrame storesDF. Choose the response that correctly fills in the numbered blanks within the code block to complete this task.
Code block:
__1__.__2__
1. storesDF
2. printSchema(“all”)
1. storesDF
2. schema
1. storesDF
2. getAs[str]
1. storesDF
2. printSchema(true)
1. storesDF
2. printSchema
Answer: E
8The code block shown below contains an error. The code block is intended to create and register a SQL UDF named “ASSESS_PERFORMANCE” using the Scala function assessPerformance() and apply it to column customerSatisfaction in the table stores. Identify the error.
Code block:
spark.udf.register(“ASSESS_PERFORMANCE”, assessPerforance)
spark.sql(“SELECT customerSatisfaction, assessPerformance(customerSatisfaction) AS result FROM stores”)
The customerSatisfaction column cannot be called twice inside the SQL statement.
Registered UDFs cannot be applied inside of a SQL statement.
The order of the arguments to spark.udf.register() should be reversed.
The wrong SQL function is used to compute column result - it should be ASSESS_PERFORMANCE instead of assessPerformance.
There is no sql() operation - the DataFrame API must be used to apply the UDF assessPerformance().
Answer: D
9The code block shown below contains an error. The code block is intended to create the Scala UDF assessPerformanceUDF() and apply it to the integer column customers1t1sfaction in Data Frame storesDF. Identify the error.
Code block:
The input type of customerSatisfaction is not specified in the udf() operation.
The return type of assessPerformanceUDF() must be specified.
The withColumn() operation is not appropriate here - UDFs should be applied by iterating over rows instead.
The assessPerformanceUDF() must first be defined as a Scala function and then converted to a UDF.
UDFs can only be applied via SQL and not through the Data Frame API.
Answer: B
10The code block shown below should create a single-column DataFrame from Scala list years which is made up of integers. Choose the response that correctly fills in the numbered blanks within the code block to complete this task.
Code block:
__1__.__2__(__3__).__4__
1. spark
2. createDataFrame
3. years
4. IntegerType