10 free sample questions from a bank of 10, with the correct answers and explanations. No signup required — start practising right now.
1Which of the following issues should a data scientist be most concerned about when generating a synthetic data set?
The data set consuming too many resources
The data set having insufficient features
The data set having insufficient row observations
The data set not being representative of the population
Answer: D
2A data scientist is performing a linear regression and wants to construct a model that explains the most variation in the data. Which of the following should the data scientist maximize when evaluating the regression performance metrics?
Accuracy
R2
p value
AUC
Answer: B
3A data scientist is building an inferential model with a single predictor variable. A scatter plot of the independent variable against the real-number dependent variable shows a strong relationship between them. The predictor variable is normally distributed with very few outliers. Which of the following algorithms is the best fit for this model, given the data scientist wants the model to be easily interpreted?
A logistic regression
An exponential regression
A linear regression
A probit regression
Answer: C
4A data scientist wants to evaluate the performance of various nonlinear models. Which of the following is best suited for this task?
AIC
Chi-squared test
MCC
ANOVA
Answer: A
5Which of the following is the layer that is responsible for the depth in deep learning?
Convolution
Dropout
Pooling
Hidden
Answer: D
6Which of the following modeling tools is appropriate for solving a scheduling problem?
One-armed bandit
Constrained optimization
Decision tree
Gradient descent
Answer: B
7Which of the following environmental changes is most likely to resolve a memory constraint error when running a complex model using distributed computing?
Converting an on-premises deployment to a containerized deployment
Migrating to a cloud deployment
Moving model processing to an edge deployment
Adding nodes to a cluster deployment
Answer: D
8A data analyst wants to save a newly analyzed data set to a local storage option. The data set must meet the following requirements:
Be minimal in size
Have the ability to be ingested quickly
Have the associated schema, including data types, stored with it
Which of the following file types is the best to use?
JSON
Parquet
XML
CSV
Answer: B
9Which of the following is a key difference between KNN and k-means machine-learning techniques?
KNN operates exclusively on continuous data, while k-means can work with both continuous and categorical data.
KNN performs better with longitudinal data sets, while k-means performs better with survey data sets.
KNN is used for finding centroids, while k-means is used for finding nearest neighbors.
KNN is used for classification, while k-means is used for clustering.
Answer: D
10A data scientist needs to:
Build a predictive model that gives the likelihood that a car will get a flat tire.
Provide a data set of cars that had flat tires and cars that did not.
All the cars in the data set had sensors taking weekly measurements of tire pressure similar to the sensors that will be installed in the cars consumers drive. Which of the following is the most immediate data concern?