CompTIA DataAI DY0-001 (V1) Exam Practice Test 3

Reviewed by Editorial Team
The ProProfs editorial team is comprised of experienced subject matter experts. They've collectively created over 10,000 quizzes and lessons, serving over 100 million users. Our team includes in-house content moderators and subject matter experts, as well as a global network of rigorously trained contributors. All adhere to our comprehensive editorial guidelines, ensuring the delivery of high-quality content.
Learn about Our Editorial Process
| By Thames
T
Thames
Community Contributor
Quizzes Created: 11201 | Total Attempts: 9,875,275
| Questions: 25 | Updated: Sep 28, 2026
Please wait...
Question 1 / 26
🏆 Rank #-- ▾
0 %
0/100
Score 0/100

1. A team compares a single decision tree against a random forest built from many trees trained on different bootstrapped samples and random feature subsets. Which two statements correctly describe the relative behavior of these two approaches? (Select two.)

Explanation

A single decision tree, especially a deep one, tends to have high variance and can easily overfit the specific training data it was built on, capturing noise along with genuine signal. A random forest combines many such trees, each trained on a different bootstrapped sample and a random subset of features, and averages their predictions, which reduces overall variance and typically improves generalization compared to any single tree. Random forests do rely on bootstrapped sampling as a core part of their construction, and while they generally require more computation than a single tree, they typically achieve better generalization performance in exchange.

Submit
Please wait...
About This Quiz
CompTIA DataAI Dy0-001 (V1) Exam Practice Test 3 - Quiz

This practice resource focuses on the CompTIA DataAI DY0-001 (V1) exam, evaluating your understanding of data analysis, artificial intelligence, and related concepts. It's designed to help learners solidify their knowledge and prepare effectively for certification. Engaging with this material will enhance your readiness and confidence for the exam.

2.

What first name or nickname would you like us to use?

You may optionally provide this to label your report, leaderboard, or certificate.

2. A computer vision system for autonomous vehicles needs to both draw a bounding box around each detected pedestrian and also determine the exact pixel-level boundary of the road surface. Which two computer vision tasks, respectively, address these two goals?

Explanation

Object detection identifies and localizes objects within an image, typically by drawing a bounding box around each detected instance, which fits identifying and marking each pedestrian's approximate location. Semantic segmentation instead classifies every individual pixel in an image according to a category, producing an exact pixel-level boundary, which is what is needed to precisely delineate the road surface rather than just a rough bounding region. Choosing the right task for each requirement matters, since a bounding box alone would not give the precise pixel-level boundary that segmentation provides.

Submit

3. A model represents each word as a dense numeric vector such that words with similar meanings end up close together in the resulting vector space, unlike a simple one-hot encoding where every word is equally distant from every other word. What does this describe?

Explanation

Word embeddings represent words as dense numeric vectors positioned in a continuous vector space such that semantically similar words end up closer together, which captures meaningful relationships between words that a sparse one-hot encoding cannot represent at all. Techniques like Word2vec and GloVe are well-known methods for learning these embeddings from large text corpora. TF-IDF instead weights term importance based on frequency statistics rather than learning a dense semantic vector space, which is a different kind of text representation entirely.

Submit

4. A supply chain team wants to minimize shipping costs subject to a set of linear constraints on warehouse capacity and delivery deadlines, where the objective function and all constraints are linear. Which classical optimization method is designed for exactly this kind of problem?

Explanation

The simplex method is a classical algorithm designed specifically for linear programming problems, where both the objective function to optimize and all constraints are linear, which matches a shipping cost minimization problem subject to linear capacity and deadline constraints. It systematically moves between vertices of the feasible region defined by the constraints to find the optimal solution efficiently. k-means, gradient boosting, and PCA instead address clustering, supervised prediction, and dimensionality reduction respectively, none of which is designed for linear programming optimization.

Submit

5. An organization keeps its most sensitive data processing on private, on-premises infrastructure for compliance reasons, while running less sensitive, more elastic workloads in a public cloud environment, with both parts of the system working together. This combined approach is called ____ deployment.

Explanation

Hybrid deployment combines on-premises infrastructure with public cloud resources, allowing an organization to keep especially sensitive workloads under tighter direct control for compliance or security reasons while still taking advantage of the cloud's elasticity and scalability for other, less sensitive workloads. This approach requires careful design of how data and processing are coordinated between the two environments, since network connectivity, latency, and security boundaries between them must all be accounted for. Hybrid deployment is often chosen specifically as a middle ground between full on-premises control and full cloud flexibility.

Submit

6. A team validates a new model's predictions against a held-out historical dataset before deployment, and separately monitors its predictions against real, live incoming data after deployment. Which two validation approaches, respectively, does this describe?

Explanation

Offline validation evaluates a model against historical, already-collected data before it is ever exposed to live traffic, which is useful for an initial assessment of expected performance under controlled conditions. Online validation instead monitors a model's actual behavior against real, live incoming data after deployment, which can reveal issues that historical data alone could not anticipate, such as genuine shifts in real-world input distributions. Relying on offline validation alone, without any online monitoring after deployment, risks missing exactly the kind of real-world drift or unexpected behavior that only live data would reveal.

Submit

7. A merged dataset combines records from three systems, one storing dates as MM/DD/YYYY, another as DD-MM-YYYY, and a third as Unix timestamps. Before analysis, what must be done to avoid misinterpreting dates?

Explanation

Standardizing date and time values to a single consistent format and time zone before analysis is essential when combining data from multiple sources that record dates differently, since ambiguous formats like MM/DD versus DD/MM can easily be misinterpreted and silently produce incorrect results. Assuming that analysis tools will automatically and correctly detect every format is a risky assumption that can lead to subtle, hard-to-detect errors, especially for dates where both interpretations produce a valid calendar date. Deleting the date column would remove potentially important information rather than solving the standardization problem.

Submit

8. A data platform team wants to be able to trace exactly which raw source and transformation steps produced a specific value in a downstream report, and also wants transformation jobs to run automatically in the correct dependency order each night. Which two concepts, respectively, address these two needs? (Select two.)

Explanation

Data lineage tracks the origin of a piece of data and every transformation applied to it on its way to a downstream report, which is exactly what is needed to trace a specific value back to its raw source and processing steps. Orchestration automates the scheduling and sequencing of pipeline jobs so that each transformation step runs only after its dependencies have completed successfully, which is what enables reliable, automated nightly processing. Confusing these two concepts would misdirect troubleshooting efforts, since a lineage problem requires tracing data transformations while a scheduling problem requires examining the orchestration configuration.

Submit

9. A pharmaceutical researcher deliberately administers a treatment to one group of participants and a placebo to another, then measures the outcome difference between groups. What type of generated data does this produce?

Explanation

Experimental data is generated through a deliberately designed intervention, such as administering a treatment to one group and a placebo to another, which allows researchers to draw stronger causal conclusions than purely observational data typically permits. This differs from administrative data, which is generated as a byproduct of routine record-keeping processes, and sensor data, which is generated by automated measurement devices without direct researcher intervention. Understanding how a specific dataset was generated is important context for correctly interpreting what conclusions the data can support.

Submit

10. A team needs to let developers test an application against realistic-looking data without exposing actual customer names, addresses, or account numbers in a lower, less secured environment. What technique addresses this?

Explanation

Data obfuscation replaces or masks sensitive values, such as real names or account numbers, with realistic-looking but fictional substitutes, preserving the structural characteristics developers need for testing without exposing actual sensitive customer data in less secured environments. This differs from full anonymization intended for external sharing, though both address protecting sensitive information, since obfuscation for internal testing purposes may have somewhat different requirements than data intended for external research partners. Feature engineering and cross-validation instead address model input construction and evaluation, unrelated to protecting sensitive data values.

Submit

11. When two models achieve similar predictive performance, a data scientist prefers the simpler one with fewer assumptions and parameters, reasoning that unnecessary complexity should be avoided unless it provides a meaningful benefit. This principle is commonly known as Occam's ____.

Explanation

Occam's razor, also called the law of parsimony, holds that among competing explanations or models with similar explanatory or predictive power, the simpler one is generally preferable, since unnecessary complexity increases the risk of overfitting and makes a model harder to interpret and maintain. This principle guides model selection decisions when a more complex model does not provide a clear, meaningful improvement over a simpler alternative. It does not mean the simplest possible model is always correct, only that added complexity should be justified by a genuine gain in performance or explanatory power.

Submit

12. A data scientist wants to cluster geographic point data where clusters may have irregular, non-circular shapes and the number of clusters is not known in advance, while also automatically identifying sparse points as noise rather than forcing them into a cluster. Which clustering algorithm fits this requirement?

Explanation

DBSCAN groups points based on density, allowing it to discover clusters of irregular, non-circular shapes without requiring the number of clusters to be specified in advance, and it explicitly labels points in sparse regions as noise rather than forcing every point into some cluster. This is a significant advantage over k-means, which assumes roughly spherical clusters and requires the number of clusters, k, to be chosen beforehand. Linear regression and PCA address prediction and dimensionality reduction respectively, neither of which is a clustering algorithm.

Submit

13. A team building a model to predict the next word in a sentence needs an architecture capable of retaining relevant information over long sequences while mitigating the vanishing gradient problem that simpler recurrent architectures struggle with. Which architecture is well suited to this?

Explanation

LSTM networks include gating mechanisms specifically designed to retain relevant information over long sequences while mitigating the vanishing gradient problem that simpler recurrent neural network architectures often struggle with over long sequences. This makes LSTMs well suited to sequential tasks like language modeling, where context from much earlier in a sequence can still matter for predicting what comes next. A basic multilayer perceptron has no inherent mechanism for handling sequential or temporal dependencies at all, making it poorly suited to this specific task.

Submit

14. A researcher wants to test whether average test scores differ significantly across three different teaching methods, using a single test that compares all three group means simultaneously rather than running multiple pairwise t-tests. Which technique is appropriate?

Explanation

ANOVA is designed specifically to compare means across three or more groups simultaneously, avoiding the inflated false positive risk that comes from running multiple separate pairwise t-tests without correction. A single t-test only compares two group means at a time, which would require several separate tests to cover all three teaching methods. A chi-squared test instead is typically used for categorical data relationships, not for comparing continuous outcome means across groups.

Submit

15. A streaming service recommends movies to a user based on the viewing patterns of other users who have similar taste, rather than analyzing the content attributes of the movies themselves. What recommendation approach does this describe?

Explanation

Collaborative filtering generates recommendations based on patterns of behavior across many users, identifying users with similar tastes and recommending items those similar users liked, rather than analyzing the content attributes of the items themselves. This differs from content-based filtering, which instead relies on the characteristics of the items, such as a movie's genre or actors, to make recommendations. Techniques like alternating least squares are commonly used to implement collaborative filtering at scale, particularly for sparse user-item interaction matrices.

Submit

16. After training a complex ensemble model, a data scientist applies a separate technique to estimate which features most influenced a specific individual prediction, without needing the model itself to be inherently interpretable. What category of technique is this?

Explanation

Post hoc explainability techniques are applied after a model has already been trained, to help explain its behavior without requiring the model itself to be inherently interpretable, such as a simple linear model would be. A local explanation specifically focuses on understanding why the model made a particular prediction for a single instance, as opposed to a global explanation, which characterizes the model's overall behavior across the entire dataset. This distinction matters because a technique providing only global explanations might not adequately explain any single unusual or high-stakes individual prediction.

Submit

17. Alongside a trained model, a data science team stores information about when it was trained, which dataset version was used, and which hyperparameters were selected, without this information being part of the model's actual predictions. This kind of supporting information is called ____.

Explanation

Metadata captures descriptive information about a model or dataset, such as training date, dataset version, and hyperparameter values, without being part of the actual predictive output itself. Maintaining thorough metadata is essential for reproducibility, auditing, and troubleshooting, since it allows a team to trace exactly how a specific model version was produced. Without accurate metadata, diagnosing why two supposedly identical model runs produced different results can become extremely difficult.

Submit

18. Before recommending a final production model, a data scientist runs a series of tests confirming the model's predictions remain stable under small input perturbations and that its performance holds across different customer segments. What is this step called?

Explanation

Specification testing verifies that a model behaves correctly and robustly under a defined set of test conditions, such as stability under small input perturbations or consistent performance across different segments of the population, before it is trusted for production use. This step provides evidence supporting a final model recommendation beyond simply reporting an aggregate accuracy number. Literature review and hyperparameter tuning instead occur earlier in the model selection process, informing which approaches to try rather than validating the final chosen model's behavior.

Submit

19. A data scientist examines a plot of residuals against fitted values from a regression model and notices a clear curved pattern rather than a random scatter around zero. What does this diagnostic plot suggest?

Explanation

A residual versus fitted plot showing a clear curved or systematic pattern, rather than a random scatter around zero, suggests the model is failing to capture some structure in the data, often indicating a non-linear relationship that a purely linear model cannot represent well. If residuals were randomly scattered with no discernible pattern, that would instead support the assumption that a linear model is appropriately specified. Addressing this might involve adding polynomial terms, transforming variables, or switching to a non-linear modeling approach.

Submit

20. A regression model's residuals show clear heteroskedasticity, with variance increasing as the predicted value increases, and the target variable is strictly positive and right-skewed. Which two considerations are relevant here? (Select two.)

Explanation

A Box-Cox transformation is a family of power transformations designed to stabilize variance and reduce skewness, and it requires the target variable to be strictly positive, which matches the scenario described here. A logarithmic transformation is a well-known special case within this family that is particularly effective for right-skewed, strictly positive data exhibiting this kind of variance pattern. Left unaddressed, heteroskedasticity can violate standard linear regression assumptions and lead to inefficient estimates and unreliable standard errors, which is precisely why addressing it matters.

Submit

21. A time series model trained on early years of stock price data performs poorly when applied to more recent years, because the statistical properties of the series, such as its mean and variance, have shifted substantially over time. What issue does this describe?

Explanation

Non-stationarity describes a time series whose statistical properties, such as mean, variance, or autocorrelation structure, change over time rather than remaining stable, which can cause a model trained on one period to generalize poorly to a different period. Many time series modeling techniques, including ARIMA, assume or require stationarity, often achieved through differencing, before they can be applied reliably. Multicollinearity and sparse data instead describe relationships between predictor variables and data density respectively, which are different issues entirely.

Submit

22. A dataset includes a column for shirt size with values small, medium, and large, which have a meaningful order but no consistent numeric spacing between them. What feature type is this?

Explanation

An ordinal variable has categories with a meaningful order, such as small, medium, and large, but the spacing or distance between categories is not necessarily consistent or numerically meaningful. This distinguishes it from a nominal variable, which has categories with no inherent order at all, such as color. Correctly identifying a variable as ordinal rather than nominal or continuous affects which encoding and modeling choices are appropriate.

Submit

23. A data scientist models a non-stationary time series by combining an autoregressive component, a differencing step to achieve stationarity, and a moving average component into a single unified model, commonly abbreviated ____.

Explanation

ARIMA combines an autoregressive component, which uses past values to predict future values, a differencing step, which transforms a non-stationary series into a stationary one by computing differences between consecutive observations, and a moving average component, which models the error term as a combination of past forecast errors. This combination makes ARIMA a flexible and widely used approach for many types of time series that are not stationary in their raw form. Extensions like SARIMA further add a seasonal component for time series with recurring seasonal patterns.

Submit

24. A recommendation system needs to measure similarity between two user preference vectors based on the angle between them rather than their absolute magnitude, while a separate application needs to measure straight-line distance between two physical GPS coordinates. Which two distance metrics, respectively, fit these two use cases? (Select two.)

Explanation

Cosine distance measures the angle between two vectors, which makes it well suited for comparing preference vectors where the direction of preference matters more than the absolute magnitude of the values. Euclidean distance measures straight-line distance in space, which is the natural and intuitive choice for physical GPS coordinates where actual geographic distance is what matters. Manhattan distance instead sums absolute differences along each axis, which is a different geometric interpretation better suited to grid-like movement rather than angle-based similarity.

Submit

25. A histogram of household income shows a long tail extending toward very high values, with most observations clustered at lower income levels. What does this pattern describe?

Explanation

Positive skewness describes a distribution with a long tail extending toward higher values, which is a very common pattern for income data, where most people earn moderate amounts but a small number of very high earners pull the tail rightward. Negative skewness would instead show a long tail extending toward lower values, with most observations clustered at higher values. Recognizing skewness is important because it can affect which statistical methods and central tendency measures, such as median versus mean, are most appropriate to summarize the data.

Submit
×
Saved
Thank you for your feedback!
View My Results
Cancel
  • All
    All (25)
  • Unanswered
    Unanswered ()
  • Answered
    Answered ()
A team compares a single decision tree against a random forest built...
A computer vision system for autonomous vehicles needs to both draw a...
A model represents each word as a dense numeric vector such that words...
A supply chain team wants to minimize shipping costs subject to a set...
An organization keeps its most sensitive data processing on private,...
A team validates a new model's predictions against a held-out...
A merged dataset combines records from three systems, one storing...
A data platform team wants to be able to trace exactly which raw...
A pharmaceutical researcher deliberately administers a treatment to...
A team needs to let developers test an application against...
When two models achieve similar predictive performance, a data...
A data scientist wants to cluster geographic point data where clusters...
A team building a model to predict the next word in a sentence needs...
A researcher wants to test whether average test scores differ...
A streaming service recommends movies to a user based on the viewing...
After training a complex ensemble model, a data scientist applies a...
Alongside a trained model, a data science team stores information...
Before recommending a final production model, a data scientist runs a...
A data scientist examines a plot of residuals against fitted values...
A regression model's residuals show clear heteroskedasticity, with...
A time series model trained on early years of stock price data...
A dataset includes a column for shirt size with values small, medium,...
A data scientist models a non-stationary time series by combining an...
A recommendation system needs to measure similarity between two user...
A histogram of household income shows a long tail extending toward...
play-Mute sad happy unanswered_answer up-hover down-hover success oval cancel Check box square blue
Alert!