CompTIA DataAI DY0-001 (V1) Exam Practice Test 2

Reviewed by Editorial Team
The ProProfs editorial team is comprised of experienced subject matter experts. They've collectively created over 10,000 quizzes and lessons, serving over 100 million users. Our team includes in-house content moderators and subject matter experts, as well as a global network of rigorously trained contributors. All adhere to our comprehensive editorial guidelines, ensuring the delivery of high-quality content.
Learn about Our Editorial Process
| By Thames
T
Thames
Community Contributor
Quizzes Created: 11201 | Total Attempts: 9,875,275
| Questions: 25 | Updated: Sep 28, 2026
Please wait...
Question 1 / 26
🏆 Rank #-- ▾
0 %
0/100
Score 0/100

1. A data scientist is comparing gradient boosting approaches for a tabular classification task. Which two statements correctly describe these techniques? (Select two.)

Explanation

Gradient boosting builds an ensemble of trees sequentially, where each new tree is trained specifically to correct the residual errors left by the current ensemble, gradually improving overall performance. XGBoost is a popular, highly optimized implementation of gradient boosting that adds regularization terms and engineering improvements for speed and scalability, making it a common choice for tabular data competitions and production systems alike. Bagging, by contrast, trains trees independently and in parallel, which is a fundamentally different approach from the sequential, error-correcting nature of boosting.

Submit
Please wait...
About This Quiz
CompTIA DataAI Dy0-001 (V1) Exam Practice Test 2 - Quiz

This practice assessment focuses on key concepts in data science as outlined in the CompTIA DataAI DY0-001 (V1) curriculum. It evaluates skills related to data analysis, machine learning, and data management, making it a valuable resource for learners preparing for certification. Enhance your understanding and readiness for the exam with... see moretargeted questions designed to reinforce essential knowledge. see less

2.

What first name or nickname would you like us to use?

You may optionally provide this to label your report, leaderboard, or certificate.

2. A financial institution wants to flag unusual transactions that deviate significantly from a customer's normal spending pattern, and separately wants to flag deliberate attempts to manipulate account balances through coordinated fake transactions. Which two specialized application areas, respectively, address these two goals? (Select two.)

Explanation

Anomaly detection focuses on identifying observations that deviate significantly from an established normal pattern, which fits flagging a transaction that is unusual for a specific customer's typical behavior. Fraud detection more specifically targets deliberate, often coordinated attempts to manipulate a system for illicit gain, which may itself use anomaly detection techniques among others, including graph analysis to uncover coordinated networks of fake accounts or transactions. Both areas commonly draw on overlapping statistical and machine-learning techniques rather than relying on a single method exclusively.

Submit

3. A system processing news articles automatically identifies and labels mentions of people, organizations, and locations within the text. What NLP task does this describe?

Explanation

Named-entity recognition identifies and classifies specific entities within text, such as people, organizations, and locations, which is exactly the task of automatically labeling these mentions within news articles. Topic modeling instead identifies latent themes across a collection of documents without necessarily labeling specific named entities, and sentiment analysis determines the emotional tone of text rather than identifying entities. Text summarization condenses content rather than tagging specific entity mentions within it.

Submit

4. A marketing team wants to continuously test several ad variants in production, dynamically shifting more traffic toward whichever variant is currently performing best, rather than running a single fixed-duration A/B test and waiting for it to conclude. What class of problem does this describe?

Explanation

A multi-armed bandit problem is specifically designed for this kind of ongoing decision-making, balancing the exploration of less-tested options against the exploitation of the option currently believed to perform best, and dynamically adjusting traffic allocation as more data arrives. This differs from a traditional fixed-duration A/B test, which splits traffic evenly and only makes a final decision once the test concludes. Bandit approaches can reduce the cost of testing by shifting traffic away from underperforming variants more quickly than waiting for a full, fixed-length experiment.

Submit

5. A team packages a trained model along with all its dependencies, runtime, and configuration into a single, portable unit that runs consistently across different environments, from a developer's laptop to a production cluster. This packaging approach is called ____.

Explanation

Containerization packages an application, including a trained model, its dependencies, and its runtime environment, into a single portable unit that behaves consistently regardless of where it is deployed, whether on a developer's laptop, a testing environment, or a production cluster. This consistency significantly reduces the classic problem of a model working in one environment but failing in another due to subtle configuration or dependency differences. Container orchestration platforms then manage scaling, scheduling, and lifecycle for many containerized services running together.

Submit

6. A deployed fraud detection model's accuracy has been gradually declining over several months, even though the model itself has not been changed. What MLOps practice is designed to detect and alert on this kind of gradual performance decay?

Explanation

Model performance monitoring continuously tracks a deployed model's metrics over time, which is designed specifically to detect gradual performance decay caused by data drift, where the input data distribution shifts, or concept drift, where the underlying relationship between inputs and outputs changes. Without this ongoing monitoring, a model's declining real-world performance might go unnoticed until it causes a significant business problem. Container orchestration and code isolation address deployment infrastructure concerns, not the detection of performance decay over time.

Submit

7. A dataset's income variable contains a few extreme values far beyond the rest of the distribution, likely due to legitimate but rare high earners rather than data entry errors. A data scientist wants to cap these extreme values at a defined percentile rather than removing them entirely. What technique does this describe?

Explanation

Winsorization caps extreme values at a defined percentile threshold, such as the 1st and 99th percentiles, rather than removing them outright, which retains the observations while limiting the disproportionate influence extreme values can have on statistics like the mean or on model training. This differs from simply removing outliers, which would discard the affected records entirely and could disregard genuinely valid extreme observations. Deciding between removal, Winsorization, or leaving values unchanged depends on whether the extreme values are more likely genuine data points or errors.

Submit

8. A data platform stores relational sales records in tables, JSON event logs with varying fields, and raw customer support call recordings. Which two storage category pairings correctly match a data type to its storage category? (Select two.)

Explanation

Structured storage suits data with a fixed, well-defined schema, such as relational sales records organized into consistent rows and columns. Semi-structured storage suits data like JSON logs that have some organizational structure but allow flexible or varying fields between records, unlike a rigid relational schema. Raw audio recordings are a classic example of unstructured data, since they have no inherent schema or fields at all, which is why pairing them with structured storage would be incorrect.

Submit

9. A team wants to enrich its internal dataset with a third-party commercial dataset but must first review the terms governing how that external data may be used, redistributed, and combined with internal data. What consideration does this represent?

Explanation

Commercial and public datasets typically come with licensing terms that govern how the data may be used, whether it can be redistributed, and whether it can be combined with other data sources, all of which must be reviewed before incorporating external data into a project. Ignoring these restrictions can create legal or compliance risk even if the data itself is technically useful. Data augmentation and winsorization instead address expanding training data and handling outliers respectively, which are unrelated to licensing terms.

Submit

10. A data science team wants to define a small set of quantifiable measures that directly track progress toward the business's stated goals, such as customer retention rate or monthly active users. What are these measures called?

Explanation

KPIs are quantifiable measures specifically chosen to track progress toward defined business goals, such as customer retention or active user counts, and they help translate a broader business objective into something concrete that can be monitored and reported over time. Choosing appropriate KPIs, rather than an overwhelming number of metrics, helps stakeholders focus on what genuinely matters for the business. Confusion matrices, eigenvalues, and loss functions are technical evaluation and modeling concepts, not business-facing KPIs.

Submit

11. As the number of features in a dataset grows very large relative to the number of observations, the data becomes increasingly sparse in the high-dimensional feature space, and distance-based methods like KNN become less meaningful because most points end up roughly equidistant from each other. This phenomenon is known as the curse of ____.

Explanation

The curse of dimensionality describes how, as the number of features grows, the volume of the feature space grows so quickly that available data becomes increasingly sparse within it, and intuitive geometric concepts like distance and density become less meaningful, since points tend to become nearly equidistant from one another. This particularly affects distance-based methods like KNN, which rely heavily on meaningful distance calculations to function well. Dimensionality reduction techniques and careful feature selection are common ways to mitigate this problem before applying distance-sensitive methods.

Submit

12. A data scientist wants to visualize high-dimensional data in two dimensions in a way that preserves local neighborhood structure and reveals natural clusters, even at the cost of distorting global distances between far-apart points. Which technique is best suited for this specific visualization goal?

Explanation

t-SNE is specifically designed for visualization and prioritizes preserving local neighborhood structure, often revealing natural clusters clearly in two or three dimensions, even though it can distort global distances between clusters that are far apart in the original high-dimensional space. PCA instead is a linear technique that preserves global variance structure and is often used for both visualization and more general dimensionality reduction, but it may not reveal nonlinear cluster structure as clearly as t-SNE for this specific purpose. Choosing between them depends on whether the goal is general-purpose dimensionality reduction or specifically visualizing local cluster structure.

Submit

13. During training, a neural network randomly deactivates a fraction of its neurons on each forward pass, forcing the network to not rely too heavily on any single neuron or small group of neurons. What regularization technique is this?

Explanation

Dropout randomly deactivates a fraction of neurons during each training pass, which prevents the network from becoming overly dependent on any specific neuron or narrow combination of neurons, encouraging more robust, distributed representations that generalize better. Batch normalization instead normalizes layer inputs to stabilize and speed up training, and early stopping halts training once validation performance stops improving, both of which are related but distinct regularization or training strategies from dropout's random deactivation approach.

Submit

14. A data scientist adds five new, mostly uninformative predictors to a regression model. Plain R2 increases slightly, but adjusted R2 decreases. Why do these two metrics disagree?

Explanation

Plain R2 mathematically cannot decrease when additional predictors are added to a model, even if those predictors are pure noise, since it simply measures the proportion of variance explained without penalty. Adjusted R2 explicitly penalizes the addition of predictors, so it can decrease when new predictors fail to improve model fit enough to offset the penalty for added complexity. This is why adjusted R2 is generally a more reliable metric for comparing models with different numbers of predictors.

Submit

15. A spam classifier calculates the probability of an email being spam by multiplying the individual probabilities of each word appearing in spam emails, assuming each word's presence is independent of the others given the class label. What algorithm does this describe?

Explanation

Naive Bayes applies Bayes' theorem while making the simplifying, naive assumption that features, such as individual words, are conditionally independent of one another given the class label, which makes the calculation tractable even for high-dimensional text data. This independence assumption is rarely perfectly true in practice, but Naive Bayes classifiers often still perform surprisingly well for tasks like spam filtering and text classification. Linear and quadratic discriminant analysis instead model the class-conditional distributions directly with different assumptions about their covariance structure, which is a different approach entirely.

Submit

16. A data scientist trains multiple different model types, such as a random forest, a gradient boosting model, and a logistic regression model, and then combines their individual predictions into a single final prediction. What is this general approach called?

Explanation

An ensemble model combines predictions from multiple individual models, often of different types or trained differently, to produce a final prediction that is typically more accurate and robust than any single model alone. This works because different models tend to make different errors, and combining them can average out some of those individual mistakes. Cross-validation and regularization instead address model evaluation and overfitting prevention, which are related but distinct concerns from combining multiple models' predictions.

Submit

17. A new analyst joining a data science team needs a reference document that defines every column in a shared dataset, including its data type, allowed values, and business meaning. This kind of reference document is called a ____.

Explanation

A data dictionary defines every field in a dataset, including its data type, valid or expected values, units, and business meaning, which is essential documentation for anyone using the dataset without having built it themselves. This differs from a data set's raw metadata, which might only capture technical properties, since a data dictionary explicitly captures business context and intended meaning as well. Maintaining an accurate, up-to-date data dictionary reduces the risk of a new analyst misinterpreting a column and drawing incorrect conclusions.

Submit

18. A stakeholder requests a model with 99.9% accuracy for a task where even human experts disagree on the correct label about 15% of the time. How should a data scientist frame this in the final recommendation?

Explanation

Distinguishing between what a stakeholder initially wants, what the business actually needs to achieve its goals, and what is realistically achievable given inherent constraints like label disagreement is an important part of justifying a final model recommendation. In this scenario, since even human experts disagree 15% of the time, an accuracy ceiling well below 99.9% may simply reflect the inherent ambiguity of the task rather than a shortcoming of the model. Clearly communicating this distinction helps set appropriate expectations rather than either overpromising or silently underdelivering.

Submit

19. A team designing a fraud detection model must ensure predictions are returned within 50 milliseconds per transaction and that the model can be retrained weekly within a fixed compute budget. What category of model design consideration do these two requirements represent?

Explanation

Time constraints, such as a maximum inference latency, and resource constraints, such as a fixed compute budget for retraining, are both design constraints that shape which model architectures and approaches are even feasible for a given use case, independent of raw predictive accuracy. A highly accurate but slow or resource-intensive model may be entirely unsuitable if it cannot meet these operational constraints. Considering constraints early in the design process avoids investing significant effort in a model that ultimately cannot be deployed under real-world limitations.

Submit

20. A dataset includes a feature ranging from 0 to 1,000,000 and another ranging from 0 to 1, and the team is preparing to train a k-nearest neighbors model that relies on distance calculations. Which two considerations are important here? (Select two.)

Explanation

Distance-based algorithms like KNN calculate distances directly from raw feature values, so a feature with a much larger numeric range will dominate the distance calculation, effectively making the smaller-range feature nearly irrelevant to the model's decisions. Scaling or normalizing features to comparable ranges before training addresses this imbalance, ensuring each feature contributes proportionally to distance calculations based on its actual informativeness rather than its raw numeric scale. Standardization, which centers and scales by standard deviation, and normalization, which rescales to a fixed range like 0 to 1, are related but distinct techniques.

Submit

21. A retail sales dataset shows a consistent, repeating spike in sales every December across multiple years. What data issue or characteristic does this pattern represent?

Explanation

Seasonality refers to a recurring pattern that repeats at consistent intervals, such as a yearly spike in December sales tied to holiday shopping. Failing to account for seasonality when modeling time series data can lead to misleading trend estimates or poor forecasts, since a model might mistake a seasonal spike for a genuine shift in the underlying trend. Time series models like SARIMA explicitly include a seasonal component to address exactly this kind of recurring pattern.

Submit

22. A data scientist wants a single visual that shows the pairwise correlation strength between every combination of numeric variables in a dataset, using color intensity to represent the strength and direction of each relationship. Which visualization fits this need?

Explanation

A correlation plot, often rendered as a heat map of a correlation matrix, displays the pairwise correlation coefficient between every combination of numeric variables at once, using color to indicate strength and direction. This gives an efficient overview of relationships across many variables simultaneously, which would be tedious to inspect one scatter plot at a time. A Sankey diagram instead visualizes flow between categories, and a violin plot visualizes the distribution shape of a single variable, neither of which shows pairwise correlation across many variables.

Submit

23. To establish that a new website design causes an increase in conversion rate, rather than merely correlates with it, a data scientist randomly assigns users to see either the old or new design and compares outcomes. This experimental design is called a randomized controlled trial, often implemented in a digital product setting as an ____ test.

Explanation

A randomized controlled trial randomly assigns subjects to a treatment or control condition, which allows a causal claim to be made about the treatment's effect since randomization balances out other confounding factors between groups on average. In a digital product context, this design is commonly called an A/B test, comparing a new variant against the existing baseline. Without randomization, an observed association between a design change and conversion rate could be explained by confounding factors rather than a true causal effect.

Submit

24. A medical test is 95% accurate at detecting a disease that affects only 1% of the population. A data scientist wants to calculate the actual probability that a patient who tested positive truly has the disease, accounting for the low base rate. Which concept is required to correctly perform this calculation?

Explanation

Bayes' rule combines the test's known accuracy with the disease's prior probability, or base rate, in the population to calculate the true posterior probability that a positive result indicates actual disease. Because the disease is rare, even a fairly accurate test can produce a surprising number of false positives relative to true positives, which is exactly the kind of counterintuitive result Bayes' rule reveals. Ignoring the base rate and relying only on the test's stated accuracy is a common and significant reasoning error in this type of problem.

Submit

25. A binary classifier for a rare disease produces a confusion matrix with 5 true positives, 2 false positives, 3 false negatives, and 990 true negatives. Which two statements about the resulting metrics are correct? (Select two.)

Explanation

Precision divides true positives by all predicted positives, which here is 5 divided by the sum of 5 true positives and 2 false positives, giving approximately 71%. Recall divides true positives by all actual positives, which here is 5 divided by the sum of 5 true positives and 3 false negatives, giving 62.5%. Given the extreme class imbalance in this scenario, accuracy alone would be misleading, since a classifier that always predicts negative would still achieve very high accuracy while completely failing to identify any positive cases.

Submit
×
Saved
Thank you for your feedback!
View My Results
Cancel
  • All
    All (25)
  • Unanswered
    Unanswered ()
  • Answered
    Answered ()
A data scientist is comparing gradient boosting approaches for a...
A financial institution wants to flag unusual transactions that...
A system processing news articles automatically identifies and labels...
A marketing team wants to continuously test several ad variants in...
A team packages a trained model along with all its dependencies,...
A deployed fraud detection model's accuracy has been gradually...
A dataset's income variable contains a few extreme values far beyond...
A data platform stores relational sales records in tables, JSON event...
A team wants to enrich its internal dataset with a third-party...
A data science team wants to define a small set of quantifiable...
As the number of features in a dataset grows very large relative to...
A data scientist wants to visualize high-dimensional data in two...
During training, a neural network randomly deactivates a fraction of...
A data scientist adds five new, mostly uninformative predictors to a...
A spam classifier calculates the probability of an email being spam by...
A data scientist trains multiple different model types, such as a...
A new analyst joining a data science team needs a reference document...
A stakeholder requests a model with 99.9% accuracy for a task where...
A team designing a fraud detection model must ensure predictions are...
A dataset includes a feature ranging from 0 to 1,000,000 and another...
A retail sales dataset shows a consistent, repeating spike in sales...
A data scientist wants a single visual that shows the pairwise...
To establish that a new website design causes an increase in...
A medical test is 95% accurate at detecting a disease that affects...
A binary classifier for a rare disease produces a confusion matrix...
play-Mute sad happy unanswered_answer up-hover down-hover success oval cancel Check box square blue
Alert!