CompTIA DataAI DY0-001 (V1) Exam Practice Test 1

Reviewed by Editorial Team
The ProProfs editorial team is comprised of experienced subject matter experts. They've collectively created over 10,000 quizzes and lessons, serving over 100 million users. Our team includes in-house content moderators and subject matter experts, as well as a global network of rigorously trained contributors. All adhere to our comprehensive editorial guidelines, ensuring the delivery of high-quality content.
Learn about Our Editorial Process
| By Thames
T
Thames
Community Contributor
Quizzes Created: 11201 | Total Attempts: 9,875,275
| Questions: 25 | Updated: Sep 28, 2026
Please wait...
Question 1 / 26
🏆 Rank #-- ▾
0 %
0/100
Score 0/100

1. A team is choosing between two ensemble tree-based approaches: one that trains many trees independently on bootstrapped samples and averages their predictions, and another that trains trees sequentially, each one correcting errors made by the previous trees. Which two terms correctly describe these two approaches, respectively? (Select two.)

Explanation

Bagging, short for bootstrap aggregation, trains many trees independently on different bootstrapped samples of the data and combines their predictions, typically by averaging or voting, which reduces variance. Boosting instead trains trees sequentially, with each new tree specifically focused on correcting the errors made by the ensemble so far, which reduces bias and often achieves strong performance but can be more prone to overfitting if not carefully tuned. Random forest is a well-known bagging-based approach, while gradient boosting and XGBoost are well-known boosting-based approaches.

Submit
Please wait...
About This Quiz
CompTIA DataAI Dy0-001 (V1) Exam Practice Test 1 - Quiz

This practice assessment focuses on key concepts and skills necessary for the CompTIA DataAI DY0-001 (V1) exam. It evaluates your understanding of data analysis, artificial intelligence, and related technologies, making it a valuable resource for aspiring professionals in the field. Engaging with this material will help reinforce your knowledge and... see moreprepare you for the certification. see less

2.

What first name or nickname would you like us to use?

You may optionally provide this to label your report, leaderboard, or certificate.

2. A team training an image classifier has a limited training set and wants to artificially expand it by creating modified versions of existing images that still represent valid, realistic examples of the same class. Which two augmentation techniques are commonly used for this purpose? (Select two.)

Explanation

Rotating and flipping images creates additional valid training examples that still represent the same class from a different, plausible orientation, effectively expanding the training set without collecting new data. Adding occlusion or cropping simulates real-world conditions where objects may be partially obscured or only partially visible, which helps the model generalize to imperfect real-world images. Replacing every pixel with a solid color, deleting training data, or reducing every image to an identical fixed image would destroy useful information rather than meaningfully augment the dataset.

Submit

3. A search engine wants to weigh words that are frequent within a specific document but rare across the overall document collection more heavily than words that appear frequently everywhere, such as common stop words. Which technique achieves this weighting?

Explanation

TF-IDF combines how frequently a term appears within a specific document with how rare that term is across the entire document collection, producing higher weights for terms that are distinctive to a document and lower weights for common terms that appear everywhere, such as typical stop words. This weighting scheme helps search and information retrieval systems prioritize genuinely distinguishing terms over generically frequent ones. One-hot encoding, PCA, and k-means serve encoding, dimensionality reduction, and clustering purposes respectively, none of which perform this specific frequency-based term weighting.

Submit

4. A logistics company wants to find the shortest possible route visiting a fixed set of delivery stops exactly once and returning to the start, which is a classic example of the traveling salesman problem. Which category of optimization does this represent?

Explanation

The traveling salesman problem is a classic example of constrained optimization, since the solution must satisfy specific constraints, namely visiting every stop exactly once and returning to the starting point, while minimizing total distance or cost. Unconstrained optimization instead searches for a maximum or minimum without such explicit structural requirements on the solution. Multi-armed bandit problems address a different class of sequential decision-making under uncertainty, not fixed-route optimization.

Submit

5. A computer vision model needs to run directly on cameras installed in remote locations with unreliable internet connectivity, rather than sending every frame to a central cloud server for inference. This deployment approach is known as ____ deployment.

Explanation

Edge deployment runs inference directly on or near the device generating the data, such as a camera or sensor, rather than requiring a round trip to a centralized cloud server for every prediction. This is especially valuable when connectivity is unreliable, when latency must be minimal, or when bandwidth constraints make continuously streaming raw data to the cloud impractical. Edge deployment often requires optimizing or compressing a model to run efficiently within the more limited compute resources typically available on edge devices.

Submit

6. A team wants to validate a newly trained recommendation model against the currently deployed model using real production traffic, routing a portion of users to each model and comparing business outcomes, before fully replacing the old model. What practice does this describe?

Explanation

Model A/B testing routes a portion of real production traffic to a new model candidate while the remainder continues to see the existing model, allowing a direct comparison of actual business outcomes under real conditions before a full rollout. This complements offline validation, which evaluates a model against historical data before it ever sees live traffic, but cannot fully capture how the model performs against real-time, real-world behavior. Container orchestration and data replication instead address deployment infrastructure and data availability, not model outcome comparison.

Submit

7. Two datasets need to be merged based on customer name, but names are inconsistently formatted between the two sources, such as 'Bob Smith' versus 'Robert Smith Jr.' for the same person. What technique helps match these records despite the inconsistency?

Explanation

A fuzzy join matches records based on approximate similarity, such as string edit distance or phonetic similarity, rather than requiring an exact match on the join key, which is well suited to handling inconsistently formatted names or other free-text identifiers across data sources. An exact key-based join would fail to match these records at all, since 'Bob Smith' and 'Robert Smith Jr.' are not identical strings. Fuzzy joins typically require careful tuning and validation of match rates, since approximate matching can introduce false matches if thresholds are set too loosely.

Submit

8. A fraud detection system needs to score transactions the instant they occur, while a separate monthly reporting job can process an entire month of accumulated data at once overnight. Which two data ingestion approaches, respectively, fit these two use cases? (Select two.)

Explanation

Streaming ingestion processes data continuously as it arrives, which is necessary for a fraud detection system that must score each transaction immediately rather than waiting for a batch to accumulate. Batching instead processes accumulated data in scheduled chunks, which is well suited to a monthly reporting job where immediate processing is unnecessary and batch efficiency is preferred. Using batching for the fraud detection use case would introduce unacceptable latency, defeating the purpose of real-time fraud scoring.

Submit

9. A team considers generating synthetic training data to supplement a small dataset of rare medical events. What is an important limitation to consider before relying heavily on this synthetic data?

Explanation

Synthetic data is only as good as the process used to generate it, and it can fail to capture rare, complex, or unexpected real-world patterns that were not anticipated or modeled during its creation, which can limit how well a model trained on it generalizes to real-world data. This is an important limitation to weigh against the potential cost and availability benefits synthetic data can offer. Models trained substantially on synthetic data should still be validated against real-world data before being trusted for high-stakes decisions.

Submit

10. A healthcare data science team needs to share a dataset containing patient records with an external research partner, while ensuring individual patients cannot be re-identified from the shared data. What practice addresses this requirement?

Explanation

Anonymizing sensitive data removes or transforms personally identifiable information so that individuals cannot reasonably be re-identified from the shared dataset, which is essential when sharing data containing PII with external parties. This is a compliance and privacy requirement that applies regardless of how useful or well-engineered the dataset's features are otherwise. Data augmentation and feature engineering address model performance and data richness, not privacy protection, and would not on their own satisfy a re-identification concern.

Submit

11. A fraud detection dataset has only 2% positive fraud cases. To address this class imbalance, a data scientist generates new synthetic minority-class examples by interpolating between existing minority-class observations, using a well-known technique abbreviated ____.

Explanation

SMOTE, the Synthetic Minority Oversampling Technique, generates new synthetic examples of the minority class by interpolating between existing minority-class observations and their nearest neighbors, rather than simply duplicating existing minority examples. This helps the model learn a more generalizable decision boundary for the minority class compared to naive oversampling, which can lead to overfitting on repeated identical examples. SMOTE is one of several mitigations for class imbalance, alongside undersampling the majority class and adjusting the model's decision threshold.

Submit

12. A data scientist runs k-means clustering with several different values of k and wants a metric that measures how similar each point is to its own cluster compared to other clusters, to help choose the best number of clusters. Which metric is designed for this purpose?

Explanation

The silhouette score measures how similar each point is to its own cluster compared to the next nearest cluster, producing a value that helps evaluate cluster cohesion and separation across different values of k. Higher average silhouette scores generally indicate better-defined clusters. R-squared and F1 score are metrics designed for regression and classification tasks respectively, not for evaluating unsupervised clustering quality.

Submit

13. A neural network's hidden layers use an activation function that outputs zero for any negative input and passes positive input through unchanged, which helps mitigate the vanishing gradient problem compared to older activation functions. Which activation function is this?

Explanation

ReLU outputs zero for negative inputs and passes positive inputs through unchanged, which is computationally simple and helps mitigate the vanishing gradient problem that older saturating activation functions like sigmoid and tanh can suffer from in deep networks. Sigmoid and tanh both saturate at their extremes, meaning their gradients approach zero for very large or very small inputs, which can slow or stall learning in deep architectures. Softmax is typically used in an output layer for multiclass classification to produce a probability distribution, not as a general hidden-layer activation function.

Submit

14. A data scientist runs an A/B test and obtains a p value of 0.03 for the difference in conversion rate between two variants, using a significance threshold of 0.05. What is the correct interpretation?

Explanation

A p value represents the probability of observing a result at least as extreme as the one obtained, assuming the null hypothesis is true, not the probability that the null hypothesis itself is true. Since 0.03 is below the conventional 0.05 threshold, the result is considered statistically significant at that threshold, meaning the observed difference is unlikely to be due to chance alone. Statistical significance does not by itself confirm practical or business significance, which requires considering effect size as well.

Submit

15. A data scientist wants a linear regression technique that not only shrinks coefficient estimates but can also drive some coefficients exactly to zero, effectively performing automatic feature selection. Which technique fits this description?

Explanation

LASSO applies an L1 penalty to the regression coefficients, which can shrink some coefficients all the way to zero, effectively removing those features from the model and performing automatic feature selection alongside regularization. Ordinary least squares applies no penalty at all and will not perform this kind of selection. Ridge regression, by contrast, uses an L2 penalty that shrinks coefficients toward zero but typically does not set them exactly to zero, which is the key distinction from LASSO.

Submit

16. A decision tree model achieves 99% accuracy on its training data but only 68% accuracy on a held-out test set. What does this pattern most likely indicate?

Explanation

A large gap between very high training accuracy and much lower test accuracy is a classic sign of overfitting, where the model has captured noise and idiosyncrasies specific to the training data rather than the underlying general pattern. Underfitting would instead show poor performance on both the training and test sets, since the model would be too simple to capture the pattern well anywhere. Reducing model complexity, applying regularization, or pruning the tree are common ways to address this.

Submit

17. A data team is preparing a public-facing dashboard and wants to ensure it is usable by colorblind viewers and screen reader users. Which two practices support this goal? (Select two.)

Explanation

Choosing colorblind-safe palettes ensures that charts remain interpretable for viewers with common forms of color vision deficiency, who may not be able to distinguish red from green in a default palette. Adding descriptive tags or alt text allows screen readers to convey chart content to users who cannot see the visual directly, which is an important accessibility consideration for public-facing reports. Relying exclusively on red and green, removing labels, or shrinking fonts would each work against accessibility rather than support it.

Submit

18. A data scientist proposes a new customer churn model and wants to demonstrate its value before recommending it for production use. What is the most appropriate first comparison to make?

Explanation

Benchmarking a new model against a simple baseline, such as a naive majority-class predictor or the conventional process currently in place, establishes whether the added complexity actually provides meaningful improvement. A model that only slightly outperforms a simple baseline may not justify the added complexity, maintenance burden, or interpretability loss it introduces. Comparing only against itself under different random seeds says nothing about whether the model adds real value over existing practice.

Submit

19. A team wants to find the combination of learning rate and tree depth that produces the best validation performance for a gradient boosting model, by exhaustively testing every combination from a predefined set of values for each parameter. What technique are they using?

Explanation

Grid search exhaustively evaluates every combination of hyperparameter values from a predefined set for each parameter, which guarantees the best combination within that grid is found but can become computationally expensive as the number of parameters or values increases. Random search instead samples a limited number of random combinations, which can be more efficient when the search space is large but does not exhaustively test every combination. Bootstrapping and cross-terms are unrelated techniques from resampling and feature engineering respectively.

Submit

20. A dataset contains a categorical variable for product color with no inherent order, and a separate ordinal variable for customer satisfaction rated low, medium, and high. Which two encoding approaches are appropriate for these two variables, respectively? (Select two.)

Explanation

One-hot encoding creates a separate binary column for each category, which avoids implying any false numeric relationship between categories that have no natural order, such as product colors. Label encoding that preserves the meaningful order of an ordinal variable, such as mapping low, medium, and high to 1, 2, and 3, retains that ordering information in a way the model can use appropriately. Applying label encoding with an arbitrary order to an unordered variable would incorrectly imply a ranking that does not exist in the underlying data.

Submit

21. In a linear regression model, two predictor variables are found to be highly correlated with each other, causing unstable and difficult-to-interpret coefficient estimates even though the model's overall predictive accuracy remains reasonable. What issue does this describe?

Explanation

Multicollinearity occurs when predictor variables are highly correlated with one another, which makes it difficult for the model to separate each variable's individual effect, leading to unstable and hard-to-interpret coefficients even if overall prediction accuracy is not severely harmed. Detecting this often involves checking the variance inflation factor or a correlation matrix among predictors. Seasonality and non-stationarity instead describe patterns over time in a single series, which is a different kind of data issue.

Submit

22. A data scientist wants a chart that displays the median, interquartile range, and potential outliers of a continuous variable in a single compact visual. Which chart type fits this need?

Explanation

A box and whisker plot compactly displays the median, the interquartile range as a box, whiskers extending to a defined range, and individual points beyond that range flagged as potential outliers, all in one visual. A Sankey diagram instead visualizes flow between categories, and a heat map visualizes the magnitude of values across two dimensions using color, neither of which is designed to summarize a single variable's distribution and outliers this way.

Submit

23. In principal component analysis, the directions of maximum variance in the data correspond to the ____ of the covariance matrix, while the amount of variance explained along each direction corresponds to the associated ____.

Explanation

Eigenvectors of the covariance matrix point in the directions along which the data varies the most, and each eigenvector's corresponding eigenvalue quantifies how much variance is captured along that direction. PCA ranks these eigenvector-eigenvalue pairs by eigenvalue magnitude and keeps the top ones as principal components, since they capture the most variance with the fewest dimensions. This is the mathematical foundation that lets PCA reduce dimensionality while retaining as much information as possible.

Submit

24. A support team wants to model the number of tickets that arrive per hour, where events occur independently at a constant average rate and the count is always a non-negative integer. Which distribution is most appropriate?

Explanation

The Poisson distribution models the count of independent events occurring at a constant average rate over a fixed interval, which fits ticket arrivals per hour well, since counts are non-negative integers and events happen independently. A normal distribution models continuous, symmetric data and is not naturally suited to modeling discrete event counts. A uniform distribution instead assumes every outcome in a range is equally likely, which does not reflect the clustering pattern typical of arrival counts.

Submit

25. A spam classifier occasionally flags legitimate emails as spam and occasionally lets actual spam through undetected. Which two statements correctly describe these two error types? (Select two.)

Explanation

A Type I error is a false positive, incorrectly rejecting a true null hypothesis, which corresponds to flagging a legitimate email as spam when it is actually not spam. A Type II error is a false negative, failing to reject a false null hypothesis, which corresponds to spam passing through undetected as if it were legitimate. These two error types generally trade off against each other, so tightening a classifier's threshold to reduce one often increases the other.

Submit
×
Saved
Thank you for your feedback!
View My Results
Cancel
  • All
    All (25)
  • Unanswered
    Unanswered ()
  • Answered
    Answered ()
A team is choosing between two ensemble tree-based approaches: one...
A team training an image classifier has a limited training set and...
A search engine wants to weigh words that are frequent within a...
A logistics company wants to find the shortest possible route visiting...
A computer vision model needs to run directly on cameras installed in...
A team wants to validate a newly trained recommendation model against...
Two datasets need to be merged based on customer name, but names are...
A fraud detection system needs to score transactions the instant they...
A team considers generating synthetic training data to supplement a...
A healthcare data science team needs to share a dataset containing...
A fraud detection dataset has only 2% positive fraud cases. To address...
A data scientist runs k-means clustering with several different values...
A neural network's hidden layers use an activation function that...
A data scientist runs an A/B test and obtains a p value of 0.03 for...
A data scientist wants a linear regression technique that not only...
A decision tree model achieves 99% accuracy on its training data but...
A data team is preparing a public-facing dashboard and wants to ensure...
A data scientist proposes a new customer churn model and wants to...
A team wants to find the combination of learning rate and tree depth...
A dataset contains a categorical variable for product color with no...
In a linear regression model, two predictor variables are found to be...
A data scientist wants a chart that displays the median, interquartile...
In principal component analysis, the directions of maximum variance in...
A support team wants to model the number of tickets that arrive per...
A spam classifier occasionally flags legitimate emails as spam and...
play-Mute sad happy unanswered_answer up-hover down-hover success oval cancel Check box square blue
Alert!