CompTIA DataAI DY0-001 (V1) Exam Practice Test 4

Reviewed by Editorial Team
The ProProfs editorial team is comprised of experienced subject matter experts. They've collectively created over 10,000 quizzes and lessons, serving over 100 million users. Our team includes in-house content moderators and subject matter experts, as well as a global network of rigorously trained contributors. All adhere to our comprehensive editorial guidelines, ensuring the delivery of high-quality content.
Learn about Our Editorial Process
| By Thames
T
Thames
Community Contributor
Quizzes Created: 11201 | Total Attempts: 9,875,275
| Questions: 25 | Updated: Sep 28, 2026
Please wait...
Question 1 / 26
🏆 Rank #-- ▾
0 %
0/100
Score 0/100

1. A market basket analysis finds that customers who buy bread also frequently buy butter. Which two statements about the metrics used to evaluate this association rule are correct? (Select two.)

Explanation

Support measures how frequently a given itemset, such as bread and butter appearing together, occurs across all transactions in the dataset, providing a sense of how common the pattern is overall. Lift compares the observed co-occurrence of the two items against what would be expected if they were purchased completely independently, with a lift value greater than 1 indicating a positive association stronger than chance alone would predict. Confidence instead measures how often butter is purchased specifically among transactions that already include bread, which is a distinct metric from support and lift, each capturing a different aspect of the association rule's strength.

Submit
Please wait...
About This Quiz
CompTIA DataAI Dy0-001 (V1) Exam Practice Test 4 - Quiz

This practice resource focuses on the CompTIA DataAI DY0-001 exam, evaluating your knowledge in data analysis, AI principles, and data management. It's essential for those preparing for a career in data science, helping you understand key concepts and skills necessary for success in the field.

2.

What first name or nickname would you like us to use?

You may optionally provide this to label your report, leaderboard, or certificate.

2. A team is designing an agent that learns to play a game by receiving rewards and penalties based on its actions over time, and separately wants a fast, rule-of-thumb approach to quickly find a reasonably good, though not necessarily optimal, solution to a complex scheduling problem. Which two approaches, respectively, address these two goals? (Select two.)

Explanation

Reinforcement learning trains an agent to learn a policy of actions by receiving rewards and penalties over time, which is exactly the mechanism needed for an agent learning to play a game through trial and error. A heuristic instead provides a practical, rule-of-thumb approach that quickly finds a reasonably good solution to a complex problem like scheduling, without the computational cost of guaranteeing a truly optimal solution, which is often an acceptable and necessary tradeoff for problems too large to solve exactly in reasonable time. These two approaches address fundamentally different problem types, sequential decision-making under reward feedback versus fast approximate problem-solving, so neither one is interchangeable with the other.

Submit

3. A text preprocessing pipeline reduces the words 'running,' 'runs,' and 'ran' all down to their common base form, but one approach does this crudely by chopping off suffixes while the other does it more carefully using vocabulary and grammatical rules to arrive at a proper dictionary base form. Which two terms describe these two approaches, respectively?

Explanation

Stemming crudely chops off word suffixes using simple rules, often producing a truncated form that may not even be a valid word, while lemmatization uses vocabulary and grammatical analysis to reduce a word to its proper dictionary base form, or lemma. Both aim to reduce word variations to a common base form to help downstream text analysis treat related word forms consistently, but lemmatization is generally more accurate at the cost of additional computational complexity. Choosing between them often depends on the specific application's need for speed versus linguistic precision.

Submit

4. A cloud provider must decide how to distribute a limited pool of GPU resources across many competing customer workloads to maximize overall utility, subject to each customer's minimum guaranteed capacity. What type of problem is this?

Explanation

Resource allocation optimization problems involve distributing a limited pool of resources, such as GPU capacity, across competing demands to maximize some overall objective, subject to constraints like minimum guaranteed capacity for each customer. This is a classic application of constrained optimization techniques within data science and operations research. Named-entity recognition, clustering, and time series forecasting instead address text analysis, grouping, and temporal prediction tasks respectively, none of which directly addresses this resource distribution problem.

Submit

5. A large-scale machine learning training job is distributed across many interconnected machines working together as a single coordinated computing resource, rather than running on a single powerful machine alone. This deployment approach is called ____ deployment.

Explanation

Cluster deployment distributes computation across many interconnected machines working together as a coordinated resource, which is often necessary for training very large models or processing datasets too large to fit or process efficiently on a single machine alone. This differs from edge deployment, which instead pushes computation out to individual, often resource-constrained devices, representing essentially the opposite end of the deployment spectrum in terms of centralization and available compute power. Choosing cluster deployment introduces its own complexity around distributing work efficiently and handling communication between the machines involved.

Submit

6. A data science team wants every code change to automatically trigger tests, retrain a model on updated data, and deploy the new model version if all checks pass, without manual intervention at each step. What practice does this represent?

Explanation

A CI/CD pipeline automates the sequence of testing, retraining, and deploying a model whenever code or data changes trigger the pipeline, removing manual, error-prone steps and enabling much faster, more consistent iteration compared to fully manual processes. Applying CI/CD principles, originally developed for general software engineering, to data science and machine learning workflows is a core MLOps practice that helps teams reliably ship model updates. A single manual deployment performed infrequently would not provide the continuous, automated feedback loop that CI/CD is specifically designed to enable.

Submit

7. A data quality audit finds that one specific data entry clerk consistently mistypes a particular field in a predictable, recurring way, while other apparent errors in the dataset appear random and isolated with no consistent pattern. What best distinguishes these two categories of error?

Explanation

A systematic error follows a consistent, recurring, and often predictable pattern, such as one specific clerk consistently mistyping a field in the same way, which means the error can potentially be identified and corrected across every affected record once the pattern is recognized. An idiosyncratic error instead appears random and isolated, with no consistent underlying pattern connecting the individual mistakes, making each one harder to detect systematically and often requiring case-by-case review. Recognizing which category an error falls into helps determine whether a targeted, pattern-based correction can be applied broadly or whether each instance needs individual investigation.

Submit

8. A data warehouse team wants recent transactional data updated every hour for near-real-time dashboards, while data older than two years should be moved to lower-cost, less frequently accessed storage. Which two concepts, respectively, describe these two practices? (Select two.)

Explanation

Refresh cycles define how frequently data is updated to remain current, which directly addresses the need for hourly updates supporting near-real-time dashboards. Archiving moves older, less frequently accessed data to lower-cost storage tiers, which reduces ongoing storage costs while still preserving the data for occasional access if needed later, addressing the two-year-plus data retention requirement. Confusing these two concepts would misdirect effort, since a refresh cycle problem concerns update frequency for current data, while an archiving problem concerns storage cost and retention for older data.

Submit

9. A manufacturing plant equips its machinery with devices that continuously record temperature, vibration, and pressure readings for later predictive maintenance analysis. What type of generated data does this represent?

Explanation

Sensor data is generated by automated measurement devices, such as those recording temperature, vibration, and pressure on manufacturing equipment, typically continuously and without direct human input at the point of generation. This differs from survey data, which is deliberately collected by asking people questions, and administrative data, which arises as a byproduct of routine record-keeping processes such as billing or registration. Understanding a dataset's generation process helps anticipate its likely quality issues, such as sensor drift or missing readings during equipment downtime, which are common challenges specific to sensor-generated data.

Submit

10. Before building any model, a data scientist meets with stakeholders to understand the underlying business problem, clarify what decision the model's output will actually inform, and identify any constraints on how the model can be used. What phase of the process does this describe?

Explanation

Requirements gathering occurs early in the process and focuses on understanding the underlying business problem, the decision the model's output will inform, and any relevant constraints, which together shape what kind of technical solution is actually appropriate. Skipping or rushing this phase risks building a technically impressive model that does not actually address the real business need or fit within practical constraints like latency or interpretability requirements. Hyperparameter tuning, deployment, and cross-validation all occur later, once the underlying problem and appropriate solution approach are already reasonably well understood.

Submit

11. To detect multicollinearity among predictors in a regression model, a data scientist calculates a statistic for each predictor that quantifies how much the variance of its coefficient estimate is inflated due to correlation with other predictors, commonly abbreviated ____.

Explanation

The variance inflation factor quantifies how much a predictor's coefficient variance is inflated due to its correlation with other predictors in the model, with higher VIF values indicating more severe multicollinearity for that specific predictor. A commonly used rule of thumb flags VIF values above 5 or 10 as concerning, though the appropriate threshold can depend on the specific modeling context. Addressing high VIF values might involve removing or combining highly correlated predictors, or applying a regularization technique that can better handle correlated features.

Submit

12. A recommendation system decomposes a large, sparse user-item rating matrix into three smaller matrices whose product approximates the original, capturing latent factors underlying user preferences and item characteristics. Which technique performs this kind of matrix decomposition?

Explanation

Singular value decomposition factors a matrix into three component matrices whose product approximates the original, which is widely used in recommendation systems to uncover latent factors underlying user preferences and item characteristics from a large, sparse user-item rating matrix. This factorization can reduce dimensionality while retaining the most important underlying structure in the data, often enabling more efficient and sometimes more accurate recommendations than working with the raw sparse matrix directly. k-nearest neighbors and confusion matrices instead serve classification and evaluation purposes, not matrix factorization.

Submit

13. A modern language model processes an entire input sequence at once, using a mechanism that lets each token weigh the relevance of every other token in the sequence directly, rather than processing tokens strictly one at a time in order. What architecture does this describe?

Explanation

Transformers use a self-attention mechanism that allows each token in a sequence to directly weigh its relevance to every other token, processing the sequence largely in parallel rather than strictly one token at a time as recurrent architectures do. This design has proven especially effective for large-scale language modeling, since it captures long-range dependencies more directly and trains more efficiently in parallel compared to sequential recurrent processing. A basic recurrent neural network instead processes tokens sequentially, which can make capturing very long-range dependencies more difficult and training less parallelizable.

Submit

14. A decision tree algorithm evaluates candidate splits using a measure of impurity that ranges from 0 for a perfectly pure node to a maximum value when classes are evenly mixed, computed as one minus the sum of squared class probabilities. Which measure is this?

Explanation

The Gini index measures node impurity as one minus the sum of squared class probabilities within that node, reaching zero when a node is perfectly pure and its maximum when classes are evenly split. Decision tree algorithms like CART commonly use the Gini index to evaluate and choose the best split at each node, though entropy-based information gain is another common alternative impurity measure. R2 and the F statistic instead apply to regression model fit and significance testing, not to evaluating classification tree splits.

Submit

15. A model predicts the probability of a binary outcome by modeling the log-odds of the outcome as a linear combination of predictors, then transforming that value back into a probability between 0 and 1. What is this linear combination of predictors, before transformation, commonly called?

Explanation

The logit is the log-odds of the outcome, modeled as a linear combination of predictors in logistic regression, and it is transformed through the logistic function to produce a probability bounded between 0 and 1. This is the core mechanism that allows logistic regression to model a binary outcome using a linear combination of predictors while still producing valid probability outputs. Related models, such as probit regression, use a different transformation function but address the same general goal of modeling binary outcomes with predictors.

Submit

16. A model designed to predict loan default accidentally includes a feature that is only recorded after the loan has already defaulted, resulting in unrealistically high test accuracy that collapses once the model is deployed against genuinely new applications. What issue does this describe?

Explanation

Data leakage occurs when information that would not actually be available at the time a real prediction needs to be made is inadvertently included in the training process, which artificially inflates offline performance metrics that then fail to hold up once the model is deployed against genuinely new, real-time applications. In this scenario, a feature recorded only after the loan already defaulted obviously would not be available before the outcome is known, making it a clear case of leakage rather than a genuinely predictive feature. Carefully auditing feature availability timing relative to the actual prediction moment is essential to catch this kind of issue before deployment.

Submit

17. A team wants every function in a shared codebase to include an inline explanation of its purpose, parameters, and return value, formatted in a way that documentation generation tools can automatically extract and render. This inline documentation convention is commonly called a ____.

Explanation

A docstring is an inline documentation convention placed directly within code, typically describing a function's purpose, its parameters, and its return value, in a format that documentation generation tools can automatically parse and render into readable reference documentation. This differs from a data dictionary, which documents the meaning of dataset fields rather than the behavior of code. Consistent docstring usage across a codebase makes it significantly easier for new team members or future maintainers to understand what a given function does without having to read through its entire implementation.

Submit

18. Before recommending a churn model for production, a data scientist reports its precision, recall, and business-relevant cost savings estimate on a final held-out test set that was not used during any part of model development or tuning. What is this step called?

Explanation

Reporting final performance measures on a held-out test set that played no role in model development or hyperparameter tuning provides an honest, unbiased estimate of how the model is expected to perform on new, unseen data, which is essential evidence to support a final model recommendation. Using a test set that was inadvertently involved in tuning decisions would produce an overly optimistic performance estimate, since the model may have been indirectly fit to that data. Feature engineering and hyperparameter tuning instead occur earlier in the workflow, before this final, unbiased evaluation step.

Submit

19. A data science team runs dozens of training experiments with different hyperparameters and architectures, and wants a systematic way to record each experiment's configuration, metrics, and artifacts so results can be compared and reproduced later. What practice addresses this need?

Explanation

Experiment tracking systematically records each training run's configuration, such as hyperparameters and architecture choices, along with resulting metrics and any saved artifacts, which allows a team to compare many experiments consistently and reproduce a specific result later if needed. Without this discipline, teams running many experiments risk losing track of which configuration produced which result, making it difficult to confidently select and reproduce the best model. Literature review instead happens earlier in the process, informing which approaches to try in the first place rather than tracking the experiments once they are running.

Submit

20. A data scientist wants to convert a continuous age variable into discrete age groups for a simpler categorical model, and separately wants to create a debt-to-income ratio feature from two existing numeric columns. Which two techniques, respectively, describe these transformations? (Select two.)

Explanation

Binning converts a continuous variable into a set of discrete categories or ranges, which is exactly what is needed to convert continuous age values into discrete age groups. Creating a ratio feature, such as debt-to-income, combines two existing numeric columns into a single new derived feature that can capture a relationship the two original columns did not directly express on their own. Geocoding instead converts location information like addresses into geographic coordinates, and pivoting reshapes data between long and wide formats, neither of which describes binning or ratio creation.

Submit

21. A data scientist tries to join daily website traffic data with monthly revenue figures directly, without first aggregating one to match the other's time resolution. What issue does this represent?

Explanation

Granularity misalignment occurs when datasets being combined have mismatched levels of resolution, such as daily versus monthly data, which must be reconciled, typically by aggregating the finer-grained data up to match the coarser one, before a meaningful join or analysis can be performed. Attempting to join mismatched granularities directly without this reconciliation step can produce nonsensical or misleading results. Multicollinearity and seasonality instead describe relationships between predictors and recurring time patterns respectively, which are unrelated to this specific structural mismatch.

Submit

22. A data scientist wants to visualize the intensity of website clicks across different regions of a webpage layout, using color to represent click density at each location. Which visualization is best suited to this?

Explanation

A heat map uses color intensity to represent the magnitude of a value across a two-dimensional space, which is exactly suited to visualizing click density across different regions of a webpage layout. A box and whisker plot instead summarizes the distribution of a single continuous variable, and a Q-Q plot compares a variable's distribution against a theoretical distribution, neither of which is designed to show spatial intensity across a layout. A Sankey diagram visualizes flow between categories, which is also a different visualization goal from spatial click density.

Submit

23. During backpropagation in a neural network, gradients of the loss function with respect to weights in earlier layers are computed by successively multiplying the derivatives of each intermediate function in the network, a calculus technique known as the ____ rule.

Explanation

The chain rule allows the derivative of a composed function to be computed as the product of the derivatives of its individual components, which is exactly what backpropagation relies on to compute how the loss changes with respect to weights in early layers, even though those weights only affect the loss indirectly through many intermediate layers. Without the chain rule, efficiently computing these gradients across many layers would not be tractable. This calculus foundation underlies the training of essentially all modern deep learning architectures.

Submit

24. A survey on income finds that the highest earners are disproportionately likely to skip the income question altogether, meaning the probability of a value being missing depends on the actual unobserved value itself. What type of missingness does this describe?

Explanation

Missing not at random occurs when the probability that a value is missing depends on the unobserved value itself, such as high earners specifically being more likely to skip an income question because of the value they would have reported. This is the most challenging type of missingness to handle correctly, since standard imputation methods that assume missingness is unrelated to the true value can introduce bias. Missing completely at random and missing at random instead describe cases where missingness is unrelated to the unobserved value, or depends only on other observed variables, respectively.

Submit

25. A team evaluating a highly imbalanced binary classifier wants a single summary metric that accounts for all four cells of the confusion matrix, remaining informative even when classes are heavily imbalanced. Which two statements about the Matthews Correlation Coefficient (MCC) are correct? (Select two.)

Explanation

MCC incorporates all four cells of the confusion matrix, true positives, true negatives, false positives, and false negatives, into a single balanced correlation-like measure between predicted and actual classifications. It ranges from negative one, indicating total disagreement, to positive one, indicating perfect prediction, with zero representing performance no better than random guessing, and it remains meaningful even under significant class imbalance where accuracy alone can be misleading. This is why MCC is often preferred over accuracy specifically for imbalanced classification problems.

Submit
×
Saved
Thank you for your feedback!
View My Results
Cancel
  • All
    All (25)
  • Unanswered
    Unanswered ()
  • Answered
    Answered ()
A market basket analysis finds that customers who buy bread also...
A team is designing an agent that learns to play a game by receiving...
A text preprocessing pipeline reduces the words 'running,' 'runs,' and...
A cloud provider must decide how to distribute a limited pool of GPU...
A large-scale machine learning training job is distributed across many...
A data science team wants every code change to automatically trigger...
A data quality audit finds that one specific data entry clerk...
A data warehouse team wants recent transactional data updated every...
A manufacturing plant equips its machinery with devices that...
Before building any model, a data scientist meets with stakeholders to...
To detect multicollinearity among predictors in a regression model, a...
A recommendation system decomposes a large, sparse user-item rating...
A modern language model processes an entire input sequence at once,...
A decision tree algorithm evaluates candidate splits using a measure...
A model predicts the probability of a binary outcome by modeling the...
A model designed to predict loan default accidentally includes a...
A team wants every function in a shared codebase to include an inline...
Before recommending a churn model for production, a data scientist...
A data science team runs dozens of training experiments with different...
A data scientist wants to convert a continuous age variable into...
A data scientist tries to join daily website traffic data with monthly...
A data scientist wants to visualize the intensity of website clicks...
During backpropagation in a neural network, gradients of the loss...
A survey on income finds that the highest earners are...
A team evaluating a highly imbalanced binary classifier wants a single...
play-Mute sad happy unanswered_answer up-hover down-hover success oval cancel Check box square blue
Alert!