CompTIA DataAI DY0-001 (V1) Exam Practice Test 5

Reviewed by Editorial Team
The ProProfs editorial team is comprised of experienced subject matter experts. They've collectively created over 10,000 quizzes and lessons, serving over 100 million users. Our team includes in-house content moderators and subject matter experts, as well as a global network of rigorously trained contributors. All adhere to our comprehensive editorial guidelines, ensuring the delivery of high-quality content.
Learn about Our Editorial Process
| By Thames
T
Thames
Community Contributor
Quizzes Created: 11201 | Total Attempts: 9,875,275
| Questions: 25 | Updated: Sep 28, 2026
Please wait...
Question 1 / 26
🏆 Rank #-- ▾
0 %
0/100
Score 0/100

1. A team is deciding between a bagging-based ensemble and a boosting-based ensemble for a tabular prediction task with noisy labels. Which two considerations are relevant to this decision? (Select two.)

Explanation

Boosting's sequential error-correction approach can make it more sensitive to noisy labels or outliers, since the algorithm may repeatedly try to fit patterns that are actually just noise rather than genuine signal. Bagging tends to be comparatively more robust to this kind of noise, since it averages predictions across many independently trained trees rather than sequentially chasing and potentially overfitting to every residual error. Neither approach is universally superior on every dataset, which is why practitioners typically evaluate both empirically on their specific data rather than assuming one always outperforms the other.

Submit
Please wait...
About This Quiz
CompTIA DataAI Dy0-001 (V1) Exam Practice Test 5 - Quiz

This quiz focuses on the CompTIA DataAI DY0-001 (V1) exam, evaluating your understanding of data analysis, machine learning concepts, and AI fundamentals. It's designed to help learners prepare effectively for the certification, ensuring they grasp essential skills and knowledge in data science. Engaging with this content will enhance your readiness... see morefor the exam and deepen your understanding of critical topics in the field. see less

2.

What first name or nickname would you like us to use?

You may optionally provide this to label your report, leaderboard, or certificate.

2. A system reads printed text from scanned invoice images to extract structured data, and a separate system follows the same detected vehicle across consecutive video frames to estimate its speed. Which two computer vision applications, respectively, address these two goals?

Explanation

Optical character recognition converts printed or handwritten text within an image into machine-readable structured text, which is exactly the task of extracting data from scanned invoice images. Tracking instead follows a detected object, such as the same vehicle, across consecutive video frames over time, which is necessary to estimate motion characteristics like speed rather than just detecting the object once in a single static frame. These are distinct computer vision tasks solving different problems, one focused on reading text from static images and the other focused on maintaining an object's identity and position across a sequence of frames.

Submit

3. A company analyzes thousands of product reviews to automatically determine whether each one expresses a positive, negative, or neutral opinion about the product. What NLP application does this describe?

Explanation

Sentiment analysis classifies the emotional tone or opinion expressed within text, such as determining whether a product review is positive, negative, or neutral, which is a widely used NLP application for understanding customer feedback at scale. Named-entity recognition instead identifies specific entities like people or organizations mentioned in text, and topic modeling identifies latent themes across a document collection, both of which are different tasks from classifying sentiment. Text summarization condenses lengthy text into a shorter form, which again addresses a different goal from determining the emotional tone of the original text.

Submit

4. A gradient-based optimization algorithm converges to a point where the objective function's slope is zero, but further exploration reveals a different point elsewhere with an even better objective value. What does this describe?

Explanation

Getting stuck at a local maximum or minimum means the algorithm has converged to a point that is optimal within its immediate neighborhood, even though a better solution exists elsewhere in the broader search space. This is a well-known challenge for optimization over non-convex objective functions, which is common in many real-world problems, including training deep neural networks. Techniques like using multiple random starting points, adding momentum, or using more sophisticated optimizers can help reduce, though not always fully eliminate, the risk of settling for a local rather than global optimum.

Submit

5. A financial institution with strict data residency regulations chooses to run its fraud detection model entirely on servers it owns and physically controls within its own data center, rather than using any external cloud provider. This deployment approach is called ____ deployment.

Explanation

On-premises deployment runs infrastructure entirely on hardware owned and physically controlled by the organization itself, which can be an important requirement for institutions facing strict data residency or regulatory requirements that limit or prohibit using external cloud providers for certain workloads. This approach trades away some of the elasticity and convenience that cloud deployment offers in exchange for direct physical control over where data resides and how it is secured. Organizations sometimes combine on-premises deployment for their most sensitive workloads with cloud deployment for less sensitive ones, which is what constitutes a hybrid deployment approach.

Submit

6. A team runs many containerized model-serving instances that need to be automatically scaled up during high demand, restarted if they fail, and load-balanced across available compute resources. What capability manages this automatically?

Explanation

Container orchestration platforms automatically manage scaling containerized workloads up or down based on demand, restarting failed containers, and distributing load across available compute resources, all without requiring manual intervention for each of these operational events. This is essential for reliably running model-serving infrastructure at scale, where manually managing many individual containers would be impractical and error-prone. A decision tree, data dictionary, and Box-Cox transformation are unrelated modeling and documentation concepts, none of which manage containerized deployment infrastructure.

Submit

7. A data scientist needs to combine a customer table with an orders table, keeping every customer record regardless of whether they have placed any orders, and filling in missing order details with nulls for customers with no orders. Which type of join accomplishes this?

Explanation

A left join keeps every record from the left table, in this case every customer, and fills in null values for any matching fields from the right table when no match exists, which is exactly what is needed to retain customers with no orders. An inner join would instead only keep customers who do have at least one matching order, dropping customers with no orders entirely, which does not meet the stated requirement. A union stacks rows from two tables with matching columns rather than joining them side by side based on a key, which is a fundamentally different operation from a join.

Submit

8. A team training a large deep learning model needs specialized hardware optimized for the highly parallel matrix operations deep learning relies on, while a separate lightweight statistical model can train adequately on general-purpose hardware. Which two statements about infrastructure requirements are correct? (Select two.)

Explanation

GPUs and TPUs are specialized hardware architectures well suited to the highly parallel matrix and tensor operations that deep learning training relies on heavily, often dramatically reducing training time compared to general-purpose CPUs for large models. A lightweight statistical model, by contrast, often trains adequately on general-purpose CPU hardware without needing specialized accelerators at all, since its computational demands are comparatively modest. Correctly sizing infrastructure to the actual model and workload avoids both under-provisioning, which slows training unnecessarily, and over-provisioning, which wastes cost on hardware capability the workload does not need.

Submit

9. A model trained exclusively on data from large enterprise customers is being considered for use in decisions about small business customers, a segment the training data does not represent well. What consideration is most relevant here?

Explanation

The relevant range of application refers to recognizing the boundaries of the population and conditions a model was actually trained on and validated against, since applying it to a meaningfully different population, such as small business customers when it was trained only on large enterprise customers, risks poor and potentially misleading performance. This is a data science operations consideration closely related to extrapolation, but specifically framed around the business context of who or what the model's insights genuinely apply to. Recognizing this limitation before deployment helps avoid confidently but incorrectly applying a model's conclusions well outside the population it was actually built to serve.

Submit

10. A data scientist is asked to recommend whether a proposed model justifies its ongoing infrastructure and maintenance costs given its expected business impact, weighing the expected financial benefit against the resources required to build and maintain it. What kind of analysis does this represent?

Explanation

A cost-benefit analysis weighs a proposed solution's expected business impact and financial benefit against the resources, including infrastructure and ongoing maintenance costs, required to build and sustain it, which helps stakeholders make an informed go or no-go decision. This kind of analysis is an important part of translating a technically feasible model into a genuinely worthwhile business recommendation, since even a highly accurate model may not be worth pursuing if its costs substantially outweigh its expected benefit. A confusion matrix, t-test, and Q-Q plot instead address model evaluation and statistical distribution comparison, which are technical tools rather than business-value assessments.

Submit

11. A model trained on house prices for homes between 1,000 and 3,000 square feet is asked to predict the price of a 10,000 square-foot home, well outside the range of its training data. Making a prediction in this unobserved region, beyond the range of the training data, is called ____, which is generally riskier than predicting within the observed range.

Explanation

Extrapolation involves making predictions outside the range of values seen in the training data, which is generally riskier than interpolation, predicting within the observed range, since the model has no direct evidence about how the relationship behaves in that unobserved region and may simply be extending a pattern that does not actually hold there. A model that fits training data well within its observed range offers no guarantee that the same relationship continues to hold far beyond that range, especially for a house 3 to 10 times larger than anything the model has seen. Recognizing when a prediction requires extrapolation is important context for judging how much confidence to place in that specific prediction.

Submit

12. A model classifies a new data point by looking at the labels of its closest neighboring points in the training data and taking a majority vote, without assuming any particular functional form for the underlying relationship between features and the outcome. What algorithm does this describe?

Explanation

KNN classifies a new point based on the majority label among its nearest neighbors in the training data, making it a non-parametric method that does not assume any specific functional form for the relationship between features and the outcome, unlike logistic regression or linear discriminant analysis, which both assume particular parametric forms. This flexibility lets KNN capture complex, non-linear decision boundaries, but it can also make the algorithm sensitive to the choice of distance metric and the value of k, along with feature scaling. KNN's need for meaningful distance calculations is also exactly why it can suffer from the curse of dimensionality in very high-dimensional feature spaces.

Submit

13. A deep neural network trains faster and more stably after a technique is applied that normalizes the inputs to each layer within a mini-batch during training, reducing internal shifts in the distribution of layer inputs as earlier layers' weights change. What technique is this?

Explanation

Batch normalization normalizes the inputs to a layer within each mini-batch during training, which reduces the internal covariate shift caused by earlier layers' weights changing during training, generally leading to faster and more stable convergence. This is a distinct technique from dropout, which instead randomly deactivates neurons to prevent overreliance on specific units, though both are commonly used together in modern deep learning architectures. Early stopping and learning rate scheduling address when to stop training and how the learning rate changes over time respectively, which are related but separate training considerations from normalizing layer inputs.

Submit

14. A data scientist evaluates a binary classifier by plotting true positive rate against false positive rate across every possible decision threshold, then summarizes overall discriminative ability as a single number between 0 and 1. What is this evaluation approach called?

Explanation

ROC curves plot the true positive rate against the false positive rate across every possible decision threshold, giving a threshold-independent view of a classifier's discriminative ability, and the area under that curve, AUC, summarizes overall performance as a single number between 0 and 1, with 1 representing perfect separation and 0.5 representing random guessing. This differs from a confusion matrix, which reflects performance at only one specific chosen threshold rather than across the full range. AUC is particularly useful for comparing classifiers independent of any single threshold choice, which matters when the ideal operating threshold may vary by business context.

Submit

15. A data scientist wants a regularized linear regression technique that combines both an L1 penalty, which can zero out coefficients for feature selection, and an L2 penalty, which handles groups of correlated predictors more gracefully than L1 alone. Which technique combines both of these penalties?

Explanation

Elastic net combines both an L1 penalty, similar to LASSO, which can drive some coefficients to exactly zero for automatic feature selection, and an L2 penalty, similar to ridge regression, which tends to handle groups of correlated predictors more gracefully by shrinking their coefficients together rather than arbitrarily selecting just one from the group. This combination can outperform either LASSO or ridge alone in situations with many correlated predictors, offering a middle ground between the two individual approaches. Ordinary least squares, by contrast, applies no penalty at all and does not perform any feature selection or handle correlated predictors specially.

Submit

16. During training, a regression model adjusts its parameters to minimize a function that quantifies the discrepancy between predicted and actual values across the training data. What is this function called?

Explanation

A loss function quantifies the discrepancy between a model's predictions and the actual observed values, and training a model fundamentally involves adjusting its parameters to minimize this function across the training data. Different tasks and modeling goals often call for different loss functions, such as mean squared error for regression tasks focused on minimizing variance in prediction error, or cross-entropy for classification tasks. The choice of loss function directly shapes what kind of errors the model is optimized to avoid, which is why selecting an appropriate loss function for the specific business problem matters.

Submit

17. When preparing a report, a data scientist tailors the level of technical detail differently for a company's chief financial officer than for a fellow data scientist on another team, recognizing that these two audiences have different backgrounds and information needs. The CFO in this example represents a business executive stakeholder, while the fellow data scientist represents a ____ stakeholder.

Explanation

Effective communication requires recognizing that different audiences, such as business executive stakeholders, business domain stakeholders, and peer or professional stakeholders, have different backgrounds, priorities, and levels of technical fluency, which should shape how results are presented to each group. A report written appropriately for a fellow data scientist, using precise technical terminology and methodological detail, would likely overwhelm or fail to resonate with a business executive whose priorities center on strategic and financial implications instead. Tailoring the message to the specific audience without changing the underlying substance of the findings is a core communication skill for translating data science results into organizational impact.

Submit

18. A team proposes replacing a manual underwriting process with a machine learning model and wants to demonstrate the model's value by comparing its outcomes directly against the existing manual process's historical decisions, not just against a naive baseline. What does this comparison represent?

Explanation

Benchmarking against the conventional process that is currently in place provides a directly relevant comparison, since it shows whether the proposed model actually improves on what the business is doing today, which is often a more meaningful and business-relevant comparison than only beating a simplistic statistical baseline. This is distinct from comparing against a naive baseline, such as always predicting the majority class, which establishes a lower bound on acceptable performance rather than demonstrating improvement over current practice. Both types of benchmarks can be useful, but comparing against the conventional process specifically speaks to the model's real-world business value.

Submit

19. Before selecting a modeling approach for a novel forecasting problem, a data scientist surveys published research and industry case studies addressing similar problems, to avoid reinventing approaches that have already been tried and evaluated by others. What step in the model design process does this represent?

Explanation

A literature review surveys existing published research and industry case studies addressing similar problems, which helps inform model selection by revealing approaches that have already been tried, along with their reported strengths and weaknesses, before committing significant effort to a specific modeling approach. This step typically occurs early in model design, well before hyperparameter tuning or experiment tracking, which both assume a general modeling approach has already been chosen. Skipping this step risks reinventing an approach that the field has already explored, potentially missing known pitfalls or better-performing alternatives.

Submit

20. A dataset contains raw street addresses that need to be converted into latitude and longitude coordinates for mapping, and separately contains sales data in a long format that needs to be reshaped into a wide format with one column per product category. Which two techniques, respectively, address these two needs? (Select two.)

Explanation

Geocoding converts raw address text into geographic coordinates, latitude and longitude, which is exactly what is needed before addresses can be plotted or analyzed spatially on a map. Pivoting reshapes data between long and wide formats, such as turning a long format with one row per product category into a wide format with one column per category, which is a data restructuring operation entirely distinct from geocoding. Confusing these two techniques would misdirect effort, since a geocoding need requires converting text addresses into coordinates while a pivoting need requires only reshaping already-numeric or categorical data.

Submit

21. A recommendation system's user-item rating matrix has millions of possible user-item pairs, but the vast majority of cells have no rating at all, since each user has only rated a small fraction of available items. What data issue does this represent?

Explanation

Sparse data describes a situation where most entries in a matrix or dataset are empty, zero, or missing, which is extremely common in recommendation systems, since any individual user typically interacts with only a tiny fraction of all available items. Specialized techniques like matrix factorization and singular value decomposition are often specifically designed to handle this kind of sparsity efficiently, since naively storing and processing a mostly empty matrix would waste significant memory and computation. Multicollinearity and seasonality instead describe predictor correlation and recurring temporal patterns, which are unrelated issues from data sparsity.

Submit

22. A data scientist examines the joint relationship among three variables simultaneously, such as how price, square footage, and neighborhood together relate to home sale outcomes, rather than examining each variable in isolation. What type of analysis is this?

Explanation

Multivariate analysis examines relationships among multiple variables simultaneously, which captures interactions and joint patterns that would be missed if each variable were only examined in isolation through univariate analysis. Univariate analysis instead focuses on the distribution and characteristics of a single variable at a time, which is a useful but more limited starting point for exploratory data analysis. Understanding multivariate relationships is often essential before building a predictive model that will ultimately need to account for how several features jointly relate to an outcome.

Submit

23. Before applying certain machine learning algorithms to high-dimensional data, a data scientist factors a matrix into a product of simpler component matrices, such as through eigendecomposition or singular value decomposition, to reveal underlying structure or reduce dimensionality. This general family of techniques is called matrix ____.

Explanation

Matrix decomposition refers to a family of techniques, including eigendecomposition and singular value decomposition, that factor a matrix into a product of simpler component matrices, which can reveal underlying structure, such as principal directions of variance, or enable more efficient computation for downstream algorithms. These decompositions form the mathematical backbone of techniques like PCA and SVD-based dimensionality reduction discussed elsewhere in the data science workflow. Understanding matrix decomposition conceptually helps clarify why these seemingly different techniques share common mathematical machinery underneath.

Submit

24. A game costs $5 to play and pays out $20 with probability 0.2 and $0 otherwise. What is the expected value of playing this game, and is it favorable to the player?

Explanation

Expected value is calculated by multiplying each possible outcome by its probability and summing the results, which here is 0.2 times $20 plus 0.8 times $0, equaling $4. Since the expected payout of $4 is less than the $5 cost to play, the game has a negative expected value for the player, making it unfavorable on average over many repeated plays. Expected value is a foundational concept for evaluating decisions under uncertainty, even though any single play could still result in the larger payout by chance.

Submit

25. A data scientist wants to measure the strength of a linear relationship between two continuous variables, and separately wants to measure the strength of a monotonic but not necessarily linear relationship between two ranked variables. Which two correlation measures, respectively, fit these two goals? (Select two.)

Explanation

Pearson correlation specifically measures the strength and direction of a linear relationship between two continuous variables, and it can understate the strength of a relationship that is strongly monotonic but not actually linear in shape. Spearman correlation instead works on ranked data and captures monotonic relationships, meaning as one variable increases the other consistently increases or decreases, without requiring that relationship to be linear. Choosing the wrong correlation measure for the underlying relationship shape can lead to underestimating how strongly two variables are actually related.

Submit
×
Saved
Thank you for your feedback!
View My Results
Cancel
  • All
    All (25)
  • Unanswered
    Unanswered ()
  • Answered
    Answered ()
A team is deciding between a bagging-based ensemble and a...
A system reads printed text from scanned invoice images to extract...
A company analyzes thousands of product reviews to automatically...
A gradient-based optimization algorithm converges to a point where the...
A financial institution with strict data residency regulations chooses...
A team runs many containerized model-serving instances that need to be...
A data scientist needs to combine a customer table with an orders...
A team training a large deep learning model needs specialized hardware...
A model trained exclusively on data from large enterprise customers is...
A data scientist is asked to recommend whether a proposed model...
A model trained on house prices for homes between 1,000 and 3,000...
A model classifies a new data point by looking at the labels of its...
A deep neural network trains faster and more stably after a technique...
A data scientist evaluates a binary classifier by plotting true...
A data scientist wants a regularized linear regression technique that...
During training, a regression model adjusts its parameters to minimize...
When preparing a report, a data scientist tailors the level of...
A team proposes replacing a manual underwriting process with a machine...
Before selecting a modeling approach for a novel forecasting problem,...
A dataset contains raw street addresses that need to be converted into...
A recommendation system's user-item rating matrix has millions of...
A data scientist examines the joint relationship among three variables...
Before applying certain machine learning algorithms to...
A game costs $5 to play and pays out $20 with probability 0.2 and $0...
A data scientist wants to measure the strength of a linear...
play-Mute sad happy unanswered_answer up-hover down-hover success oval cancel Check box square blue
Alert!