CDSP logo
Focused certification exam prep
Start practice

CDSP Exam Domains 2026: Complete Guide to All 7 Content Areas

TL;DR
  • CDSP (Certified Data Science Practitioner, CertNexus) exam DSP-210 covers 7 domains; weights are published as ranges, not fixed percentages.
  • Performing exploratory data analysis is the largest domain at 25-36%, so it deserves the biggest share of your study hours.
  • ETL (17-25%) and Building models (19-27%) together can account for nearly half the exam.
  • Testing models, Operationalizing the pipeline, and Communicating findings are each small, but they are still scored content.

How the DSP-210 Blueprint Is Built

The Certified Data Science Practitioner credential from CertNexus is assessed through exam DSP-210, delivered by Pearson VUE either at a test center or via OnVUE online proctoring. DSP-210 launched in September 2024 and replaced the earlier DSP-110, which was retired on February 28, 2025. The current blueprint is version 1.6, last modified February 3, 2025, and it organizes everything you can be tested on into seven domains.

What makes this blueprint different from many vendor exams is its shape. It follows the lifecycle of a real data science project: you define the problem, get the data, explore it, build models, test them, deploy the pipeline, and communicate results. Reading the domains in order is effectively reading a project plan. That is good news for candidates with hands-on experience, because the exam rewards people who have actually carried a project from a vague business question to a delivered result.

If you are still deciding whether to pursue the credential at all, start with What Is CDSP Certification? and the CDSP requirements overview. There are no formal prerequisites or mandatory training, though programming, statistics, and data-handling competence are recommended. This guide assumes you have decided to sit the exam and want to understand exactly what each content area demands.

Domain Weight Ranges at a Glance

CertNexus publishes each domain as a percentage range rather than a single number. That matters for planning: the ranges overlap enough that you cannot compute an exact question count per domain. Treat them as relative priorities, not quotas.

DomainPublished WeightRelative Priority
1. Defining the need to be addressed through the application of data science7-9%Low-moderate
2. Extracting, Transforming, and Loading Data17-25%High
3. Performing exploratory data analysis25-36%Highest
4. Building models19-27%High
5. Testing models4-7%Low
6. Operationalizing the pipeline5-8%Low
7. Communicating findings4-7%Low
Reading the ranges correctly: The three heavy domains (ETL, EDA, and model building) can together make up the majority of scored items. Even at their lower bounds they sum to well over half of the exam. That is why a candidate who is strong on those three but ignores the smaller domains can still pass, while a candidate who is weak on EDA almost certainly cannot.

Domain 1: Defining the Need to Be Addressed Through the Application of Data Science (7-9%)

This is the "before you touch any data" domain. It tests whether you can translate a business situation into a problem that data science can actually address, and whether you can recognize when it cannot. Questions here tend to be scenario-driven: a stakeholder describes a pain point and you must identify the right framing.

What candidates must understand

Problem framing is the skill being measured, not any particular tool.

  • Distinguishing a business question from an analytical question, and restating one as the other
  • Identifying whether a problem calls for prediction, classification, clustering, or simple description
  • Recognizing what data would be needed and whether it plausibly exists
  • Defining success in terms the stakeholder cares about, not just a model metric
  • Spotting constraints such as privacy, data availability, and feasibility early

Why this small domain trips people up

Candidates from heavily technical backgrounds often skim this domain because it feels soft. But the questions reward judgment: choosing the framing that matches the stated goal, even when a more sophisticated technique is tempting. A common trap is selecting an answer that sounds more advanced when the scenario calls for something simpler. At 7-9% you will see a modest number of these items, but they are some of the most winnable points on the exam if you read carefully.

Domain 2: Extracting, Transforming, and Loading Data (17-25%)

ETL is the second-largest block of the exam after EDA, and it reflects a reality every working practitioner knows: most project time is spent getting data into usable shape. Expect questions about acquiring data from different sources, reshaping it, handling quality problems, and storing it for downstream use.

Core ETL competencies

Think in terms of the full path from raw source to analysis-ready dataset.

  • Extracting data from files, databases, and other structured or semi-structured sources
  • Combining datasets through joins and merges, and understanding how join choices change row counts
  • Cleaning: missing values, duplicates, inconsistent formats, and outliers
  • Transforming: type conversion, encoding categorical variables, scaling, and feature derivation
  • Reshaping data between wide and long formats
  • Loading results into a form suitable for analysis or modeling

Programming fluency shows up here first

Because the credential recommends programming and data-handling competence, this is where that recommendation becomes concrete. You should be comfortable reading short code snippets and predicting their output, or choosing which operation accomplishes a described transformation. If your data wrangling skills rely on a graphical tool, practice reading the equivalent code operations. The exam is closed-book and proctored, so you cannot look up syntax during the session.

Key Takeaway

Do not treat ETL as the "boring" domain. At 17-25% it is worth more than Domains 1, 5, 6, and 7 combined at their upper bounds, and the concepts (joins, missing data, encoding) also feed directly into the model-building questions later.

Domain 3: Performing Exploratory Data Analysis (25-36%)

This is the heart of the CDSP exam and its largest domain by a clear margin. Whatever else you do, make sure your EDA knowledge is solid. A quarter to over a third of the exam sits here, which means this domain alone can decide a close result.

What EDA covers on the exam

EDA blends statistics, visualization, and interpretation. Questions ask you to read a situation and decide what the data is telling you.

  • Descriptive statistics: central tendency, spread, skewness, and what each reveals about a distribution
  • Choosing appropriate visualizations for a given variable type or relationship
  • Understanding distributions and recognizing when assumptions such as normality are violated
  • Correlation versus causation, and the limits of what a correlation coefficient can tell you
  • Detecting outliers and deciding how to treat them
  • Hypothesis testing concepts, including p-values, significance, and common misinterpretations
  • Feature selection and dimensionality reduction as outgrowths of exploration
  • Identifying data quality problems that only become visible once you look at the data

Interpretation beats memorization

The most common EDA mistake is memorizing definitions without being able to apply them. A question might describe a dataset with a long right tail and ask which summary statistic is most appropriate, or present a scenario in which a significant p-value is being misread. You need to reason, not recall. Work through real datasets and force yourself to articulate what each plot or statistic tells you and what it does not.

Statistics is the hidden prerequisite: CertNexus lists statistical competence as recommended background, and this domain is why. If your statistics is rusty, expect to spend extra time on distributions, hypothesis testing logic, and the meaning of confidence in an estimate before you go deeper into modeling. For a sense of how this affects overall difficulty, see How Hard Is the CDSP Exam?

Domain 4: Building Models (19-27%)

The third heavy domain covers the machine learning work most people associate with data science. You are expected to know which family of algorithm fits which type of problem, how to prepare data for training, and how to train and tune models responsibly.

Model-building essentials

Expect to choose methods based on problem type and data characteristics.

  • Supervised learning: regression for continuous targets and classification for categorical ones
  • Unsupervised learning: clustering and dimensionality reduction
  • Splitting data into training, validation, and test sets and why leakage between them is dangerous
  • Feature engineering and its effect on model performance
  • Overfitting and underfitting, and the levers used to address each, such as regularization
  • Hyperparameter tuning and the role of cross-validation
  • Ensemble approaches and when combining models helps
  • Handling imbalanced classes

Match the method to the problem

The exam favors applied reasoning over mathematical derivation. You are more likely to be asked which approach suits a described scenario than to derive a gradient. Know the strengths and weaknesses of each model family, what assumptions they make, and what symptoms indicate a problem such as high variance. Because Domain 4 builds on EDA and ETL, weaknesses upstream will cost you here as well.

Domain 5: Testing Models (4-7%)

Testing models is among the smallest domains, but it contains concepts that are easy to confuse and therefore easy to get wrong. It is about evaluating whether a model is actually any good, and for what purpose.

Evaluation concepts to master

Metric choice is the central skill.

  • Regression metrics such as mean squared error and R-squared, and how to interpret them
  • Classification metrics: accuracy, precision, recall, F1, and the confusion matrix
  • ROC curves and AUC as threshold-independent measures
  • Why accuracy misleads on imbalanced data
  • Validating generalization versus simply measuring training performance
  • Comparing candidate models fairly

The trap here is metric selection. A scenario will describe the cost of different errors, such as a missed positive being far worse than a false alarm, and you must pick the metric that reflects that. Because this domain is only 4-7%, it is tempting to under-study it, but the concepts overlap with Domain 4 and the questions are very learnable.

Domain 6: Operationalizing the Pipeline (5-8%)

This domain asks what happens after a model works in a notebook. It tests your understanding of moving a data science solution into a repeatable, maintainable process.

Operational topics

The emphasis is on conceptual understanding of deployment and maintenance.

  • Packaging a workflow so that it can be run repeatedly and consistently
  • Monitoring model performance over time and recognizing drift
  • Deciding when to retrain or update a model
  • Documenting the pipeline so others can maintain it
  • Considering scale, latency, and integration with existing systems

Candidates who have only worked in exploratory notebooks sometimes find this domain unfamiliar. You do not need deep engineering expertise, but you do need to understand why a model that performed well at launch can degrade and what a responsible handoff looks like.

Domain 7: Communicating Findings (4-7%)

The final domain closes the loop: a result nobody understands has no value. Questions ask how to present analytical outcomes to the audience that needs them.

Communication skills tested

Audience awareness is the common thread.

  • Selecting visualizations that convey a finding clearly to a non-technical audience
  • Translating model output into business implications and recommendations
  • Stating limitations, uncertainty, and assumptions honestly
  • Avoiding misleading charts or overclaiming from the evidence
  • Tailoring the level of technical detail to the stakeholder

These items usually reward the answer that is clearest and most honest rather than the most technically impressive. They are generally approachable, and at 4-7% they are a reasonable place to collect dependable points.

Sequencing Your Prep by Domain

Rather than studying in the order the domains are listed, consider ordering by dependency and weight. The heavy domains build on each other, so a logical sequence gives you compounding returns.

Weeks 1-2

Data handling and EDA foundations

  • Work through ETL operations in code: joins, cleaning, encoding, reshaping
  • Begin EDA with descriptive statistics and distribution reading, since both domains share this groundwork
  • Skim Domain 1 problem framing while you are still warm on business context
Weeks 3-4

Deep EDA and hypothesis testing

  • Spend the most time here: EDA is up to 36% of the exam
  • Practice interpreting plots and statistical output on real datasets
  • Revisit weak statistics topics before they compound into modeling errors
Weeks 5-6

Modeling, then testing

  • Cover supervised and unsupervised methods, tuning, and overfitting
  • Study evaluation metrics immediately afterward so Domain 5 reinforces Domain 4
Week 7

Operations, communication, and full review

  • Cover Domains 6 and 7, which are lighter and quick to absorb
  • Take timed practice sets and revisit whichever domain scores lowest

For a broader plan that accounts for your own background, the CDSP study guide walks through preparation end to end, and the CDSP cheat sheet is useful for a final-days review of key facts. Once you have covered the material, test yourself under realistic conditions with the CDSP practice tests to see which domains still need work.

Exam Mechanics That Shape How You Study

Understanding the delivery format helps you prepare appropriately. DSP-210 has 90 total questions, of which 75 are scored and 15 are unscored. You will not know which are which, so every question deserves a genuine effort. The appointment is two hours, but that includes a 5-minute agreement and a 5-minute tutorial, which leaves roughly 110 minutes for the questions themselves. That works out to a little over a minute per question if you spread your time evenly.

The exam is closed-book and proctored. CertNexus materials differ on the question style: the live exam page describes multiple choice and multiple response, while blueprint v1.6 describes multiple choice and single response. This discrepancy has not been clarified by the issuer, so prepare for both. Practice reading each question carefully for instructions such as how many answers to select. Whether personal calculators are permitted, and whether delivery is adaptive, has not been verified, so confirm current policies with Pearson VUE or CertNexus before test day rather than assuming.

The passing standard is 72% or 69% depending on the equated form you receive. Note that this is a cut score, not a pass rate. For details, see the CDSP passing score breakdown and the discussion of the CDSP pass rate. On cost, the exam voucher is USD $367.50, an optional digital guide is $103.95, and a guide-plus-voucher bundle is $424.31; the full picture is in the CDSP certification cost guide. Scheduling details are covered in the CDSP exam dates article.

Think about validity from day one: The credential is valid for 3 years. To maintain it you can either earn 90 CECs and pay a $150 renewal fee, or pass the current examination before expiry. Because the blueprint was revised between DSP-110 and DSP-210, it is worth keeping your skills current enough that re-examination remains a viable path.

Frequently Asked Questions

Which CDSP domain is weighted most heavily?

Performing exploratory data analysis (Domain 3) is the largest at 25-36%. Building models (19-27%) and Extracting, Transforming, and Loading Data (17-25%) follow. Together these three make up the bulk of the scored content.

Are the domain weights exact percentages?

No. CertNexus publishes each domain as a range, so you cannot calculate an exact number of questions per domain. Use the ranges to prioritize study time rather than to predict precise question counts.

Can I skip the smaller domains and still pass?

It is risky. Testing models, Operationalizing the pipeline, and Communicating findings are each only 4-8%, but together they are still a meaningful slice, and their questions are often very learnable. Treat them as efficient points rather than optional extras.

Do I need prior experience before attempting the exam?

There are no formal prerequisites or mandatory training. However, programming, statistics, and data-handling competence are recommended, and the exam's scenario style rewards hands-on experience. See the CDSP requirements page for more on eligibility.

Where can I find a one-page summary of the key facts?

The CDSP cheat sheet condenses the must-know details, and the ROI analysis helps you decide whether the investment makes sense for your career before you commit to preparing.

Ready to pass your CDSP exam?

Put this into practice with free CDSP questions across every exam domain.