Data quality problems are rarely solved by deleting rows that look strange. The current CompTIA Data+ DA0-002 blueprint includes data acquisition and preparation, analysis, governance, automated data quality monitoring, profiling, quality metrics, testing, and data health checks. The practical goal is to determine whether data is fit for the decision it is supposed to support.
Validation checks known rules; profiling discovers what is actually in the data. Those activities work together with lineage, metadata, and reporting. Candidates coming from broader data fundamentals such as DP-900 should think beyond syntax: a dataset can be technically queryable and still be misleading because values are incomplete, duplicated, stale, inconsistent, or joined incorrectly.
Profile the dataset before deciding how to clean it
Profiling summarizes structure and distribution before transformation: row counts, null rates, distinct values, minimum and maximum values, frequencies, patterns, lengths, and type consistency are common starting points. Tools such as pandas for data analysis can make those checks repeatable across large datasets rather than relying on manual spreadsheet inspection. In data quality, validation, and profiling, profile the dataset before deciding how to clean it is useful only when the evidence supports the chosen control rather than when the feature merely exists. Profile a representative snapshot and record the extraction time so later comparisons use the same data boundary.
Look for unexpected categories, impossible ranges, sudden cardinality changes, unusual text patterns, and columns whose observed values do not match their documented meaning. A customer-status column with 98 spellings may indicate free-text drift, multiple source systems, or an undocumented code migration. For profile the dataset before deciding how to clean it, the technician should capture an observable result and use that result to decide the next step instead of relying on habit. Do not normalize values until the business meaning of each variation is understood.
Compare profiles across partitions, source systems, regions, or time periods to identify local defects that disappear in aggregate metrics. A global null rate of one percent can hide a new source that is missing a critical field in every record. Within data quality, validation, and profiling, this profile the dataset before deciding how to clean it decision should narrow uncertainty by comparing the starting state with the expected state after one justified action. Data quality should be inspected at the granularity where decisions and pipelines actually operate.
Define quality dimensions in business terms
Completeness asks whether required values are present; validity asks whether they conform to allowed rules; accuracy asks whether they represent reality; uniqueness asks whether one entity is represented more than intended. Consistency examines agreement across fields or systems, while timeliness asks whether data is fresh enough for the use case. The define quality dimensions in business terms workflow is strongest when another technician can reconstruct why the action followed from the evidence and what would have changed the decision. Not every dataset needs the same target for every dimension.
Translate quality expectations into measurable rules tied to business use. An optional middle name can be incomplete without harming analysis, while a missing transaction amount may make a financial record unusable. Scope matters during define quality dimensions in business terms: a technically correct action applied to the wrong user, host, data set, or policy can create a second incident. Avoid generic goals such as ‘100% clean data’ when the requirement is actually a set of field-specific tolerances.
Assign ownership for the rule and the data domain so a failed check has somewhere to go. Technical teams can detect an invalid product code, but the business data owner may be the only person who can decide whether the code should be corrected, retired, or mapped. For define quality dimensions in business terms, treat the observed result as evidence; if it contradicts the working explanation, revise the explanation before changing more of the environment. Quality without ownership becomes a dashboard of unresolved warnings.
Validate types, ranges, formats, and domain rules
Type validation confirms that a field can be interpreted as the intended number, date, boolean, identifier, or category rather than being stored merely as text. Range checks can reject impossible values such as negative quantities where the business model forbids them. A safe validate types, ranges, formats, and domain rules implementation leaves a verification path so rollback, escalation, or peer review can occur without guessing what was changed. Preserve rejected records or error details instead of silently coercing values that fail parsing.
Format rules help with structured identifiers, postal codes, dates, account numbers, and other values whose syntax matters. Regular expressions can identify patterns, but they do not prove the value is authentic or semantically valid. Good validate types, ranges, formats, and domain rules practice separates a visible symptom from its underlying cause, because removing one error is not enough when the condition that produced it remains. Separate syntax validation from reference checks against authoritative systems.
Cross-field rules validate relationships: an end date should not precede a start date, a refunded transaction may require a refund reason, and a country may constrain valid region codes. These checks often catch errors that column-level profiling misses. During validate types, ranges, formats, and domain rules, each step should either confirm or eliminate a plausible cause so later investigators do not inherit an unexplained sequence of changes. Document the business rule so future analysts know why a record failed.
Handle missing, duplicate, and outlier data deliberately
Missing values can mean unknown, not applicable, not collected, not yet available, or lost during processing. Those meanings have different consequences, so replacing every null with zero or an average can distort analysis. Document the handle missing, duplicate, and outlier data deliberately decision point clearly enough that a temporary workaround cannot quietly become an undocumented permanent configuration. Use imputation only when the method and its effect on downstream interpretation are understood.
Duplicate detection should define what makes two records the same entity or event. Exact row duplication is easy to detect, while near-duplicate customer records may require identifiers, normalized names, addresses, and time context. Do not delete duplicates until you know whether repeated rows represent an error or legitimate repeated transactions.
Outliers can be data-entry mistakes, sensor failures, rare but valid events, or the most important observations in the dataset. Use statistical and domain context together before capping, excluding, or transforming them. Where controls overlap during handle missing, duplicate, and outlier data deliberately, verify which layer owns the behavior before editing settings because one control can mask the effect of another. Record exclusions so a surprising result can be traced back to the cleaning decision that changed the population.
Validate joins and referential relationships
Many data-quality failures appear only after tables are joined: unmatched keys, one-to-many explosions, duplicated reference rows, stale identifiers, and different code systems can all distort counts. Compare row counts and match rates before and after the join. In data quality, validation, and profiling, validate joins and referential relationships is useful only when the evidence supports the chosen control rather than when the feature merely exists. A query that runs successfully can still create a numerically wrong dataset.
Referential integrity checks ask whether foreign keys point to valid parent records and whether orphaned data is expected. In analytical environments, historical facts may legitimately reference a dimension member that has been retired, so the resolution should preserve history without hiding the mismatch. For validate joins and referential relationships, the technician should capture an observable result and use that result to decide the next step instead of relying on habit. Define how unknown or late-arriving members are represented.
Normalize key types, whitespace, casing, and encoding before blaming the source for failed matches. A numeric identifier stored as text in one system and integer in another can create apparent missing data. Within data quality, validation, and profiling, this validate joins and referential relationships decision should narrow uncertainty by comparing the starting state with the expected state after one justified action. Keep transformation logic visible so future pipeline changes do not reintroduce the mismatch.
Use lineage and metadata to locate the source of defects
Lineage shows where data originated, which transformations changed it, and which reports or models consume the result. When a field suddenly becomes null, lineage helps determine whether the source stopped sending it, an ingestion mapping changed, or a transformation dropped it. The use lineage and metadata to locate the source of defects workflow is strongest when another technician can reconstruct why the action followed from the evidence and what would have changed the decision. Fix defects at the earliest responsible layer when possible instead of patching every downstream report.
Metadata such as owner, refresh schedule, source of truth, definitions, units, and allowed values makes quality rules interpretable. A temperature of 30 is valid or invalid depending on unit and context. Scope matters during use lineage and metadata to locate the source of defects: a technically correct action applied to the wrong user, host, data set, or policy can create a second incident. Keep the data dictionary synchronized with the actual pipeline so documentation does not become a competing source of truth.
Automate health checks and watch for drift
Automated checks can track null rate, distinct count, schema, row volume, freshness, distribution, failed constraints, and reference match rates on every pipeline run. Set thresholds that reflect normal variability rather than alerting on every small change. A safe automate health checks and watch for drift implementation leaves a verification path so rollback, escalation, or peer review can occur without guessing what was changed. The monitoring should identify which dataset, partition, rule, and run introduced the failure.
Data drift is a change in distribution or relationship over time and may reflect real business change rather than bad data. Compare the new pattern with source changes, seasonality, product launches, and collection methods before labeling it an error. Good automate health checks and watch for drift practice separates a visible symptom from its underlying cause, because removing one error is not enough when the condition that produced it remains. Escalate unexplained drift because it can invalidate models or reports even when every row passes simple validation.
Test transformations and reports, not just raw inputs
Unit tests can validate deterministic transformation logic, while integration and requirement tests confirm that multiple stages together produce the expected result. User acceptance testing checks whether business users receive the correct interpretation and workflow, not merely whether the pipeline completed. Document the test transformations and reports, not just raw inputs decision point clearly enough that a temporary workaround cannot quietly become an undocumented permanent configuration. A green ingestion job does not prove the report is correct.
Visualization practices associated with Power BI data-analysis skills also require verifying totals, filters, relationships, and aggregation level before presentation. A polished dashboard can amplify bad data if calculations or joins are wrong. Reconcile critical metrics against a trusted source before publishing changes.
Turn quality findings into an owned remediation process
A quality report should state the rule, affected records, severity, business impact, source owner, first observed time, and recommended action. Trend repeated failures separately from one-time anomalies so chronic source problems receive structural fixes. In data quality, validation, and profiling, turn quality findings into an owned remediation process is useful only when the evidence supports the chosen control rather than when the feature merely exists. Do not measure success only by the number of rules passed; measure whether important decisions are protected from bad data.
Synthetic records and test fixtures, including controlled dummy data for testing, can help validate edge cases without corrupting production data. Use examples that cover missing values, boundary conditions, duplicates, malformed inputs, and legitimate unusual records. For turn quality findings into an owned remediation process, the technician should capture an observable result and use that result to decide the next step instead of relying on habit. Keep test data clearly separated from production so it does not become a new quality defect.