Check
Record row and column counts before cleaning
Use this student checklist before submitting a data cleaning, EDA, machine learning, or visualization assignment.
The checklist is a pre-analysis review for common data-quality issues such as missing values, duplicates, data types, outliers, and undocumented transformations.
Record row and column counts before cleaning
Profile missingness by variable
Check unique identifiers and duplicate logic
Validate dates, categories, and numeric ranges
A customer dataset can contain the same person more than once legitimately, so duplicate detection should inspect the relevant key and context rather than automatically deleting repeated rows.
Answers are kept specific to this page so students can check requirements, method, files, and limitations without reading repeated site-wide text.
No. First determine whether they are true duplicates or valid repeated observations. Removing legitimate repeats can change the analysis.
The choice depends on why values are missing, how many are affected, variable type, and the intended analysis. Deleting or imputing without justification can bias results.
Incorrect date, categorical, or numeric types can break calculations, create wrong sorting, or cause models to treat values incorrectly.
Use it only as the assignment permits. Many courses also require the formula, working, units, assumptions, or written interpretation.
Send the assignment brief, dataset, deadline, tool requirement, and grading rubric. A clear quote can be shared after reviewing the exact task.