DataAudit is an R package for systematic data-quality
auditing and validation.
It provides tools for detecting common data-quality problems, defining reusable validation rules, generating structured audit reports, and calculating an overall data-quality score.
DataAudit supports checks for missing values, blank values, duplicates, infinite values, constant and near-zero variance variables, outliers, ranges, categories, data types, patterns, dates, identifiers, lengths, character case, whitespace, sequences, dependencies, uniqueness, and cross-variable consistency.
You can install DataAudit from GitHub with:
# install.packages("pak")
pak::pak("vinodhpmd/DataAudit")Then load the package:
library(DataAudit)Create a dataset containing several data-quality problems:
dat <- data.frame(
ID = c(
"A001",
"A002",
"A002",
"BAD"
),
Species = c(
"Dog",
"Horse",
"",
"Cat"
),
Age = c(
5,
40,
-2,
NA
),
stringsAsFactors = FALSE
)
dat## ID Species Age
## 1 A001 Dog 5
## 2 A002 Horse 40
## 3 A002 -2
## 4 BAD Cat NA
Use audit_data() for a general first-pass audit:
report <- audit_data(
dat,
include_nzv = FALSE
)
report## DataAudit Report
## ================
##
## Rows: 4
## Variables: 3
## Variables with issues: 2
## Duplicated rows: 0
## Total issues: 2
## Rule violations: 0
## Overall issues: 2
##
## Variables with issues
## ---------------------
## Variable Class Missing MissingPercent Blank Infinite Unique Constant
## Species character 0 0 1 0 4 FALSE
## Age numeric 1 25 0 0 3 FALSE
## NearZeroVariance Issue
## FALSE TRUE
## FALSE TRUE
##
## No duplicated rows detected.
Detailed results are available through:
report$overview## Metric Value
## 1 Rows 4
## 2 Variables 3
## 3 Missing values 1
## 4 Blank values 1
## 5 Infinite values 0
## 6 Duplicated rows 0
## 7 Constant variables 0
## 8 Near-zero variance variables 0
## 9 Variables with issues 2
report$variables## Variable Class Missing MissingPercent Blank Infinite Unique
## ID ID character 0 0 0 0 3
## Species Species character 0 0 1 0 4
## Age Age numeric 1 25 0 0 3
## Constant NearZeroVariance Issue
## ID FALSE FALSE FALSE
## Species FALSE FALSE TRUE
## Age FALSE FALSE TRUE
report$duplicates## Row Duplicate
## 1 1 FALSE
## 2 2 FALSE
## 3 3 FALSE
## 4 4 FALSE
report$summary## Rows Variables VariablesWithIssues DuplicateRows TotalIssues RuleViolations
## 1 4 3 2 0 2 0
## OverallIssues
## 1 2
Specific checks can also be performed independently.
missing_check(dat)## Variable Total Missing Complete MissingPercent CompletePercent
## 1 ID 4 0 4 0 100
## 2 Species 4 0 4 0 100
## 3 Age 4 1 3 25 75
blank_check(dat)## Variable TextVariable Missing Blank BlankPercent HasBlank
## ID ID TRUE 0 0 0 FALSE
## Species Species TRUE 0 1 25 TRUE
## Age Age FALSE 1 0 0 FALSE
duplicate_check(dat)## Row Duplicate
## 1 1 FALSE
## 2 2 FALSE
## 3 3 FALSE
## 4 4 FALSE
range_check(
dat,
variable = "Age",
min = 0,
max = 30
)## Row Value BelowMinimum AboveMaximum OutOfRange
## 1 1 5 FALSE FALSE FALSE
## 2 2 40 FALSE TRUE TRUE
## 3 3 -2 TRUE FALSE TRUE
## 4 4 NA FALSE FALSE FALSE
category_check(
dat,
variable = "Species",
allowed = c(
"Dog",
"Cat"
)
)## Row Value Missing ValidCategory InvalidCategory
## 1 1 Dog FALSE TRUE FALSE
## 2 2 Horse FALSE FALSE TRUE
## 3 3 FALSE FALSE TRUE
## 4 4 Cat FALSE TRUE FALSE
Reusable validation rules can be defined with
audit_rules():
rules <- audit_rules(
range = list(
Age = c(0, 30)
),
category = list(
Species = c(
"Dog",
"Cat"
)
),
unique = "ID",
required = "Species",
pattern = list(
ID = "^A[0-9]{3}$"
)
)Apply the rules directly:
rule_results <- apply_audit_rules(
dat,
rules
)
rule_results## DataAudit Rule Validation
## =========================
##
## Status: Violations detected
## Total violations: 8
##
## Rule summary
## ------------
## Rule Violations
## Range 2
## Category 2
## Unique 2
## Required 1
## Pattern 1
##
## Variables with rule violations
## ------------------------------
## - Range: Age
## - Category: Species
## - Unique: ID
## - Required: Species
## - Pattern: ID
Custom validation rules can be incorporated directly into the general audit:
report <- audit_data(
dat,
include_nzv = FALSE,
rules = rules
)
report## DataAudit Report
## ================
##
## Rows: 4
## Variables: 3
## Variables with issues: 2
## Duplicated rows: 0
## Total issues: 2
## Rule violations: 8
## Overall issues: 10
##
## Variables with issues
## ---------------------
## Variable Class Missing MissingPercent Blank Infinite Unique Constant
## Species character 0 0 1 0 4 FALSE
## Age numeric 1 25 0 0 3 FALSE
## NearZeroVariance Issue
## FALSE TRUE
## FALSE TRUE
##
## No duplicated rows detected.
##
## Custom rule validation
## ----------------------
## Rule Violations
## Range 2
## Category 2
## Unique 2
## Required 1
## Pattern 1
##
## Total rule violations: 8
The report now combines automatic auditing and custom validation:
report$summary## Rows Variables VariablesWithIssues DuplicateRows TotalIssues RuleViolations
## 1 4 3 2 0 2 8
## OverallIssues
## 1 10
report$rules## DataAudit Rule Validation
## =========================
##
## Status: Violations detected
## Total violations: 8
##
## Rule summary
## ------------
## Rule Violations
## Range 2
## Category 2
## Unique 2
## Required 1
## Pattern 1
##
## Variables with rule violations
## ------------------------------
## - Range: Age
## - Category: Species
## - Unique: ID
## - Required: Species
## - Pattern: ID
A compact summary can be obtained with:
summary(report)## $dataset
## Rows Variables VariablesWithIssues DuplicateRows TotalIssues RuleViolations
## 1 4 3 2 0 2 8
## OverallIssues
## 1 10
##
## $variables_with_issues
## Variable Class Missing MissingPercent Blank Infinite Unique
## Species Species character 0 0 1 0 4
## Age Age numeric 1 25 0 0 3
## Constant NearZeroVariance Issue
## Species FALSE FALSE TRUE
## Age FALSE FALSE TRUE
##
## $duplicate_rows
## [1] Row Duplicate
## <0 rows> (or 0-length row.names)
##
## $rules
## Rule Violations
## 1 Range 2
## 2 Category 2
## 3 Unique 2
## 4 Required 1
## 5 Pattern 1
##
## attr(,"class")
## [1] "summary.DataAuditReport" "list"
Calculate an overall data-quality score with:
score <- audit_score(report)
score## DataAudit Quality Score
## =======================
##
## Score: 54.55 / 100
## Quality: Poor
##
## Dataset
## -------
## Rows: 4
## Variables: 3
## Data cells: 12
##
## Issues
## ------
## Automatic issues: 2
## Rule violations: 8
## Overall issues: 10
The score ranges from 0 to 100, with higher scores indicating fewer detected data-quality issues relative to dataset size.
DataAudit classifies scores as:
The score is intended as a compact audit summary and should be interpreted alongside the underlying diagnostics and the intended use of the dataset.
DataAudit currently provides functions for:
Raw data
|
v
audit_data()
|
+----> Automatic checks
|
+----> audit_rules()
|
v
Rule validation
|
v
DataAuditReport
|
+----> summary()
|
+----> audit_score()
Vinodh Kumar Obli Rajendran
ORCID: 0000-0002-7232-4122
Keerthi Aaradhana
DataAudit is released under the MIT License.