ItemRest 1.0.0 identifies candidate item sets under loading and
numerical screening criteria. It does not establish an optimal
measurement solution. The search considers combinations of flagged
items, reassesses remaining items, and stops a branch when no removable
problems remain or the model cannot be estimated. Factor count is fixed
at the discovery baseline. Identical item sets are evaluated once. A
bounded search reports that it is incomplete. The original correlation
matrix and observation counts are computed once per dataset. Each
retained set uses a submatrix, with its own positive-definiteness check
and any explicitly requested smoothing. Validation and bootstrap samples
compute their own matrices. Full psych fits can be kept with
store_fits=TRUE; the default keeps the loadings, factor
correlations, reliability, and diagnostics.
Automatic parallel analysis retains the
parallel_method="fa" default;
parallel_method="pc" uses full-matrix component
eigenvalues. For ordinal data, independent item permutations preserve
category margins and missing positions, and each reference dataset uses
the same correlation estimator. The suggested count may differ from a
normal Pearson reference. These are reference samples, not additional
recomputations of the real-data matrix in the removal search.
This example has two correlated factors, a weak item, a cross-loading item, and one explicitly reversed item. It checks implementation; a single simulation does not establish methodological performance or substantively valid deletion.
set.seed(5102026)
n <- 600L
F1 <- rnorm(n)
F2 <- .3 * F1 + sqrt(1 - .3^2) * rnorm(n)
d <- data.frame(
A1 = F1 + rnorm(n, sd=.5), A2 = F1 + rnorm(n, sd=.5),
A3 = F1 + rnorm(n, sd=.5), A4 = F1 + rnorm(n, sd=.5),
B1 = F2 + rnorm(n, sd=.5), B2 = F2 + rnorm(n, sd=.5),
B3 = F2 + rnorm(n, sd=.5), B4 = -F2 + rnorm(n, sd=.5),
Weak = rnorm(n), Cross = .6 * F1 + .6 * F2 + rnorm(n, sd=.5)
)
split <- itemrest_split(d, n_factors=2, seed=5102026,
keys=c(B4=-1), retain_items=c("A1", "B1"),
item_reasons=c(A1="Domain A anchor", B1="Domain B anchor"),
max_solutions=500, verbose=FALSE)
result <- split$discovery
print(result)
#>
#> ItemRest: Candidate Solutions
#> Factor count: 2 (fixed at baseline); ordering: none
#> Seed: 5102026
#> Search: complete within threshold-driven scope
#> No-removal baseline:
#> Solution_ID Total_Explained_Var Analysis_Status Candidate
#> S00001 0.7280208 ok FALSE
#> Diagnostics
#> low_loading_items; cross_loading_items
#> Candidate solutions:
#> Solution_ID Baseline Removed_Items N_Remaining Total_Explained_Var
#> S00004 FALSE Cross-Weak 8 0.8094965
#> Analysis_Status Candidate Diagnostics
#> ok TRUE
#> Candidates require content review and independent evaluation. Explained variance across different item sets is descriptive.
#> Let algorithms be your compass, not your captain.Without a supplied seed, the local execution date is encoded as
DDMMYYYY. For 5 October 2026 this means numeric seed 5102026 and label
05102026. Use the recorded numeric seed to reproduce the
run on a different date. The caller RNG state is restored. Settings
include thresholds, keys, and the number of parallel-analysis
replications (100 by default), along with R/package versions.
result$removal_summary[, c("Solution_ID", "Baseline", "Removed_Items",
"Analysis_Status", "Candidate", "Problem_Items", "Diagnostics")]
#> Solution_ID Baseline Removed_Items Analysis_Status Candidate Problem_Items
#> 1 S00001 TRUE None ok FALSE Cross-Weak
#> 2 S00002 FALSE Cross ok FALSE Weak
#> 3 S00003 FALSE Weak ok FALSE Cross
#> 4 S00004 FALSE Cross-Weak ok TRUE None
#> Diagnostics
#> 1 low_loading_items; cross_loading_items
#> 2 low_loading_items
#> 3 cross_loading_items
#> 4
result$search
#> $evaluated
#> [1] 4
#>
#> $complete
#> [1] TRUE
#>
#> $limit_reached
#> [1] FALSE
#>
#> $failed
#> [1] 0
#>
#> $skipped
#> [1] 0
#>
#> $termination
#> [1] "completed"
#>
#> $scope
#> [1] "threshold_reachable_item_sets"
#>
#> $max_solutions
#> [1] 500
result$content_decisions
#> Item Protected Reason
#> 1 A1 TRUE Domain A anchor
#> 2 A2 FALSE <NA>
#> 3 A3 FALSE <NA>
#> 4 A4 FALSE <NA>
#> 5 B1 TRUE Domain B anchor
#> 6 B2 FALSE <NA>
#> 7 B3 FALSE <NA>
#> 8 B4 FALSE <NA>
#> 9 Weak FALSE <NA>
#> 10 Cross FALSE <NA>The no-removal baseline is always retained and appears first in the full table. Failure or underidentification of one subset does not abort other branches. Numerical diagnostics distinguish inadmissible variances, nonpositive definite correlations, explicit smoothing, reported/nonreported convergence, high factor correlations, and too few qualifying items per factor. Default screening requires three nonflagged primary items per factor. A pattern loading above one requires review and is not by itself classified as a Heywood case under oblique rotation. Only qualifying, admissible solutions without review flags enter the candidate table. These rules do not establish validity or independent stability.
Listwise deletion (default) occurs once on all baseline items,
retaining the same observations across subsets. Pairwise handling
records pair-count ranges and uses the minimum pair count as
conservative EFA sample size. Missing-data bias requires separate
consideration. pd_action="smooth" is an explicit opt-in;
smoothed solutions are recorded and require manual review.
No winner is selected and default ordering is discovery order. Optional ordering by removal count or explained variance is presentation. Explained variance is mean model communality computed using L Phi L’, so oblique correlations are included. Values from different retained sets have different denominators and are not evidence that one instrument is superior.
result$solution_details[["S00001"]]$assessment$factor_reliability
#> Factor N_Items Items Raw_Alpha Standardized_Alpha
#> 1 ULS2 4 B1-B2-B3-B4 0.9456989 0.9458077
#> 2 ULS1 4 A1-A2-A3-A4 0.9429409 0.9429681
#> Analysis_Correlation_Alpha Omega_Total
#> 1 0.9458077 0.9458179
#> 2 0.9429681 0.9430077
result$removal_summary[, c("Solution_ID", "Cronbachs_Alpha",
"Standardized_Alpha", "Analysis_Correlation_Alpha", "Omega_Total")]
#> Solution_ID Cronbachs_Alpha Standardized_Alpha Analysis_Correlation_Alpha
#> 1 S00001 0.8850927 0.8798170 0.8798170
#> 2 S00002 0.8548714 0.8479526 0.8479526
#> 3 S00003 0.9081655 0.9083550 0.9083550
#> 4 S00004 0.8839382 0.8840359 0.8840359
#> Omega_Total
#> 1 0.9433957
#> 2 0.9309639
#> 3 0.9631775
#> 4 0.9568569Factor reports use qualifying nonflagged primary assignments; low/cross-loading items are unassigned for subscale reliability. Global alpha describes all retained items and is insufficient for a multidimensional scale. Raw alpha uses observed covariances; standardized alpha uses Pearson correlations. Selected-correlation alpha follows the analysis correlation matrix, which may combine ordinal and continuous correlations under qgraph automatic detection. These measures are explicitly distinguished, as described in the psych alpha documentation.
Model omega total is common variance of a standardized unit-weighted
sum divided by its model-implied total variance. It includes all common
factors, also within factor subscales, and is not omega hierarchical or
proof of unidimensionality. Scoring direction is specified using
keys; ItemRest does not automatically reverse items. Signed
and absolute loading ranges are reported separately.
split$validation$validation_summary[, c("Solution_ID", "Candidate",
"Analysis_Status", "Problem_Items", "N_Obs")]
#> Solution_ID Candidate Analysis_Status Problem_Items N_Obs
#> S00001 S00001 FALSE ok Cross-Weak 180
#> S00004 S00004 TRUE ok None 180
split$validation$solution_details[["S00001"]]$factor_congruence
#> ULS2 ULS1
#> ULS2 0.99726923 0.04224237
#> ULS1 0.04736451 0.99821327Holdout EFA refits the same retained sets without selecting more
items. The full discovery/holdout factor-congruence matrix is supplied
for matching and inspection. Disjoint row indices are returned by
itemrest_split. Alternatively, use
itemrest_validate(result, independent_data). The caller
must establish that observations are independent.
if (requireNamespace("lavaan", quietly=TRUE)) {
cfa <- itemrest_validate(result, d[split$validation_rows, ], method="cfa", seed=5102026)
cfa$validation_summary
}
#> Solution_ID Baseline Analysis_Status Error_Message Warnings Messages
#> S00001 S00001 TRUE ok
#> S00004 S00004 FALSE ok
#> Converged Admissible Max_Factor_Correlation Review_Required N_Obs
#> S00001 TRUE TRUE 0.3259867 FALSE 180
#> S00004 TRUE TRUE 0.2790009 FALSE 180
#> CFI TLI RMSEA SRMR CFI_Robust TLI_Robust
#> S00001 0.9228228 0.8978537 0.14300871 0.11743398 0.9223361 0.8972096
#> S00004 0.9972073 0.9958844 0.03341418 0.03052968 0.9966085 0.9950021
#> RMSEA_Robust CFI_Scaled TLI_Scaled RMSEA_Scaled Estimator Missing_Method
#> S00001 0.14343421 0.9196104 0.8936020 0.14503399 MLR listwise
#> S00004 0.03681991 0.9964340 0.9947449 0.03763912 MLR listwiseCFA freezes discovery primary assignments with zero cross-loadings and correlated factors. It uses MLR for continuous data and WLSMV for specified/detected ordinal items. With pairwise discovery handling, continuous CFA uses FIML; this is labelled in the output rather than presented as the same missing-data estimator. Fit indices do not become automatic validity decisions. See the lavaan categorical-data documentation.
boot <- itemrest_bootstrap(result, d[split$discovery_rows, ], n_boot=20, seed=5102026)
boot$item_stability
#> Item Retained_In_Any_Candidate Retained_In_All_Candidates
#> 1 A1 1.00 1.00
#> 2 A2 1.00 1.00
#> 3 A3 1.00 1.00
#> 4 A4 1.00 1.00
#> 5 B1 1.00 1.00
#> 6 B2 1.00 1.00
#> 7 B3 1.00 1.00
#> 8 B4 1.00 1.00
#> 9 Weak 0.00 0.00
#> 10 Cross 0.25 0.25
#> Mean_Retention_Among_Candidates Mean_Removal_Among_Candidates
#> 1 1.00 0.00
#> 2 1.00 0.00
#> 3 1.00 0.00
#> 4 1.00 0.00
#> 5 1.00 0.00
#> 6 1.00 0.00
#> 7 1.00 0.00
#> 8 1.00 0.00
#> 9 0.00 1.00
#> 10 0.25 0.75
#> Retained_In_Any_Candidate_All_Attempts
#> 1 1.00
#> 2 1.00
#> 3 1.00
#> 4 1.00
#> 5 1.00
#> 6 1.00
#> 7 1.00
#> 8 1.00
#> 9 0.00
#> 10 0.25
#> Retained_In_All_Candidates_All_Attempts Mean_Retention_All_Attempts
#> 1 1.00 1.00
#> 2 1.00 1.00
#> 3 1.00 1.00
#> 4 1.00 1.00
#> 5 1.00 1.00
#> 6 1.00 1.00
#> 7 1.00 1.00
#> 8 1.00 1.00
#> 9 0.00 0.00
#> 10 0.25 0.25
boot$solution_stability
#> Solution_ID Candidate_Frequency Candidate_Frequency_All_Attempts
#> 1 S00001 0.00 0.00
#> 2 S00002 0.00 0.00
#> 3 S00003 0.25 0.25
#> 4 S00004 0.75 0.75
table(boot$replicates$Status)
#>
#> complete
#> 20Every bootstrap replicate reruns the threshold-driven search with the discovery factor count fixed. Frequencies distinguish retention in any candidate, in all candidates, and average retention among candidates. Each complete replication has equal weight; candidate-free replications count as zero for any/all retention. Failed/incomplete searches are reported and excluded from denominators. Exact-set frequencies describe stability without selecting a rank-based winner. These frequencies do not replace independent validation or give probabilities of validity.
The complete executable script ships at
examples/reproducible_workflow.R in the installed package.
Content-based decisions and independent replication remain necessary
before using any candidate as a measurement instrument.