Candidate Item Sets, Diagnostics, and Independent Evaluation

ItemRest 1.0.0 identifies candidate item sets under loading and numerical screening criteria. It does not establish an optimal measurement solution. The search considers combinations of flagged items, reassesses remaining items, and stops a branch when no removable problems remain or the model cannot be estimated. Factor count is fixed at the discovery baseline. Identical item sets are evaluated once. A bounded search reports that it is incomplete. The original correlation matrix and observation counts are computed once per dataset. Each retained set uses a submatrix, with its own positive-definiteness check and any explicitly requested smoothing. Validation and bootstrap samples compute their own matrices. Full psych fits can be kept with store_fits=TRUE; the default keeps the loadings, factor correlations, reliability, and diagnostics.

Automatic parallel analysis retains the parallel_method="fa" default; parallel_method="pc" uses full-matrix component eigenvalues. For ordinal data, independent item permutations preserve category margins and missing positions, and each reference dataset uses the same correlation estimator. The suggested count may differ from a normal Pearson reference. These are reference samples, not additional recomputations of the real-data matrix in the removal search.

Simulated discovery and holdout observations

This example has two correlated factors, a weak item, a cross-loading item, and one explicitly reversed item. It checks implementation; a single simulation does not establish methodological performance or substantively valid deletion.

set.seed(5102026)
n <- 600L
F1 <- rnorm(n)
F2 <- .3 * F1 + sqrt(1 - .3^2) * rnorm(n)
d <- data.frame(
  A1 = F1 + rnorm(n, sd=.5), A2 = F1 + rnorm(n, sd=.5),
  A3 = F1 + rnorm(n, sd=.5), A4 = F1 + rnorm(n, sd=.5),
  B1 = F2 + rnorm(n, sd=.5), B2 = F2 + rnorm(n, sd=.5),
  B3 = F2 + rnorm(n, sd=.5), B4 = -F2 + rnorm(n, sd=.5),
  Weak = rnorm(n), Cross = .6 * F1 + .6 * F2 + rnorm(n, sd=.5)
)
split <- itemrest_split(d, n_factors=2, seed=5102026,
  keys=c(B4=-1), retain_items=c("A1", "B1"),
  item_reasons=c(A1="Domain A anchor", B1="Domain B anchor"),
  max_solutions=500, verbose=FALSE)
result <- split$discovery
print(result)
#> 
#> ItemRest: Candidate Solutions
#> Factor count: 2 (fixed at baseline); ordering: none 
#> Seed: 5102026 
#> Search: complete within threshold-driven scope 
#> No-removal baseline:
#>  Solution_ID Total_Explained_Var Analysis_Status Candidate
#>       S00001           0.7280208              ok     FALSE
#>                             Diagnostics
#>  low_loading_items; cross_loading_items
#> Candidate solutions:
#>  Solution_ID Baseline Removed_Items N_Remaining Total_Explained_Var
#>       S00004    FALSE    Cross-Weak           8           0.8094965
#>  Analysis_Status Candidate Diagnostics
#>               ok      TRUE            
#> Candidates require content review and independent evaluation. Explained variance across different item sets is descriptive.
#> Let algorithms be your compass, not your captain.

Without a supplied seed, the local execution date is encoded as DDMMYYYY. For 5 October 2026 this means numeric seed 5102026 and label 05102026. Use the recorded numeric seed to reproduce the run on a different date. The caller RNG state is restored. Settings include thresholds, keys, and the number of parallel-analysis replications (100 by default), along with R/package versions.

Baseline and diagnostics

result$removal_summary[, c("Solution_ID", "Baseline", "Removed_Items",
  "Analysis_Status", "Candidate", "Problem_Items", "Diagnostics")]
#>   Solution_ID Baseline Removed_Items Analysis_Status Candidate Problem_Items
#> 1      S00001     TRUE          None              ok     FALSE    Cross-Weak
#> 2      S00002    FALSE         Cross              ok     FALSE          Weak
#> 3      S00003    FALSE          Weak              ok     FALSE         Cross
#> 4      S00004    FALSE    Cross-Weak              ok      TRUE          None
#>                              Diagnostics
#> 1 low_loading_items; cross_loading_items
#> 2                      low_loading_items
#> 3                    cross_loading_items
#> 4
result$search
#> $evaluated
#> [1] 4
#> 
#> $complete
#> [1] TRUE
#> 
#> $limit_reached
#> [1] FALSE
#> 
#> $failed
#> [1] 0
#> 
#> $skipped
#> [1] 0
#> 
#> $termination
#> [1] "completed"
#> 
#> $scope
#> [1] "threshold_reachable_item_sets"
#> 
#> $max_solutions
#> [1] 500
result$content_decisions
#>     Item Protected          Reason
#> 1     A1      TRUE Domain A anchor
#> 2     A2     FALSE            <NA>
#> 3     A3     FALSE            <NA>
#> 4     A4     FALSE            <NA>
#> 5     B1      TRUE Domain B anchor
#> 6     B2     FALSE            <NA>
#> 7     B3     FALSE            <NA>
#> 8     B4     FALSE            <NA>
#> 9   Weak     FALSE            <NA>
#> 10 Cross     FALSE            <NA>

The no-removal baseline is always retained and appears first in the full table. Failure or underidentification of one subset does not abort other branches. Numerical diagnostics distinguish inadmissible variances, nonpositive definite correlations, explicit smoothing, reported/nonreported convergence, high factor correlations, and too few qualifying items per factor. Default screening requires three nonflagged primary items per factor. A pattern loading above one requires review and is not by itself classified as a Heywood case under oblique rotation. Only qualifying, admissible solutions without review flags enter the candidate table. These rules do not establish validity or independent stability.

Listwise deletion (default) occurs once on all baseline items, retaining the same observations across subsets. Pairwise handling records pair-count ranges and uses the minimum pair count as conservative EFA sample size. Missing-data bias requires separate consideration. pd_action="smooth" is an explicit opt-in; smoothed solutions are recorded and require manual review.

No winner is selected and default ordering is discovery order. Optional ordering by removal count or explained variance is presentation. Explained variance is mean model communality computed using L Phi L’, so oblique correlations are included. Values from different retained sets have different denominators and are not evidence that one instrument is superior.

Factor reliability and scoring direction

result$solution_details[["S00001"]]$assessment$factor_reliability
#>   Factor N_Items       Items Raw_Alpha Standardized_Alpha
#> 1   ULS2       4 B1-B2-B3-B4 0.9456989          0.9458077
#> 2   ULS1       4 A1-A2-A3-A4 0.9429409          0.9429681
#>   Analysis_Correlation_Alpha Omega_Total
#> 1                  0.9458077   0.9458179
#> 2                  0.9429681   0.9430077
result$removal_summary[, c("Solution_ID", "Cronbachs_Alpha",
  "Standardized_Alpha", "Analysis_Correlation_Alpha", "Omega_Total")]
#>   Solution_ID Cronbachs_Alpha Standardized_Alpha Analysis_Correlation_Alpha
#> 1      S00001       0.8850927          0.8798170                  0.8798170
#> 2      S00002       0.8548714          0.8479526                  0.8479526
#> 3      S00003       0.9081655          0.9083550                  0.9083550
#> 4      S00004       0.8839382          0.8840359                  0.8840359
#>   Omega_Total
#> 1   0.9433957
#> 2   0.9309639
#> 3   0.9631775
#> 4   0.9568569

Factor reports use qualifying nonflagged primary assignments; low/cross-loading items are unassigned for subscale reliability. Global alpha describes all retained items and is insufficient for a multidimensional scale. Raw alpha uses observed covariances; standardized alpha uses Pearson correlations. Selected-correlation alpha follows the analysis correlation matrix, which may combine ordinal and continuous correlations under qgraph automatic detection. These measures are explicitly distinguished, as described in the psych alpha documentation.

Model omega total is common variance of a standardized unit-weighted sum divided by its model-implied total variance. It includes all common factors, also within factor subscales, and is not omega hierarchical or proof of unidimensionality. Scoring direction is specified using keys; ItemRest does not automatically reverse items. Signed and absolute loading ranges are reported separately.

Holdout EFA and optional CFA

split$validation$validation_summary[, c("Solution_ID", "Candidate",
  "Analysis_Status", "Problem_Items", "N_Obs")]
#>        Solution_ID Candidate Analysis_Status Problem_Items N_Obs
#> S00001      S00001     FALSE              ok    Cross-Weak   180
#> S00004      S00004      TRUE              ok          None   180
split$validation$solution_details[["S00001"]]$factor_congruence
#>            ULS2       ULS1
#> ULS2 0.99726923 0.04224237
#> ULS1 0.04736451 0.99821327

Holdout EFA refits the same retained sets without selecting more items. The full discovery/holdout factor-congruence matrix is supplied for matching and inspection. Disjoint row indices are returned by itemrest_split. Alternatively, use itemrest_validate(result, independent_data). The caller must establish that observations are independent.

if (requireNamespace("lavaan", quietly=TRUE)) {
  cfa <- itemrest_validate(result, d[split$validation_rows, ], method="cfa", seed=5102026)
  cfa$validation_summary
}
#>        Solution_ID Baseline Analysis_Status Error_Message Warnings Messages
#> S00001      S00001     TRUE              ok                                
#> S00004      S00004    FALSE              ok                                
#>        Converged Admissible Max_Factor_Correlation Review_Required N_Obs
#> S00001      TRUE       TRUE              0.3259867           FALSE   180
#> S00004      TRUE       TRUE              0.2790009           FALSE   180
#>              CFI       TLI      RMSEA       SRMR CFI_Robust TLI_Robust
#> S00001 0.9228228 0.8978537 0.14300871 0.11743398  0.9223361  0.8972096
#> S00004 0.9972073 0.9958844 0.03341418 0.03052968  0.9966085  0.9950021
#>        RMSEA_Robust CFI_Scaled TLI_Scaled RMSEA_Scaled Estimator Missing_Method
#> S00001   0.14343421  0.9196104  0.8936020   0.14503399       MLR       listwise
#> S00004   0.03681991  0.9964340  0.9947449   0.03763912       MLR       listwise

CFA freezes discovery primary assignments with zero cross-loadings and correlated factors. It uses MLR for continuous data and WLSMV for specified/detected ordinal items. With pairwise discovery handling, continuous CFA uses FIML; this is labelled in the output rather than presented as the same missing-data estimator. Fit indices do not become automatic validity decisions. See the lavaan categorical-data documentation.