--- title: "Build a Mini Reference Library" author: "Win Cowger" date: "`r Sys.Date()`" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Build a Mini Reference Library} %\VignetteEncoding{UTF-8} %\VignetteEngine{knitr::rmarkdown} --- ```{r, include = FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>", warning = FALSE ) data.table::setDTthreads(2) ``` ```{r setup} library(OpenSpecy) ``` This example combines a few files bundled with OpenSpecy into a small reference library. Real library builds usually need larger lookup tables and more curation, but the same helper functions apply. ## Read And Combine Spectra `c_spec()` and `build_lib()` default to the widest range represented by the source spectra at resolution 6. Values outside each source's original range are kept as `NA`, so useful spectral regions are not discarded. This compact vignette example uses the shared overlapping range to keep the rendered output small. ```{r mini-library-read} mini_files <- c( read_extdata("raman_hdpe.csv"), read_extdata("ftir_ldpe_soil.asp"), read_extdata("raman_atacamit.spc") ) mini_sources <- lapply(mini_files, read_any) mini_sources <- lapply(mini_sources, function(x) { x$metadata$intensity_units <- "absorbance" attr(x, "intensity_unit") <- "absorbance" x }) mini_raw <- c_spec(mini_sources, range = "common", res = 6) check_OpenSpecy(mini_raw) dim(mini_raw$spectra) mini_raw$metadata[, "file_name", with = FALSE] ``` ## Create And Fill A Lookup Lookup templates help users see which metadata values need curation. When `path` is not supplied, the template is returned as a `data.table`; users can write it to CSV by supplying `path`. Lookup joins use exact values, so edit the template until the key column matches the object metadata. ```{r mini-library-template} template <- make_lib_lookup_template( mini_raw, columns = "file_name", add = c("material", "material_type") ) template ``` For this small example, create the lookup table directly in R. The same shape could come from a CSV edited outside R. ```{r mini-library-join} lookup <- data.table::data.table( file_name = basename(mini_files), library_type = c("example", "example", "example"), spectrum_type = c("raman", "ftir", "raman"), material = c("hdpe", "ldpe in soil", "atacamite") ) hierarchy <- data.table::data.table( material = c("hdpe", "ldpe in soil", "atacamite"), material_class = c("polyethylene", "polyethylene", "copper mineral"), material_type = c("plastic", "plastic", "mineral") ) join_lib_metadata(mini_raw, lookup, by = "file_name", require_complete = TRUE)$metadata[ , c("file_name", "library_type", "spectrum_type", "material"), with = FALSE ] ``` ## Build A Mini Library `build_lib()` can run the ordinary lookup and material hierarchy joins before applying named recipes. Recipe names become names in the returned list. Empty recipes keep the merged spectra unchanged; other recipe lists are passed to `process_spec()`. Missing values and processing attributes are handled automatically. Metadata column names are also cleaned to lowercase underscore names. Known aliases are coalesced using an editable lookup table. Variants that differ only by underscores or one terminal plural `s` match automatically. Before sources are merged, `build_lib()` also converts declared reflectance and transmittance spectra to absorbance by default. A nonempty `attr(x, "intensity_unit")` is the primary truth for the whole object; otherwise, `metadata$intensity_units` is evaluated spectrum by spectrum. Unknown or missing units are left unchanged with a warning. Use `convert_intensity = FALSE` when unit handling has already been completed outside the builder. The normal input is one or more file paths; one `OpenSpecy` or a list of `OpenSpecy` objects supports sources already loaded in memory. An RDS path may store either one object or a list of objects. Other formats are read with `read_any()`. Large same-axis source lists are prepared in bulk, which is much faster for legacy RDS files that store many one-spectrum objects. Named stages and elapsed time are reported while the library is built; use `progress = FALSE` when quiet output is preferred. Supplying `restrict_range_args` triggers the existing `restrict_range()` operation before deduplication and recipes; multiple retained ranges can exclude a known silent region without custom workflow code. ```{r metadata-name-lookup} name_lookup <- lib_metadata_name_lookup( project_code = c("campaign id", "study code"), regex = list(instrument_mode = "^method_[0-9]+$") ) name_lookup[ canonical_name %in% c("material_color", "number_of_accumulations") ] lib_clean_name(c("User Name", "Laser (%)", "Method...3")) ``` Named arguments add exact aliases to the defaults, while `regex` adds patterns evaluated against cleaned names. Overlapping regex patterns produce an error that identifies the source column and matching rules. Pass the result as `metadata_name_lookup`. Set `clean_metadata_values = TRUE` in `build_lib()` or `clean_values = TRUE` in `lib_clean_metadata()` to lowercase, trim, and ASCII-normalize character metadata values before joins. Ordinary and hierarchical joins run whenever their corresponding lookup input is non-`NULL`. `spectrum_identity` receives one additional deterministic cleanup: recognizable paths are reduced to their basename and trailing file extensions supported by `read_any()` are removed. Numeric OPUS extensions include `.10` and any other terminal period followed only by digits. Exact lookup keys receive the same cleanup. Each built library records changed values and counts in its `spectrum_identity_cleanup_report` attribute. Keep flexible class patterns in a separate table and apply them afterward with `predict_class_reference()`. Automatic ordinary lookups use the single shared column with overlapping values and unique lookup keys. Lookups with no usable shared key are skipped with a message; lookups with multiple usable shared keys are treated as ambiguous, so wrap a lookup as `list(lookup = table, by = "key")` when an explicit key is needed. Named `by` vectors map metadata names to different lookup names. Canonical source metadata is standardized before these external joins. Every spectrum receives `library_name` from a populated `organization`, otherwise from `user_name`; this is the canonical source-library grouping used by build assessments. A blank `organization` is still filled from the reviewed internal `user_name` alias for the existing type lookup. The older lookup-level `fallback_by` field is deprecated. Set `fill_only = TRUE` to fill blank values such as `library_type` or `spectrum_type` without replacing populated source metadata. ```{r mini-library-build} mini_libs <- build_lib( mini_files, recipes = list( raw = list(), derivative = list( conform_spec = FALSE, smooth_intens = TRUE, smooth_intens_args = list(window = 15, derivative = 1), make_rel = TRUE ), nobaseline = list( conform_spec = FALSE, smooth_intens = FALSE, subtr_baseline = TRUE, make_rel = TRUE ) ), metadata_lookups = lookup, material_hierarchy = hierarchy, clean_metadata_values = TRUE, convert_intensity = FALSE, assess = TRUE, dedupe = FALSE ) names(mini_libs) check_OpenSpecy(mini_libs$raw) check_OpenSpecy(mini_libs$derivative) attr(mini_libs$derivative, "derivative_order") attr(mini_libs$nobaseline, "baseline") mini_libs$raw$metadata[ , .(file_name, material, material_class, material_type, sn, assessment_flag, assessment_checks) ] ``` `prune_lib()` performs the spectrum-supported cleanup used before medoid or model creation. With `cross_class = TRUE`, it first flags correlations strictly above `cross_class_threshold` between different reviewed classes within each source library. The spectrum with the most active wrong-class neighbors is removed first and scores are recalculated; adjacent maximum-score ties remove both endpoints. A second pass then uses the number of independent opposing libraries, removing the higher-evidence endpoint or both on a tie. Generic and unclassified labels are excluded. The audit records phase, view, component, identities/classes/libraries, active degree or independent evidence, decision round, threshold, and reason. Within the same spectral-technique pool, `other` may then be reassigned to the nearest established class, `other plastic` only to a candidate whose `material_type` is `plastic`, and `other material` only to `organic matter` or `mineral`. The matched `material_type` is copied with the class. It then evaluates material classes from largest to smallest using bounded correlation blocks. FTIR and NIR share a candidate pool; Raman is separate; 2200--2420 cm^-1^ is excluded by default. After generic reassignment, each `spectrum_type` and material-class group must contain at least `min_n` spectra across the complete input database; `min_n` is not applied separately to each source library. Smaller resolved groups are reassigned as a whole when an eligible destination exists and removed only when no destination exists; groups exactly at the threshold are retained whole and larger groups are never reduced below it. Constant or unmatched spectra in eligible classes are retained, and deterministic IDs resolve correlation ties. Use `return = "report"` for the retained IDs, frozen class schedule, reassignment correlations, spectrum-level removals, and `excluded_classes`, which reports observed support and the number of additional or reassigned spectra needed. End-to-end builds combine these class rows in `assessments$cleanup$summary` with the recipe name. `build_lib(prune = ...)` maps recipe names to `prune_lib()` argument lists, so each selected spectral representation is pruned independently. Resolved class/type groups below `min_n` are reassigned as a whole to the most-correlated established class in the same technique pool and material type; they are dropped only when no eligible correlated destination exists. The pruning assessment records the destination, mean class correlation, action, and reason. In the official workflow, an otherwise unresolved identity is temporarily labeled `other`. The default `remove_other = TRUE` removes blank identities and the unresolved literal `other` class before quality control. Reviewed broad `other plastic` and `other material` categories stay in the reference libraries and remain eligible for constrained reassignment during derivative/no-baseline pruning. The builder keeps reviewed identifiers, source metadata, prior labels, reasons, actions, and typed before/after counts in `assessments$cleanup$summary`. The separate `assessments$cleanup$dropped_spectrum_identities` table contains only the sorted distinct source identities removed anywhere in cleanup. Set `remove_other = FALSE` to retain generic rows for the constrained semisupervised `prune_lib()` reassignment described above. ```{r mini-library-prune} pruned <- prune_lib( mini_libs$derivative, min_n = 1, return = "report", progress = FALSE ) pruned$summary pruned$schedule pruned$excluded_classes ``` The composable default leaves the new first pass off. Enable it explicitly for a library carrying canonical `library_name` metadata; the official derivative and no-baseline recipes do this by default: ```{r cross-class-prune-policy, eval=FALSE} prune_lib( reference_library, cross_class = TRUE, cross_class_threshold = 0.9 ) ``` ## Official Reference-Library Workflow The version-controlled [`workflows/OpenSpecy_reference_library.R`](https://github.com/wincowgerDEV/OpenSpecy-package/blob/main/workflows/OpenSpecy_reference_library.R) script retains the source and output path finders and passes those paths to one `build_lib()` call. The complete visual flow, including checkpoint reuse and old/new assessments, is maintained in [`.specify/memory/build-lib-diagram.html`](https://github.com/wincowgerDEV/OpenSpecy-package/blob/main/.specify/memory/build-lib-diagram.html). Canonically named exact-class, regex-class, library-type, material-hierarchy, known-bad-ID, material-form-regex, and common-use tables live under `workflows/data/`. Raw-data corrections are completed externally before this workflow runs. The source-library paths and output directory are explicit inputs. When `workflow_data` is not supplied, `build_lib()` looks for the helper tables under `data/` beside the calling script, then under `data/` or `workflows/data/` in the current working directory. Calling `build_lib()` with no `x` now stops with an actionable error instead of guessing external paths. The official build derives `library_name` from organization first and user name second, fills a missing organization from `user_name` before one source-type lookup, and asserts complete library/spectrum-type coverage. Per-recipe retention rows in `assessments$cleanup$summary` report counts at preparation, exclusion, deduplication, broad-category review, quality control, pruning, post-transform, and final partition stages. A completely dropped source names the first empty stage and its reason. The exact class lookup runs first; `predict_class_reference()` then evaluates the separate regex table only for blank materials. Exact/regex overlaps are audited without overwriting exact values, and conflicting regex predictions stop for review. `material_class` describes chemistry rather than physical form. Every reviewed standard plastic class begins with `poly` for discoverability (`other plastic` is the explicit catch-all exception); polyethylene and polypropylene remain separate, and `polyhydroxy(meth)acrylates` retains its chemically meaningful optional group notation. Poly(styrene-butadiene) rubber and poly(ethylene-propylene-diene) rubber have distinct classes. Paint binder labels such as acrylic, alkyd, urethane, nitrocellulose, and vinyl combinations remain in `spectrum_identity`, `paint` remains in `material_form`, and the hierarchy maps each reviewed binder to its relevant chemical polymer family. Remaining uncertain identities follow the selected generic-row review policy. The official workflow also standardizes two enrichment fields. It normalizes and concatenates all atomic metadata values once per spectrum, then evaluates the reviewable `material_form_regex.csv`. A uniquely matching existing `material_form` takes precedence; otherwise the full metadata row is searched. Cross-category matches remain `NA` and are retained in the clash audit. The controlled vocabulary currently covers paint, rubber, hard plastic, fiber, pellet, film plastic, fragment, foam, and sphere/bead. `common_use_reference.csv` is a material-class lookup for the proposed predominant global ultimate end-market. Quantitative rows retain consumer, industrial, and unresolved mass shares: `consumer` and `industrial` require more than 50%, while `mixed` requires at least 25% on each side, at least 75% classified, and neither side above 50%. Where comparable mass shares do not exist, a qualitative proposal is allowed only with a cited application source, retrieval date, and explicit review note; share cells remain blank. Consumer endpoints include products used directly by individuals, such as passenger tires, household goods, clothing, and patient-facing healthcare products. Classes with substantial consumer and industrial endpoints use `mixed`. Ambiguous catch-alls and poorly supported specialty groups remain `NA`. Commodity shares use the [OECD 2019 polymer-by-application mass dataset](https://data-explorer.oecd.org/vis?df%5Bag%5D=OECD.ENV.EEI&df%5Bid%5D=DSD_PU_2019%40DF_PU_2019); PTFE uses a separate material-specific end-use source recorded in the CSV. Coverage, provenance, matched evidence, and clashes are retained with the upstream build assessments. Derivative and nobaseline spectra pass three quality gates before pruning: FTIR CO2 is flattened when the 2200--2420 maximum divided by the 2420--2550 silent-region maximum is greater than two; high-tail detection ignores NA padding, trims each spectrum on its finite support, and drops failed corrections; and finite running SNR below two (or unavailable SNR) is removed. Same-library majority closure then precedes independent-library evidence, generic reassignment, and conventional top-match pruning. After rounding and type partitioning, the full and model-range views are closed again. Raw remains unpruned. Every spectrum excluded by cross-class closure is preserved in the review quarantine. Each completed library, medoid, model, and assessment component is written with a SHA-256 manifest under `output_dir/checkpoints`. `reuse = TRUE` loads it only when the source/lookup signatures, arguments, package version, and builder contract still match. Validated output is promoted into a versioned release directory, which contains the seven legacy RDS names, explicit algorithm-named model RDS files, one aggregate build RDS, and a release manifest. Existing versioned payloads are immutable: a different hash stops promotion. The release also contains `quarantined_spectra.rds`, a versioned bundle of valid `OpenSpecy` objects by recipe/type, row-level removal metadata, the long conflict audit, and build provenance. Its review copy and CSV mirrors are checkpointed under `output_dir/review` before downstream model and comparison stages. The returned object has four top-level elements: `libraries`, `medoids`, `models`, and `assessments`. Libraries are recipe-first and then keyed by `ftir`, `raman`, or `nir`: full Raman uses 200--4000, FTIR uses 400--4000, and NIR uses 4000--12000 cm^-1^. FTIR/Raman medoids and models use 800--3200; the NIR identification interval is derived from the longest region with at least 90% finite coverage (falling back to available coverage) and recorded on the object. All-blank metadata columns are dropped independently by type. The enrichment columns are retained even when every value for one type is missing; other all-blank metadata columns are dropped independently by type. Each spectrum must observe at least 10% of its type-specific identification axis to enter medoid selection or model fitting. PAM operates on a temporary relative matrix where each spectrum's missing values are filled by that spectrum's finite mean. Selected medoid IDs are then pulled from the original range-restricted object, so published medoids retain their original `NA` positions. Groups with more than 3,000 spectra use five deterministic 1,000-spectrum PAM samples; each candidate set is scored against the complete group with the same correlation distance, avoiding a full oversized dissimilarity matrix. `models` is algorithm-first (`logistic_regression` or `random_forest`), then recipe and type. Logistic regression retains derivative/nobaseline medoid training and fills restored gaps with the finite mean at each wavenumber across training medoids. Random forests use the complete eligible raw, derivative, and nobaseline type libraries, relative normalization, wavenumber-mean filling, inverse-frequency balanced case sampling, 500 probability trees, permutation importance, and all available ranger threads. Balanced sampling improves minority representation in each bootstrap without also applying a class-vote correction. Because this resampling can make out-of-bag accuracy optimistic, treat OOB results as fit diagnostics. The release assessment separately applies each deployed model to its complete corresponding source dataset. `train_spec_model()` exposes both engines; `build_model_lib()` remains a compatible wrapper. Each model stores its filler, support audit, training diagnostics, and one tidy `tests` table. Logistic lambda selection remains maximum out-of-fold macro class accuracy with its calibrated alpha, no-intercept, grouped multinomial, and weighting settings unchanged. Full assessment uses every candidate and downloaded legacy artifact. Candidate and legacy artifacts independently use a seeded approximately ten-percent class/type-stratified sample. Rows sharing a physical identifier or exact spectral content are one group and cannot cross training and test. This source-local policy tolerates taxonomy evolution without fuzzy cross-version class matching; its results describe each artifact on its own source population and are not paired causal deltas. Held-out groups are removed from full reference libraries before their identification assessment. Medoid artifacts are instead tested as deployed: each medoid library identifies every spectrum in its complete corresponding processed library (for example, `medoid_derivative.rds` searches `derivative.rds`). Existing logistic and random-forest artifacts likewise identify their complete corresponding source libraries once; assessment does not select new medoids or call `train_spec_model()`. `match_spec()` uses stored fillers for partial spectra. Macro class accuracy is the primary metric, followed by evaluated coverage, overall accuracy, and clear old/new `assess_spec(report = "all")` shifts. The review surface contains at most ten tables nested under `cleanup`, `ref_lib`, `medoid`, `model`, and `functionality`. Old and new values occupy adjacent columns; accuracy, misidentification counts, absolute diagnostic correlations, and warning/error rate shifts are sorted from largest to smallest. Pass rows are omitted from quality shifts. Machine-level split and row-test evidence is available through `attr(reference_library_build$assessments, "evidence")`. Published accuracy tables contain aggregate overall and macro metrics, not per-class accuracy rows. A versioned release writes all global cleanup, quality, pruning, model, and comparison evidence to `assessments.rds`; its library, medoid, and model files contain only runtime data and prediction state. The accompanying `reference_library_build.rds` is a lightweight index that can be supplied to `rebuild_lib_artifacts()`. The smaller 1,000-spectrum run in the manual benchmark is development-only and is not part of `build_lib()` or final acceptance evidence. The standard call therefore has this shape: ```{r official-call-shape, eval=FALSE} reference_library_build <- build_lib( files, output_dir = output_dir, previous_library_dir = "system", remove_other = TRUE ) ``` After upstream libraries have already completed, rebuild only the downstream artifacts into a new explicit output location: ```{r downstream-rebuild-shape, eval=FALSE} updated_reference_library_build <- rebuild_lib_artifacts( completed_build_dir, output_dir = downstream_output_dir, previous_library_dir = "system" ) ``` `completed_build_dir` may instead be a `reference_library_build.rds`, a `libraries.rds` checkpoint, or the corresponding in-memory build object. The script is excluded from package builds but remains available in the GitHub repository so library releases can be reviewed and reproduced.