mintyr mintyr package logo

CRAN status CRAN total downloads Dev Version Lifecycle: experimental R-CMD-check GitHub last commit

mintyr turns “many groups × many variables” data into analysis-ready pieces — and back into files — with data.table speed and no surprises.

A lot of real-world analysis follows the same loop: take a wide table, split it by trait and group, fit something on every piece (often with cross-validation), then write each piece or result to its own file or sheet. mintyr packages that loop into a small set of consistent functions:

 files ──► import ──► reshape & nest ──► cross-validate / summarise ──► export ──► files
          (xlsx/csv)   (by trait × group)      (per piece)          (txt/csv/xlsx)

It was born in animal breeding, where this loop is everywhere (multi-trait, multi-breed, multi-farm data feeding ASReml-R, HIBLUP or DMU), but nothing in it is specific to genetics.

Who is it for?

Field Typical use
Animal & plant breeding Per-trait / per-breed models, genetic correlations across traits or environments, input files for HIBLUP or DMU
Multi-site / multi-batch experiments The same analysis for every site × variable, with results written back per site
Grouped machine learning Reproducible (stratified) k-fold CV run separately within every subgroup
Reporting Read dozens of Excel workbooks, clean them in one table, write them back to the original file/sheet layout
Any grouped analysis Pairwise correlations between many variables, top/bottom-N% summaries per group

Installation

# From CRAN
install.packages("mintyr")

# Development version from GitHub
pak::pak("tony2015116/mintyr")

Quick start

Fit a model for every trait within every group, using 4-fold cross-validation, and export each subset to its own folder:

library(mintyr)
library(data.table)

# 1. Wide -> long, nested by trait (`name`) and group (`am`)
nested <- w2l_nest(mtcars, cols = c("mpg", "qsec"), by = "am")
nested
#>    name am               data
#> 1:  mpg  1 <data.table[13x9]>
#> 2:  mpg  0 <data.table[19x9]>
#> 3: qsec  1 <data.table[13x9]>
#> 4: qsec  0 <data.table[19x9]>

# 2. Reproducible 4-fold CV inside every trait x group
cv <- nest_cv(nested, v = 4, seed = 2026)

# 3. Predictive ability per fold, then averaged
cv[, r := mapply(function(tr, va) {
  fit <- lm(value ~ wt + hp, data = tr)
  cor(predict(fit, va), va$value)
}, train, validate)]
cv[, .(mean_r = round(mean(r), 3)), by = .(name, am)]
#>    name am mean_r
#> 1:  mpg  1  0.816
#> 2:  mpg  0  0.881
#> 3: qsec  1  0.884
#> 4: qsec  0  0.858

# 4. One file per trait x group: <path>/<name>/<am>/data.txt
export_nest(nested, path = file.path(tempdir(), "by_trait"))

What’s inside

Import and export — a lossless round trip

Function What it does
import_xlsx() Read many workbooks and sheets into one table, with excel_name / sheet_name source columns; optional parallel reading (workers)
import_csv() Read many CSV/TXT files with fread(), with a source-file column
export_xlsx() Write a table back to the original file/sheet layout, or a list of tables into one workbook or one file each
export_nest() Write every nested table to <path>/<group1>/<group2>/<column>.txt
export_list() Write every element of a named list to its own file (names may contain / for sub-folders)

Reshape and nest

Function What it does
w2l_nest() / w2l_split() Wide → long, then nest (list-column) or split (named list) by variable and group
c2p_nest() All pairs (or triples, …) of variables, renamed to value1, value2, … for uniform pairwise analysis
r2p_nest() Same variable across levels of another column (e.g. one column per environment), aligned by an ID

Cross-validation and summaries

Function What it does
split_cv() / nest_cv() Repeated, optionally stratified k-fold CV for a list of tables or a nested table; returns fold indices and (optionally) the subsets
top_perc() Statistics of the top / bottom X% per variable and group
format_digits() Format numeric columns for reports (decimals, percentages)
get_path_info() Extract file names or path segments, cross-platform

Examples by use case

Pairwise correlations between many variables, per group

pairs <- c2p_nest(mtcars, cols = c("mpg", "hp", "wt"), by = "am")
pairs[, .(r = sapply(data, function(d) cor(d$value1, d$value2))), by = .(pairs, am)]

The same trait in different environments (e.g. genotype × environment)

# growth: one row per animal x farm
r2p_nest(growth, names_from = "farm", cols = c("adg", "backfat"), id = "animal")
# -> per trait: animal | farm1 | farm2, ready for a cross-environment correlation

Clean many Excel files and put them back where they came from

dt <- import_xlsx(list.files("raw", pattern = "\\.xlsx$", full.names = TRUE))
dt <- dt[!is.na(weight)]                  # any data.table processing
export_xlsx(dt, path = "cleaned")         # same file names, same sheet names

Input files for command-line breeding software

export_nest(nested, path = "hiblup_runs", na = "NA")       # HIBLUP
export_nest(nested, path = "dmu_runs", na = "-9999",       # DMU: numeric missing code,
            col.names = FALSE)                             #      no header

Design principles

Upgrading from 0.1.x

Version 0.2.0 unifies argument names (cols2l, cols2bind, trait, nest_dt, group_cols, export_path, rbind, … → cols, data, by, path, combine, …); the old names are no longer accepted. Cross-validation no longer depends on rsample: the splits column is replaced by train_idx / validate_idx. See NEWS for the complete old → new table.

Cheat sheet

mintyr package quick reference guide and cheatsheet

Acknowledgments

AI assistants helped turn the initial ideas behind mintyr into code, refine its structure and write its documentation.