mintyr turns “many groups x many variables” data into analysis-ready pieces and back into files. A typical analysis follows one loop:

 files --> import --> reshape & nest --> cross-validate / summarise --> export --> files

Each step has its own article:

Step Functions Article
Import and export import_xlsx(), import_csv(), export_xlsx(), export_nest(), export_list() vignette("import-and-export")
Reshape and nest w2l_nest(), w2l_split(), c2p_nest(), r2p_nest() vignette("reshape-and-nest")
Cross-validation split_cv(), nest_cv() vignette("cross-validation")
Descriptive statistics desc_stats(), top_perc(), format_digits() vignette("descriptive-statistics")
Utilities get_path_info(), mintyr_example() vignette("utilities")
library(mintyr)
library(data.table)
#> 
#> Attaching package: 'data.table'
#> The following object is masked from 'package:base':
#> 
#>     %notin%

The loop in five steps

1. Import. Several workbooks become one table; excel_name and sheet_name record where every row came from.

files <- mintyr_example(mintyr_examples("xlsx_test"))
raw <- import_xlsx(files)
head(raw)
#>    excel_name sheet_name  col1   col2   col3
#>        <char>     <char> <num> <char> <lgcl>
#> 1: xlsx_test1     Sheet1     4      d  FALSE
#> 2: xlsx_test1     Sheet1     5      f   TRUE
#> 3: xlsx_test1     Sheet1     6      e   TRUE
#> 4: xlsx_test1     Sheet2     1      a   TRUE
#> 5: xlsx_test1     Sheet2     2      b  FALSE
#> 6: xlsx_test1     Sheet2     3      c   TRUE

2. Describe. A report table per group, with a total row.

desc_stats(mtcars, cols = c("mpg", "hp", "wt"), by = "cyl",
           fmt = "{mean} ± {sd}", total = TRUE, shape = "wide")
#>       cyl          mpg             hp          wt
#>    <char>       <char>         <char>      <char>
#> 1:      4 26.66 ± 4.51  82.64 ± 20.93 2.29 ± 0.57
#> 2:      6 19.74 ± 1.45 122.29 ± 24.26 3.12 ± 0.36
#> 3:      8 15.10 ± 2.56 209.21 ± 50.98 4.00 ± 0.76
#> 4:  Total 20.09 ± 6.03 146.69 ± 68.56 3.22 ± 0.98

3. Reshape and nest. One row per trait and group, the data in a list-column.

nested <- w2l_nest(mtcars, cols = c("mpg", "qsec"), by = "am")
nested
#>      name    am               data
#>    <char> <num>             <list>
#> 1:    mpg     1 <data.table[13x9]>
#> 2:    mpg     0 <data.table[19x9]>
#> 3:   qsec     1 <data.table[13x9]>
#> 4:   qsec     0 <data.table[19x9]>

4. Cross-validate inside every piece. Reproducible 4-fold CV; a model per fold, then the mean predictive ability per trait and group.

cv <- nest_cv(nested, v = 4, seed = 2026)
cv[, r := mapply(function(tr, va) {
  fit <- lm(value ~ wt + hp, data = tr)
  cor(predict(fit, va), va$value)
}, train, validate)]
cv[, .(mean_r = round(mean(r), 3)), by = .(name, am)]
#>      name    am mean_r
#>    <char> <num>  <num>
#> 1:    mpg     1  0.816
#> 2:    mpg     0  0.881
#> 3:   qsec     1  0.884
#> 4:   qsec     0  0.858

5. Export. One file per trait and group, e.g. as input for HIBLUP or DMU.

out <- file.path(tempdir(), "by_trait")
files <- export_nest(nested, path = out)
#> [ export_nest ] Auto-selected nested columns: data
#> [ export_nest ] Auto-selected grouping columns: name, am
#> [ export_nest ] Export complete. 4 file(s) written to: /tmp/Rtmp8EhuAs/by_trait
basename(dirname(files))
#> [1] "1" "0" "1" "0"
unlink(out, recursive = TRUE)

Design principles

  • No side effects: input objects are never modified by reference.
  • No silent data loss: file names are sanitised, and functions stop instead of overwriting files or returning ambiguous results.
  • Consistent arguments: data, cols, by, out_type, path mean the same thing in every function.
  • Light dependencies: data.table, readxl, writexl and base R.