mintyr turns “many groups x many variables” data into analysis-ready pieces and back into files. A typical analysis follows one loop:
files --> import --> reshape & nest --> cross-validate / summarise --> export --> files
Each step has its own article:
| Step | Functions | Article |
|---|---|---|
| Import and export |
import_xlsx(), import_csv(),
export_xlsx(), export_nest(),
export_list()
|
vignette("import-and-export") |
| Reshape and nest |
w2l_nest(), w2l_split(),
c2p_nest(), r2p_nest()
|
vignette("reshape-and-nest") |
| Cross-validation |
split_cv(), nest_cv()
|
vignette("cross-validation") |
| Descriptive statistics |
desc_stats(), top_perc(),
format_digits()
|
vignette("descriptive-statistics") |
| Utilities |
get_path_info(), mintyr_example()
|
vignette("utilities") |
library(mintyr)
library(data.table)
#>
#> Attaching package: 'data.table'
#> The following object is masked from 'package:base':
#>
#> %notin%1. Import. Several workbooks become one table;
excel_name and sheet_name record where every
row came from.
files <- mintyr_example(mintyr_examples("xlsx_test"))
raw <- import_xlsx(files)
head(raw)
#> excel_name sheet_name col1 col2 col3
#> <char> <char> <num> <char> <lgcl>
#> 1: xlsx_test1 Sheet1 4 d FALSE
#> 2: xlsx_test1 Sheet1 5 f TRUE
#> 3: xlsx_test1 Sheet1 6 e TRUE
#> 4: xlsx_test1 Sheet2 1 a TRUE
#> 5: xlsx_test1 Sheet2 2 b FALSE
#> 6: xlsx_test1 Sheet2 3 c TRUE2. Describe. A report table per group, with a total row.
desc_stats(mtcars, cols = c("mpg", "hp", "wt"), by = "cyl",
fmt = "{mean} ± {sd}", total = TRUE, shape = "wide")
#> cyl mpg hp wt
#> <char> <char> <char> <char>
#> 1: 4 26.66 ± 4.51 82.64 ± 20.93 2.29 ± 0.57
#> 2: 6 19.74 ± 1.45 122.29 ± 24.26 3.12 ± 0.36
#> 3: 8 15.10 ± 2.56 209.21 ± 50.98 4.00 ± 0.76
#> 4: Total 20.09 ± 6.03 146.69 ± 68.56 3.22 ± 0.983. Reshape and nest. One row per trait and group, the data in a list-column.
nested <- w2l_nest(mtcars, cols = c("mpg", "qsec"), by = "am")
nested
#> name am data
#> <char> <num> <list>
#> 1: mpg 1 <data.table[13x9]>
#> 2: mpg 0 <data.table[19x9]>
#> 3: qsec 1 <data.table[13x9]>
#> 4: qsec 0 <data.table[19x9]>4. Cross-validate inside every piece. Reproducible 4-fold CV; a model per fold, then the mean predictive ability per trait and group.
cv <- nest_cv(nested, v = 4, seed = 2026)
cv[, r := mapply(function(tr, va) {
fit <- lm(value ~ wt + hp, data = tr)
cor(predict(fit, va), va$value)
}, train, validate)]
cv[, .(mean_r = round(mean(r), 3)), by = .(name, am)]
#> name am mean_r
#> <char> <num> <num>
#> 1: mpg 1 0.816
#> 2: mpg 0 0.881
#> 3: qsec 1 0.884
#> 4: qsec 0 0.8585. Export. One file per trait and group, e.g. as input for HIBLUP or DMU.
out <- file.path(tempdir(), "by_trait")
files <- export_nest(nested, path = out)
#> [ export_nest ] Auto-selected nested columns: data
#> [ export_nest ] Auto-selected grouping columns: name, am
#> [ export_nest ] Export complete. 4 file(s) written to: /tmp/Rtmp8EhuAs/by_trait
basename(dirname(files))
#> [1] "1" "0" "1" "0"
unlink(out, recursive = TRUE)data,
cols, by, out_type,
path mean the same thing in every function.data.table,
readxl, writexl and base R.