desc_stats() computes descriptive statistics for numeric columns, overall
or by group, in one tidy table (one row per group x variable, or one row
per group in wide shape). The statistic sets follow
rstatix::get_summary_stats(). An optional total row gives the
statistics of every variable over all records; results can be returned
as numbers or as report-ready text such as "26.66 ± 4.51".
desc_stats(
data,
cols = NULL,
by = NULL,
type = "common",
stats = NULL,
total = FALSE,
total_label = "Total",
min_n = 1L,
probs = c(0, 0.25, 0.5, 0.75, 1),
conf_level = 0.95,
na_rm = TRUE,
digits = NULL,
fmt = NULL,
labels = NULL,
shape = "long",
sep = "_",
out_type = "dt"
)A data.frame or data.table. It is never modified.
Numeric columns to summarise, as names or indices. Default
NULL: every numeric column not in by.
Grouping column(s), as names or indices. Default NULL. Must not
be named variable or value.
Preset set of statistics (ignored when stats or fmt is
given):
"full"n, min, max, median, q1, q3, iqr, mad, mean, sd, se, ci, cv, skew, kurt
"common" (default)n, min, max, median, iqr, mean, sd, se, ci
"robust"n, median, iqr
"five_number"n, min, q1, median, q3, max
"mean_sd", "mean_se"n, mean and sd / se
"mean_ci"n, mean, ci_low, ci_high
"median_iqr", "median_mad"n, median and iqr / mad
"quantile"n and the quantiles given by probs
"mean", "median"n and the mean / median
Optional character vector of statistics, overriding type:
any of "n", "n_miss", "min", "max", "mean", "median",
"q1", "q3", "iqr", "mad", "sd", "se", "ci", "ci_low",
"ci_high", "cv", "skew", "kurt" and "quantile", in the order
wanted.
Logical. If TRUE, add totals following one rule: across
groups the statistics are recomputed on all records, across variables
only counts are added up (different variables are not one quantity).
Total row (with by): an extra group whose by columns hold
total_label; every variable is summarised over all records, so its
n is the sum of the group n.
Total variable / column (with two or more variables, when
n or n_miss is requested or a fmt template uses only counts):
an extra variable
total_label, last, whose n / n_miss are the sums over the
variables in each row; its other statistics are NA, and in wide
shape its empty columns are dropped. A count-only wide table thus
gets row and column totals, as janitor::adorn_totals().
Without by only the total variable is added. Default FALSE.
Label of the total row and the total variable. Default
"Total". Must not already occur in a by column or as a variable name
(or label).
Minimum number of (non-missing) values a group needs for its
statistics to be reported. Groups with fewer values keep n / n_miss
but all other statistics are NA, so that e.g. the SD of a farm with
two animals is not mistaken for a reliable figure. Default 1.
Probabilities for "quantile"; output columns are named
q0, q25, q2.5, ... Default c(0, 0.25, 0.5, 0.75, 1).
Confidence level of ci. Default 0.95.
Logical. If TRUE (default) missing values are removed
before computing statistics, and n counts the non-missing values. If
FALSE, any missing value makes the statistics NA and n counts all
values.
NULL (default, no rounding), one non-negative integer for
all variables, or a named vector per variable, optionally with one
unnamed default: c(2, adg = 0, bf = 1) rounds adg to 0, bf to 1
and every other variable to 2 decimals (variables without a value are
not rounded). The counts n / n_miss are integers. With fmt, these
are the decimals of placeholders without their own {stat:d} (default
2).
NULL (default) or a character vector of templates that turn
the statistics into report-ready text, e.g. "{mean} ± {sd}". Each
template becomes one text column, replacing the numeric statistics;
name the elements to name the columns (default: the template without
braces, e.g. "mean ± sd"). Placeholders:
{stat} — any statistic listed in stats (except
"quantile"), with digits decimals (counts without decimals).
{stat:d} — with d decimals, e.g. {mean:1}.
{stat:d\%} — multiplied by 100 and followed by \%, e.g.
{cv:1\%}.
{q2.5}, {q97.5}, ... — percentiles ({q1} / {q3} are the
quartiles).
The statistics needed by the templates are computed automatically;
type, stats and probs are then ignored. Missing values are shown
as "NA" (e.g. the SD of a single record).
NULL (default) or a named character vector of display
names for the variables, e.g. c(adg = "ADG (g)", bf = "Backfat (mm)").
Applied to the variable column (long shape) or the column names (wide
shape). Variables without a label keep their name.
"long" (default): one row per group x variable, one column
per statistic. "wide": one row per group (total row last), columns
<variable><sep><statistic> in variable order, as
in a typical report table. With a single statistic or template (e.g.
stats = "mean", fmt = "{mean} ± {sd}") the columns are simply named
after the variables. Both shapes show the same numbers.
Separator between variable and statistic in wide column names.
Default "_".
"dt" (default) for a data.table, "df" for a
data.frame.
A data.table (or data.frame). Long shape: the by columns,
variable and one column per statistic (or per fmt template, as
text). Wide shape: the by columns followed by one column per
variable x statistic (or template).
Definitions: q1 / q3 and quantile use stats::quantile()
(type 7), iqr = q3 - q1, mad is stats::mad() (scaled by 1.4826),
se = sd / sqrt(n), ci is the half-width of the t-based confidence
interval of the mean (qt((1 + conf_level) / 2, n - 1) * se), and
ci_low / ci_high are its bounds (mean -/+ ci),
cv = sd / mean as a ratio (shown in percent with fmt = "{cv:1\%}"),
skew is the adjusted Fisher-Pearson skewness G1 and kurt the excess
kurtosis G2, as in SAS, SPSS and Excel's SKEW() / KURT() (NA with
fewer than 3 / 4 values).
The CV is only meaningful for strictly positive data and is NA when a
group contains values <= 0. Statistics that are undefined for a group
(e.g. sd with one value) are NA.
Computation: statistics are computed column by column on the input table
(no reshaping to long format, so memory use stays close to the input
size). n, mean, sd, min, max and median use data.table's
optimised grouped functions, quantiles are read off once-sorted groups,
so tens of thousands of groups (sires, litters, pens) are summarised in
about a second per million records.
Rows are ordered by variable (in the order of cols),
then by group (in the original order of the by values: numeric order,
factor levels, alphabetical for text), with the total row last. With a
total row, by columns that are not factors or character are returned as
character; factors gain the label as their last level.
top_perc() for statistics of the top / bottom X% per group.
# Example 1: Common statistics for every numeric column of iris
desc_stats(iris)
#> variable n min max median iqr mean sd se
#> <char> <int> <num> <num> <num> <num> <num> <num> <num>
#> 1: Sepal.Length 150 4.3 7.9 5.80 1.3 5.843333 0.8280661 0.06761132
#> 2: Sepal.Width 150 2.0 4.4 3.00 0.5 3.057333 0.4358663 0.03558833
#> 3: Petal.Length 150 1.0 6.9 4.35 3.5 3.758000 1.7652982 0.14413600
#> 4: Petal.Width 150 0.1 2.5 1.30 1.5 1.199333 0.7622377 0.06223645
#> ci
#> <num>
#> 1: 0.13360085
#> 2: 0.07032302
#> 3: 0.28481463
#> 4: 0.12298004
# Example 2: Mean and SD by group, with a total row over all records
desc_stats(
iris,
by = "Species", # Grouping column
type = "mean_sd", # Preset: n, mean, sd
total = TRUE, # n = sum of the groups, mean / sd of all records
digits = 2 # Round the statistics
)
#> Species variable n mean sd
#> <fctr> <char> <int> <num> <num>
#> 1: setosa Sepal.Length 50 5.01 0.35
#> 2: versicolor Sepal.Length 50 5.94 0.52
#> 3: virginica Sepal.Length 50 6.59 0.64
#> 4: Total Sepal.Length 150 5.84 0.83
#> 5: setosa Sepal.Width 50 3.43 0.38
#> 6: versicolor Sepal.Width 50 2.77 0.31
#> 7: virginica Sepal.Width 50 2.97 0.32
#> 8: Total Sepal.Width 150 3.06 0.44
#> 9: setosa Petal.Length 50 1.46 0.17
#> 10: versicolor Petal.Length 50 4.26 0.47
#> 11: virginica Petal.Length 50 5.55 0.55
#> 12: Total Petal.Length 150 3.76 1.77
#> 13: setosa Petal.Width 50 0.25 0.11
#> 14: versicolor Petal.Width 50 1.33 0.20
#> 15: virginica Petal.Width 50 2.03 0.27
#> 16: Total Petal.Width 150 1.20 0.76
#> 17: setosa Total 200 NA NA
#> 18: versicolor Total 200 NA NA
#> 19: virginica Total 200 NA NA
#> 20: Total Total 600 NA NA
#> Species variable n mean sd
#> <fctr> <char> <int> <num> <num>
# Example 3: Count table with row and column totals
# (counts are the only statistic that adds up across variables)
desc_stats(
iris,
by = "Species",
stats = "n",
total = TRUE, # Total row and Total column
shape = "wide"
)
#> Species Sepal.Length Sepal.Width Petal.Length Petal.Width Total
#> <fctr> <int> <int> <int> <int> <int>
#> 1: setosa 50 50 50 50 200
#> 2: versicolor 50 50 50 50 200
#> 3: virginica 50 50 50 50 200
#> 4: Total 150 150 150 150 600
# Without groups the total adds up the counts of all variables
desc_stats(iris, type = "mean_sd", total = TRUE, digits = 2)
#> variable n mean sd
#> <char> <int> <num> <num>
#> 1: Sepal.Length 150 5.84 0.83
#> 2: Sepal.Width 150 3.06 0.44
#> 3: Petal.Length 150 3.76 1.77
#> 4: Petal.Width 150 1.20 0.76
#> 5: Total 600 NA NA
# Example 3b: Report table - groups in rows, traits in columns, "mean ± sd"
desc_stats(
mtcars,
cols = c("mpg", "hp", "wt"),
by = "cyl",
fmt = "{mean} ± {sd}", # Report-ready text
total = TRUE,
shape = "wide" # One row per group, one column per trait
)
#> cyl mpg hp wt
#> <char> <char> <char> <char>
#> 1: 4 26.66 ± 4.51 82.64 ± 20.93 2.29 ± 0.57
#> 2: 6 19.74 ± 1.45 122.29 ± 24.26 3.12 ± 0.36
#> 3: 8 15.10 ± 2.56 209.21 ± 50.98 4.00 ± 0.76
#> 4: Total 20.09 ± 6.03 146.69 ± 68.56 3.22 ± 0.98
# Example 4: Several templates, per-placeholder decimals, CV in percent and
# percentiles
desc_stats(
mtcars,
cols = c("mpg", "wt"),
by = "am",
fmt = c(N = "{n}",
"Mean ± SD" = "{mean:1} ± {sd:1}",
"CV" = "{cv:1%}",
"Median [P2.5, P97.5]" = "{median:1} [{q2.5:1}, {q97.5:1}]"),
total = TRUE
)
#> am variable N Mean ± SD CV Median [P2.5, P97.5]
#> <char> <char> <char> <char> <char> <char>
#> 1: 0 mpg 19 17.1 ± 3.8 22.4% 17.3 [10.4, 23.7]
#> 2: 1 mpg 13 24.4 ± 6.2 25.3% 22.8 [15.2, 33.4]
#> 3: Total mpg 32 20.1 ± 6.0 30.0% 19.2 [10.4, 32.7]
#> 4: 0 wt 19 3.8 ± 0.8 20.6% 3.5 [2.8, 5.4]
#> 5: 1 wt 13 2.4 ± 0.6 25.6% 2.3 [1.5, 3.4]
#> 6: Total wt 32 3.2 ± 1.0 30.4% 3.3 [1.6, 5.4]
#> 7: 0 Total 38 <NA> <NA> <NA>
#> 8: 1 Total 26 <NA> <NA> <NA>
#> 9: Total Total 64 <NA> <NA> <NA>
# Example 5: Quantiles, returned as a data.frame
desc_stats(iris, cols = 1:2, type = "quantile",
probs = c(0.05, 0.5, 0.95), out_type = "df")
#> variable n q5 q50 q95
#> 1 Sepal.Length 150 4.600 5.8 7.255
#> 2 Sepal.Width 150 2.345 3.0 3.800