desc_stats() computes descriptive statistics for numeric columns, overall or by group, in one tidy table (one row per group x variable, or one row per group in wide shape). The statistic sets follow rstatix::get_summary_stats(). An optional total row gives the statistics of every variable over all records; results can be returned as numbers or as report-ready text such as "26.66 ± 4.51".

desc_stats(
  data,
  cols = NULL,
  by = NULL,
  type = "common",
  stats = NULL,
  total = FALSE,
  total_label = "Total",
  min_n = 1L,
  probs = c(0, 0.25, 0.5, 0.75, 1),
  conf_level = 0.95,
  na_rm = TRUE,
  digits = NULL,
  fmt = NULL,
  labels = NULL,
  shape = "long",
  sep = "_",
  out_type = "dt"
)

Arguments

data

A data.frame or data.table. It is never modified.

cols

Numeric columns to summarise, as names or indices. Default NULL: every numeric column not in by.

by

Grouping column(s), as names or indices. Default NULL. Must not be named variable or value.

type

Preset set of statistics (ignored when stats or fmt is given):

"full"

n, min, max, median, q1, q3, iqr, mad, mean, sd, se, ci, cv, skew, kurt

"common" (default)

n, min, max, median, iqr, mean, sd, se, ci

"robust"

n, median, iqr

"five_number"

n, min, q1, median, q3, max

"mean_sd", "mean_se"

n, mean and sd / se

"mean_ci"

n, mean, ci_low, ci_high

"median_iqr", "median_mad"

n, median and iqr / mad

"quantile"

n and the quantiles given by probs

"mean", "median"

n and the mean / median

stats

Optional character vector of statistics, overriding type: any of "n", "n_miss", "min", "max", "mean", "median", "q1", "q3", "iqr", "mad", "sd", "se", "ci", "ci_low", "ci_high", "cv", "skew", "kurt" and "quantile", in the order wanted.

total

Logical. If TRUE, add totals following one rule: across groups the statistics are recomputed on all records, across variables only counts are added up (different variables are not one quantity).

  • Total row (with by): an extra group whose by columns hold total_label; every variable is summarised over all records, so its n is the sum of the group n.

  • Total variable / column (with two or more variables, when n or n_miss is requested or a fmt template uses only counts): an extra variable total_label, last, whose n / n_miss are the sums over the variables in each row; its other statistics are NA, and in wide shape its empty columns are dropped. A count-only wide table thus gets row and column totals, as janitor::adorn_totals().

Without by only the total variable is added. Default FALSE.

total_label

Label of the total row and the total variable. Default "Total". Must not already occur in a by column or as a variable name (or label).

min_n

Minimum number of (non-missing) values a group needs for its statistics to be reported. Groups with fewer values keep n / n_miss but all other statistics are NA, so that e.g. the SD of a farm with two animals is not mistaken for a reliable figure. Default 1.

probs

Probabilities for "quantile"; output columns are named q0, q25, q2.5, ... Default c(0, 0.25, 0.5, 0.75, 1).

conf_level

Confidence level of ci. Default 0.95.

na_rm

Logical. If TRUE (default) missing values are removed before computing statistics, and n counts the non-missing values. If FALSE, any missing value makes the statistics NA and n counts all values.

digits

NULL (default, no rounding), one non-negative integer for all variables, or a named vector per variable, optionally with one unnamed default: c(2, adg = 0, bf = 1) rounds adg to 0, bf to 1 and every other variable to 2 decimals (variables without a value are not rounded). The counts n / n_miss are integers. With fmt, these are the decimals of placeholders without their own {stat:d} (default 2).

fmt

NULL (default) or a character vector of templates that turn the statistics into report-ready text, e.g. "{mean} ± {sd}". Each template becomes one text column, replacing the numeric statistics; name the elements to name the columns (default: the template without braces, e.g. "mean ± sd"). Placeholders:

  • {stat} — any statistic listed in stats (except "quantile"), with digits decimals (counts without decimals).

  • {stat:d} — with d decimals, e.g. {mean:1}.

  • {stat:d\%} — multiplied by 100 and followed by \%, e.g. {cv:1\%}.

  • {q2.5}, {q97.5}, ... — percentiles ({q1} / {q3} are the quartiles).

The statistics needed by the templates are computed automatically; type, stats and probs are then ignored. Missing values are shown as "NA" (e.g. the SD of a single record).

labels

NULL (default) or a named character vector of display names for the variables, e.g. c(adg = "ADG (g)", bf = "Backfat (mm)"). Applied to the variable column (long shape) or the column names (wide shape). Variables without a label keep their name.

shape

"long" (default): one row per group x variable, one column per statistic. "wide": one row per group (total row last), columns <variable><sep><statistic> in variable order, as in a typical report table. With a single statistic or template (e.g. stats = "mean", fmt = "{mean} ± {sd}") the columns are simply named after the variables. Both shapes show the same numbers.

sep

Separator between variable and statistic in wide column names. Default "_".

out_type

"dt" (default) for a data.table, "df" for a data.frame.

Value

A data.table (or data.frame). Long shape: the by columns, variable and one column per statistic (or per fmt template, as text). Wide shape: the by columns followed by one column per variable x statistic (or template).

Details

Definitions: q1 / q3 and quantile use stats::quantile() (type 7), iqr = q3 - q1, mad is stats::mad() (scaled by 1.4826), se = sd / sqrt(n), ci is the half-width of the t-based confidence interval of the mean (qt((1 + conf_level) / 2, n - 1) * se), and ci_low / ci_high are its bounds (mean -/+ ci), cv = sd / mean as a ratio (shown in percent with fmt = "{cv:1\%}"), skew is the adjusted Fisher-Pearson skewness G1 and kurt the excess kurtosis G2, as in SAS, SPSS and Excel's SKEW() / KURT() (NA with fewer than 3 / 4 values). The CV is only meaningful for strictly positive data and is NA when a group contains values <= 0. Statistics that are undefined for a group (e.g. sd with one value) are NA.

Computation: statistics are computed column by column on the input table (no reshaping to long format, so memory use stays close to the input size). n, mean, sd, min, max and median use data.table's optimised grouped functions, quantiles are read off once-sorted groups, so tens of thousands of groups (sires, litters, pens) are summarised in about a second per million records.

Rows are ordered by variable (in the order of cols), then by group (in the original order of the by values: numeric order, factor levels, alphabetical for text), with the total row last. With a total row, by columns that are not factors or character are returned as character; factors gain the label as their last level.

See also

top_perc() for statistics of the top / bottom X% per group.

Examples

# Example 1: Common statistics for every numeric column of iris
desc_stats(iris)
#>        variable     n   min   max median   iqr     mean        sd         se
#>          <char> <int> <num> <num>  <num> <num>    <num>     <num>      <num>
#> 1: Sepal.Length   150   4.3   7.9   5.80   1.3 5.843333 0.8280661 0.06761132
#> 2:  Sepal.Width   150   2.0   4.4   3.00   0.5 3.057333 0.4358663 0.03558833
#> 3: Petal.Length   150   1.0   6.9   4.35   3.5 3.758000 1.7652982 0.14413600
#> 4:  Petal.Width   150   0.1   2.5   1.30   1.5 1.199333 0.7622377 0.06223645
#>            ci
#>         <num>
#> 1: 0.13360085
#> 2: 0.07032302
#> 3: 0.28481463
#> 4: 0.12298004

# Example 2: Mean and SD by group, with a total row over all records
desc_stats(
  iris,
  by = "Species",                 # Grouping column
  type = "mean_sd",               # Preset: n, mean, sd
  total = TRUE,                   # n = sum of the groups, mean / sd of all records
  digits = 2                      # Round the statistics
)
#>        Species     variable     n  mean    sd
#>         <fctr>       <char> <int> <num> <num>
#>  1:     setosa Sepal.Length    50  5.01  0.35
#>  2: versicolor Sepal.Length    50  5.94  0.52
#>  3:  virginica Sepal.Length    50  6.59  0.64
#>  4:      Total Sepal.Length   150  5.84  0.83
#>  5:     setosa  Sepal.Width    50  3.43  0.38
#>  6: versicolor  Sepal.Width    50  2.77  0.31
#>  7:  virginica  Sepal.Width    50  2.97  0.32
#>  8:      Total  Sepal.Width   150  3.06  0.44
#>  9:     setosa Petal.Length    50  1.46  0.17
#> 10: versicolor Petal.Length    50  4.26  0.47
#> 11:  virginica Petal.Length    50  5.55  0.55
#> 12:      Total Petal.Length   150  3.76  1.77
#> 13:     setosa  Petal.Width    50  0.25  0.11
#> 14: versicolor  Petal.Width    50  1.33  0.20
#> 15:  virginica  Petal.Width    50  2.03  0.27
#> 16:      Total  Petal.Width   150  1.20  0.76
#> 17:     setosa        Total   200    NA    NA
#> 18: versicolor        Total   200    NA    NA
#> 19:  virginica        Total   200    NA    NA
#> 20:      Total        Total   600    NA    NA
#>        Species     variable     n  mean    sd
#>         <fctr>       <char> <int> <num> <num>

# Example 3: Count table with row and column totals
# (counts are the only statistic that adds up across variables)
desc_stats(
  iris,
  by = "Species",
  stats = "n",
  total = TRUE,                   # Total row and Total column
  shape = "wide"
)
#>       Species Sepal.Length Sepal.Width Petal.Length Petal.Width Total
#>        <fctr>        <int>       <int>        <int>       <int> <int>
#> 1:     setosa           50          50           50          50   200
#> 2: versicolor           50          50           50          50   200
#> 3:  virginica           50          50           50          50   200
#> 4:      Total          150         150          150         150   600

# Without groups the total adds up the counts of all variables
desc_stats(iris, type = "mean_sd", total = TRUE, digits = 2)
#>        variable     n  mean    sd
#>          <char> <int> <num> <num>
#> 1: Sepal.Length   150  5.84  0.83
#> 2:  Sepal.Width   150  3.06  0.44
#> 3: Petal.Length   150  3.76  1.77
#> 4:  Petal.Width   150  1.20  0.76
#> 5:        Total   600    NA    NA

# Example 3b: Report table - groups in rows, traits in columns, "mean ± sd"
desc_stats(
  mtcars,
  cols = c("mpg", "hp", "wt"),
  by = "cyl",
  fmt = "{mean} ± {sd}",          # Report-ready text
  total = TRUE,
  shape = "wide"                  # One row per group, one column per trait
)
#>       cyl          mpg             hp          wt
#>    <char>       <char>         <char>      <char>
#> 1:      4 26.66 ± 4.51  82.64 ± 20.93 2.29 ± 0.57
#> 2:      6 19.74 ± 1.45 122.29 ± 24.26 3.12 ± 0.36
#> 3:      8 15.10 ± 2.56 209.21 ± 50.98 4.00 ± 0.76
#> 4:  Total 20.09 ± 6.03 146.69 ± 68.56 3.22 ± 0.98

# Example 4: Several templates, per-placeholder decimals, CV in percent and
# percentiles
desc_stats(
  mtcars,
  cols = c("mpg", "wt"),
  by = "am",
  fmt = c(N = "{n}",
          "Mean ± SD" = "{mean:1} ± {sd:1}",
          "CV" = "{cv:1%}",
          "Median [P2.5, P97.5]" = "{median:1} [{q2.5:1}, {q97.5:1}]"),
  total = TRUE
)
#>        am variable      N  Mean ± SD     CV Median [P2.5, P97.5]
#>    <char>   <char> <char>     <char> <char>               <char>
#> 1:      0      mpg     19 17.1 ± 3.8  22.4%    17.3 [10.4, 23.7]
#> 2:      1      mpg     13 24.4 ± 6.2  25.3%    22.8 [15.2, 33.4]
#> 3:  Total      mpg     32 20.1 ± 6.0  30.0%    19.2 [10.4, 32.7]
#> 4:      0       wt     19  3.8 ± 0.8  20.6%       3.5 [2.8, 5.4]
#> 5:      1       wt     13  2.4 ± 0.6  25.6%       2.3 [1.5, 3.4]
#> 6:  Total       wt     32  3.2 ± 1.0  30.4%       3.3 [1.6, 5.4]
#> 7:      0    Total     38       <NA>   <NA>                 <NA>
#> 8:      1    Total     26       <NA>   <NA>                 <NA>
#> 9:  Total    Total     64       <NA>   <NA>                 <NA>

# Example 5: Quantiles, returned as a data.frame
desc_stats(iris, cols = 1:2, type = "quantile",
           probs = c(0.05, 0.5, 0.95), out_type = "df")
#>       variable   n    q5 q50   q95
#> 1 Sepal.Length 150 4.600 5.8 7.255
#> 2  Sepal.Width 150 2.345 3.0 3.800