Statistics in the command window#
Every name on this page is typed at the command window (or passed to
ICoreBlocks --console "<line>") and answers what MATLAB's Statistics and
Machine Learning Toolbox answers for the same arguments — same parameter
order, same defaults, same tie and edge rules. This is the largest of the
toolbox families here: 190 names, covering the distributions, the
descriptive statistics, the classical tests, regression, and the
numeric-in/numeric-out half of clustering and multivariate analysis.
What is here is everything that takes numbers and answers numbers. What is
NOT here is everything that answers an OBJECT — fitlm, fitcsvm,
makedist, cvpartition and the rest of the model classes — because this
console has no object types. The last section says what to reach for instead.
The families, and what to reach for#
| You want to | Use |
|---|---|
| Summarise a sample | mean, median, std, var, range, iqr, mad, prctile, quantile, zscore, skewness, kurtosis, moment, bounds |
| Summarise it robustly or by group | trimmean, geomean, harmmean, grpstats, tabulate, crosstab, tiedrank |
| Evaluate a distribution | <name>pdf, <name>cdf, <name>inv, <name>rnd, <name>stat, <name>fit for norm, unif, exp, poiss, bino, t, chi2, f, gam, beta, logn, wbl, rayl, ev, gev, gp, nbin, geo, hyge |
| …by name instead of by spelling | pdf(name, x, …), cdf, icdf, random(name, …), mle(x, "Distribution", name) |
| Evaluate a multivariate density | mvnpdf, mvncdf, mvnrnd, mvtpdf, mvtrnd, wishrnd, iwishrnd |
| Test a mean | ttest (one sample or PAIRED), ttest2 (two samples), ztest |
| Test a variance or a shape | vartest, vartest2, chi2gof, kstest, kstest2, lillietest, jbtest, adtest, runstest |
| Compare several groups | anova1, anova2, kruskalwallis, ranksum, signrank, signtest, friedman, multcompare |
| Fit a linear model | regress, robustfit, stepwisefit, glmfit/glmval, nlinfit with nlparci/nlpredci |
| Fit a penalised one | lasso, ridge, plsregress |
| Measure association | corr, partialcorr, cov, corrcov, tiedrank |
| Reduce dimensions | pca, pcacov, pcares, factoran, canoncorr, cmdscale, mdscale, procrustes, rotatefactors |
| Cluster | kmeans, kmedoids, linkage, cluster, clusterdata, cophenet, inconsistent, silhouette, optimalleaforder |
| Measure distance | pdist, pdist2, squareform, mahal, knnsearch, rangesearch |
| Classify and score | classify, confusionmat, perfcurve |
| Resample | randsample, datasample, bootstrp, bootci, jackknife |
| Look at a distribution empirically | ecdf, ksdensity |
The nineteen distribution families#
Each family has the same five or six spellings, and they mean the same thing in
every one: pdf the density, cdf the cumulative probability, inv the
quantile, rnd a DRAW, stat the mean and VARIANCE as [m, v], and fit the
maximum-likelihood estimate with its confidence interval. <name>cdf takes a
trailing "upper" for the upper tail computed exactly rather than as 1 - p.
| Family | Parameters | Spellings |
|---|---|---|
| Normal | mu, sigma | normpdf, normcdf, norminv, normrnd, normstat, normfit, normlike |
| Uniform | a, b | unifpdf, unifcdf, unifinv, unifrnd, unifstat, unifit |
| Exponential | mu, which is the MEAN and not a rate | exppdf, expcdf, expinv, exprnd, expstat, expfit |
| Poisson | lambda | poisspdf, poisscdf, poissinv, poissrnd, poisstat, poissfit |
| Binomial | n, p | binopdf, binocdf, binoinv, binornd, binostat, binofit |
| Student t | v | tpdf, tcdf, tinv, trnd |
| Chi-square | v | chi2pdf, chi2cdf, chi2inv, chi2rnd |
| F | v1, v2 | fpdf, fcdf, finv, frnd |
| Gamma | a, b — b is the SCALE, so the mean is a*b | gampdf, gamcdf, gaminv, gamrnd, gamstat, gamfit |
| Beta | a, b | betapdf, betacdf, betainv, betarnd, betastat |
| Lognormal | mu, sigma of the LOG | lognpdf, logncdf, logninv, lognrnd, lognstat, lognfit |
| Weibull | A, B — the SCALE first | wblpdf, wblcdf, wblinv, wblrnd, wblstat, wblfit |
| Rayleigh | b | raylpdf, raylcdf, raylinv, raylrnd, raylstat, raylfit |
| Extreme value | mu, sigma | evpdf, evcdf, evinv, evrnd, evstat, evfit |
| Generalized extreme value | k, sigma, mu | gevpdf, gevcdf, gevinv, gevrnd, gevstat |
| Generalized Pareto | k, sigma, theta | gppdf, gpcdf, gpinv, gprnd, gpstat |
| Negative binomial | r, p | nbinpdf, nbincdf, nbininv, nbinrnd, nbinstat |
| Geometric | p | geopdf, geocdf, geoinv, geornd, geostat |
| Hypergeometric | M, K, N | hygepdf, hygecdf, hygeinv, hygernd, hygestat |
Four families answer a fit this console does not have — beta, negative
binomial and the two generalized ones — and those refuse by name; their
densities, cumulatives, quantiles and draws are all here. mle(x, "Distribution", name)
reaches the fits it does have under one name.
A worked line or two#
The normal family, and a sample summarised:
>>> [normpdf(1), normcdf(1), norminv(0.975)]
[[0.241971, 0.841345, 1.95996]] # Matrix of Double
>>> x = [2.1, 3.4, 1.9, 4.2, 2.8, 3.1, 2.5, 3.9]; [mean(x), std(x), prctile(x, 75), iqr(x)]
[[2.9875, 0.82191, 3.65, 1.35]] # Matrix of Double
A one-sample t-test against a hypothesised mean, and the same data fitted with its exact intervals:
>>> x = [2.1, 3.4, 1.9, 4.2, 2.8, 3.1, 2.5, 3.9]; [h, p, ci] = ttest(x, 2.5)
h = 0 # Integer
p = 0.137319 # Double
ci = [[2.30037, 3.67463]] # Matrix of Double
>>> x = [2.1, 3.4, 1.9, 4.2, 2.8, 3.1, 2.5, 3.9]; [muhat, sigmahat, muci, sigmaci] = normfit(x)
muhat = 2.9875 # Double
sigmahat = 0.82191 # Double
muci = [[2.30037], [3.67463]] # Matrix of Double
sigmaci = [[0.543426], [1.67281]] # Matrix of Double
Regression, correlation with its p-value, and principal components:
>>> X = [1 2; 2 3; 3 5; 4 4; 5 7]; y = [2; 3; 5; 4; 8]; regress(y, [ones(5,1), X])
[[-0.6], [-0.0444444], [1.22222]] # Matrix of Double
>>> x = [1 2 3 4 5]; y = [2 4 5 4 5]; [rho, pval] = corr(x', y')
rho = 0.774597 # Double
pval = 0.124027 # Double
>>> X = [1 2; 2 3; 3 5; 4 4; 5 7]; [coeff, score, latent] = pca(X); latent'
[[5.91469, 0.285306]] # Matrix of Double
Distances and clustering, with the start pinned so the answer is reproducible:
>>> D = pdist([1 1; 2 2; 4 5]); [D; squareform(D)]
[[1.41421, 5, 3.60555], [0, 1.41421, 5], [1.41421, 0, 3.60555], [5, 3.60555, 0]] # Matrix of Double
>>> X = [1 1; 1.5 2; 3 4; 5 7; 3.5 5; 4.5 5; 3.5 4.5]; [idx, C] = kmeans(X, 2, "Start", [1 1; 5 7])
idx = [[1], [1], [2], [2], [2], [2], [2]] # Matrix of Double
C = [[1.25, 1.5], [3.9, 5.1]] # Matrix of Double
Every transcript on this page is real output, captured with
ICoreBlocks --console "<line>" on 2026-09-05 at commit bd974745.
Ten things that surprise people#
ttest(x, y)with two VECTORS is the PAIRED test. The same subjects measured twice,n - 1degrees of freedom. The unpaired one isttest2. Nothing in either call says which is which, and they answer different p-values for the same pair.pdf("normal", 1)is NaN wherenormpdf(1)is 0.242. The generic, string-named forms read an omitted parameter as ZERO, not as the family's default — MATLAB'spdf.mdoesif nargin < 3, a = 0; end, and a normal density with sigma 0 is NaN.randomdoes not zero-fill at all; it forwards its arguments, sorandom("normal")is an error on both sides:``
>>> [pdf("normal", 1), normpdf(1)] [[NaN, 0.241971]] # Matrix of Double``wblpdf's parameters are (SCALE, SHAPE) —wblpdf(x, A, B)has A the scale and B the exponent, the reverse of the (k, λ) most textbooks write.ridge's default answer is not on your data's scale and has no intercept.ridge(y, X, k)andridge(y, X, k, 1)are the same call: both answer p coefficients for the centred, SCALED design. Onlyridge(y, X, k, 0)divides back by the column deviations and prepends the intercept — so the two forms differ in SIZE as well as in value.robustfitADDS a constant column;regressnever does. Two neighbouring functions with opposite conventions. Pass"off"asrobustfit's fifth argument to stop it.silhouette's default metric is SQUARED Euclidean, not Euclidean. It is the one place in this toolbox where a squared distance is the default, and it moves every number.runstest's default reference is the MEAN, not the median, and values exactly equal to it are DROPPED before any run is counted.cluster(Z, "MaxClust", k)cuts by DISTANCE, not by the inconsistency its documentation calls the default. MATLAB'scluster.mrewrites the criterion whenever no cutoff was given and none was named; reproduced here.lasso'sBcomes back in ASCENDING lambda order, on the ORIGINAL scale of X, and the sequence can be shorter than you asked for. The fit runs downward from lambdaMax and reverses at the end, and it stops early once it passesDFmaxor its error falls below a thousandth of the null model's. It also standardises with the POPULATION deviation whereridgeuses the sample one.- Anything that DRAWS cannot be compared to MATLAB and does not pretend
to be.
normrnd,random,bootstrp,datasample,kmeanswithout a"Start"— MATLAB's stream is a Mersenne Twister and this console's is not, so the same seed gives different numbers and no seed makes them agree. The translation is sound; the values are the sampler's. Pin the start, or compare a statistic of the draw rather than the draw.
What is not here#
Refused, with the reason in the message — the common shape is a result that MATLAB returns as a STRUCT or a CELL, which this console has no kind for:
anova1's table and stats,chi2gof's third output,runstest's third,robustfit's stats,lasso'sFitInfo— every one of these refuses by name rather than being quietly absent, and says which rule it met.corr's Spearman and Kendall p-values (the correlations themselves are here). MATLAB reaches those two through a permutation distribution in compiled code.nlpredci's simultaneous OBSERVATION interval — its Scheffé parameter is chosen by an unpublished rule.rotatefactorson an ALREADY-rotated loading matrix, which restarts MATLAB from a random orthogonal matrix.runstest(..., "ud"), which needs a shipped data file.
Not known to the console at all — typing one of these answers
unknown function '<name>':
| If you reach for | Use |
|---|---|
fitlm, fitglm, fitnlm, stepwiselm | regress, glmfit/glmval, nlinfit, stepwisefit — the numeric forms of the same fits |
fitcsvm, fitctree, fitcknn, fitcnb, fitrgp, TreeBagger | classify, knnsearch, confusionmat, perfcurve. There is no model object to train and predict from here |
makedist, prob.* distribution objects | the <name>pdf/cdf/inv/rnd families, or the string-named pdf/cdf/icdf/random |
cvpartition, crossval, kfoldLoss | randsample and datasample to build the folds yourself |
fitlme, fitglme, fitrm, anovan, bayesopt | nothing — mixed effects and named-factor designs are outside this console's numeric surface |
distributionFitter, classificationLearner, regressionLearner | the same names typed directly; there are no interactive tools here |
Four names on this console are core MATLAB's rather than the toolbox's, and
so carry no toolbox tag even though MATLAB puts them here: rms, prctile,
quantile and iqr. They work the same way; the tag only says whether a
licence would be needed on the other side.
Where to look next#
- Command glossary — console commands, verbs, functions — every name the console answers, grouped by toolbox, with each one's arguments.
- The command window — the command engine for a user — the engine itself: variables, multiple outputs, scripts, and running a line headlessly.
- Curve fitting in the command window —
fitand its model library, for fitting a curve rather than testing a hypothesis. - Numerics — what the solver will and will not do — what a double can carry, and where these answers stop being exact.