What it means
Share of observations in category k: divide its count by the total number of observations. Always between 0 and 1.
Components
- sample size
- category/class/variable indices
- absolute frequency
- relative frequency
279 entries from the Bocconi Statistics 30001 course — every formula plus the R and UBStats commands you need at the computer — each with a plain-English explanation and a breakdown of what every symbol or argument stands for. Built to help you prepare efficiently and drill until it sticks.
279 results
What it means
Share of observations in category k: divide its count by the total number of observations. Always between 0 and 1.
Components
What it means
The same share expressed on a 0-100 scale; just multiply the relative frequency by 100.
Components
What it means
Counting every category once must return the whole sample: a useful check that no observation was lost or double-counted.
Components
What it means
Shares of all categories must add up to the whole (1). If they don't, a frequency is wrong.
Components
What it means
Length of a grouped interval: upper endpoint minus lower endpoint. Needed whenever classes have unequal size.
Components
What it means
Height of a histogram bar: relative frequency spread over the class width, so unequal classes stay comparable by area.
Components
What it means
Inverting the density: the area of a histogram bar (height x width) is the class's relative frequency.
Components
What it means
Adds up shares from the first class to class k, i.e. the proportion of data at or below that class.
Components
What it means
Function giving the fraction of observations less than or equal to x; the empirical version of a CDF, non-decreasing from 0 to 1.
Components
What it means
The balance point of the data: add all values and divide by how many there are.
Components
What it means
When data are grouped, weight each class value by its count instead of listing all observations.
Components
What it means
Same weighted average written with shares, so it is directly a weighted mean with weights summing to 1.
Components
What it means
Deviations above and below the mean cancel exactly, which is why we square them to measure spread.
Components
What it means
Rearranged mean: knowing the mean and n gives the total. Handy for combining or updating groups.
Components
What it means
Crudest spread measure: distance between largest and smallest value. Very sensitive to outliers.
Components
What it means
Spread of the middle 50% of the data. Robust to outliers, unlike the range.
Components
What it means
Boxplot rule: points below this threshold are flagged as (mild) outliers.
Components
What it means
Boxplot rule: points above this threshold are flagged as (mild) outliers.
Components
What it means
Average squared distance from the population mean, computed over all N units.
Components
What it means
Computationally faster form: mean of squares minus square of the mean.
Components
What it means
Average squared deviation using n-1 (Bessel's correction) so the estimator is unbiased for the population variance.
Components
What it means
Same value from raw sums: sum of squares minus n times the squared mean, divided by n-1.
Components
What it means
Square root of the population variance, back in the original units of the variable.
Components
What it means
Square root of the sample variance; the usual reported measure of spread.
Components
What it means
Relative spread: standard deviation per unit of mean. Lets you compare variability across different scales or units.
Components
What it means
Average product of paired deviations over the population: positive when the two variables tend to move together.
Components
What it means
Mean of the products minus product of the means; faster from raw sums.
Components
What it means
Sample version with n-1 in the denominator, so it is unbiased for the population covariance.
Components
What it means
Raw-sum form of sample covariance, convenient when you are given sums from a table.
Components
What it means
Covariance rescaled by both standard deviations, giving a unit-free measure of linear association in [-1,1].
Components
What it means
Sample version of the correlation coefficient; measures strength and direction of a linear relationship only.
Components
What it means
Correlation can never leave [-1,1]; the extremes mean perfectly linear relationships, 0 means no linear association.
Components
What it means
The PMF: probability that a discrete random variable takes exactly the value x.
Components
What it means
Every probability lies between 0 (impossible) and 1 (certain).
Components
What it means
The PMF must exhaust all possible outcomes, so its values add to 1.
Components
What it means
Cumulative distribution function: probability of being at or below x. Non-decreasing, from 0 to 1.
Components
What it means
Long-run average of X: each value weighted by its probability.
Components
What it means
Expected squared distance from the mean, weighted by probabilities.
Components
What it means
Variance equals the mean of the square minus the square of the mean. Fastest route in most exercises.
Components
What it means
A single trial with two outcomes, success (1) with probability p and failure (0).
Components
Showing 40 / 279