P-Value: Statistics Study Notes
October 10, 2026
📊 Understanding the P-Value in Statistical Hypothesis Testing
- Core Concepts: Definition, calculation methods, and theoretical foundations of p-values
- Interpretation & Usage: Proper application, significance levels (), and scientific interpretation
- Distributions: Behavior under simple null hypotheses and composite hypotheses
- Controversies & Misuse: Common misconceptions, ASA statements, and ongoing debates
- Practical Examples: Coin flipping experiments and sequential testing
- Historical Development: From Arbuthnot and Laplace to Pearson and Fisher
- Related Statistical Indices: E-values, q-values, Probability of Direction, S-values, and second-generation p-values
💡 Introduction and Context
In null-hypothesis significance testing, the p-value is the probability of obtaining test results at least as extreme as the result actually observed, under the assumption that the null hypothesis is correct.
- A very small p-value indicates that such an extreme observed outcome would be highly unlikely if the null hypothesis were true.
- Despite being a standard practice in academic publications across quantitative fields, the misinterpretation and misuse of p-values are widespread, making them a major topic of discussion in mathematics and metascience.
American Statistical Association (ASA) Statements
- 2016 Formal Statement: The ASA clarified that:
- P-values do not measure the probability that the studied hypothesis is true, nor the probability that data were produced by random chance alone.
- A p-value (or statistical significance) does not measure the size of an effect or the importance of a result.
- A p-value does not provide a good measure of evidence regarding a model or hypothesis without context or other evidence.
- 2019 Replicability Statement: A subsequent ASA task force concluded that "p-values and significance tests, when properly applied and interpreted, increase the rigor of the conclusions drawn from data."
🔍 Basic Concepts
In statistics, any conjecture concerning the unknown probability distribution of a collection of random variables representing observed data is called a statistical hypothesis.
- Null Hypothesis Test: A test where a single hypothesis is stated to see whether it is tenable, without investigating other specific hypotheses.
- Null Hypothesis (): The default hypothesis under which a specific property does not exist. Typically, it states that some parameter (such as a correlation or difference between means) in populations of interest is zero.
Evaluating Statistical Significance
- Data is often reduced to a single numerical test statistic , whose marginal probability distribution connects to the main research question.
- The p-value quantifies the statistical significance of the observed value of .
- Smaller p-values provide stronger evidence against the null hypothesis. Rejecting implies there is sufficient evidence against it.
| Alternative Interpretation Scenarios | Implication when () is rejected |
|---|---|
| (i) Mean Deviation | The mean of is not . |
| (ii) Variance Deviation | The variance of is not . |
| (iii) Distribution Deviation | is not normally distributed. |
The more independent observations one has, the more accurate the test will be, increasing the precision in determining deviations (e.g., showing a mean is not zero). However, this increases the need to evaluate the real-world or scientific relevance of the deviation.
📐 Definition and Interpretation
Mathematical Definitions
Let be an observed test-statistic from an unknown distribution . The p-value represents the prior probability of observing a value at least as "extreme" as if the null hypothesis is true:
- One-sided right-tail test:
- One-sided left-tail test:
- Two-sided test (general):
- Two-sided test (symmetric distribution about zero):
Significance Levels ()
- Type I Error: The error of rejecting a true null hypothesis, considered critical to avoid.
- Significance Level (): A preassigned threshold (e.g., or ) set by researchers before examining data to bound the probability of a Type I error.
- Originally proposed by Ronald Fisher in 1925, setting means rejecting if the p-value falls below this threshold. P-values from independent datasets can be combined using methods like Fisher's combined probability test.
Distribution of P-Values
- Under a Simple Hypothesis: If fixes the probability distribution of precisely and continuously, the p-value is uniformly distributed between 0 and 1 when the null hypothesis is true.
- Variability: Repeating the same test independently with fresh data yields a different p-value each time.
- P-curve: A collection of significant p-values from multiple studies on the same subject used to assess research reliability and detect publication bias or p-hacking.
Distribution for Composite Hypotheses
- Simple Hypothesis: Parameter value is assumed to be a single number.
- Composite Hypothesis: Parameter value is given by a set of numbers (e.g., ).
- When is composite, the exact distribution of the test statistic may vary across possible parameter values. To handle this, the p-value is defined by taking the least favorable null-hypothesis case (typically on the border between the null and alternative hypotheses). This ensures that a significance test at level maintains a maximum Type I error rate of .
🛠️ Calculation of P-Values
Computing a p-value requires three components:
- A formulated null hypothesis.
- A chosen test statistic (scalar function of all observations, e.g., -statistic, -statistic, -statistic) and decision of a one-tailed or two-tailed test.
- The underlying observational data.
Common Statistical Tests
- Z-test: For means of a normal distribution with known variance.
- T-test: Based on Student's t-distribution for means when variance is unknown.
- F-test: Based on the F-distribution for variances.
- Chi-squared test (Pearson's): For categorical/discrete data using normal approximations via the Central Limit Theorem for large samples.
Computational Evolution: While computing test statistics is straightforward, calculating sampling distributions and cumulative distribution functions (CDFs) was historically done using printed tables via interpolation or extrapolation (pioneered by Fisher via inverse CDF quantile functions). Today, this is computed programmatically using statistical software.
🪙 Practical Example: Testing Coin Fairness
An experiment is conducted to determine if a coin is fair or biased, resulting in 14 heads out of 20 total flips.
Test Parameters
- Null Hypothesis (): The coin is fair, , tosses are independent.
- Test Statistic: Total number of heads.
- Alpha Level (): .
- Observation (): 14 heads out of 20 flips.
Calculation Steps
- One-tailed p-value (right-tail favoring heads):
- Two-tailed p-value (symmetric binomial distribution):
Conclusion
Because the p-value () exceeds the alpha threshold (), the data falls within the range expected 95% of the time for a fair coin. Therefore, the null hypothesis is not rejected. (Note: Had there been 15 heads, the p-value would be , leading to rejection of ).
Optional Stopping Complication
Sequential testing changes p-value calculations. If an experiment uses optional stopping rules (e.g., stopping early if extreme outcomes occur vs. planning a fixed 6 flips), the definition of "at least as extreme" shifts, making p-values deeply dependent on the experimenter's original stopping intentions.
⚠️ Misuse and Scientific Debates
The ASA and broader scientific communities highlight several prominent criticisms regarding p-values:
- Automatic Acceptance: Accepting alternative hypotheses for any without contextual backing (study design, measurement quality, external evidence).
- Misunderstanding Probability: Mistaking the p-value for the probability that the null hypothesis is true, or assuming it infers population parameters from samples.
- Proposed Alternatives: Some statisticians advocate abandoning p-values in favor of confidence intervals, likelihood ratios, or Bayes factors, though feasibility is heavily debated.
- Alternative Interpretations: Suggestions include removing fixed thresholds entirely to treat p-values as continuous indices of evidence, or reporting prior probabilities of real effects to gauge false-positive risks.
📜 Historical Milestones
- 1710 (John Arbuthnot): Studied London birth records showing excess male births over 82 years. Calculated a p-value of ( in ) using a sign test, concluding "Art, not Chance" governed births—marking the first recorded significance test.
- 1770s (Pierre-Simon Laplace): Analyzed nearly half a million births using a parametric binomial test, concluding male birth excess was a real, unexplained effect.
- Karl Pearson: Formally introduced the p-value (notated as capital ) via Pearson's chi-squared test.
- Ronald Fisher (1925, 1935): Popularized the p-value in Statistical Methods for Research Workers and The Design of Experiments. He established the standard threshold (1 in 20 chance) and introduced landmark designs like the lady tasting tea experiment. He also contrasted his exact p-value approach with the Neyman-Pearson "Acceptance Procedure" decision framework.
📊 Related Statistical Indices
| Index Name | Alternative Name / Definition | Primary Purpose |
|---|---|---|
| E-value | Robust p-value alternative / Expect value | Deals with optional continuation of experiments; represents the expected number of extreme test statistics under . |
| q-value | Positive false discovery rate analog | Used in multiple hypothesis testing to maintain statistical power while minimizing false positives. |
| Probability of Direction () | Bayesian numerical equivalent | Proportion of the posterior distribution matching the median's sign (50%–100%), representing certainty of effect direction. |
| Second-generation p-values | Extended p-value concept | Ignores extremely small, practically irrelevant effect sizes when evaluating significance. |
| S-value | Surprisal value () | Converts p-values into a logarithmic scale to intuitively measure how "surprised" one should be by a result. |