---
title: "P-Value Calculator — Z-Test and T-Test, One-Tailed and Two-Tailed"
description: "Calculate p-values from z-scores or t-scores for one-tailed and two-tailed tests. Explains the 0.05 threshold, reproducibility debate, and how US/EU/UK research reporting standards differ."
url: "https://worldcalculators.org/calculators/p-value/"
canonical: "https://worldcalculators.org/calculators/p-value/"
---

# P-Value Calculator

Calculate p-values from z-scores or t-scores for one-tailed and two-tailed tests. Explains the 0.05 threshold, reproducibility debate, and how US/EU/UK research reporting standards differ.

## Quick answer

The reproducibility crisis (or replication crisis) refers to findings that many published scientific results fail to replicate when repeated. A 2015 study in Science reproduced only 36% of 100 psychology studies. Key contributors: p-hacking (testing multiple hypotheses and only reporting p < 0.05), HARKing (hypothesising after results are known), publication bias (journals favouring significant results). Major efforts to address this include pre-registration of studies (AsPredicted.org, OSF) and the use of Bayesian methods. The UK Medical Research Council and US NIH both now require pre-registration.

## Frequently Asked Questions

The reproducibility crisis (or replication crisis) refers to findings that many published scientific results fail to replicate when repeated. A 2015 study in Science reproduced only 36% of 100 psychology studies. Key contributors: p-hacking (testing multiple hypotheses and only reporting p < 0.05), HARKing (hypothesising after results are known), publication bias (journals favouring significant results). Major efforts to address this include pre-registration of studies (AsPredicted.org, OSF) and the use of Bayesian methods. The UK Medical Research Council and US NIH both now require pre-registration.
Physics (particle physics at CERN): requires 5-sigma (p < 0.0000003) for a "discovery" — the Higgs boson announcement used this threshold. This is because: (1) results can be checked against theory precisely, (2) experiments run for years producing massive datasets, (3) a wrong announcement would be a major setback. Biology and psychology: 0.05 threshold, partly because effect sizes are smaller and data noisier. Genomics uses genome-wide significance p < 5×10⁻⁸ to correct for testing ~1 million genetic variants simultaneously (Bonferroni correction).

Source: https://worldcalculators.org/calculators/p-value/
Markdown mirror of the HTML page. Prefer this URL for RAG; the interactive calculator still lives on the HTML page.
