Read This First
If this page feels abrupt, start here
These links provide the wider frame, earlier distinction, or branch map that makes the current page easier to enter.
-
What is Induction?
Start here if the current page feels compressed: What is Induction? gives the broader frame before the argument narrows into the present pressure.
-
Philosophy of Science Branch Guide
If this page feels abrupt, start with the Philosophy of Science branch guide so the wider map is visible before the close reading begins.
Read This Next
If the page clicked, continue here
These are not just nearby pages. They are the strongest next moves if you want the pressure of this page to keep unfolding.
-
Inductive Density
Inductive Density keeps the same branch pressure in view but turns it from a different angle.
-
The Problem of Induction
The Problem of Induction keeps the same branch pressure in view but turns it from a different angle.
-
The Notion of Laws
The Notion of Laws keeps the same branch pressure in view but turns it from a different angle.
Prompt 1: Is the arbitrary threshold for statistical significance usually set at a 0.05 p-value the optimal way to indicate a significant relationship between two variables? Have other metrics for significance been proposed?
Why the 0.05 threshold is not the whole story
P-Value Issues matters here as a line of inquiry, not just a topic label. The section should narrow the reader's attention toward the tension that actually needs investigation.
By the end of the section, the reader should be able to restate P-Value Issues in plain language, identify an easy misuse of it, and say what would count as a stronger reason for or against the view.
- State the clearest version of P-Value Issues before testing it.
- Ask what evidence, example, or argument would genuinely change the reader's judgment.
- Notice where a familiar phrase is doing more work than the reasoning beneath it.
- Keep the neighboring concepts visible so the page does not collapse different questions together.
Critiques of the 0.05 Threshold The main critique is that the 0.05 threshold can be arbitrary and may not suit all research contexts. It can lead to overemphasizing the importance of a specific p-value (e.g., 0.049 is deemed significant, while 0.051 is not) without considering the effect size, research design, data quality, or the broader context of the study. Additionally, the reliance on the 0.05 threshold may contribute to issues like p-hacking (manipulating data to achieve a p-value below the significance threshold) and publication bias (the tendency to only publish studies with significant results).
Alternatives and Supplements to the P-value Adjusting the P-value Threshold: Some fields and studies may use more stringent thresholds (e.g., 0.01 or 0.001) to claim statistical significance, especially in contexts where the cost of a false positive is high. Confidence Intervals: Presenting confidence intervals alongside p-values provides a range of plausible values for the effect size, offering more information about the precision of the estimate. Bayesian Statistics: Bayesian methods incorporate prior knowledge along with the current data to compute the probability of a hypothesis. This approach can provide a more nuanced understanding of the evidence in favor or against a hypothesis. Effect Size and Power Analysis: Emphasizing effect size and conducting power analysis can help researchers understand the practical significance of their findings, not just whether they are statistically significant. Pre-registration of Studies: Pre-registering hypotheses, methods, and analyses before data collection can help mitigate selective reporting and p-hacking, emphasizing the robustness of findings over the attainment of a specific p-value.
Adjusting the P-value Threshold Some fields and studies may use more stringent thresholds (e.g., 0.01 or 0.001) to claim statistical significance, especially in contexts where the cost of a false positive is high.
Confidence Intervals Presenting confidence intervals alongside p-values provides a range of plausible values for the effect size, offering more information about the precision of the estimate.
Bayesian Statistics Bayesian methods incorporate prior knowledge along with the current data to compute the probability of a hypothesis. This approach can provide a more nuanced understanding of the evidence in favor or against a hypothesis.
Effect Size and Power Analysis Emphasizing effect size and conducting power analysis can help researchers understand the practical significance of their findings, not just whether they are statistically significant.
Pre-registration of Studies Pre-registering hypotheses, methods, and analyses before data collection can help mitigate selective reporting and p-hacking, emphasizing the robustness of findings over the attainment of a specific p-value.
Recent Discussions and Recommendations Some scientific journals and associations have encouraged moving beyond the rigid adherence to the 0.05 threshold. For example, the American Statistical Association (ASA) published a statement in 2016 outlining the limitations of p-values and cautioning against their misuse. In 2019, a special issue of “The American Statistician” discussed alternatives to the null hypothesis significance testing (NHST) framework and the 0.05 cutoff, advocating for a more holistic approach to evaluating evidence.
Arbitrary nature The choice of 0.05 has historical background rather than statistical justification. Other values like 0.01 or 0.10 could be equally valid depending on the research context and potential consequences of errors.
Focus on binary outcome It dichotomizes results into “significant” or “not significant,” neglecting the gradient of evidence strength within each category. A p-value of 0.051 is practically indistinguishable from 0.049, yet classified differently.
Misinterpretations Misconceptions abound, like equating p-value with the probability of the null hypothesis being true or the effect size of the relationship.
Confidence intervals Provide a range of plausible values for the effect size, offering nuanced information beyond a binary “significant” label.
Effect size measures Directly quantify the magnitude of the observed relationship, aiding in interpreting its practical significance. Examples include Cohen’s d, correlation coefficients, or risk ratios.
Bayesian analysis Shifts the focus from null hypothesis testing to estimating the probability of the alternative hypothesis being true given the data, incorporating prior knowledge and avoiding strict p-value thresholds.
Research question and objectives Are you primarily interested in rejecting the null hypothesis or understanding the nature and strength of the relationship?
Sample size and power Smaller samples are more prone to misleading p-values. Considering power helps ensure your study has a reasonable chance of detecting a real effect.
Field-specific conventions and considerations Some disciplines may have established norms for significance levels or prefer specific effect size measures.
Prompt 2: Explain how “confidence intervals” work.
Clarifying Definition and Interpretation
P-Value Issues matters here as a line of inquiry, not just a topic label. The section should narrow the reader's attention toward the tension that actually needs investigation.
By the end of the section, the reader should be able to restate P-Value Issues in plain language, identify an easy misuse of it, and say what would count as a stronger reason for or against the view.
- State the clearest version of P-Value Issues before testing it.
- Ask what evidence, example, or argument would genuinely change the reader's judgment.
- Notice where a familiar phrase is doing more work than the reasoning beneath it.
- Keep the neighboring concepts visible so the page does not collapse different questions together.
Applications CIs are widely used in research to assess the reliability of an estimate. They are particularly useful in health sciences for estimating effect sizes, differences between groups, and association measures.
Limitations The interpretation of confidence intervals is sometimes misunderstood. A 95% CI does not mean that there is a 95% probability that the interval contains the true parameter value in a frequentist sense. Instead, it reflects the proportion of such intervals that would contain the parameter if the experiment were repeated under the same conditions.
Calculate the point estimate This could be a mean, median, proportion, or any other relevant statistic based on your sample.
Choose a confidence level This is the percentage of times you want your interval to capture the true population parameter. Common choices are 90%, 95%, and 99%.
Calculate the margin of error This value represents the amount of uncertainty around your point estimate. It depends on the sample size, variability within the sample, and the chosen confidence level. Higher confidence levels lead to wider margins of error and vice versa.
Construct the interval Add and subtract the margin of error from your point estimate. This gives you the lower and upper bounds of your confidence interval.
Prompt 3: Comment on how reliance on a confidence interval rather than a binary p-value would reduce the distortive effects of the pressure to find a p-value significance.
Why confidence intervals beat binary p-value thinking
The live issue is whether P-Value Issues can carry the argumentative weight being placed on it. That means separating the strongest support from the rhetorical packaging around it.
By the end of the section, the reader should be able to restate P-Value Issues in plain language, identify an easy misuse of it, and say what would count as a stronger reason for or against the view.
Calculate the point estimate This could be a mean, median, proportion, or any other relevant statistic based on your sample.
Choose a confidence level This is the percentage of times you want your interval to capture the true population parameter. Common choices are 90%, 95%, and 99%.
Calculate the margin of error This value represents the amount of uncertainty around your point estimate. It depends on the sample size, variability within the sample, and the chosen confidence level. Higher confidence levels lead to wider margins of error and vice versa.
Construct the interval Add and subtract the margin of error from your point estimate. This gives you the lower and upper bounds of your confidence interval.
- CIs provide a range of plausible values for the parameter being estimated, offering a richer context for interpretation.
- P-values encourage a binary view of research findings: This dichotomy can be misleading, as it ignores the continuous nature of evidence.
- P-hacking involves selectively reporting results or manipulating data analysis until statistically significant results are found.
- The obsession with p-values can lead to publication bias, where studies with significant results are more likely to be published than those without.
- Emphasizing CIs can promote more robust research practices, such as pre-registration of studies and specifying confidence levels a priori.
- By providing a clearer picture of the uncertainty and variability associated with estimates, CIs can help in assessing the replicability of findings.
Prompt 4: Why is the replacement of p-values with confidence intervals so difficult?
Clarifying Values with confidence intervals so difficult
values with confidence intervals so difficult matters here as a line of inquiry, not just a topic label. The section should narrow the reader's attention toward the tension that actually needs investigation.
By the end of the section, the reader should be able to restate values with confidence intervals so difficult in plain language, identify an easy misuse of it, and say what would count as a stronger reason for or against the view.
- P-values have been deeply ingrained in the statistical methodology of many fields for decades.
- There is a widespread misunderstanding of both p-values and CIs among researchers.
- The scientific publishing industry and peer review processes have historically emphasized p-values as the criterion for statistical significance and publication worthiness.
- P-values provide a simple, if not simplistic, binary outcome that can be easily interpreted as “significant” or “not significant.” This simplicity is appealing for making quick decisions about research findings, even if it reduces the complexity of the data to a misleading.
- While many statisticians and researchers advocate for the use of CIs over p-values, there is no universal agreement on the best alternative approach.
- Any significant change in scientific practice faces resistance due to the human tendency to stick with known and trusted methods.
What ties this page together.
A good route is to identify the strongest version of the idea, then test where it needs qualification, evidence, or a neighboring concept.
The main pressure comes from treating a useful distinction as final, or treating a local insight as if it solved more than it actually solves.
Read this page as part of the wider Philosophy of Science branch: the prompts point inward to the topic, but they also point outward to neighboring questions that keep the topic honest.
- Question 1: What is the primary critique of using a p-value threshold of 0.05 for determining statistical significance?
- Question 3: Why is it difficult to replace p-values with confidence intervals in research practice?
- Question 4: How do confidence intervals help in understanding the replicability of findings?
- Which distinction inside P-Value Issues is easiest to miss when the topic is explained too quickly?
- What is the strongest charitable reading of this topic, and what is the strongest criticism?
Deep Understanding Quiz Check your understanding of P-Value Issues
This quiz checks whether the main distinctions and cautions on the page are clear. Choose an answer, read the feedback, and click the question text if you want to reset that item.
Future Branches
Where this page naturally expands
Nearby pages in the same branch include Inductive Density, The Problem of Induction, The Notion of Laws, and Demarcation for Scientific Laws; those links are not decorative, but suggested continuations where the pressure of this page becomes sharper, stranger, or more usefully contested.