Fisher’s Exact Test produces a p-value that tells you the probability of seeing your observed data, or something more extreme, if there were truly no association between the two variables in your table. A small p-value (typically below 0.05) suggests the pattern you found is unlikely to be a fluke, while a large one means you cannot rule out chance. That sounds straightforward, but interpreting the result well requires knowing a few things the raw number does not tell you, from whether a one-sided or two-sided version was used, to why the test is notoriously conservative, to what happens when your p-value and confidence interval seem to contradict each other.
What the P-Value Is Actually Measuring
Fisher’s Exact Test works by looking at a two-by-two table and asking: given the row and column totals we have, how likely is this particular arrangement of counts (and every arrangement more extreme) under the assumption that the two variables are unrelated? Instead of relying on a large-sample approximation the way a chi-squared test does, it calculates the probability directly from the hypergeometric distribution, which simply counts all the possible ways the data could have landed in the four cells of the table while keeping the margins fixed.1PubMed Central. Two-stage case–control designs for rare genetic variants This is why it is called “exact” — it does not estimate probabilities; it enumerates them.
Because of this, the p-value you get is not an approximation. It is the precise probability under the null hypothesis. When someone reports “Fisher’s Exact Test, p = 0.032,” they are saying: if the two groups truly had no difference in the outcome, there is about a 3.2 percent chance of seeing an association this strong or stronger purely by chance. The smaller the p-value, the harder it is to blame the result on random variation alone.
One thing the p-value does not tell you is how large the difference is. A tiny p-value can come from a huge effect in a small sample or a trivial effect in a massive sample. To get at the size of the association, you need a measure like the odds ratio, relative risk, or the simple difference in proportions.2PubMed Central. Exact confidence intervals for the relative risk and the odds ratio Always pair your p-value with an effect-size measure so readers know whether the association is clinically or practically meaningful, not just statistically detectable.
One-Sided Versus Two-Sided and Why It Matters More Than You Think
A persistent source of confusion in interpreting Fisher’s Exact Test is whether the reported p-value comes from a one-sided or a two-sided version. A one-sided test asks whether the outcome in one group is specifically higher (or specifically lower) than in the other. A two-sided test asks whether the groups differ in either direction. The two-sided p-value will always be at least as large as the one-sided one, so choosing the wrong version can make a result look more or less significant than it should.
A review of six major medical journals found that researchers frequently reported Fisher’s Exact Test without specifying whether they used a one-sided or two-sided version, creating the potential for misinterpretation of statistical significance.3JAMA. The Inexact Use of Fisher’s Exact Test in Six Major Medical Journals If you are reading a paper that simply says “Fisher’s Exact Test, p = 0.04” without clarifying directionality, you cannot be sure that the result would survive a two-sided threshold. And if you are running the test yourself, the default in most statistical software may not be the version you intended. Always check and always report which version you used.
As a rule of thumb, use a two-sided test unless you had a strong, pre-specified reason before collecting data to expect a difference in one specific direction. Using a one-sided test after looking at the data to see which direction the difference fell is a form of p-hacking that inflates false-positive rates.
Why “Exact” Does Not Mean “Best”
The word “exact” gives the test an air of mathematical perfection, which leads many researchers to assume it is always the right choice for small samples. In reality, Fisher’s Exact Test is widely recognized as being conservative. That means its actual false-positive rate tends to sit below the nominal level you set. If you set your threshold at 5 percent, the test’s true Type I error rate may hover around 2 or 3 percent. You are less likely to declare a difference when one does not exist, which sounds like a good thing, but it comes at a cost: you are also less likely to detect a real difference. In practical terms, the test under-rejects, so you may miss genuine associations.4PubMed. A revisit to contingency table and tests of independence: bootstrap is preferred to Chi-square approximations as well as Fisher’s exact test
The conservatism stems from the discrete nature of the test. Because the hypergeometric distribution can only produce certain probability values, the test cannot always get its actual significance level to match your chosen threshold exactly. It has to round down to the nearest attainable level, which leaves unused alpha on the table. This is not a bug per se — it follows directly from the mathematics — but it does mean the test has lower statistical power than some alternatives, especially with small or unbalanced samples.5Wiley StatsRef: Statistics Reference Online. Fisher’s exact test including power and mid‐P value
Research comparing several approaches to analyzing two-by-two tables found that when the smallest expected cell count is at least one, an alternative chi-squared formulation (the “N-1” chi-squared test) performed well, while Fisher’s Exact Test was preferred only when expected counts dropped below that threshold.6PubMed. Chi-squared and Fisher-Irwin tests of two-by-two tables with small sample recommendations So the common advice to “just use Fisher’s when the sample is small” is an oversimplification. The test earns its keep primarily when expected frequencies are very small — think single digits — and other tests’ approximations break down.
The Mid-P Correction and How It Changes Your Results
One popular way to deal with the conservatism problem is the mid-p adjustment, sometimes called Lancaster’s mid-p correction. Instead of counting the full probability of the observed table, the mid-p approach counts only half of that probability, then adds the probabilities of all more extreme tables. This nudges the p-value downward, bringing the test’s actual significance level closer to the nominal one you set.
Simulation studies have found that the mid-p version of Fisher’s test consistently outperforms the standard version in terms of power while keeping the false-positive rate close to the stated level.7PubMed. Using Lancaster’s mid-P correction to the Fisher’s exact test for adverse impact analyses Research comparing recommended tests for two-by-two tables has confirmed that Fisher’s test with the mid-p adjustment produces results close to those of unconditional exact tests, which do not fix the margins and tend to be more powerful.8PubMed. Recommended tests for association in 2 x 2 tables
A systematic examination of the mid-p test across a wide range of sample sizes showed that its actual significance levels tend to be closer to the nominal level than those of various classical alternatives, for both one-sided and two-sided tests.9PubMed Central. A quasi-exact test for comparing two binomial proportions If your software offers a mid-p option (R does, for example), it is worth considering, particularly when your standard Fisher’s p-value is close to the significance boundary and you suspect the conservatism is the reason it does not cross.
There is a trade-off: the mid-p value is technically no longer “exact” in the strict sense, and it can occasionally exceed the nominal level slightly. But for most practical purposes, it strikes a better balance between controlling false positives and not missing real effects.
When the P-Value and Confidence Interval Disagree
One of the more disorienting things that can happen when interpreting Fisher’s Exact Test is getting a significant p-value while the confidence interval for the odds ratio includes one (meaning “no effect”), or vice versa. In large-sample tests, these two pieces of output are mathematically guaranteed to agree. With Fisher’s Exact Test, they can conflict.
The reason is that the standard exact confidence interval for the odds ratio is not actually the inversion of the two-sided Fisher’s test. It is built by inverting a different procedure: one that rejects when either one-sided Fisher test rejects at half the significance level. These are subtly different tests, and because the underlying distribution is discrete, they can produce contradictory conclusions. Research on this problem has led to the development of “matching” confidence intervals that are specifically constructed to agree with the two-sided Fisher p-value, resolving the discrepancy.10PubMed Central. Confidence intervals that match Fisher’s exact or Blaker’s exact tests
If you encounter this conflict in your own analysis, the practical takeaway is: do not panic, and do not assume one output is “wrong.” Instead, check which method your software used to construct the confidence interval. If your software has the option to compute matching intervals (the R package exact2x2 does this), switch to those so the p-value and interval tell a consistent story. If that option is not available, report both results transparently and note the discrepancy rather than selectively presenting whichever one supports your conclusion.
Small Changes in Data, Large Swings in the P-Value
Another underappreciated quirk is how sensitive the two-tailed Fisher’s Exact p-value can be to tiny changes in the data. A systematic evaluation of 920 pairs of similar contingency tables found that the two-tailed exact p-value is extremely sensitive to small perturbations: in one illustrative case, a one percent increase in the denominator of one group produced a 32 percent drop in the exact p-value, despite changing the actual treatment success rate by only about 0.1 percent.11PubMed Central. Sensitivity of Fisher’s exact test to minor perturbations in 2 x 2 contingency tables
This sensitivity is not a random glitch. It happens because the two-tailed version has to decide which tables in the opposite tail to include, and small changes in the data can shift which tables qualify, causing the p-value to jump. The same research found that simply doubling the one-tailed exact p-value gives a more consistent measure of how strong the evidence is, avoiding the erratic swings of the standard two-tailed calculation.
The practical implication: if your two-tailed Fisher’s p-value is sitting right around your significance threshold, be cautious. Run a sensitivity check by slightly varying your table entries to see whether the conclusion holds. If adding or removing one or two observations flips your result from significant to non-significant (or back), the evidence is probably too fragile to hang a strong conclusion on, regardless of what the exact p-value says.
Boschloo’s Test and Other Alternatives
If Fisher’s Exact Test is conservative and the mid-p correction is not truly exact, is there a better option? For many situations, yes. Boschloo’s test uses the p-value from Fisher’s test as its test statistic but then evaluates significance using an unconditional framework rather than conditioning on the margins. The result is a test that is uniformly more powerful than Fisher’s: it never has less power, and in many configurations it has substantially more.12Biometrics. A Cautionary Note on Exact Unconditional Inference for a Difference Between Two Independent Binomial Proportions
Work on the design of randomized clinical trials has reinforced this point. Because Boschloo’s test is always at least as powerful as Fisher’s, trials designed around it require smaller sample sizes to detect the same effect, which matters for cost, ethics, and feasibility.13PubMed. Design of randomized clinical trials with a binary endpoint: Conditional versus unconditional analyses of a two-by-two table
So why does Fisher’s test remain the default? Inertia, mostly. Fisher’s test has been around since the 1930s, it is available in every major software package, reviewers and regulatory agencies recognize it instantly, and many textbooks still recommend it without caveat. Boschloo’s test is computationally heavier (though that matters less every year) and less widely known. If you are in a field where reviewers expect Fisher’s, you may need to report it for convention while noting the Boschloo result alongside it. If you have the freedom to choose, Boschloo’s test is the stronger option for a standard two-by-two comparison of independent proportions.
Reporting the Test Properly
Good reporting is part of good interpretation. Researchers have argued that when hypothesis testing is the chosen framework, authors should compute Fisher-exact p-values from the actual randomization procedure rather than relying on asymptotic approximations, and should also display the shape of the null randomization distribution so readers can assess the result visually.14PubMed Central. When possible, report a Fisher-exact P value and display its underlying null randomization distribution In practice, here is what a well-reported Fisher’s Exact Test result should include:
- Directionality: State whether the test was one-sided or two-sided. If one-sided, state the predicted direction.
- Exact p-value: Report the actual p-value (e.g., p = 0.038), not just “p < 0.05.” Rounding to a threshold discards information.
- Effect measure: Report the odds ratio, relative risk, or difference in proportions alongside the p-value. The p-value alone says nothing about effect size.
- Confidence interval: Ideally, use a matching confidence interval so it agrees with the test. Specify the method used.
- Table or counts: Show the actual two-by-two table, or at minimum the group sizes and event counts, so readers can verify or reanalyze.
Leaving out any of these elements forces readers to guess at aspects of your analysis, which is exactly the ambiguity the JAMA review flagged decades ago. Transparency in reporting also lets other researchers apply a mid-p correction or Boschloo’s test to your data if they want to, which is a form of scientific generosity.
Stratified Tables and Larger Designs
Fisher’s Exact Test is most commonly associated with the simple two-by-two table, but the underlying logic extends to more complex designs. When you have a potential confounding variable, you can stratify the data into multiple two-by-two tables (one per stratum) and use a stratified version of Fisher’s test. This approach has been developed for situations with rare cell frequencies where the more familiar Cochran-Mantel-Haenszel test may not hold up well, and sample-size formulas for planning such studies have been proposed.15PubMed Central. Stratified Fisher’s exact test and its sample size calculation
When counts are truly sparse — imagine comparing adverse event rates for a rare side effect across treatment arms — stratified Fisher’s test can be more trustworthy than pooling everything into a single table and hoping the large-sample assumptions hold.
Fisher’s Test in Genomics and Pathway Analysis
Outside clinical trials, one of the biggest modern uses of Fisher’s Exact Test is in genomics, particularly for pathway and gene-set enrichment analysis. When a researcher performs an experiment that produces a long list of genes (say, genes whose expression changed after a treatment), they want to know whether that list is enriched for genes belonging to known biological pathways. The standard approach is to construct a two-by-two table for each pathway: genes in the pathway versus not, differentially expressed versus not. Fisher’s Exact Test then asks whether the overlap is larger than what you would expect by chance.16PubMed Central. Pathway enrichment analysis and visualization of omics data using g:Profiler, GSEA, Cytoscape and EnrichmentMap
This is a setting where the test’s conservatism is arguably less harmful, because researchers are running thousands of tests simultaneously (one per pathway) and need to control the overall false-positive rate through multiple-testing corrections like the Benjamini-Hochberg procedure. Starting from a test that is already conservative provides an extra layer of caution. The most common databases queried in this context are gene ontology biological processes and molecular pathways from curated repositories.17Nature Communications. Integrative pathway enrichment analysis of multivariate omics data
If you are interpreting Fisher’s Exact Test results in a genomics context, the key difference from a single clinical-trial comparison is scale. A p-value of 0.03 for one pathway out of 5,000 tested is meaningless without adjustment. Always look at the corrected p-value (sometimes labeled “adjusted p,” “q-value,” or “FDR”) rather than the raw Fisher’s p-value when evaluating enrichment results.
Odds Ratio Bias With Small Counts
When your two-by-two table contains very small counts or zero cells, the odds ratio estimate you get alongside Fisher’s Exact Test can be biased. Zero cells are a particular headache: they make the standard odds ratio either zero or undefined, depending on which cell is empty. A common quick fix is to add 0.5 to every cell (the Haldane correction), but this introduces its own bias. Research on the conditional distribution underlying Fisher’s test has explored computing the expected odds ratio after excluding the zero-cell configurations from sampling, revealing that the bias of the odds ratio estimate can be quantified more carefully than the naive corrections suggest.18PubMed Central. Bias of Odds Ratio Estimate in Fisher’s Exact Test
For the non-specialist, the lesson is simple: when your table has cells with counts of zero or low single digits, treat the odds ratio with extra skepticism. The p-value from Fisher’s test may still be valid, but the point estimate of the odds ratio can be misleading. Report the confidence interval for the odds ratio and pay attention to its width. A confidence interval stretching from 0.4 to 25 may technically exclude 1 (and thus agree with a significant p-value), but it tells you very little about the true size of the effect.
The Tea-Tasting Origin and What It Reveals About the Logic
The test traces back to a story that sounds like a parlor game. One afternoon in the 1920s, the statistician Ronald Fisher offered a cup of tea to his colleague Muriel Bristol, who refused it, insisting she could tell by tasting whether the milk or the tea had been poured into the cup first. Fisher was skeptical, but another colleague suggested they test her claim on the spot, and Bristol proceeded to identify the cups correctly at a rate that surprised the group.19Social Science Research Network. The Lady Tasting Tea
The experiment that followed became Fisher’s vehicle for developing the exact test. The setup was a perfect two-by-two table: cups prepared milk-first versus tea-first, and Bristol’s guesses of milk-first versus tea-first. The question was whether her correct identifications exceeded what luck alone would produce, given a fixed number of cups in each category. The total number of cups was small enough that a large-sample approximation would not be trustworthy, so Fisher worked out the exact probabilities by hand. This is the same logic the test uses today in every software package: enumerate all possible tables with the same margins, compute each table’s probability under the null, and sum the probabilities of all tables as extreme or more extreme than the observed one.
Understanding this origin story clarifies a conceptual point that trips people up: the test conditions on the margins being fixed. In the tea experiment, the number of milk-first and tea-first cups was set by the experimenter, and the number of guesses in each direction was fixed by Bristol. Both margins were truly fixed. In many modern applications, only one margin (or neither) is truly fixed, and conditioning on both can make the test unnecessarily conservative. This is the core reason Boschloo’s test and other unconditional alternatives exist — they relax the fixed-margin assumption when it does not reflect how the data were actually generated.