Correlation coefficient
Compute the Pearson and Spearman correlation between two paired data sets, with r², the regression slope, an optional Kendall τ, a strength reading and the scatter data.
Runs in your browserEvery computation happens in your browser — your data never leaves this device.
Input
r, ρ and τ are computed locally; Spearman uses average ranks for ties and Kendall uses the τ-b tie correction.
Result
0.60000.60002.20000.73795What this tool does
- Check whether two measurements move together — height and weight, ad spend and revenue — before you invest in a regression.
- Switch to Spearman ρ when the relation is monotonic but not linear, or when outliers would dominate Pearson r: it only looks at ranks and handles ties.
- Use Kendall τ for small samples and rank data (1–5 survey answers, judge rankings): it measures the share of concordant pairs and behaves better with many ties.
- Build an appendix or a report: grab r, r², the regression slope and intercept in one pass, and copy the scatter data into your charting tool.
Example
Input
X series 1, 2, 3, 4, 5; Y series 2, 4, 5, 4, 5; "Also compute Kendall τ" enabled
Output
Sample size 5, Pearson r = 0.7746, r² = 0.6000, regression slope = 0.6000, regression intercept = 2.2000, Spearman ρ = 0.7379, Kendall τ = 0.6708, strength: strong positive correlation
r² = 0.6 means 60% of the variation in Y is explained by the linear relation with X. Spearman ρ comes out slightly lower because Y contains two 4s and two 5s, and average ranks flatten those positions.
Frequently asked questions
Which coefficient should I use?
Use Pearson r when the relation is roughly linear and there are no influential outliers. Use Spearman ρ for ranks, ties, or a monotonic-but-curved relation such as exponential growth. Use Kendall τ when the sample is small (n < 30) or heavily tied, where its confidence intervals behave better. Compute all three and compare; if they disagree, trust ρ or τ.
A high r means X causes Y, right?
No. A correlation only describes how two variables move together. A third factor can drive both — temperature raises both ice-cream sales and swimming-pool visits — which produces a strong r with no causal link. Causal claims need an experiment or an identification strategy.
What if the two series have different lengths?
The tool stops with "X and Y must have the same length". Correlation needs paired observations: an extra value on one side has no partner. If a pair has a missing value, delete both members of that pair before pasting.
Why does a constant series break the calculation?
Its variance is 0, so the formula divides by zero. A series that never changes has no linear relation with anything, so the tool reports "variance is 0, the correlation is undefined" instead of returning NaN.
Why is there a size limit on Kendall τ?
τ-b compares every pair, so it is O(n²): 2000 points means about two million comparisons. Beyond the cap the tool refuses rather than freezing the page; at that size Spearman ρ is more than enough.
Keywords:correlation coefficientpearsonspearmankendall taur squaredscatter data相关系数皮尔逊斯皮尔曼肯德尔线性相关散点