Benchmark results builder
Turn several benchmark runs into one table: mean, median, min/max, standard deviation, speed-up versus the baseline and ranking — copyable as Markdown or plain text.
Runs in your browserEvery computation happens in your browser — your data never leaves this device.
Result table
| Rank | Name | Runs | Mean | Median | Min | Max | Std dev | Relative |
|---|---|---|---|---|---|---|---|---|
| 2 | loop-sum | 3 | 122 | 122 | 120 | 124 | 1.63 | 1.24× (+23.65%) |
| 1 | array-map(baseline) | 3 | 98.67 | 99 | 96 | 101 | 2.05 | 1.00× (+0.00%) |
| 3 | json-parse | 3 | 209.67 | 210 | 205 | 214 | 3.68 | 2.13× (+112.50%) |
Copyable output
| Rank | Name | Mean | Median | Min | Max | Std dev | Relative | | --- | --- | --- | --- | --- | --- | --- | --- | | 1 | array-map (baseline) | 98.67 | 99 | 96 | 101 | 2.05 | 1.00× (+0.00%) | | 2 | loop-sum | 122 | 122 | 120 | 124 | 1.63 | 1.24× (+23.65%) | | 3 | json-parse | 209.67 | 210 | 205 | 214 | 3.68 | 2.13× (+112.50%) |
1. array-map | Mean 98.67 | Median 99 | Min 96 | Max 101 | Std dev 2.05 | Relative 1.00× (+0.00%) 2. loop-sum | Mean 122 | Median 122 | Min 120 | Max 124 | Std dev 1.63 | Relative 1.24× (+23.65%) 3. json-parse | Mean 209.67 | Median 210 | Min 205 | Max 214 | Std dev 3.68 | Relative 2.13× (+112.50%)
What this tool does
- Paste several benchmark runs and get mean, median, min/max, standard deviation, speed-up versus the baseline and a ranking — no spreadsheet formulas required.
- Drop the Markdown table straight into a README or an issue when writing a performance comparison.
- Decide whether an optimisation actually helped: compare the new implementation against the old one and check whether the ratio exceeds the measurement noise (the standard deviation tells you how big that noise is).
- Tidy up multiple load-test rounds and use the median and extremes to spot results skewed by an outlier.
Example
Input
loop-sum 120 124 122 array-map 96 99 101 json-parse 210 205 214
Output
| Rank | Name | Mean | Median | Min | Max | Std dev | Relative | | --- | --- | --- | --- | --- | --- | --- | --- | | 1 | array-map (baseline) | 98.67 | 99 | 96 | 101 | 2.05 | 1.00× (+0.00%) | | 2 | loop-sum | 122 | 122 | 120 | 124 | 1.63 | 1.24× (+23.65%) | | 3 | json-parse | 209.67 | 210 | 205 | 214 | 3.68 | 2.13× (+112.50%) |
This is the real output with the defaults (baseline = best entry, lower is better, 2 decimals). Ratios are always ≥ 1 here, so the slowest entry takes 2.13× as long as the fastest.
Frequently asked questions
Should the baseline be the first entry or the best one?
It depends on the story you want to tell. With the best entry, every ratio is ≥ 1 and reads as “how many times slower than the fastest”, which is handy for picking a winner. With the first entry you compare against the original implementation in input order, which is what you want when measuring an improvement (ratios can be below 1 and show as −x%).
Is the standard deviation population or sample?
Population by default (divide by n); enabling “Use sample standard deviation” switches to n−1, which is slightly larger for the same data. With only three measurements the gap is noticeable, so state which one you used in a report. With a single measurement the sample value cannot be computed and is left blank.
Can the values carry units?
Yes. `12.4ms`, `1.2GB/s`, `320ops/s` and `95%` all parse — only the leading number is used. The unit is treated as a comment and never converted, so mixing ms and s gives a wrong ratio; normalise units before pasting. Scientific notation such as `1e3` works too.
Why do two entries share a rank?
Because ranking compares exact means: identical means share a rank and the next entry skips a position. More common in practice is a difference that is smaller than the standard deviation — that will not tie, but it should not be read as a real difference either.
Is my data uploaded?
No. Parsing, statistics and table generation all run in your browser, the page makes no network requests, and the benchmark data you paste is never logged.
Keywords:benchmark基准测试性能对比markdown 表格标准差speedupmeasures跑分