比较基准结果
🌐 Comparing benchmark results
node:bench 并不把某个基准指定为基线,也不会对运行结果做通过/失败的比较。它展示原始样本、稳定的基准标识、参数和标签,以便比较策略可以保留在更高级的工具中。工具可以使用 benchId 来匹配兼容源布局中相同的声明和参数,并使用标签或自身的元数据来识别基线。
比较工具应该保留原始采样率,并验证执行计划和相关环境细节是否可比。适当的分析取决于实验设计和数据分布。例如,独立样本可能使用Welch t检验或基于秩的检验,而刻意配对的观察则需要配对分析。工具在测试多个基准时,还应考虑效应量、不确定性和校正。node:perf_hooks 中的通用 <Histogram> 统计可以支持这种分析,但运行器不会选择方法或显著性阈值。
🌐 Comparison tools should retain the raw sample rates and verify that execution
plans and relevant environment details are comparable. The appropriate analysis
depends on the experimental design and distribution. For example, independent
samples might use Welch's t-test or a rank-based test, while observations that
were deliberately paired require paired analysis. Tools should also consider
effect sizes, uncertainty, and correction when testing multiple benchmarks.
The general-purpose <Histogram> statistics in node:perf_hooks can support such
analysis, but the runner does not select a method or significance threshold.