Node.js v26.11.1 文档


基准测试运行器#>

🌐 Benchmark runner

稳定性: 1.0 - 早期开发

源代码: lib/bench.js

node:bench 模块支持在当前进程中定义和运行 JavaScript 基准测试,并在新的子进程中运行一个基准测试文件。该模块仅在使用 --experimental-bench 标志启动 Node.js 时可用,并且只能通过 node: 方案导入:

🌐 The node:bench module supports defining and running JavaScript benchmarks in the current process, and running one benchmark file in a fresh child process. The module is only available when Node.js is started with the --experimental-bench flag and can only be imported with the node: scheme:

import { bench, suite } from 'node:bench';const { bench, suite } = require('node:bench');

示例基准#>

🌐 Example benchmark

把以下内容保存为 benchmark.mjs:

🌐 Save the following as benchmark.mjs:

import { bench, suite } from 'node:bench';

suite('URL', () => {
  const input = 'https://example.com/a?b=c';

  bench('construct', {
    samples: 30,
    params: { input: 'short' },
  }, (b) => {
    const operations = 10_000;
    let totalLength = 0;

    b.start();
    for (let i = 0; i < operations; i++) {
      totalLength += new URL(input).href.length;
    }
    b.end(operations);

    if (totalLength !== operations * input.length) {
      throw new Error('Unexpected URL result');
    }
  });
}); 

从命令行运行基准测试:

🌐 Run the benchmark from the command line:

node --experimental-bench --bench benchmark.mjs 

基准测试会按照声明的顺序依次执行。声明的基准测试会自动安排。可以在声明的同一轮中调用 run() 来使用事件流或设置过滤。如果自动安排的运行失败且没有调用 run(),进程的退出代码会被设置为 1。

🌐 Benchmarks are executed serially in declaration order. Declared benchmarks are scheduled automatically. Call run() during the same turn as the declarations to consume the event stream or configure filtering. If an automatically scheduled run fails and run() was not called, the process exit code is set to 1.

测量模型#>

🌐 Measurement model

每次热身和测量样本都会使用一个新的 BenchContext 调用基准函数一次。该函数必须要么正好调用一次 context.start() 和 context.end(operations),要么正好调用一次 context.record(sample) 来提供一个外部测量的样本。在 start() 之前的设置和在 end() 之后的清理不在测量区域内。返回 Promise 的函数会被等待。

🌐 Each warmup and measured sample invokes the benchmark function once with a fresh BenchContext. The function must either call context.start() and context.end(operations) exactly once, or call context.record(sample) exactly once to provide an externally measured sample. Setup before start() and cleanup after end() are outside the measured region. Promise-returning functions are awaited.

默认情况下,每次样本调用之间都会有一次事件循环。嵌入式运行器可以使用 yieldBetweenSamples 来禁用这一点。运行器按顺序执行基准测试,但它不提供进程隔离。进程中的其他工作、JIT 编译、垃圾回收、CPU 频率变化以及系统负载都可能影响结果。在比较结果时保留原始样本,并且遇到数据噪声大或分布偏斜时,应调查原因,而不是把置信区间当作通过/不通过的标准。

🌐 By default, an event loop turn occurs between sample invocations. An embedded runner can disable this using yieldBetweenSamples. The runner executes benchmarks serially, but it does not provide process isolation. Other work in the process, JIT compilation, garbage collection, CPU frequency changes, and system load can all affect results. Keep raw samples when comparing results and investigate noisy or skewed distributions rather than treating a confidence interval as a pass/fail threshold.

测量完整性#>

🌐 Measurement integrity

一个统计上一致的结果并不能证明基准测试测量了预期的工作。优化的运行时可能会去掉结果未被使用的工作,或者将其专门化得比被建模的工作负载更窄。框架和循环开销也可能在操作过短时占据主导地位。为了减少这些风险:

🌐 A statistically consistent result does not prove that a benchmark measured the intended work. An optimizing runtime can remove work whose result is unused or specialize it more narrowly than the workload being modeled. Framework and loop overhead can also dominate operations that are too short. To reduce these risks:

  • 让通过测量工作产生的值在测量区间外也可观察,例如通过验证从每个结果得出的汇总。仅通过未使用的本地计算传递它们是不够的。
  • 在每个样本中执行足够的操作,以分摊固定定时器读取和对 context.start() 与 context.end() 的调用。如果循环记录相对于一次操作而言影响显著,那么每次迭代批量处理多个操作,并报告总操作次数。
  • 检查原始 samples 数据,看看是否有显示预热不足或优化分级的趋势,是否存在与垃圾回收一致的停顿,以及是否有多峰分布。
  • 用一个独立的基准形状来验证令人惊讶的结果,这个形状以不同的方式完成相同的工作。

node:bench 不会强制特定的优化状态,也不会推断引擎是否消除了某些工作。这类控制和诊断是特定于运行时的,并且带有启发性,不能替代对基准工作负载的验证。

动态采样和可变批次#>

🌐 Dynamic sampling and variable batches

在测量样本期间调用 context.done() 会在该样本之后完成基准测试。这允许更高级的工具将 samples 视为最大值并实现动态采样策略。

🌐 Calling context.done() during a measured sample completes the benchmark after that sample. This allows a higher-level tool to treat samples as a maximum and implement a dynamic sampling policy.

不同样本之间的操作次数可能会有所不同。汇总统计将每个样本的 rate 视为同等权重的一次观测。特别是,summary.mean 是每个样本速率的算术平均值。它不是按以下方式计算的合并吞吐量:

🌐 The number of operations can differ between samples. Summary statistics treat each sample's rate as one equally weighted observation. In particular, summary.mean is the arithmetic mean of the per-sample rates. It is not the pooled throughput calculated as:

1_000_000_000 * sum(sample.operations) / sum(sample.duration_ns) 

当样本持续时间不同的时候,这两个值可能会有所不同,因为合并吞吐量会按每个样本的持续时间对每个样本速率进行加权。一个高级工具如果要更改批量大小,就应该选择与其分析相匹配的汇总方式。它可以从原始的 samples 计算合并吞吐量;操作计数应该作为 bigint 值求和,因为它们的总和可能超过 Number.MAX_SAFE_INTEGER,尽管每个计数本身不可能超过。

🌐 The two values can differ when sample durations vary because pooled throughput weights each per-sample rate by its duration. A higher-level tool that varies batch sizes should choose the aggregation that matches its analysis. It can calculate pooled throughput from the raw samples; operation counts should be summed as bigint values because their total can exceed Number.MAX_SAFE_INTEGER even though each count cannot.

比较基准结果#>

🌐 Comparing benchmark results

node:bench 并不把某个基准指定为基线,也不会对运行结果做通过/失败的比较。它展示原始样本、稳定的基准标识、参数和标签,以便比较策略可以保留在更高级的工具中。工具可以使用 benchId 来匹配兼容源布局中相同的声明和参数,并使用标签或自身的元数据来识别基线。

比较工具应该保留原始采样率,并验证执行计划和相关环境细节是否可比。适当的分析取决于实验设计和数据分布。例如,独立样本可能使用Welch t检验或基于秩的检验,而刻意配对的观察则需要配对分析。工具在测试多个基准时,还应考虑效应量、不确定性和校正。node:perf_hooks 中的通用 <Histogram> 统计可以支持这种分析,但运行器不会选择方法或显著性阈值。

🌐 Comparison tools should retain the raw sample rates and verify that execution plans and relevant environment details are comparable. The appropriate analysis depends on the experimental design and distribution. For example, independent samples might use Welch's t-test or a rank-based test, while observations that were deliberately paired require paired analysis. Tools should also consider effect sizes, uncertainty, and correction when testing multiple benchmarks. The general-purpose <Histogram> statistics in node:perf_hooks can support such analysis, but the runner does not select a method or significance threshold.

可重复使用的长条垫#>

🌐 Reusable runners

模块级声明函数使用共享运行器并自动调度它。更高级的工具可以创建独立的、显式启动的运行器,如下所示:

🌐 The module-level declaration functions use a shared runner and schedule it automatically. Higher-level tools can create isolated, explicitly started runners instead:

import { createRunner } from 'node:bench';

const runner = createRunner({ yieldBetweenSamples: false });

runner.bench('example', { samples: 100 }, (b) => {
  const operations = chooseOperationCount();
  b.start();
  runOperations(operations);
  const sample = b.end(operations);

  if (hasEnoughData(sample)) b.done();
});

for await (const record of runner.run()) {
  // Consume structured benchmark records.
} 

每个运行器都有独立的声明、钩子、过滤和输出。不像模块级声明,在一个明确的运行器上创建基准测试并不会安排执行。这允许包收集声明并稍后启动它们。调用明确运行器的run()函数会阻止额外的声明,而第二次调用run()则会报错。

🌐 Each runner has independent declarations, hooks, filtering, and output. Unlike the module-level declarations, creating a benchmark on an explicit runner does not schedule execution. This allows packages to collect declarations and start them later. Calling the explicit runner's run() function prevents additional declarations and a second call to run() is an error.

命令行运行器#>

🌐 Command-line runner

--bench 标志用于运行一个或多个指定的基准测试文件或通配符模式:

🌐 The --bench flag runs one or more explicit benchmark files or glob patterns:

node --experimental-bench --bench benchmark.mjs
node --experimental-bench --bench --bench-reporter=json 'benchmarks/**/*.js' 

文件会被排序并按顺序执行。默认的 --bench-isolation=process 模式会在每个子进程中运行每个文件,并生成一个汇总摘要。结构化事件会直接传递给父进程,不会进行 JSON 转换,可以保留 BigInt 类型的持续时间、错误和参数值。子进程写入的 stdout 和 stderr 会作为诊断记录输出,因此不会破坏报告器的输出。

🌐 Files are sorted and executed serially. The default --bench-isolation=process mode runs each file in a separate child process and emits one aggregate summary. Structured events are transferred to the parent without JSON conversion, preserving BigInt durations, errors, and parameter values. Child writes to stdout and stderr are emitted as diagnostic records so they do not corrupt reporter output.

--bench-isolation=none 会把所有文件导入到运行器进程中。这种模式启动开销较低,但模块、堆和进程状态会在文件间共享,而且用户的输出和错误会与报告器共用 stdout 和 stderr。

工作线程隔离不是命令行模式。每个新创建的 <Worker> 都有一个独立的 V8 隔离区、JavaScript 堆和事件循环,通常启动开销比子进程低。重用工作线程可以保留它的模块和堆状态。工作线程还共享 libuv 的全进程线程池,并且可以共享进程全局的本地或附加状态,所以它们并不能提供像进程隔离那样的边界。

🌐 Worker-thread isolation is not a CLI mode. Each newly constructed <Worker> has a separate V8 isolate, JavaScript heap, and event loop, typically with lower startup cost than a child process. Reusing a worker preserves its module and heap state. Workers also share libuv's process-wide thread pool and can share process-global native or addon state, so they do not provide the same boundary as process isolation.

更高级的工具可以通过在 worker 内加载基准代码来实验 worker 隔离,在那里进行测量,传输结构化的样本数据,并将其传给 context.record()。报告的 duration_ns 可以在 worker 捕获两个时间戳时排除消息传输。工具应该明确识别 worker 模块和工作负载。它们不应该将任意函数或闭包转换成字符串来在隔离环境间移动,因为闭包无法与其原始词法环境重建。

🌐 Higher-level tools can experiment with worker isolation by loading benchmark code inside a worker, measuring there, transferring structured sample data, and passing it to context.record(). The reported duration_ns can exclude message transport when the worker captures both timestamps. Tools should identify worker modules and workloads explicitly. They should not stringify arbitrary functions or closures to move them between isolates, because closures cannot be reconstructed with their original lexical environment.

传递给 --bench 的基准文件应该声明基准,但不得调用 run()。CLI 支持 --bench-name-pattern、--bench-samples、--bench-warmup、--bench-reporter 和 --bench-reporter-destination。详情请参见 命令行选项文档。

🌐 Benchmark files passed to --bench should declare benchmarks but must not call run(). The CLI supports --bench-name-pattern, --bench-samples, --bench-warmup, --bench-reporter, and --bench-reporter-destination. See the command-line options documentation for details.

通过 --require 或 --import 传递的预加载模块不应该声明基准。这样的声明不会与入口文件关联,而且它们的 entryFile 值是 null。它们的 fileRunId 用来标识发生这些操作的运行器或子执行进程。在进程隔离的情况下,预加载模块会被评估,并且它的声明会在每个基准子进程中运行一次。

🌐 Preload modules passed through --require or --import should not declare benchmarks. Such declarations are not associated with an entry file and have an entryFile value of null. Their fileRunId identifies the runner or child execution in which they occurred. With process isolation, a preload is evaluated and its declarations run once for every benchmark child process.

基准报告器#>

🌐 Benchmark reporters

内置报告器可以从仅限 Scheme 的 node:bench/reporters 模块中获取:

🌐 The built-in reporters are available from the scheme-only node:bench/reporters module:

import { json, spec } from 'node:bench/reporters';const { json, spec } = require('node:bench/reporters');

Reporter 的值可以直接传递给 stream.compose():

🌐 Reporter values can be passed directly to stream.compose():

import { bench, run } from 'node:bench';
import { spec } from 'node:bench/reporters';
import process from 'node:process';

bench('example', (b) => {
  b.start();
  doWork();
  b.end(1);
});

run().compose(spec).pipe(process.stdout); 

spec 报告器会缓冲结果,并输出一个简洁的表格,包含样本数量、平均速率、平均值的 95% 置信区间、中位速率以及警告。修改系数超过 5% 会显示为 noisy,绝对偏度超过 1 会显示为 skewed。具体的可读格式可能会有所变化。

🌐 The spec reporter buffers results and outputs a concise table containing the sample count, mean rate, 95% confidence interval for the mean, median rate, and warnings. A coefficient of variation above 5% is reported as noisy, and an absolute skewness above 1 is reported as skewed. The exact human-readable format is subject to change.

json 报告器会将每个生命周期记录作为换行分隔的 JSON 输出。BigInt 值,包括 duration_ns,会被编码为十进制字符串。错误会使用它们的 name、message、stack、code、cause 和 errors 属性表示。根据 JSON 的要求,非有限数字会被编码为 null。

🌐 The json reporter emits every lifecycle record as newline-delimited JSON. BigInt values, including duration_ns, are encoded as decimal strings. Errors are represented using their name, message, stack, code, cause, and errors properties. As required by JSON, non-finite numbers are encoded as null.

自定义报告器使用相同的组合契约。它们可以是 stream.compose() 接受的转换或函数。组合后的可读流可以传输到任意可写的目标:

🌐 Custom reporters use the same composition contract. They can be transforms or functions accepted by stream.compose(). The composed readable can be piped to any writable destination:

import { run } from 'node:bench';
import process from 'node:process';

async function* names(source) {
  for await (const { type, data } of source) {
    if (type === 'bench:complete') {
      yield `${data.name}\n`;
    }
  }
}

run().compose(names).pipe(process.stdout); 

createRunner([options])#>

  • options <Object>
    • yieldBetweenSamples <boolean> 在样本回调之间安排事件循环轮次。禁用此功能也会防止基于计时器的中止信号在同步回调之间触发。基准测试超时仍会根据单调截止时间进行检查。默认值: true。
  • 返回:<Object> 一个独立的基准测试运行器,包含有界的 after、afterEach、before、beforeEach、bench、describe、run 和 suite 函数。

创建一个显式启动的基准测试运行器。通过一个运行器进行的声明不会与通过另一个运行器或模块级函数进行的声明相互影响。调用返回的 run() 函数来启动运行器并获得它的 BenchmarksStream。

🌐 Creates an explicitly started benchmark runner. Declarations made through one runner do not interact with declarations made through another runner or through the module-level functions. Call the returned run() function to start the runner and obtain its BenchmarksStream.

每个运行器只能启动一次。它的 run() 函数接受与模块级别的 run() 相同的选项。run({ yieldBetweenSamples }) 会覆盖传递给 createRunner() 的值。

🌐 Each runner can be started once. Its run() function accepts the same options as the module-level run(). run({ yieldBetweenSamples }) overrides the value passed to createRunner().

bench([name][, options], fn)#>

  • name <string> 基准名称。默认: fn 的 name 属性,或者当 fn 没有名称时为 '<anonymous>'。
  • options <Object>
    • diagnosticChannels <Array> 字符串诊断通道名称,通过并集去重并从包含的测试套件继承。数组中的符号值会被自动忽略。默认值: []。
    • only <boolean> 当任何基准或包含的套件设置了 only 时,层级中没有 only 的基准会被跳过。默认值: false。
    • params <Object> 字符串、有限数字或布尔类型元数据,用于标识这个基准配置。在构建稳定的基准身份时,参数键会被排序。**默认值:**一个空对象。
    • samples <number> 测量回调调用的最大次数。必须是一个正的 32 位无符号整数。基准测试可以通过调用 context.done() 提前结束。默认值: 30。
    • signal <AbortSignal> 允许中止此基准测试。
    • skip <boolean> | <string> 如果为真,该基准测试将被跳过。一个字符串会被包含在结果中作为跳过的原因。默认值:false。
    • tags <string[]> 与基准相关的标签。标签会被转换为小写、去重,并通过包含的测试套件以并集的方式继承。默认值: []。
    • timeout <number> 基准测试失败的毫秒数。默认值: Infinity。
    • warmup <number> 在测量样本之前未报告的回调调用次数。必须是32位无符号整数。默认值: 0。
  • fn <Function> | <AsyncFunction> 基准函数。它接收一个 BenchContext。
  • 返回:<Promise> 在顶层基准测试完成后,会用基准结果填充,或者在套件中声明 undefined 时立即填充。

热身调用使用与测量样本相同的回调和计时契约,但它们的样本会被丢弃。发生异常、拒绝、超时、终止、缺少计时调用或重复计时调用都会停止当前基准测试。后续基准测试仍会继续运行。

🌐 Warmup invocations use the same callback and timing contract as measured samples, but their samples are discarded. An exception, rejection, timeout, abort, missing timing call, or duplicate timing call stops the current benchmark. Later benchmarks continue to run.

在超时或中止之后,运行器会暂时等待异步基准工作的完成,然后再继续。如果它仍然处于挂起状态,所有后续被选中运行的基准测试都会失败而不会运行,以免它们的测量与该工作重叠。

🌐 After a timeout or abort, the runner briefly waits for asynchronous benchmark work to settle before continuing. If it remains pending, all later benchmarks that were selected to run fail without running so that their measurements cannot overlap with that work.

对于每次热身和测量回调,运行器都会订阅配置好的诊断通道。每次发布都会排入一个上下文诊断,其 message 为 { name, message },包含通道名称字符串和发布的消息。当回调完成或被中止时,订阅会被移除。

🌐 For each warmup and measured callback, the runner subscribes to the configured diagnostics channels. Each publication queues a context diagnostic whose message is { name, message }, containing the string channel name and the published message. Subscriptions are removed when the callback settles or is aborted.

超时或中止无法打断同步的 JavaScript,也不会强制取消忽略 context.signal 的异步工作。

🌐 A timeout or abort cannot interrupt synchronous JavaScript and does not forcibly cancel asynchronous work that ignores context.signal.

benchId 是基于声明源文件、层级测试套件和基准名称,以及规范化参数生成的。对于同一源位置的重复运行,它是稳定的,但嵌入的源值在不同的检出根目录、模块格式、操作系统或路径大小写之间并未标准化。

🌐 The benchId is based on the declaration source file, hierarchical suite and benchmark names, and canonicalized parameters. It is stable for repeated runs from the same source location, but the embedded source value is not normalized across checkout roots, module formats, operating systems, or path casing.

执行范围是单独表示的。runId 标识一次逻辑运行,而 fileRunId 标识该运行中的文件运行器或子执行。entryFile 字段记录导致声明的入口文件导入,而由预加载模块做出的声明则为 null。因此,当入口文件使用共享声明助手时,相同的 benchId 可以出现在多个 fileRunId 下。在一个文件执行范围内多次声明相同的 benchId 会报错,而不是合并样本。

🌐 Execution scope is represented separately. A runId identifies one logical run, while fileRunId identifies a file runner or child execution within that run. The entryFile field records which entry-file import caused a declaration and is null for declarations made by preload modules. The same benchId can therefore occur under multiple fileRunId values when entry files use a shared declaration helper. Declaring the same benchId more than once within one file execution scope reports an error rather than merging the samples.

bench.skip([name][, options], fn)#>

bench(name, { ...options, skip: true }, fn) 的简写。

🌐 Shorthand for bench(name, { ...options, skip: true }, fn).

bench.only([name][, options], fn)#>

bench(name, { ...options, only: true }, fn) 的简写。

🌐 Shorthand for bench(name, { ...options, only: true }, fn).

suite([name][, options], fn)#>

  • name <string> 套件名称。默认值: fn 的 name 属性,或者当 fn 没有名称时为 '<anonymous>'。
  • options <Object>
    • diagnosticChannels <Array> 字符串诊断通道名称会被嵌套的测试套件和基准继承。数组中的符号值会被静默忽略。默认值: []。
    • only <boolean> 选择此测试套件中嵌套的所有基准。默认值: false。
    • skip <boolean> | <string> 跳过此测试套件中所有嵌套的基准测试。 默认值: false。
    • tags <string[]> 嵌套测试套件和基准继承的标签。 默认值: []。
  • fn <Function> | <AsyncFunction> 一个声明嵌套套件、基准测试和钩子的函数。
  • 返回:<Promise> 当顶层套件完成时,或者在另一个套件中声明时会立即与 undefined 一起完成。

当收集声明时,测试套件函数会运行。在基准执行开始之前,会等待返回 Promise 的套件函数完成。

🌐 Suite functions run while declarations are collected. Promise-returning suite functions are awaited before benchmark execution begins.

describe([name][, options], fn)#>

suite() 的别名。

🌐 Alias for suite().

before(fn)#>

注册一个钩子,在当前测试套件的基准测试开始前只运行一次。

🌐 Registers a hook that runs once before the benchmarks in the current suite.

after(fn)#>

注册一个钩子,在当前测试套件的基准测试之后只运行一次。

🌐 Registers a hook that runs once after the benchmarks in the current suite.

beforeEach(fn)#>

  • fn <Function> | <AsyncFunction> 这个钩子函数。它接收一个包含基准测试的 name、params 和 signal 的对象。

注册一个钩子,在当前套件中的每个完整的逻辑基准之前运行一次。它不会在每个样本之前运行。每个样本的设置应该放在基准函数中,在 context.start() 或 context.record() 之前。

🌐 Registers a hook that runs once before each complete logical benchmark in the current suite. It does not run before every sample. Per-sample setup belongs in the benchmark function before context.start() or context.record().

afterEach(fn)#>

  • fn <Function> | <AsyncFunction> 这个钩子函数。它接收一个包含基准测试的 name、params 和 signal 的对象。

注册一个钩子,在当前套件中的每个完整的逻辑基准测试之后运行一次。它不会在每个样本之后运行。每个样本的清理应该在基准函数里在 context.end() 或 context.record() 之后进行。

🌐 Registers a hook that runs once after each complete logical benchmark in the current suite. It does not run after every sample. Per-sample cleanup belongs in the benchmark function after context.end() or context.record().

run([options])#>

  • options <Object>
    • namePattern <string> | <RegExp> 只运行完整层级名称与模式匹配的基准测试。字符串值会被解释为 JavaScript 正则表达式。
    • samples <number> 覆盖每个基准测试的最大测量回调调用次数。必须是正的 32 位无符号整数。
    • signal <AbortSignal> 允许中途终止正在进行的基准测试。
    • warmup <number> 覆盖每个基准的未报告预热回调调用次数。必须是32位无符号整数。
    • yieldBetweenSamples <boolean> 在采样回调之间安排事件循环的一个周期。默认值: true,或者对于显式运行器,使用传递给 createRunner() 的值。
  • 返回:BenchmarksStream

返回用于进程内基准运行的对象模式事件流。在声明基准的同一轮中调用 run(),在自动执行开始之前。如果不需要返回的流,可以选择不调用 run()。由 createRunner() 创建的显式运行器不会自动运行,因此其 run() 函数可以稍后调用。

🌐 Returns the object-mode event stream for the in-process benchmark run. Call run() during the same turn in which benchmarks are declared, before automatic execution begins. Calling run() is optional when the returned stream is not needed. An explicit runner created by createRunner() does not run automatically, so its run() function may be called later.

import { bench, run } from 'node:bench';

bench('example', { samples: 3 }, (b) => {
  b.start();
  doWork();
  b.end(1);
});

for await (const { type, data } of run()) {
  if (type === 'bench:complete' && data.error === undefined) {
    console.log(data.name, data.summary.mean);
  }
} 

runFile(path[, options])#>

  • path <string> | <Buffer> | <URL> 一个基准模块的路径。
  • options <Object>
    • env <Object> 子进程环境。属性值必须是字符串或 undefined。这会替代父环境,而不是扩展它。默认值: process.env 的快照。
    • execArgv <string[]> Node.js 子进程的命令行选项。这是替换继承的选项,而不是扩展它们。基准测试运行器的选项、位置参数以及选择其他执行模式的选项不允许使用。默认值: 从当前进程继承的兼容选项。
    • signal <AbortSignal> 在中止时终止子进程。
  • 返回:BenchmarksStream

在一个新的子进程中准确运行一个基准模块,并返回它的对象模式事件流。当调用 runFile() 时,相对路径 path 会从当前工作目录解析。path 不会被当作通配符处理。除非信号被中止或流在启动前被销毁,否则每次调用都会使用一个新的子进程。输入发现、排序、并发、重试和多文件调度仍然由调用方负责。

🌐 Runs exactly one benchmark module in a fresh child process and returns its object-mode event stream. A relative path is resolved from the current working directory when runFile() is called. path is not interpreted as a glob. Unless the signal is aborted or the stream is destroyed before startup, every call uses a new child. Input discovery, ordering, concurrency, retries, and multi-file scheduling remain the caller's responsibility.

当启用权限模型时,调用者必须对 path 拥有文件系统读取权限,并且有创建子进程的权限。

🌐 When the Permission Model is enabled, the caller must have file system read access to path and permission to create child processes.

记录使用先进的子进程序列化,保留支持的结构化值,例如 bigint 和错误。子进程写入 stdout 和 stderr 会变成 'bench:diagnostic' 记录。权限失败、模块加载错误、子进程异常退出或取消操作也会发出错误诊断,并生成一个终端 'bench:summary',其 success 属性为 false;这些执行失败不会让流出错。如果模块评估在声明基准测试之后失败,这些声明仍会在失败的总结之前执行。

🌐 Records use advanced child process serialization, preserving supported structured values such as bigint and errors. Child writes to stdout and stderr become 'bench:diagnostic' records. A permission failure, module loading error, abnormal child exit, or cancellation also emits an error diagnostic and produces a terminal 'bench:summary' whose success property is false; these execution failures do not error the stream. If module evaluation fails after declaring benchmarks, those declarations still run before the unsuccessful summary.

env、有效的继承选项,以及明确提供的 execArgv 会在调用 runFile() 时被复制。运行器会移除 NODE_OPTIONS,替换与 IPC 相关的环境变量,并设置它自己的子上下文、运行身份和文件身份变量,这些会覆盖 env 中同名的属性。子 Node.js 选项要通过 execArgv 传递,而不是 NODE_OPTIONS。标准的 child_process 环境传播仍然适用,包括 NODE_V8_COVERAGE、权限模型选项以及必需的 z/OS 变量。在子进程启动前中止 signal 会生成一个 AbortError 诊断,而不会启动子进程。在执行过程中中止则会向子进程发送 SIGTERM,如果子进程没有退出,则会升级为 SIGKILL。销毁返回的流遵循相同的终止流程。

类:BenchContext#>

🌐 Class: BenchContext

一个 BenchContext 实例会传递给每次基准测试调用。每次预热和测量样本都会创建一个新实例。

🌐 An instance of BenchContext is passed to every benchmark invocation. A new instance is created for every warmup and measured sample.

context.index#>

当前 context.phase 中的零基调用索引。预热样本和测量样本有各自独立的索引序列。

🌐 The zero-based invocation index within the current context.phase. Warmup and measured samples have separate index sequences.

context.name#>

基准名称

🌐 The benchmark name.

context.params#>

基准的规范化参数元数据。

🌐 The benchmark's canonicalized parameter metadata.

context.phase#>

当前的样本阶段。未报告的热身调用是 'warmup',测量的调用是 'measurement'。

🌐 The current sample phase. It is 'warmup' for an unreported warmup invocation and 'measurement' for a measured invocation.

context.signal#>

当基准测试被中止、超时或完成时触发的中止信号。

🌐 An abort signal that is triggered when the benchmark is aborted, times out, or finishes.

context.start()#>

使用 process.hrtime.bigint() 开始测量区域。多次调用 start() 会出错。

🌐 Starts the measured region using process.hrtime.bigint(). Calling start() more than once is an error.

context.end(operations[, options])#>

  • operations <number> 已完成操作的数量。必须是一个正的安全整数。
  • options <Object>
    • detail <any> 额外的可结构化克隆示例数据。使用 CLI 进程隔离时,它还必须支持高级子进程序列化。
  • 返回:<Object> 样本的 operations、duration_ns、计算得到的 rate,以及可选的克隆 detail。

结束测量区域。在验证 operations 之前会捕获结束时间戳。在 start() 之前调用 end()、调用多次,或者记录零时长样本都是错误的。如果提供了 detail,它会在捕获结束时间戳后被克隆,所以克隆时间是在测量区域之外的。

🌐 Ends the measured region. The end timestamp is captured before operations is validated. Calling end() before start(), calling it more than once, or recording a zero-duration sample is an error. When provided, detail is cloned after the end timestamp is captured, so cloning time is outside the measured region.

context.record(sample)#>

  • sample <Object>
    • operations <number> 已完成操作的数量。必须是一个正的安全整数。
    • duration_ns <bigint> 外部测量的正持续时间,以纳秒为单位,不超过 Number.MAX_SAFE_INTEGER。
    • detail <any> 额外的可结构化克隆示例数据。使用 CLI 进程隔离时,它还必须支持高级子进程序列化。
  • 返回:<Object> 归一化样本,包括它计算出的 rate 和可选的克隆 detail。

记录由另一个时钟或执行环境进行的测量。当一个高级工具在工作线程中测量工作并且需要排除消息传输时间时,这很有用。在一个回调中,record() 与 start() 和 end() 互斥,并且必须恰好调用一次。

🌐 Records a measurement made by another clock or execution environment. This is useful when a higher-level tool measures work in a worker and needs to exclude message transport from the duration. record() is mutually exclusive with start() and end() within one callback and must be called exactly once.

context.diagnostic(message[, options])#>

  • message <any> 一个可结构化克隆的诊断值。使用 CLI 进程隔离时,它还必须支持高级子进程序列化。
  • options <Object>
    • level <string> 要么是 'info',要么是 'warning'。默认: 'info'。
    • detail <any> 额外的可结构化克隆诊断数据。使用 CLI 进程隔离时,它还必须被高级子进程序列化支持。
  • 返回:<undefined>

将与当前基准测试、阶段和样本索引相关的诊断排队。多个诊断会保留调用顺序。它们会在样本回调完成后、该样本的 'bench:sample' 事件之前发出。即使热身样本不会发出,热身诊断也会被发出。在回调失败之前排队的诊断会在失败的 'bench:complete' 事件之前发出,并且它们本身不会导致基准测试失败。如果在回调完成之前超时或中止发生,排队的诊断可能不会被发出。

🌐 Queues a diagnostic associated with the current benchmark, phase, and sample index. Multiple diagnostics preserve call order. They are emitted after the sample callback settles and before that sample's 'bench:sample' event. Warmup diagnostics are emitted even though warmup samples are not. Diagnostics queued before a callback failure are emitted before the failed 'bench:complete' event and do not themselves cause the benchmark to fail. If a timeout or abort wins before the callback settles, queued diagnostics might not be emitted.

消息和细节会同步克隆。选项也会同步验证。因此,在 context.start() 和 context.end() 之间调用 diagnostic() 会把这些工作算入测量的时间。无效的参数或者不可克隆的消息或细节会违反示例协议。

🌐 The message and detail are cloned synchronously. Options are also validated synchronously. Calling diagnostic() between context.start() and context.end() therefore includes that work in the measured duration. Invalid arguments or an uncloneable message or detail violate the sample contract.

context.done()#>

请求在当前测量的样本之后完成基准测试。回调仍然必须调用 start() 和 end(),或者 record()。在预热调用期间调用 done() 是错误的。如果 done() 没有被调用,配置的 samples 值仍然是最大测量调用次数。

🌐 Requests successful benchmark completion after the current measured sample. The callback must still call either start() and end(), or record(). Calling done() during a warmup invocation is an error. The configured samples value remains the maximum number of measured invocations if done() is not called.

类:BenchmarksStream#>

🌐 Class: BenchmarksStream

BenchmarksStream 是一个对象模式的 <stream.Readable>。每条生命周期记录既会作为命名事件被发出,也会以 { type, data } 的形式在流中可用。

事件会按执行顺序发出:

🌐 The events are emitted in execution order:

  • 'bench:plan'
  • 'bench:start'
  • 'bench:sample'
  • 'bench:complete'
  • 'bench:diagnostic'
  • 'bench:summary'

命名事件的有效负载、可读记录以及基准完成值都是独立的快照。通过一种传递机制收到的值的更改不会影响通过其他机制收到的值。和其他<EventEmitter>事件一样,同一个命名事件的多个监听器会收到相同的事件负载。通过<SharedArrayBuffer>引用的内存仍然是共享的,遵循结构化克隆语义。

🌐 Named event payloads, readable records, and benchmark completion values are independent snapshots. Mutating a value received through one delivery mechanism does not change values received through the others. As with other <EventEmitter> events, multiple listeners for the same named event receive the same event payload. Memory referenced through a <SharedArrayBuffer> remains shared, following structured clone semantics.

一旦消费者开始读取,运行器会遵守流的对象模式高水位标记,并在消费者比生产者慢时在记录之间等待。这些等待发生在采样计时结束后,记录不会被丢弃。快照的创建和传输等待不计入基准超时。在可读消费开始之前,记录会累积在标准可读缓冲区中,并包含在 readableLength 中。这可以防止未读流和只使用命名事件的消费者产生死锁,但缓冲区可能会无限增长。不需要可读记录的仅命名事件消费者应调用 stream.resume() 来丢弃它们。销毁流会停止可读传输,但不会取消基准执行,因此基准完成的承诺仍会结算。自动调度的模块级运行会在内部清空其流。

🌐 Once a consumer starts reading, the runner honors the stream's object-mode high-water mark and waits between records when the consumer is slower than the producer. These waits occur after sample timing has ended, and records are not dropped. Snapshot creation and delivery waits are excluded from benchmark timeout accounting. Before readable consumption starts, records accumulate in the standard readable buffer and are included in readableLength. This keeps an unread stream and a consumer using only named events from deadlocking, but the buffer can grow without bound. A named-event-only consumer that does not need readable records should call stream.resume() to discard them. Destroying the stream stops readable delivery but does not cancel benchmark execution, so benchmark completion promises still settle. Automatically scheduled module-level runs drain their stream internally.

通过进程隔离,每个子进程发送的记录只有在父进程接受后才会收到确认。子进程在收到确认之前不会发送额外的记录,这样当报告程序比较慢时,进程间通信的传输就会被限制。

🌐 With process isolation, each record sent by a child is acknowledged only after the parent has accepted it. A child sends no additional record until it receives that acknowledgement, bounding the IPC relay when a reporter is slow.

每个基准范围的事件都包含 runId、fileRunId、entryFile、benchId、parentId 和 namePath。runId 和 fileRunId 是不透明的,并且在每次运行时都会变化。entryFile 标识导致声明的顶层基准文件,而 file 标识声明本身的源位置。parentId 基于包含它的测试套件的源文件和层级名称路径。

🌐 Every benchmark-scoped event contains runId, fileRunId, entryFile, benchId, parentId, and namePath. runId and fileRunId are opaque and change between runs. entryFile identifies the top-level benchmark file whose loading caused the declaration, while file identifies the source location of the declaration itself. parentId is based on the containing suite's source file and hierarchical name path.

在异步套件声明完成后,进程内运行器会按照声明顺序为收集到的每个基准测试发出一个 'bench:plan' 事件。该运行器的所有计划会在其套件钩子或基准测试回调执行之前发出。使用进程隔离时,文件会在不同的子进程中运行,所以后面的文件的计划会在前一个子进程完成之后才发出。没有隔离时,所有文件共享一个运行器,它们的计划会在任何基准测试执行之前发出。计划数据包含基准测试范围的身份、位置、标签以及 基准测试结果 中描述的参数,同时还包括:

🌐 After asynchronous suite declarations settle, an in-process runner emits one 'bench:plan' event for every benchmark it collected, in declaration order. All plans from that runner are emitted before its suite hooks or benchmark callbacks run. With process isolation, files run in separate children, so plans for a later file are emitted after an earlier child has completed. With no isolation, all files share one runner and their plans are emitted before any benchmark executes. Plan data contains the benchmark-scoped identity, location, tags, and parameters described in benchmark result, together with:

  • diagnosticChannels <string[]> 在每次回调期间订阅的继承字符串通道名称。
  • samples <number> 在运行级别覆盖之后,测量回调调用的有效最大次数。
  • warmup <number> 运行级别覆盖后未报告的热身回调调用的有效次数。
  • timeout <number> | <null> 超时时间(毫秒),如果没有配置超时则为 null。
  • yieldBetweenSamples <boolean> 事件循环轮次是否安排在采样回调之间。
  • selected <boolean> 在应用 skip、only 和 namePattern 选择后,该基准是否有资格运行。执行仍可能因重复声明、套件构建、钩子、终止或其他运行时失败而被阻止。
  • skip <boolean> | <string> 当 selected 是 false 时,明确的跳过值或选择原因,例如 'only' 或 'name pattern'。

该计划包含执行者已知的执行设置。运行时版本、操作系统、处理器和其他环境元数据故意留给报告工具和更高级的工具收集。

🌐 The plan contains execution settings known to the runner. Runtime version, operating system, processor, and other environment metadata are intentionally left for reporters and higher-level tools to collect.

'bench:complete' 数据包含一个 基准测试结果。失败的结果有一个额外的 error 属性,并且可能包含错误发生前记录的样本。跳过的结果有一个额外的 skip 属性,以及一个空的 samples 数组。'bench:diagnostic' 报告加载、套件和钩子错误,以及公共上下文诊断。上下文诊断包含基准范围内的身份字段、phase、index、message、level、源位置,以及可选的 detail。'bench:summary' 包含整体的 runId、fileRunId、entryFile、success、counts、duration_ns 和 file 属性。fileRunId、entryFile 和 file 是 <string> | <null>;当汇总多个文件时,它们是 null。

示例结果#>

🌐 Sample result

每个测量样本都有以下属性:

🌐 Each measured sample has the following properties:

  • operations <number> 传递给 context.end() 或 context.record() 的正操作计数。
  • duration_ns <bigint> 以纳秒为单位的测量时间。
  • rate <number> 每秒操作次数。
  • detail <any> 可选的克隆样本详情。

基准测试结果#>

🌐 Benchmark result

一个完成的基准测试结果包含:

🌐 A completed benchmark result contains:

  • runId <string> 不透明的逻辑运行身份。
  • fileRunId <string> 不透明的文件运行器或子执行身份。
  • entryFile <string> | <null> 导致此声明的顶层文件。
  • benchId <string> 在相同源布局中的稳定声明身份。
  • parentId <string> | <null> 包含套件身份的稳定版本。
  • name <string> 基准名称
  • namePath <string[]> 分层套件和基准名称。
  • file <string> 声明源文件。
  • line <number> 源线路。
  • column <number> 源列
  • tags <string[]> 继承的规范标签
  • params <Object> 规范参数元数据。
  • samples <Object[]> 按测量调用顺序的精确测量样本。
  • summary <Object>
    • mean <number> 每个样本速率的等权算术平均,而不是所有操作和时间段的总吞吐量。
    • median <number> 每个样本的中位率
    • min <number> 每个样本的最小速率。
    • max <number> 每个样本的最大速率。
    • stddev <number> 比率的总体标准差
    • coefficientOfVariation <number> stddev / mean。
    • confidenceInterval <Object> 均值速率的95%学生t置信区间,具有lower和upper属性。
    • medianConfidenceInterval <Object> 中位数发生率的95%非参数置信区间,基于lower和upper属性。
    • skewness <number> 缩放速率直方图的偏度。
Node.js 中文网 - 粤ICP备13048890号