Skip to content

@workflow/next 4.0.9 dev: eager builder queues one full rebuild per watchpack batch of its start-up scan #4445

Description

@yoavaviram

Summary

In @workflow/next 4.0.9, the eager dev builder (dist/builder-eager.js) queues one full rebuild for every watchpack aggregated event and never merges queued work. At next dev start, watchpack reports its own initial scan of the working directory in several batches. The builder skips only the first batch, so every later batch is read as newly added files and queues its own full rebuild. Those full rebuilds (whole-app discovery plus both esbuild bundles) then run back to back inside the dev server process, for minutes on a medium-sized project.

Versions

  • @workflow/next 4.0.9 (dev / watch mode only)
  • next 16.2.12 (next dev)
  • Node 24.14.0, Windows 11

What the builder does

In getNextBuilderEager(), watch mode:

  1. watcher.watch({ directories: [workingDir], startTime: 0 }) starts watchpack over the whole working directory.
  2. Watchpack walks the tree and emits aggregated several times while the initial scan is still running.
  3. The handler compares the current time-info entries with the previous ones. The first event is skipped (isInitial). Every later event contains files that were not in the previous snapshot, so addedFiles is non-empty.
  4. For each such event it calls enqueue(async () => fullRebuild()). enqueue chains onto a promise queue with no deduplication, so N batches mean N full rebuilds, run one after another.

A queued rebuild that has not started yet re-globs and re-reads the tree when it does start (fullRebuild calls getInputFiles()), so every rebuild after the first pending one repeats the same work.

Observed effect

Measured by hooking the watcher's aggregated event on a project of about 2,700 files in the working directory:

  • The initial scan arrived as 6 to 15 aggregated batches per start (for example 29, 971, 623, 342 added files in successive batches).
  • Each batch queued one full rebuild. Each full rebuild took 10 to 78 s depending on machine load. On one cold start, 15 batches queued 15 full rebuilds, which ran back to back for about three minutes after start.
  • While the queue ran, other work in the same process slowed sharply. Local-world reads of .next/workflow-data took 100 to 280 ms each inside the dev server, against 1.3 to 1.6 ms for the same files from a standalone script. There was no event-loop freeze (no timer gap over 2.7 s), so this is contention, not blocking.
  • An identical workflow scenario (the same number of storage reads and flow deliveries each time) took 18.8 s once the queue had drained, against 32.3 s, 62.2 s and 84.3 s while it was still running. Median flow delivery went from 236 ms to 453 to 1,700 ms.

Reproduction

  1. A Next.js app using withWorkflow, with a few thousand files under the project directory (source, tests, docs and so on).
  2. Remove .next and run next dev.
  3. Count the rebuilds: the rebuild log lines repeat many times with no source change, or add a counter in the aggregated handler and in fullRebuild. There is one full rebuild per aggregated batch after the first.
  4. Start a workflow run during that window and compare its timing with a run after the queue has drained.

Suggested fix

Keep at most one pending rebuild. A batch that arrives while a rebuild is waiting merges into it (a full rebuild wins over a modified-only one). The pending slot clears when the rebuild starts, so any change that arrives after it has started still queues exactly one more, and no change is lost. This only touches watch mode, so next build is unaffected.

With this change applied locally to 4.0.9, a cold start with 6 batches ran 2 rebuilds instead of 6, and the queue drained about a minute before the first workflow run. The workflow-data read and delivery timings above returned to their drained values (4.2 ms mean read, 236 ms median delivery, 18.8 s for the scenario).

--- a/dist/builder-eager.js
+++ b/dist/builder-eager.js
@@ -150,6 +150,8 @@ async function getNextBuilderEager() {
                     return normalizedEntries;
                 };
                 let rebuildQueue = Promise.resolve();
+                // The one rebuild waiting to start: 'full' | 'modified' | null.
+                let pendingRebuild = null;
                 const enqueue = (task) => {
                     rebuildQueue = rebuildQueue.then(task).catch((error) => {
                         console.error('Failed to process file change', error);
@@ -264,14 +266,25 @@ async function getNextBuilderEager() {
                     for (const added of addedFiles) {
                         knownFiles.add(added);
                     }
+                    // A rebuild that has not started yet reads the tree when it starts, so one
+                    // pending rebuild covers every change that arrives before it; a change after
+                    // it starts queues one more.
+                    const kind = addedFiles.length > 0 || removedFiles.length > 0 ? 'full' : 'modified';
+                    if (pendingRebuild) {
+                        if (kind === 'full') {
+                            pendingRebuild = 'full';
+                        }
+                        return;
+                    }
+                    pendingRebuild = kind;
                     enqueue(async () => {
-                        if (addedFiles.length > 0 || removedFiles.length > 0) {
+                        const run = pendingRebuild;
+                        pendingRebuild = null;
+                        if (run === 'full') {
                             await fullRebuild();
                             return;
                         }
-                        if (modifiedFiles.length > 0) {
-                            await rebuildExistingFiles();
-                        }
+                        await rebuildExistingFiles();
                     });
                 });
                 watcher.watch({

Notes on newer versions

From reading the published code only (not measured): in 4.1.13 the watcher starts with startTime: Date.now(), which may stop the initial scan from being reported as added files. The aggregated handler there still calls enqueue(async () => { await fullRebuild(); }) once per event with no merging, so a burst of events (a branch switch, a formatter run, a code generator) would still queue one full rebuild per batch. Open PRs #3333 and #3470 appear to rework this area on main (debounced rebuilds, one follow-up rebuild on overlap). If those land and cover the 4.x line, this issue can be closed against them.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions