Repository navigation
Conversation
Add `src/cfg/wto.h` with `WeakTopologicalOrdering` (`WTO`) and `WTOWorklist` built on top of `DomTree`. In a reducible CFG ordered in reverse postorder, every cycle is a natural loop headed by a block that dominates all blocks in the cycle, allowing a Bourdoncle Weak Topological Ordering to be constructed directly from the dominator tree and natural loops of the CFG. Include unit tests in `test/gtest/wto.cpp` and `TODO` comments noting follow-on optimizations.
Replace RPOQueue with WTOWorklist in ConstraintAnalysis and RedundantSetElimination so that loops stabilize before flow values propagate to downstream blocks. This avoids quadratic/cubic blowups on functions with sequential loops while also speeding up general workloads. Benchmark results across 16 WebAssembly modules (3 iterations): - --constraint-analysis: - esbuild.wasm: 381.60s -> 7.59s (-98.0%, 50.3x speedup) - 15 non-esbuild modules geomean: 1.564s -> 1.487s (-4.9%) - 15 non-esbuild modules total: 72.00s -> 57.24s (-20.5%) - All 16 modules geomean: 2.205s -> 1.646s (-25.4%) - All 16 modules total: 453.60s -> 64.83s (-85.7%) - --rse: - esbuild.wasm: >600s (timeout) -> 1.62s (>370x speedup) - 15 non-esbuild modules total: 26.40s -> 25.47s (-3.5%) - All 16 modules geomean: N/A -> 0.922s (total: 27.09s)
Read each basic block's reverse-postorder index from contents.index in DomTree instead of allocating and populating an unordered_map<BasicBlock*, Index>, and skip self-loop backedges immediately with predIndex >= index. Update OnceReduction and test/example/domtree.cpp to initialize contents.index, and remove the redundant index initialization loop in WeakTopologicalOrdering. Benchmark results across 16 WebAssembly modules (3 iterations, interleaved): - --constraint-analysis: - Geomean: 1.646s -> 1.576s (-4.3%) - Total time: 64.83s -> 62.82s (-3.1%) - --rse: - Geomean: 0.922s -> 0.859s (-6.8%) - Total time: 27.09s -> 25.40s (-6.3%)
During reverse-RPO natural loop discovery in WeakTopologicalOrdering, collapse each discovered loop body into its header using union-find with path compression, and skip over already-collapsed inner loops when walking immediate dominators in dominates(). This prevents outer loops from re-traversing inner loop bodies, bounding natural loop discovery to O(E alpha(N)) instead of O(N * depth) on deeply nested loops. Benchmark results across 16 WebAssembly modules (3 iterations, interleaved): - --constraint-analysis: - Geomean: 1.576s -> 1.564s (-0.7%) - Total time: 62.82s -> 62.40s (-0.7%) - --rse: - Geomean: 0.859s -> 0.853s (-0.7%) - Total time: 25.40s -> 25.06s (-1.3%)
When the CFG has no backedges (checked via CFGWalker::loopTops), evaluate queued blocks in a single reverse-postorder pass in WTOWorklist::run without constructing DomTree or WeakTopologicalOrdering. Benchmark results across 16 WebAssembly modules (3 iterations, interleaved): - --constraint-analysis: - Geomean: 1.610s -> 1.606s (-0.2%) - Total time: 64.36s -> 64.05s (-0.5%; dart_essentials: 2.92s -> 2.59s, -11.6%) - --rse: - Geomean: 0.870s -> 0.871s (+0.1%) - Total time: 26.10s -> 26.03s (-0.3%; dart_essentials: 2.00s -> 1.71s, -14.4%)
Replace the recursive std::variant/Cycle representation of WeakTopologicalOrdering with a single contiguous std::vector<Entry> where cycle ends store the header block and jump target index. This avoids per-cycle vector allocations and recursive evaluation in WTOWorklist::run. Benchmark results across 16 WebAssembly modules (3 iterations, interleaved): - --constraint-analysis: - Geomean: 1.606s -> 1.592s (-0.9%) - Total time: 64.05s -> 63.67s (-0.6%; faster on 11/16 modules) - --rse: - Geomean: 0.871s -> 0.852s (-2.2%) - Total time: 26.03s -> 25.39s (-2.5%; faster on 12/16 modules)
# Conflicts: # src/cfg/wto.h
# Conflicts: # src/cfg/wto.h
# Conflicts: # src/cfg/wto.h
# Conflicts: # test/gtest/wto.cpp
kripken
reviewed
Oct 8, 2026
| // end of the cycle headed by `block` (`cycleTarget` is the entry index of the | ||
| // cycle header). | ||
| struct Entry { | ||
| BasicBlock* block = nullptr; |
Member
There was a problem hiding this comment.
Suggested change
| BasicBlock* block = nullptr; | |
| BasicBlock* block; |
This made me think it was optional, and I don't see a benefit to this default?
tlively
force-pushed
the
wto-fast-paths
branch
from
October 9, 2026 03:33
7513840 to
5acacd3
Compare
kripken
approved these changes
Oct 9, 2026
kripken
left a comment
Member
There was a problem hiding this comment.
Thanks, comments look very clear now!
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Replace the recursive std::variant/Cycle representation of
WeakTopologicalOrdering with a single contiguous std::vector
where cycle ends store the header block and jump target index. This
avoids per-cycle vector allocations and recursive evaluation in
WTOWorklist::run.
Benchmark results across 16 WebAssembly modules (3 iterations,
interleaved):