Repository navigation
CI perf: Gradle e2e shards unbalanced — b-p leg is the merge-queue critical path (~3 min/merge-group run) #1171
Description
Activity
- addedci-perfCI / merge-queue performance finding (profiler routine)CI / merge-queue performance finding (profiler routine)
on Oct 8, 2026 mikolalysenko commented
on Oct 8, 2026 CollaboratorAuthorMore actions[agent] Triaged as
priority:p3(CI-only). No open PR references this yet; it is not a duplicate of the other CI-perf reports (#1170–#1178 each target a different workflow cost).
Generated by Claude Code
mikolalysenko commented
on Oct 9, 2026 CollaboratorAuthorMore actionsFresh numbers from the profiler run at 2026-10-09 04:16 UTC. Sample: 23 successful
CImerge_group runs since 18:00.Gradle e2e shard run time, p50 / p90 (n=22 each):
shard p50 p90 8.14.3/21 gradle_hosted_b13.2 13.8 9.8.0/21 _b11.7 12.1 7.6.6/17 --skip11.3 11.9 8.14.3/21 --skip10.8 11.2 7.6.6/17 _b10.1 12.9 6.9.4/11 --skip9.8 10.3 6.9.4/11 _b9.4 10.6 9.8.0/21 --skip8.9 9.2 - Imbalance: about 4.3 min between the slowest and fastest shard at p50, unchanged since this issue was filed.
- Critical path: a Gradle shard was the last job before
ci-okin 6 of 23 runs. Those runs were mostly between 19:00 and 21:00, while the queue was saturated: shard queue waits were 3.3–6.7 min. Since 22:50,coverage-docker (sbt)→coverage-mergehas been last instead (CI perf: coverage-docker (sbt) — 5.6-min image rebuild warms 3 unused toolchains; merge-queue critical path (~2.5 min/merge-group run) #1225). - After CI perf: coverage-docker (sbt) — 5.6-min image rebuild warms 3 unused toolchains; merge-queue critical path (~2.5 min/merge-group run) #1225: shards end at about 17–19 min, so Gradle becomes the critical path again. Rebalancing so that every shard is at most ~11 min would save about 2 min per merge_group run.
Generated by Claude Code
mikolalysenko commented
on Oct 11, 2026 CollaboratorAuthorMore actionsCI profiler, 2026-10-11 12:17 UTC: resolved by #1375 (lean per-package-manager gate, merged 2026-10-09 21:29).
The 8 unbalanced Gradle shards no longer run on PRs or in the merge queue. Lean scope runs one Gradle e2e leg (
e2e_gradle_discovery_build e2e_gradle_agent_build e2e_redirect_gradle_build …) at p50 4.4 / p90 4.7 min (n=29 merge_group runs in the last 24h). The leg was the last job beforeci-okin 1 of 30 merge_group runs, against 6 of 23 when this issue was refreshed. The merge-queue critical path is nowe2e-build-windows(#1385), and a whole merge_group run takes p50 9.6 min, where one shard alone used to take 13.2. The full-shard matrix runs only nightly, on dispatch and in full scope, where shard balance doesn't gate merges. Closing as completed.
Generated by Claude Code
Measurement
Merge queue. After Cut merge-group CI from ~46 to ~20 min: shard Gradle e2e and test legs, skip test-release in queue, cancel orphaned runs #1133 merged (2026-10-08 18:55Z), merge_group CI reaches
ci-okin 22–25 min. In 8 of the 9 non-cancelled runs since then, the last job to finish beforeci-okwas a Gradle leg ofe2e (ubuntu-latest, e2e_redirect_gradle_build | e2e_gradle_discovery_build …). Examples:PRs. Since 10-08 02:16Z, an
e2e (ubuntu-latest, …)leg was the last job beforeci-okin 94 of 111 successful PR CI runs. PR push →ci-okp50 is 41 min and p90 53 min.Per-leg times. "Run e2e tests" step in three merge_group runs (37828932401, 37828929076, 37822515034), per Gradle line:
hosted_[b-p]Job links: 8.14.3 b-p, 784 s, 7.6.6 catch-all, 694 s, 8.14.3 b-p, 803 s, 7.6.6 b-p, 750 s.
Where the time goes
Breakdown of job 113492360128, 817 s in total:
Slowest tests:
hosted_fallback_snippet_compiles_kotlin(8.14.3)hosted_config_cache_second_row(8.14.3)hosted_vendored_takeover_and_eject(7.6.6)hosted_multiproject_buildsrc_includebuildhosted_stale_lock_failshosted_tamper_failsMost other hosted tests take 140–190 s.
In the agent leg, three suites (12 s, 200 s, 373 s) run back-to-back in a
forloop, so each suite's tail is serialized.Root cause
The shards from #1133 split by test-name prefix (ci.yml around lines 1240–1255, enforced by
scripts/test_ci_gradle_prefixes.py), not by measured duration. Per Gradle line the total is about 1,900 s across 4 legs, which would be about 480 s each if balanced, but the heaviest leg takes 700–800 s.Each test is slow by design: it uses a fresh
GRADLE_USER_HOMEand--no-daemon(gradle_build_common/mod.rs:283). Caching~/.gradlewould not help.Proposed fix
hosted_config_cache_second_rowandhosted_fallback_snippet_compiles_kotlinfrom b-p, andhosted_vendored_takeover_and_ejectfrom the catch-all. Update the skip/filter lists andscripts/test_ci_gradle_prefixes.pyso that every test still runs in exactly one leg.&+wait, or onecargo testinvocation with multiple--test) instead of a sequentialforloop. That saves about 1 min on that leg.--test-threads=8for the two heaviest Gradle legs. The tests are CPU-bound, so this is an estimated ~6 min off the leg.Expected saving
ci-oklatency on Gradle-heavy runs.e2e-buildis done) also shortens this path. With both changes, the Gradle legs stay the long pole, so this remains additive.Coverage and risk
test_ci_gradle_prefixes.pymust keep asserting full and disjoint coverage.ci-okandclippyare unchanged.Effort
S
ROI
ROI = critical-path saving × confidence / effort. Critical-path minutes are weighted at 1.0 per minute.