Skip to content

runtests: Add parallel mode - #4648

Draft
hdiethelm wants to merge 7 commits into
LinuxCNC:masterfrom
hdiethelm:tests_parallel_v2
Draft

hdiethelm wants to merge 7 commits into
LinuxCNC:masterfrom
hdiethelm:tests_parallel_v2

Conversation

@hdiethelm

@hdiethelm hdiethelm commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

This adds a parallel execution mode for runtests and uses it in CI.

The linuxcnc instances are isolated using bwrap, so this works without #2722

Solves: #4588

Additionally, I fixed the issue that ctrl-c did not work with runtests.

ToDo:

  • Properly check if all tests pass and also fail if they should
  • Check the gui test CI artifacts

TBD if an issue:

  • There is one issue: with -v, the stdout and stderr are not in order any more
  • The output is not in order of the tests. --keep-order would allow to keep the order. However: parallel: Warning: No more file handles. with many threads

@hdiethelm

Copy link
Copy Markdown
Contributor Author

Bummer: https://github.com/containers/bubblewrap/releases
--overlay-src is only available in 0.11.0 so it needs #4477 and won't work on debian bookworm.

@BsAtHome

BsAtHome commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

It is still a useful tool for running the tests on the local dev machine. Once CI is upgraded we can at least run trixie and sid on parallel tests.

@BsAtHome

BsAtHome commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

When you propagate -n in the WORKER_OPT, then you need to add bwrap options to map file stdout and stderr in the test's directory. Otherwise they will get lost when the overlay is removed.

To easily add extra options, you may want to make a list of bwrap options like:

BWRAPOPTS=(
    "--ro-bind" "/" "/"
    "--dev" "/dev"
    "--tmpfs" "/tmp"
    "--tmpfs" "/var/tmp"
    "--overlay-src" "$HOME" "--tmp-overlay" "$HOME"
    "--overlay-src" "$TOPDIR" "--tmp-overlay" "$TOPDIR"
    "--bind" "$TOPDIR/tests" "$TOPDIR/tests"
    "--unshare-ipc"
    "--unshare-pid"
    "--unshare-net"
    "--proc" "/proc"
    "--die-with-parent"
)
...
        CMD="bwrap ${BWRAPOPTS[*]} -- scripts/runtests ${WORKER_OPT[*]} -w {}"

@hdiethelm

hdiethelm commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

It is still a useful tool for running the tests on the local dev machine. Once CI is upgraded we can at least run trixie and sid on parallel tests.

Might be there are options... Backports, i could manually create overlays. Or just build the new bwrap in ci... ;-)

As long as one job takes longer, the whole CI run stays constant.

Comment thread scripts/runtests.in Outdated
Comment on lines +420 to +426
c) CLEAN_ONLY=1; WORKER_OPT+=(-c) ;;
n) NOCLEAN=1 ; WORKER_OPT+=(-n) ;;
u) NOSUDO=true; WORKER_OPT+=(-u) ;;
v) VERBOSE=1; WORKER_OPT+=(-v) ;;
s) STOP=1; WORKER_OPT+=(-s) ;;
p) PRINT=1; WORKER_OPT+=(-p) ;;
d) export ENABLE_CRASHDUMPS=1; WORKER_OPT+=(-d) ;;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Propagating -c does not make sense. It is resolved immediately below.
Propagating -s does not make sense because each test runs as a singular instance that always stops. However, parallel can be instructed to stop processing files when one process fails. That is where this option should be redirected to.
Propagating -d probably requires extra bwrap options to ensure the crash dump is accessible.

@hdiethelm

Copy link
Copy Markdown
Contributor Author

When you propagate -n in the WORKER_OPT, then you need to add bwrap options to map file stdout and stderr in the test's directory. Otherwise they will get lost when the overlay is removed.

This option binds the whole test folder, so the results are directly written to the host's filesystem:
--bind \"$TOPDIR/tests\" \"$TOPDIR/tests\"

To easily add extra options, you may want to make a list of bwrap options like:

BWRAPOPTS=(
    "--ro-bind" "/" "/"
    "--dev" "/dev"
    "--tmpfs" "/tmp"
    "--tmpfs" "/var/tmp"
    "--overlay-src" "$HOME" "--tmp-overlay" "$HOME"
    "--overlay-src" "$TOPDIR" "--tmp-overlay" "$TOPDIR"
    "--bind" "$TOPDIR/tests" "$TOPDIR/tests"
    "--unshare-ipc"
    "--unshare-pid"
    "--unshare-net"
    "--proc" "/proc"
    "--die-with-parent"
)
...
        CMD="bwrap ${BWRAPOPTS[*]} -- scripts/runtests ${WORKER_OPT[*]} -w {}"

Thanks, looks better.

@BsAtHome

BsAtHome commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

When you propagate -n in the WORKER_OPT, then you need to add bwrap options to map file stdout and stderr in the test's directory. Otherwise they will get lost when the overlay is removed.

This option binds the whole test folder, so the results are directly written to the host's filesystem: --bind \"$TOPDIR/tests\" \"$TOPDIR/tests\"

That is a problem. Running parallel tests may have side effects. Some tests share files that are not meant to be accessed/changed in parallel.

@hdiethelm

hdiethelm commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

When you propagate -n in the WORKER_OPT, then you need to add bwrap options to map file stdout and stderr in the test's directory. Otherwise they will get lost when the overlay is removed.

This option binds the whole test folder, so the results are directly written to the host's filesystem: --bind \"$TOPDIR/tests\" \"$TOPDIR/tests\"

That is a problem. Running parallel tests may have side effects. Some tests share files that are not meant to be accessed/changed in parallel.

Hmm, are you sure? Each test is in a separate folder and I guess it should not write files in an other folder.

I could bind a temporary folder for each test, copy all content in and at the end, copy the results together. But this would be cumbersome. Might be there is an overlay option to do something similar, just more efficient.

@BsAtHome

BsAtHome commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

This option binds the whole test folder, so the results are directly written to the host's filesystem: --bind \"$TOPDIR/tests\" \"$TOPDIR/tests\"

That is a problem. Running parallel tests may have side effects. Some tests share files that are not meant to be accessed/changed in parallel.

Hmm, are you sure? Each test is in a separate folder and I guess it should not write files in an other folder.
I could bind a temporary folder for each test.

Yes, some tests share stuff. Most often they are in a sub-subdirectory of tests. For example, there are written variable files or intermediaries. The only two files we know of that should move out of the overlay are stdout and stderr. The rest should remain private.

@hdiethelm

hdiethelm commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

This option binds the whole test folder, so the results are directly written to the host's filesystem: --bind \"$TOPDIR/tests\" \"$TOPDIR/tests\"

That is a problem. Running parallel tests may have side effects. Some tests share files that are not meant to be accessed/changed in parallel.

Hmm, are you sure? Each test is in a separate folder and I guess it should not write files in an other folder.
I could bind a temporary folder for each test.

Yes, some tests share stuff. Most often they are in a sub-subdirectory of tests. For example, there are written variable files or intermediaries. The only two files we know of that should move out of the overlay are stdout and stderr. The rest should remain private.

Do you have an example of such a test? For the gui tests, there are also some images needed.

I need to see how to move files out of overlays with bwrap.

Just rebased on top of #4477 to see how well it works but in CI, some tests fail. And in CI+Docker, there are still permission issues.

Any clue where:
+SET_TERM_COND termCond=2, tolerance=0.001
arrives from? I had that also locally but then it just disappeared.

But 14min down to 1m45 would be quite an improvement. In CI, doc's are anyway the longest running process, so it would not decrease the overall runtime, just the worker usage.

Locally, it would be nice anyway.

@hdiethelm
hdiethelm force-pushed the tests_parallel_v2 branch 3 times, most recently from 08005c8 to f21f2ce Compare October 8, 2026 21:30
@hdiethelm

Copy link
Copy Markdown
Contributor Author

Sorry about the many CI runs. Locally all works but in CI not. However, I found a bug:
$HOME/.tool.mmap must be available for the ui-smoke tests to pass. This is probably created by an earlier tests normally but in parallel mode, this is isolated, that's why they fail: #4656

There are also other strange issues but they also appear only in CI, like touchy_postgui.hal:7: Pin 'touchy_test.mpg' does not exist or scripts/runtests: line 102: result: Value too large for defined data type. I was not able to reproduce them until now, not locally and also not in a docker container.

Due to it will anyway not work in CI until it is updated until #4477 is merged, it would probably make sense to get this working in CI later and have parallel mode only for local testing.

I found a way to get the newly created files out of the overlay, so everything can be isolated. See last commit. It is a bit cumbersome and but it seams to work.

@hdiethelm

Copy link
Copy Markdown
Contributor Author

@grandixximo @BsAtHome
Anyone of you have an idea why this fails?
https://github.com/LinuxCNC/linuxcnc/actions/runs/37963039420/job/113930409249?pr=4648
The important hints are:
touchy_postgui.hal:7: Pin 'touchy_test.mpg' does not exist
and
+SET_TERM_COND termCond=2, tolerance=0.001

Reducing the number of jobs from 32 to 8 reduces the failed tests and sometimes also no fail except the GCC G71 issue. Here, only one test fails with no obvious reason: https://github.com/LinuxCNC/linuxcnc/actions/runs/37964686002/job/113936125329?pr=4648

Parallel runtests with installed debian packages are still broken, I am on it, looks like a path issue from my side.

@hdiethelm
hdiethelm force-pushed the tests_parallel_v2 branch 3 times, most recently from ba9f388 to 58c9f9e Compare October 9, 2026 22:13
@hdiethelm

hdiethelm commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor Author

So, after some messing around with github-ci / docker and so on:

  • Less jobs give less failures in CI. Could be:
    • Tests have issues if CPU is to busy
    • Something is not isolated
  • The Debian package tests have issues
    • chmod 0777 $(find tests/ -type d) results in "Value too large for defined data type", chown -R testrunner:testrunner tests solves this
    • Tests needing sudo for halcompile fail due to sudo inside bwrap does not work. The tests would need to be run as root but this is blocked by linuxcnc. One example: tests/module-loading/rtapi-app-main-fails
  • Crashdumps will most probably also not work

Looks like getting it to run in CI with Debian packages is something between a lot of effort to not possible.

In CI for the rest of the jobs looks possible but needs #4477

So I moved the CI integration to an other branch for later: https://github.com/hdiethelm/linuxcnc-fork/tree/tests_parallel_v2_ci

I think parallel tests are still useful for quick local testing. In CI, it is anyway less useful due to the total runtime is set by the doc build.

grandixximo and others added 3 commits October 10, 2026 09:27
The first stat.poll() in a process tried tool_mmap_user() once. If task
had not created ~/.tool.mmap yet, a static flag latched and every later
poll() returned NULL without setting an exception, so Python raised
SystemError forever and a new linuxcnc.stat() did not help. A stale
~/.tool.mmap from an earlier run hid it; isolated test runs and a fresh
HOME failed every time.

Retry the attach on each poll until it succeeds and leave initialized at
0 meanwhile, so tool_table and toolinfo report no data instead of
touching a NULL mapping. poll() warns once, and says when the mmap
became available, so the warning does not read as a failure (hdiethelm
asked for that after hunting the cause in CI logs).

tool_mmap_user() warns once and names the file and the step that
failed: open or fstat with the errno, or a file still shorter than the
mapping. The creator truncates before extending, and mapping the file in
that window would SIGBUS on first access, so a short file counts as
absent. BsAtHome asked for the reason and the size_t comparison.

Drop the drive.py comment that described the SystemError as normal
startup noise.

Fixes LinuxCNC#4656
@hdiethelm

Copy link
Copy Markdown
Contributor Author

So, squashed, a few bugs fixed and rebased on top of #4658

What I tested so far:

time scripts/runtests -p -n on base branch (tool-mmap-poll-retry)

Runtest: 327 tests run, 327 successful, 0 failed, 3 skipped, 0 shmem errors

real	13m9.683s
user	1m12.772s
sys	0m54.528s

time scripts/runtests -p -n -j 32 on this branch

Runtest parallel: 330 tests 330 tests run, 327 successful, 0 failed, 3 skipped, 0 shmem errors

real	1m12.718s
user	0m2.309s
sys	0m1.289s

The contents of the tests folders after the tests look quite similar, no files missing but there are differences like pid / times and so on. So after the round trough the overlayfs, all files arrive correctly in tests.

At least on my PC, it passes most of the time with:
for ((i=4;i<64;i+=4)) ; do time scripts/runtests -j $i -p || break ; done
Once, I got:

--- /home/hannes/linuxcnc-src/tests/hal-show/expected	2026-10-10 00:25:14.508323629 +0200
+++ /home/hannes/linuxcnc-src/tests/hal-show/result	2026-10-10 02:03:30.283284848 +0200
@@ -2,7 +2,7 @@
 Owner   Type  Dir                 Value  Name
     10  bool  IN                  FALSE  conv-bool-uint.0.in <== net-conv-bool-uint.0.in
     16  real  IN                      0  conv-real-sint.0.in <== net-conv-real-sint.0.in
-    19  sint  IN                      0  conv-sint-real.0.in <== net-conv-sint-real.0.in
+    18  sint  IN                      0  conv-sint-real.0.in <== net-conv-sint-real.0.in
     22  sint  IN                      0  conv-sint-uint.0.in <== net-conv-sint-uint.0.in
     13  uint  IN     0x0000000000000000  conv-uint-bool.0.in <== net-conv-uint-bool.0.in
     25  uint  IN     0x0000000000000000  conv-uint-sint.0.in <== net-conv-uint-sint.0.in
@@ -38,7 +38,7 @@
 Owner   Type  Dir                 Value  Name
     10  bool  IN                   TRUE  conv-bool-uint.0.in <== net-conv-bool-uint.0.in
     16  real  IN           2.147484e+09  conv-real-sint.0.in <== net-conv-real-sint.0.in
-    19  sint  IN            -2147483648  conv-sint-real.0.in <== net-conv-sint-real.0.in
+    18  sint  IN            -2147483648  conv-sint-real.0.in <== net-conv-sint-real.0.in
     22  sint  IN   -9223372036854775808  conv-sint-uint.0.in <== net-conv-sint-uint.0.in
     13  uint  IN     0x00000000FFFFFFFF  conv-uint-bool.0.in <== net-conv-uint-bool.0.in
     25  uint  IN     0xFFFFFFFFFFFFFFFF  conv-uint-sint.0.in <== net-conv-uint-sint.0.in
*** /home/hannes/linuxcnc-src/tests/hal-show: FAIL: result differed from expected

I have the feeling that some tests are just a bit flaky. Now that you can run the tests really fast, it just surfaces. I will run the original tests in a loop for a day or so to see if I have similar issues. But it might also be that something is not properly isolated and there are rare races.

One issue: parallel mode only works in run_in_place without -u due to sudo is never used with run_in_place. An option to fix this issue would be to run all sudo tests sequential after running the other tests when run_in_place is not active. Then it would even work in CI + Debian package.

@grandixximo

Copy link
Copy Markdown
Contributor

@BsAtHome maybe cleanup all the uneccesary sudo from the CI First? I can look into that if you give the go ahead :-)

@hdiethelm

Copy link
Copy Markdown
Contributor Author

@BsAtHome maybe cleanup all the uneccesary sudo from the CI First? I can look into that if you give the go ahead :-)

I think this test for example needs sudo when run with debian package installed: tests/module-loading/rtapi-app-main-fails
Without sudo, you can not install a hal component. And if it is not installed, you can not load it due to security reasons.

So as long as one test needs sudo in this case, either no parallel or the sequential/parallel fix.

Can you test if this branch also works on your PC? Due to all the issues I had in CI, it is well possible that there are also issues on other PC's.

@grandixximo

Copy link
Copy Markdown
Contributor

Tested on my machine: Debian 13, 8 cores, 15 GB, bwrap 0.12.0, GNU parallel 20240222, the whole run nested inside my own bwrap sandbox under a 6 GB memory cap.

  • sequential -p tests: 327 run, 327 ok, 3 skipped, 13m38s
  • -j 4: 327 ok, 0 failed, 3 skipped, 3m32s
  • -j 8: 327 ok, 0 failed, 3 skipped, 1m49s
  • -j 16: 327 ok, 0 failed, 3 skipped, 1m20s

I did not hit the hal-show owner id flake in these runs.

Failure detection works: with a bogus line appended to tests/hal-show/expected, -j 8 exits 1, prints the diff and lists the test under Failed:, and the tree is clean afterwards.

One difference from sequential mode: without -n, a failed test's result and stderr are lost with the tmp-overlay, while sequential only removes them on success. Maybe always use the result overlay and copy back the failed tests (all of them with -n)?

Nit: the summary reads 330 tests 330 tests run, 327 successful, ... 3 skipped, while sequential prints 327 tests run ... 3 skipped.

@BsAtHome

Copy link
Copy Markdown
Contributor

@BsAtHome maybe cleanup all the uneccesary sudo from the CI First? I can look into that if you give the go ahead :-)

Yes, all unnecessary sudo should be eliminated. Generally, you wouldn't allow sudo to run during local runs or when building a package. That would be a potential security risk.
The only required exception we have is the setsuid/setcap in the build process. But that you have to run separately on your local machine because you do not need it for RIP.

@BsAtHome

Copy link
Copy Markdown
Contributor
--- /home/hannes/linuxcnc-src/tests/hal-show/expected	2026-10-10 00:25:14.508323629 +0200
+++ /home/hannes/linuxcnc-src/tests/hal-show/result	2026-10-10 02:03:30.283284848 +0200
@@ -2,7 +2,7 @@
...
-    19  sint  IN                      0  conv-sint-real.0.in <== net-conv-sint-real.0.in
+    18  sint  IN                      0  conv-sint-real.0.in <== net-conv-sint-real.0.in
...

This is clearly a race condition. It shows the order of component creation differs. The component ID increments by one on every comp_id = hal_init(compname) call and halcmd creates a component. There is a sequence of events while forking off sub processes that is wrapped in hal_exit()/hal_init() (and the underlying rtapi_init()/rtapi_exit()). That can definitely disturb the ordering when the hal_init() call is stalled for a moment and the child gets to create a component before the halcmd parent process re-registers itself.

The solution is to ignore the component ID in the test result comparison. The race cannot be solved easily because there are multiple calls to rtapi_init() (including one in the hal_lib_init() code path), and you cannot really control the ordering. All calls to rtapi_init() act on one and the same shared memory counter.

This allows to run tests in parallel using bwarp for isolation of the
linuxcnc processes
@hdiethelm

Copy link
Copy Markdown
Contributor Author

Tested on my machine: Debian 13, 8 cores, 15 GB, bwrap 0.12.0, GNU parallel 20240222, the whole run nested inside my own bwrap sandbox under a 6 GB memory cap.

* sequential `-p tests`: 327 run, 327 ok, 3 skipped, 13m38s

* `-j 4`: 327 ok, 0 failed, 3 skipped, 3m32s

* `-j 8`: 327 ok, 0 failed, 3 skipped, 1m49s

* `-j 16`: 327 ok, 0 failed, 3 skipped, 1m20s

I did not hit the hal-show owner id flake in these runs.

Failure detection works: with a bogus line appended to tests/hal-show/expected, -j 8 exits 1, prints the diff and lists the test under Failed:, and the tree is clean afterwards.

Nice, thanks for testing. Yes, this tests are mostly flaky in CI with 32 jobs and 4 cpu's. Locally, I only have a fail ~1 in 10 runs, even if I go up to 128 threads. I need to test this in a CPU limited VM or on my slow CNC pc.

Right now, I run it with 16 cores (9950X) and 96GB RAM. I was lucky to upgrade my PC before AI hit... :-)

One difference from sequential mode: without -n, a failed test's result and stderr are lost with the tmp-overlay, while sequential only removes them on success. Maybe always use the result overlay and copy back the failed tests (all of them with -n)?

The reason behind this was, that the CI failed and I expected the different overlay mounts to be an issue. Looks like this was not the case, changed back to copy back always.

Nit: the summary reads 330 tests 330 tests run, 327 successful, ... 3 skipped, while sequential prints 327 tests run ... 3 skipped.

Yes, the reason behind was consistency check during developing. I added the check in code and changed the output to be equal. Now it shows:
Runtest parallel: 327 tests run, 327 successful, 0 failed, 3 skipped, 0 shmem errors

tests/halcompile/userspace-count-names uses sudo but does not declare it
@hdiethelm

Copy link
Copy Markdown
Contributor Author

So while this is still true: Looks like getting it to run in CI with Debian packages is something between a lot of effort to not possible.
There is an easy way around: Just not run sudo tests in parallel mode. There are only a few that don't take long. Pushed.

@hdiethelm

Copy link
Copy Markdown
Contributor Author

Finally, I got everything running in CI: #4665

Just the G71 tests fail due to the GCC issue.

There was indeed an issue with overlaying directory's to /tmp inside docker for test data recovery due to it does not support an overlay mount inside an overlay mount. Fixed with --tmpfs /tmp:rw,size=512m.

It cuts down the runtime from 16-19 min to 7-8 min for the jobs switched to parallel except for bookworm where bwrap is to old.

So I wait until #4658 is merged and #4663 has settled and then call this ready. I solved the sudo problem, so #4663 is not absolutely necessary. But if it is decided to remove sudo tests, it should go in first.

Still open from my side: Check for flaky tests

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants