fix(proxy): retry with backoff on transient accept errors instead of exiting#2369
fix(proxy): retry with backoff on transient accept errors instead of exiting#2369politerealism wants to merge 4 commits into
Conversation
…exiting The proxy accept loop unconditionally broke on any accept() error, permanently killing the proxy while the sandbox continued to report Ready. Replace the break with a sleep-and-continue pattern: EMFILE/ENFILE errors get exponential backoff (100ms to 3.2s) to let file descriptors drain, and all other accept errors get a fixed 100ms backoff matching the existing metadata_server and edge_tunnel patterns. Closes NVIDIA#2337 Signed-off-by: Sean Burdine <sburdine@nvidia.com> Signed-off-by: politerealism <burdcat17@gmail.com>
|
/ok to test 8d0099f |
- Use libc::EMFILE / libc::ENFILE instead of raw errno values; promote libc from dev-dependency to regular dependency - Extract accept_backoff() and is_fd_exhaustion_error() into testable helpers - Fix unreachable 5s cap: bump exponent limit from min(6) to min(7) so the 5_000ms ceiling is reachable (100 * 2^6 = 6400, capped to 5000) - Add 6 unit tests covering exponential progression, counter reset, saturation at cap, EMFILE/ENFILE detection, and non-FD error rejection Signed-off-by: Sean Burdine <sburdine@nvidia.com> Signed-off-by: politerealism <burdcat17@gmail.com>
|
Follow-up filed: #2372 — the SSH server accept loop ( |
johntmyers
left a comment
There was a problem hiding this comment.
gator-agent
PR Review Status
Validation: This is a concentrated network-supervisor fix for the reproducible, linked bug in #2337.
Head SHA: 629fd7d4cef2348aea55394598c47d2600df1ccc
Review findings:
- Warning — missing regression coverage: The new unit tests cover errno classification and backoff arithmetic, but they never execute the accept loop or prove that it accepts a new connection after EMFILE/ENFILE pressure clears. Please add an isolated Unix subprocess/integration test that lowers
RLIMIT_NOFILE, induces the accept failure, releases a descriptor, and verifies a subsequent proxied connection succeeds. Subprocess isolation avoids changing the test runner's process-wide limit. - See the inline warning about distinguishing terminal listener errors from retryable failures.
Thanks @politerealism, I also checked the SSH accept-loop gap you called out in #2372. That is appropriately tracked as separate follow-up work and does not expand or block this focused proxy fix.
Docs: Fern docs are not needed because this changes internal recovery and operational logging, not a user-facing command, configuration, API, or workflow.
Next state: gator:in-review while these findings await an author update. The sandbox-network scope will require test:e2e after review feedback is resolved.
Three-way error classification for the accept loop: transient errors (EMFILE, ECONNABORTED) retry with backoff, terminal errors (EBADF, EINVAL, ENOTSOCK) exit immediately, and unknown errors exit after 5 consecutive failures. Adds unit tests for the classifier and error handler, plus a subprocess-isolated integration test that validates EMFILE recovery by lowering RLIMIT_NOFILE. Signed-off-by: Quinn Burdine <sburdine@redhat.com> Signed-off-by: politerealism <burdcat17@gmail.com>
johntmyers
left a comment
There was a problem hiding this comment.
gator-agent
PR Review Status
Validation: This remains a concentrated network-supervisor fix for the reproducible, linked bug in #2337.
Head SHA: abfe555af0e45201a1d52e57643622aae401ceb8
Review findings:
- Resolved: the new classification exits immediately for
EBADF,EINVAL, andENOTSOCK, bounds unknown errors to five retries, and resets both counters after a successful accept. - Warning — production retry coverage remains incomplete: see the inline finding on the new FD-exhaustion test.
Thanks @politerealism, I checked the error-classification and subprocess-test changes in your new commit. The terminal-listener concern is resolved, but the requested production-path regression is not yet covered because the test uses a separate blocking listener and can pass without observing EMFILE.
Docs: Fern docs are not needed because this changes internal recovery and operational logging, not a user-facing command, configuration, API, or workflow.
Next state: gator:in-review while the remaining regression-test finding awaits an author update. The sandbox-network scope will require test:e2e after review feedback is resolved.
We can ignore this and accept the new unit and integration tests. cc @gator-agent |
|
/ok to test abfe555 |
|
Label |
Signed-off-by: Quinn Burdine <sburdine@redhat.com> Signed-off-by: politerealism <burdcat17@gmail.com>
|
/ok to test 39705af |
johntmyers
left a comment
There was a problem hiding this comment.
gator-agent
PR Review Status
Validation: This remains a concentrated network-supervisor fix for the reproducible, linked bug in #2337.
Head SHA: 39705af9f2d2b7732c13a0b606be0898082badba
Review findings:
- Warning: the Linux transient accept-error allowlist is incomplete, so repeated retryable errors can still terminate the unobserved proxy task.
- Warning: the FD-exhaustion test can accept the earlier queued connection after dropping it, rather than the fresh connection it later validates.
Thanks @politerealism, I checked your explanation that the unit and integration tests divide the retry coverage. Thanks @johntmyers; I treated your maintainer waiver as resolving the earlier request to exercise the exact production async wiring. The two inline findings below are distinct issues found in the current cumulative diff.
Docs: Fern docs are not needed because this changes internal recovery and operational logging, not a user-facing command, configuration, API, or workflow.
E2E: test:e2e is already applied, and /ok to test was posted for this exact head SHA.
Next state: gator:in-review while these findings await an author update.
| AcceptErrorClass::Transient | ||
| } | ||
| Some(libc::EBADF | libc::EINVAL | libc::ENOTSOCK) => AcceptErrorClass::Terminal, | ||
| _ => AcceptErrorClass::Unknown, |
There was a problem hiding this comment.
gator-agent
Warning: Tokio consumes WouldBlock, but other OS errors reach this classifier. Linux documents ENETDOWN, EPROTO, ENOPROTOOPT, EHOSTDOWN, ENONET, EHOSTUNREACH, EOPNOTSUPP, and ENETUNREACH as accept errors reliable servers should retry like EAGAIN; ENOBUFS and ENOMEM also indicate resource pressure. These currently become Unknown and terminate the unobserved proxy task after six failures, recreating the stale-Ready failure mode (CWE-755). Add Linux-specific transient cases, use backoff for ENOBUFS/ENOMEM, and add classification tests. See Linux accept(2).
|
|
||
| // Release FDs. | ||
| held_fds.clear(); | ||
| drop(backlog_conn); |
There was a problem hiding this comment.
gator-agent
Warning: On Linux, EMFILE is returned while reserving the output FD before the pending connection is dequeued. backlog_conn can therefore remain first in the accept queue; dropping it and then accepting once may return that closed connection instead of client, making the final read fail or hang. Keep backlog_conn alive and use it to verify recovery, or accept/drain it before creating and validating the fresh connection.
Summary
breakin the proxy accept loop with a sleep-and-continue pattern so transient errors (EMFILE, ENFILE, etc.) no longer permanently kill the proxymetadata_serverandedge_tunnelpatternsRelated Issue
Closes #2337
Changes
crates/openshell-supervisor-network/src/proxy.rsbreakon accept error withtokio::time::sleep(backoff).awaitraw_os_error()(EMFILE=24, ENFILE=23)Testing
cargo check -p openshell-supervisor-network— compiles cleancargo test -p openshell-supervisor-network— all tests passChecklist