Skip to content

fix: exit with _exit after a requested shutdown (low priority) - #197

Merged
solderzzc merged 1 commit into
SharpAI:mainfrom
CodeAndCanvas728:pr/shutdown-clean-exit
Sep 26, 2026
Merged

solderzzc merged 1 commit into
SharpAI:mainfrom
CodeAndCanvas728:pr/shutdown-clean-exit

Conversation

@CodeAndCanvas728

Copy link
Copy Markdown
Contributor

Low priority. This only affects stopping the server while a request is in flight. Normal operation is unaffected.

Problem

Both shutdown handlers (SIGTERM and SIGINT) emit exiting{reason:"requested"} and then call Darwin.exit(0). exit() runs C++ static destructors and atexit handlers while an inference thread may still be inside an MLX GPU eval. So stopping the server mid-generation crashes right after it announces a clean shutdown:

  • Crash: SIGSEGV in mlx::core::fast::CustomKernel::eval_gpu, dereferencing its kernel unordered_map after it's destroyed. One case was SIGABRT instead.
  • Exit status: 139 or 134.
  • Crash report: a .ips file for every stop.

A daemon following the Engine Protocol sees exiting{requested}, then a crash, for what should be a clean stop.

Change

Flush stdout and stderr, then call _exit(0). Nothing needs those destructors at that point, and emitEvent has already flushed the exiting event. The helper is exitAfterShutdownRequest().

Verification

Repro: Qwen3.6-35B-A3B with --stream-experts --ssd-prefetch. Start a streaming 2000-token generation, then send SIGTERM 25 s in.

Exit status Crash report exiting event
Before 3/3 crashed (139, 139, 134) one .ips per run emitted
After 3/3 exit 0 none emitted

SwiftLMTests: 194 passed.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Nn7Du4yB5jtocihUNxxBGs

Low priority: only affects stopping the server while a request is in flight.

The shutdown handlers called Darwin.exit(0), which runs C++ static
destructors and atexit handlers while an inference thread may still be
inside an MLX GPU eval. Sending SIGTERM mid-generation therefore crashed
right after the exiting{reason:"requested"} event (SIGSEGV in
CustomKernel::eval_gpu's kernel map, or SIGABRT), leaving a crash report
and a non-zero exit status for what the Engine Protocol reports as a
clean stop.

Flush stdout/stderr and _exit(0) instead. Nothing needs those destructors
at this point, and the exiting event is already flushed by emitEvent.

Repro (Qwen3.6-35B-A3B, --stream-experts --ssd-prefetch): SIGTERM 25 s
into a streaming 2000-token generation. Before: 3/3 crashed (exit 139/134,
one .ips each). After: 3/3 exit 0, no crash report, exiting event present.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn7Du4yB5jtocihUNxxBGs
@solderzzc
solderzzc merged commit 48f253c into SharpAI:main Sep 26, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants