Skip to content

fix: demote resident params to disk residency when memory reclamation fails - #2119

Merged
leejet merged 2 commits into
leejet:masterfrom
losewayy:demote-resident-params
Oct 10, 2026
Merged

leejet merged 2 commits into
leejet:masterfrom
losewayy:demote-resident-params

Conversation

@losewayy

@losewayy losewayy commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

Summary

When ModelManager::ensure_compute_backend_capacity exhausts its normal tiers (other runners' workspace, preferred eviction order, global LRU eviction), states resident on the compute device (params_backend == compute_backend) remain unevictable because they have no reloadable backing copy. This adds a final tier: demote eligible resident states to ResidencyMode::Disk in LRU order, which only unlocks evictability — release_params_storage_blocks still decides what is actually freed, and the unchanged disk-resident load path reloads them on next use.

Demotion is skipped for protected, pinned, staged, unloaded, required, view, split-buffer, LoRA-fused, and non-file-backed states, and it never runs when capacity already fits. States that end up not being released get their residency reverted, so permanent disk residency applies only to tensors that actually entered the disk lifecycle — afterwards they behave identically to native --params-backend ...=disk params.

Fixes #2118. On a 12 GB card the Qwen Image 2.1 Q6_K DiT (5.6 GB) previously stayed GPU-resident while the Wan-family VAE needed a 7.8 GB decode workspace, forcing 49 tiles at 256px (15.9 s) instead of a full-frame decode (4.4 s).

Test evidence (RTX 5070 Ti Laptop 12 GB, CUDA, master f89d9b1)

Scenario Before After
Qwen Image 2.1 @ 1024x1024 VAE decode 15.9 s, 49 tiles @ 256px ~3.7-4.5 s, full frame
Wan2.2 TI2V-5B @ 480x832x33f decode 366.9 s, 84 tiles @ 112x128 103.7 s, 15 tiles @ 240x256
Wan2.2 TI2V-5B end-to-end 402.9 s 136.4 s
Qwen Image 2.1 @ 832x480 (no pressure) 2.06 s full frame unchanged, no demotion
SD 1.5 512x512 / img2img / hires-fix normal unchanged
Oversized decode (1280x1280, ~13 GB peak) tiled still tiled (correct fallback)

Reload path verified with sd-server: after the first request evicts the demoted DiT, the second request reloads it from disk in ~2.3 s and samples normally; subsequent decodes evict only as much as needed (LRU, block-granular).

Notes for reviewers

  • Demotion is a state flip only; no new memory-management primitives are introduced.
  • A demoted state that is evicted and later reloaded is indistinguishable from a native disk-resident state, including the pre-existing limitation that a loaded disk-resident tensor cannot move params backends in assign_compute_backend (this already applies to --params-backend disk today).
  • Eviction remains block-granular: a block frees only when all of its states are eligible, so blocks mixing eligible and ineligible tensors simply stay resident.

Summary by CodeRabbit

  • Bug Fixes
    • Improved recovery when compute-device memory is insufficient by reclaiming eligible, reloadable tensors before reporting failure. Protected or otherwise ineligible tensors remain untouched.
    • When this recovery provides enough capacity, the operation can continue instead of immediately failing. If capacity remains insufficient, the existing failure behavior is preserved.

@coderabbitai

coderabbitai Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 0df3ea12-7243-4178-ab3a-9a609336d492

📥 Commits

Reviewing files that changed from the base of the PR and between 6842cd0 and cd4c797.


📒 Files selected for processing (1)
  • src/model_manager.cpp

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 9 remain after this review.



📝 Walkthrough

Walkthrough

The compute-capacity check now has a final reclamation pass for eligible parameter tensors resident on the requested non-CPU backend. It temporarily marks them disk-resident, retries eviction, and restores their residency as needed.

Changes

Compute capacity reclamation

Layer / File(s) Summary
Demote and restore resident states
src/model_manager.cpp
The capacity check temporarily marks eligible loaded parameter states as disk-resident before retrying eviction. Eligibility excludes views, CPU or mismatched backends, pinned or staged tensors, required or protected states, mmap-backed or split-buffer-backed states, states without source metadata, and states with an applied LoRA epoch. The method restores demoted states that remain loaded after capacity is reached, or restores all demoted states if capacity remains insufficient.

Priority: ➖ Normal

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix · Severity of issue fixed: Medium

Suggested reviewers: leejet


Merge Risk: ⚪ Minimal · up to cd4c7

The change lets memory-constrained runs reclaim resident model memory instead of falling back to slower tiled decoding. No actionable merge-blocking risk was identified.

Architecture Summary

Architecture risk: 🔵 Low · up to cd4c7

The change affects 1 system.

Changed systems: src

Architecture concerns
No architecture-level concerns identified.

Review details

Systems and components

  • observed — src (service) was modified; 1 changed file maps to changed impact.

Before / after behavior

  • observed — Modified behavior in src/model_manager.cpp: After the existing eviction attempts, the method gathers states held by mmap storage blocks and temporarily marks eligible resident tensors as disk-resident before retrying eviction. Eligibility excludes views, CPU or mismatched backends, pinned or staged tensors, protected or required states, mmap-backed or split-buffer-backed states, empty sources, and states with an applied LoRA epoch. Successful capacity reclamation restores the demoted states that remain loaded to parameter-backend residency; if no attempt succeeds, all demotions are restored before the existing capacity failure handling continues.

Pre-merge checks | Passed 6 | Failed 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 1 functions across 1 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (6 passed)
Check name Status Explanation
Title check Passed The title clearly and concisely describes the main change: demoting resident parameters to disk residency when normal memory reclamation fails.
Description check Passed The description provides a complete summary, motivation, implementation details, exclusions, test evidence, reload-path verification, and performance results. The related issue is included in the summ…
Linked Issues check Passed The pull request directly references and fixes issue #2118. The implementation matches the issue objective of making reloadable resident parameters evictable under memory pressure.
Out of Scope Changes check Passed The change is focused on one memory-reclamation path in src/model_manager.cpp. No unrelated files, public declarations, or alternative issue directions are included.
Linked Issues check Passed The change satisfies issue #2118. After ordinary eviction fails, ensure_compute_backend_capacity scans global_candidates and demotes eligible resident states to ResidencyMode::Disk. It excludes …
Out of Scope Changes check Passed The whole-PR diff changes only src/model_manager.cpp. The added code directly supports issue #2118 by making eligible resident parameters evictable under capacity pressure. It adds no unrelated API,…

  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR

  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

auto-fit produces worse plans than --offload-to-cpu on mid-range GPUs: resident weights are non-evictable and starve VAE decode

2 participants