Repository navigation
feat: --no-vision loads a vision-capable checkpoint as text-only - #206
Merged
Merged
Conversation
VLM auto-detection loads checkpoints that ship a vision_config (Qwen3.5/3.6 MoE among them) through the VLM factory, and the prompt cache is skipped for every VLM-loaded model (skipPromptCache includes isVLM). A text-only coding agent on such a checkpoint therefore re-prefills its whole history on every turn. --no-vision skips auto-detection and loads the text LLM, so the hybrid prompt cache applies. On a 16 GB M2 with a 35B-A3B qwen3_5_moe and --stream-experts, turn 2 of an agent session (14.7k tokens) drops from 166.6 s to 31.7 s and turn 3 (16.5k) from 189.0 s to 29.7 s. --vision and --no-vision are mutually exclusive. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014J4GsqznbKNX8tcKzNRxEr
…SharpAI#202 Gemma 4 VLM text-only requests are now cached, so the flag only matters for other VLM families. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Member
|
Thanks for this. I pushed two small changes to the branch (AI-assisted, Claude Code), plus a merge of current
🤖 Generated with Claude Code |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
VLM auto-detection loads any checkpoint that ships a
vision_configthrough the VLM factory. Qwen3.5/3.6 MoE checkpoints do this even when they're used purely for text. Every VLM-loaded model skips the prompt cache (skipPromptCache = isMultimodalRequest || params.kvBits != nil || isVLM), so a text-only coding agent on such a checkpoint re-prefills its whole history on every turn.--no-visionskips auto-detection and loads the text LLM, so the existing hybrid prompt cache applies. It cannot be combined with--vision, and when both are given argument parsing rejects them. The flag only affects auto-detection; with it unset, behaviour is unchanged.Measured
Setup: 16 GB M2, a 35B-A3B
qwen3_5_moe(4-bit),--stream-experts --ssd-prefetch --prefill-size 2048. The workload is a three-turn agent session, where each turn appends to the previous messages.--no-visionPrompt cache HIT (hybrid): 12680/14657Prompt cache HIT (hybrid): 14650/16504Known interaction:
--ctx-sizeturns off the hybrid cache--ctx-sizeset, Qwen3.5 full-attention layers becomeRotatingKVCache, becausenewCache→makeAttentionKVCachehonoursmaxKVSize.hybridCacheBoundaryonly acceptsMambaCacheplus plainKVCacheSimple, so the cache is silently skipped.--no-vision --ctx-size 32768: no hits on turns 2 and 3 (166.6 s and 189.0 s).RotatingKVCachethat hasn't wrapped yet, or at least to log why the cache was skipped. To cap prompt length without losing the cache, feat: --max-prompt-tokens rejects over-long prompts before prefill #205's--max-prompt-tokensis the alternative.Testing
swift test --filter SwiftLMTests: 207 tests, 0 failures. NewNoVisionFlagTestscover parsing and the mutual exclusion with--vision.qwen3_5_moe reports vision support, but --no-vision was given; loading as a text-only LLM.The cache hits are shown above.🤖 Generated with Claude Code