Skip to content

feat: --no-vision loads a vision-capable checkpoint as text-only - #206

Merged
solderzzc merged 3 commits into
SharpAI:mainfrom
CodeAndCanvas728:pr/no-vision
Oct 4, 2026
Merged

solderzzc merged 3 commits into
SharpAI:mainfrom
CodeAndCanvas728:pr/no-vision

Conversation

@CodeAndCanvas728

@CodeAndCanvas728 CodeAndCanvas728 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Why

VLM auto-detection loads any checkpoint that ships a vision_config through the VLM factory. Qwen3.5/3.6 MoE checkpoints do this even when they're used purely for text. Every VLM-loaded model skips the prompt cache (skipPromptCache = isMultimodalRequest || params.kvBits != nil || isVLM), so a text-only coding agent on such a checkpoint re-prefills its whole history on every turn.

--no-vision skips auto-detection and loads the text LLM, so the existing hybrid prompt cache applies. It cannot be combined with --vision, and when both are given argument parsing rejects them. The flag only affects auto-detection; with it unset, behaviour is unchanged.

Measured

Setup: 16 GB M2, a 35B-A3B qwen3_5_moe (4-bit), --stream-experts --ssd-prefetch --prefill-size 2048. The workload is a three-turn agent session, where each turn appends to the previous messages.

Turn Prompt Auto-detected VLM --no-vision
1 12.7k 146.5 s (cold) 145.1 s (cold)
2 14.7k 178.6 s, full re-prefill 31.7 s: Prompt cache HIT (hybrid): 12680/14657
3 16.5k not run (the VLM path never uses the cache) 29.7 s: Prompt cache HIT (hybrid): 14650/16504

Known interaction: --ctx-size turns off the hybrid cache

  • With --ctx-size set, Qwen3.5 full-attention layers become RotatingKVCache, because newCache → makeAttentionKVCache honours maxKVSize.
  • hybridCacheBoundary only accepts MambaCache plus plain KVCacheSimple, so the cache is silently skipped.
  • I measured --no-vision --ctx-size 32768: no hits on turns 2 and 3 (166.6 s and 189.0 s).
  • This PR doesn't change that. One possible follow-up is to accept a RotatingKVCache that hasn't wrapped yet, or at least to log why the cache was skipped. To cap prompt length without losing the cache, feat: --max-prompt-tokens rejects over-long prompts before prefill #205's --max-prompt-tokens is the alternative.

Testing

  • swift test --filter SwiftLMTests: 207 tests, 0 failures. New NoVisionFlagTests cover parsing and the mutual exclusion with --vision.
  • Live: the server logs qwen3_5_moe reports vision support, but --no-vision was given; loading as a text-only LLM. The cache hits are shown above.
  • Built with Xcode 27.0 (27A266a). The flag is documented in the README flags table.

🤖 Generated with Claude Code

CodeAndCanvas728 and others added 3 commits October 4, 2026 17:00
VLM auto-detection loads checkpoints that ship a vision_config (Qwen3.5/3.6
MoE among them) through the VLM factory, and the prompt cache is skipped for
every VLM-loaded model (skipPromptCache includes isVLM). A text-only coding
agent on such a checkpoint therefore re-prefills its whole history on every
turn. --no-vision skips auto-detection and loads the text LLM, so the hybrid
prompt cache applies.

On a 16 GB M2 with a 35B-A3B qwen3_5_moe and --stream-experts, turn 2 of an
agent session (14.7k tokens) drops from 166.6 s to 31.7 s and turn 3 (16.5k)
from 189.0 s to 29.7 s.

--vision and --no-vision are mutually exclusive.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014J4GsqznbKNX8tcKzNRxEr
…SharpAI#202

Gemma 4 VLM text-only requests are now cached, so the flag only matters for
other VLM families.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@solderzzc

Copy link
Copy Markdown
Member

Thanks for this. I pushed two small changes to the branch (AI-assisted, Claude Code), plus a merge of current main:

  1. The "reports vision support, but --no-vision was given" note is now gated on !self.audio. Before, --audio --no-vision printed that note even though the audio path was what actually decided the load.
  2. Refreshed the README row and --help text: after fix: use the prompt cache for Gemma 4 text requests on VLM loads #202 the prompt cache is no longer skipped for Gemma 4 VLM text-only requests. The flag still matters for other VLM families (Qwen-VL / Qwen3.5 etc.), so the wording now says "most VLM-loaded models".

NoVisionFlagTests passes locally on the merged head. I have not merged this; waiting on CI and the maintainer.

🤖 Generated with Claude Code

@solderzzc
solderzzc merged commit 2a85ed2 into SharpAI:main Oct 4, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants