Repository navigation
feat(core): continue responses after output token limits - #53876
Open
rekram1-node wants to merge 6 commits into
Open
rekram1-node wants to merge 6 commits into
rekram1-node wants to merge 6 commits into
Conversation
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Continue responses that finish with
lengthwhen no local tool results already drive another request. Preserve partial output and add this synthetic user-role instruction:Allow at most two consecutive automatic continuations. Read only the three most recent user/assistant messages, including the current response. If all three ended with
length, preserve the response and report an output-limit error. This applies whether the responses contain text, reasoning, or tools. New user input or a different finish reason breaks the streak; intervening synthetic messages and tool results do not. No separate nudge counter or metadata is needed, and the bounded history check survives compaction and restart.Failed tool input gets one error tool result containing recovery guidance and an input excerpt capped at 2,048 characters, with the original length and a truncation marker when needed:
length: explain that the call was not executed; ask for smaller tool calls with complete arguments.Wait for the finish reason before choosing the malformed-input error wording. Completed tool calls still execute eagerly, and failed arguments remain
{}rather than being repaired or executed. Local tool results drive the next request without a separate synthetic nudge, but their assistant responses still count toward the consecutive output-limit cap. Provider-executed tools do not suppress text completion.Reuse the existing within-step request loop. Transport retries, compaction policy, and public APIs remain unchanged.
Verification
bun test test/session-runner.test.ts test/session-step.test.ts test/session-runner-tool-events.test.ts: 257 passed.bun test test/session-runner-message.test.ts: 29 passed.bun typecheckinpackages/core: passed.bun run check: passed (repository lint warnings remain).