feat(pstack): make /architect designs resist agent mistakes - #495
Merged
Merged
Conversation
…d change them Co-authored-by: lauren <poteto@users.noreply.github.com>
…d hand-synced list red flags Co-authored-by: lauren <poteto@users.noreply.github.com>
Co-authored-by: lauren <poteto@users.noreply.github.com>
cursor
Bot
force-pushed
the
lauren/architect-agent-proof-designs-cfe8
branch
from
October 3, 2026 22:49
69945c2 to
c3cecc6
Compare
poteto
marked this pull request as ready for review
October 3, 2026 22:49
creedants
added a commit
to creedants/pstack-t3
that referenced
this pull request
Oct 4, 2026
Syncs `vendor/pstack` to upstream 0.15.9 ([cursor/plugins@e43c7ee](cursor/plugins@e43c7ee)), which brings in cursor/plugins#495 and #496. ## Upstream changes - `$architect` assumes the next contributor is an agent that sees only the files it opened. `design-red-flags.md` adds four flags: split ownership, two ways to do one task, importable internals, and hand-synced lists. - The Perf issue playbook replaces its eight strategy families with seven ordered performance mantras and stops at the first one that meets the target. - Hillclimb orders perf hypotheses by those mantras. `benchmark-checklist` updates its cross-reference. ## Re-ported overrides - `t3/overrides/architect/SKILL.md` - `t3/overrides/poteto-mode/playbooks/perf-issue.md` - `t3/overrides/poteto-mode/playbooks/hillclimb.md` Each takes upstream's new text and keeps its T3 role wiring. The lock is updated. ## Verification - `python3 scripts/build.py` passes with no override drift. - `python3 -m unittest discover -s tests -v` passes (36 tests).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
The next contributor to a design is usually an agent. It sees only the files it opened, copies the nearest example, and takes the shortest path that compiles.
/architectscreened candidates for depth and leakage. It did not screen for shapes where a change that looks right in one file is wrong for the repo. This change adds that screen at design time. It complements/correct, which fixes the same mistake classes in a repo that already exists.What changed
architect/SKILL.md, Phase B: judge each candidate as if an agent makes the next change. Prefer the design where a change that looks right from one file is right for the whole repo. The screen sentence no longer copies the red-flag names, so a new red flag is a change to one file.architect/references/design-red-flags.md: add four red flags in the existing format. They are split ownership, two ways to do one task, importable internals, and hand-synced list./correctmerge (feat(pstack): add /correct skill #494).Verification
node scripts/validate-plugins.mjspasses.pstack,/correct, PR numbers, agent ids, and other skill namessource,credit,adapt,based on,inspired,notes from,thanks,author, and URLse.g.,for example,such as,likelint,export,private,enum,schema,registry,codegen,barrel,package.json,single sourceuser,order,status,event,table,queue,webhook,endpointhuman,person,developer,operatorEval: old skill against new skill
Candidates are Opus 5.5 only. The judges are GPT-5.6 Sol and Fable 5.1, and they do not see the skill version. Odango EAP is not in the model list for this run. The first version of this PR used Gemini and Kimi as judges. All of that judging is redone.
Task. A small TypeScript order backend asks for outgoing webhooks. The seed code invites all four flaws:
export *that re-exports internalsnotifyCustomersender with four callersRuns. There were three
/architectruns for each skill version, assigned at random. Each run had an Opus 5.5 parent and two Opus 5.5 runners. The parent read its skill version, read the repo, wrote the rubric, screened the candidates, picked a base, grafted, and wrote the synthesized design. A harness spawned the runners and the cross-judge so that each subagent had a fixed model. Both versions had the same limits: two runners per run (the skill's minimum), a size cap, and no compile or tool runs. A first attempt with three unbounded runners per run hit the tool time limit on 16 of 18 runners, so I discarded it.Two synthesis rounds used the same candidates:
Scoring. GPT-5.6 Sol and Fable 5.1 each scored all 12 synthesized designs. The copies they scored had the synthesis notes and the red-flag names removed. Each flaw scores 0 to 2. Each quality item scores 1 to 5.
Across both rounds, the flaw total is 5.08 for the old skill and 6.00 for the new skill. Flaw totals per run (Sol/Fable):
My spot check of the 12 designs:
dbfrom the root.Possible regression. On needless layers, GPT-5.6 Sol gave all 12 designs 5. Fable gave all six new designs 4, and gave the old designs a mean of 4.5. Its reasons for the new designs were small forwarding pieces: an aggregator object, a list method that repeats another, and a second query path. The runtime file count does not show a consistent difference (one-off migration scripts are not counted). I did not change the skill for this. To check it, test whether the importable-internals red flag pushes designs toward forwarding index files.
Limits.
dbexport as a second way to emit an event. Fable mostly does not./workspaceby default. That run used the new version, so the text it could see was the same.