Skip to content

feat(pstack): make /architect designs resist agent mistakes - #495

Merged
poteto merged 3 commits into
mainfrom
lauren/architect-agent-proof-designs-cfe8
Oct 3, 2026
Merged

poteto merged 3 commits into
mainfrom
lauren/architect-agent-proof-designs-cfe8

Conversation

@poteto

@poteto poteto commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Why

The next contributor to a design is usually an agent. It sees only the files it opened, copies the nearest example, and takes the shortest path that compiles. /architect screened candidates for depth and leakage. It did not screen for shapes where a change that looks right in one file is wrong for the repo. This change adds that screen at design time. It complements /correct, which fixes the same mistake classes in a repo that already exists.

What changed

  • architect/SKILL.md, Phase B: judge each candidate as if an agent makes the next change. Prefer the design where a change that looks right from one file is right for the whole repo. The screen sentence no longer copies the red-flag names, so a new red flag is a change to one file.
  • architect/references/design-red-flags.md: add four red flags in the existing format. They are split ownership, two ways to do one task, importable internals, and hand-synced list.
  • Bump pstack from 0.15.7 to 0.15.8. The branch is rebased on the /correct merge (feat(pstack): add /correct skill #494).

Verification

  • node scripts/validate-plugins.mjs passes.
  • Leak grep on the 26 added lines. Every group has zero hits:
    • company and product names
    • model families and slugs
    • internal names: author handles, pstack, /correct, PR numbers, agent ids, and other skill names
    • attribution words: source, credit, adapt, based on, inspired, notes from, thanks, author, and URLs
    • example markers: e.g., for example, such as, like
    • technique names: lint, export, private, enum, schema, registry, codegen, barrel, package.json, single source
    • domain nouns: user, order, status, event, table, queue, webhook, endpoint
    • words for the human: human, person, developer, operator
    • em dashes and curly quotes

Eval: old skill against new skill

Candidates are Opus 5.5 only. The judges are GPT-5.6 Sol and Fable 5.1, and they do not see the skill version. Odango EAP is not in the model list for this run. The first version of this PR used Gemini and Kimi as judges. All of that judging is redone.

Task. A small TypeScript order backend asks for outgoing webhooks. The seed code invites all four flaws:

  • an email job status that both the worker and the admin routes write
  • an event-type list copied by hand into a type union, an admin array, a SQL CHECK, and the docs
  • a root export * that re-exports internals
  • a legacy notifyCustomer sender with four callers

Runs. There were three /architect runs for each skill version, assigned at random. Each run had an Opus 5.5 parent and two Opus 5.5 runners. The parent read its skill version, read the repo, wrote the rubric, screened the candidates, picked a base, grafted, and wrote the synthesized design. A harness spawned the runners and the cross-judge so that each subagent had a fixed model. Both versions had the same limits: two runners per run (the skill's minimum), a size cap, and no compile or tool runs. A first attempt with three unbounded runners per run hit the tool time limit on 16 of 18 runners, so I discarded it.

Two synthesis rounds used the same candidates:

  • Round 1. The arena cross-judge was Gemini 3.8 Flash. I kept these designs and scored them again with the new judges.
  • Round 2. The arena cross-judge was GPT-5.6 Sol. A fresh Opus 5.5 parent did the synthesis again and could not read round 1. No Gemini, Kimi, or Grok model took part in this round.

Scoring. GPT-5.6 Sol and Fable 5.1 each scored all 12 synthesized designs. The copies they scored had the synthesis notes and the red-flag names removed. Each flaw scores 0 to 2. Each quality item scores 1 to 5.

Mean of both judges R1 old R1 new R2 old R2 new
Split ownership (0-2) 1.50 2.00 2.00 2.00
Two ways to do one task (0-2) 1.33 1.50 1.17 1.50
Importable internals (0-2) 1.00 1.00 1.00 1.00
Hand-synced lists (0-2) 1.00 1.67 1.17 1.33
Flaw total (0-8) 4.83 6.17 5.33 5.83
Interface depth (1-5) 4.83 5.00 5.00 5.00
Needless layers (1-5) 4.67 4.50 4.83 4.50
Requirements coverage (1-5) 4.00 4.83 4.83 4.50
Runtime webhook files (my count) 4.0 3.7 3.3 4.3

Across both rounds, the flaw total is 5.08 for the old skill and 6.00 for the new skill. Flaw totals per run (Sol/Fable):

  • Round 1: old 4/6, 4/6, 3/6; new 6/6, 5/6, 7/7.
  • Round 2: old 4/6, 4/6, 6/6; new 4/6, 6/7, 6/6.

My spot check of the 12 designs:

  • Deep imports. A deep import of webhook internals fails the build in 2 of 6 old designs and in 6 of 6 new designs. The judges gave all 12 designs 1 on this flaw, because each design still re-exports the existing db from the root.
  • New hand-synced lists. Five of 6 old designs add SQL CHECK lists for new fields with no test. Three of 6 new designs add no unchecked list. Two of them add a drift test, and one leaves the CHECK out and says it would be a second hand-synced list.
  • Split ownership. Every round-2 design has one writer for each field. GPT-5.6 Sol found a second writer in all three round-1 old designs and in none of the round-1 new designs.

Possible regression. On needless layers, GPT-5.6 Sol gave all 12 designs 5. Fable gave all six new designs 4, and gave the old designs a mean of 4.5. Its reasons for the new designs were small forwarding pieces: an aggregator object, a list method that repeats another, and a second query path. The runtime file count does not show a consistent difference (one-off migration scripts are not counted). I did not change the skill for this. To check it, test whether the importable-internals red flag pushes designs toward forwarding index files.

Limits.

  • Three runs per version in each round is a small sample, and both rounds use the same candidates.
  • The judges disagree most on the "two ways" flaw. GPT-5.6 Sol counts the root db export as a second way to emit an event. Fable mostly does not.
  • One round-2 parent reported that a search went to /workspace by default. That run used the new version, so the text it could see was the same.
  • The eval files are not in the diff.
Open in Web Open in Cursor 

cursoragent and others added 3 commits October 3, 2026 22:48
…d change them

Co-authored-by: lauren <poteto@users.noreply.github.com>
…d hand-synced list red flags

Co-authored-by: lauren <poteto@users.noreply.github.com>
Co-authored-by: lauren <poteto@users.noreply.github.com>
@cursor
cursor Bot force-pushed the lauren/architect-agent-proof-designs-cfe8 branch from 69945c2 to c3cecc6 Compare October 3, 2026 22:49
@poteto
poteto marked this pull request as ready for review October 3, 2026 22:49
@poteto
poteto merged commit a586282 into main Oct 3, 2026
2 checks passed
creedants added a commit to creedants/pstack-t3 that referenced this pull request Oct 4, 2026
Syncs `vendor/pstack` to upstream 0.15.9
([cursor/plugins@e43c7ee](cursor/plugins@e43c7ee)),
which brings in cursor/plugins#495 and #496.

## Upstream changes
- `$architect` assumes the next contributor is an agent that sees only
the files it opened. `design-red-flags.md` adds four flags: split
ownership, two ways to do one task, importable internals, and
hand-synced lists.
- The Perf issue playbook replaces its eight strategy families with
seven ordered performance mantras and stops at the first one that meets
the target.
- Hillclimb orders perf hypotheses by those mantras.
`benchmark-checklist` updates its cross-reference.

## Re-ported overrides
- `t3/overrides/architect/SKILL.md`
- `t3/overrides/poteto-mode/playbooks/perf-issue.md`
- `t3/overrides/poteto-mode/playbooks/hillclimb.md`

Each takes upstream's new text and keeps its T3 role wiring. The lock is
updated.

## Verification
- `python3 scripts/build.py` passes with no override drift.
- `python3 -m unittest discover -s tests -v` passes (36 tests).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants