0
Fork 0
mirror of https://github.com/obra/superpowers.git synced 2026-09-25 22:09:05 +00:00

Compare commits

...

226 commits

Author SHA1 Message Date
Jesse Vincent
6b2b4ae883 inline-eval: gated plan-set run under the directory format: boundary script ran 2/2, plan set kept consistent with the code 2/2, continuation 2/2 2026-09-18 17:16:12 -07:00
Jesse Vincent
f8abe60781 inline-eval: directory-form planning on both designs (same volume, 34-36 files for the game spec) and its plan executes 9/9 2026-09-18 15:31:59 -07:00
Jesse Vincent
501ff76a96 inline-eval: directory-form plan executes the same as the file form (inline 9/9, SDD 6/6, same cost) 2026-09-18 14:50:18 -07:00
Jesse Vincent
40969a5cd0 inline-eval: plan-boundary micro-test script arm (same ceiling) 2026-09-18 13:29:04 -07:00
Jesse Vincent
821ab4ead2 inline-eval: plan-boundary micro-test is a ceiling (every arm fixes the mismatch in a fresh session); the gate is judged by the full run 2026-09-18 13:20:05 -07:00
Jesse Vincent
a69d394d73 inline-eval: D1 variant recorded (plan as a directory: header file plus one file per task) 2026-09-18 13:02:36 -07:00
Jesse Vincent
8fe695fe43 inline-eval: ledgerlite-dir fixture (terse plan as a directory of task files) 2026-09-18 13:02:00 -07:00
Jesse Vincent
6f1319ba8c inline-eval: C4 variant (plan-boundary script gate before each next plan) and its micro-test arm 2026-09-18 12:57:22 -07:00
Jesse Vincent
2c40a20b9d inline-eval: plan-set execution interviews (a boundary check must be a named, gated step) 2026-09-18 12:55:31 -07:00
Jesse Vincent
2ef85de579 inline-eval: plan-boundary micro-test repo snapshot (engine built under Advance, plan 2 says Tick) 2026-09-18 12:52:24 -07:00
Jesse Vincent
c7ac1480d9 inline-eval: plan-boundary micro-test rig (ruling slot vs boundary scan for keeping later plans true) 2026-09-18 12:52:20 -07:00
Jesse Vincent
4501db0ccc inline-eval: plan-set execution: continuation 2/2, remaining-plans slot 2/2, keep-later-plans-true via ruling slot 0/2 (edit 1/2, index only) 2026-09-18 12:49:04 -07:00
Jesse Vincent
1fd0be9245 inline-eval: plan-set trap fixture: plans 1-2 consistently say Tick 2026-09-18 11:27:29 -07:00
Jesse Vincent
b0333fb132 inline-eval: C3 plan set (5 plans in 3.4k lines); plan-1 execution names remaining plans 2/2, continuation untested by that prompt 2026-09-18 11:27:14 -07:00
Jesse Vincent
f833d3edd5 inline-eval: plan-set execution fixture with a planted cross-plan naming conflict 2026-09-18 11:27:04 -07:00
Jesse Vincent
c525789d3e inline-eval: INLINE_EVAL_PROMPT override 2026-09-18 11:26:30 -07:00
Jesse Vincent
850db12a74 inline-eval: combined variant (ledgerlite 475-551 lines, executes 9/9); C3 plan-set variant queued 2026-09-18 10:11:15 -07:00
Jesse Vincent
1ec1bd405b inline-eval: C3 variant recorded (all plans up front with a Plan Set section; rulings name the later plans they touch and edit them; executors run the set) 2026-09-18 09:28:24 -07:00
Jesse Vincent
d55489a186 inline-eval: C2 variant recorded (C1 + Remaining Plans header section and closing-report slot in both executors) 2026-09-18 09:21:28 -07:00
Jesse Vincent
3764ca4319 inline-eval: C1 combined writing-plans variant recorded (recipe with verification step, proportion, first-plan, reader framing, housekeeping) 2026-09-18 09:11:16 -07:00
Jesse Vincent
b9cb1f90f1 inline-eval: replacement one-line-tests cosmic rep (2,257 lines, 17 min) 2026-09-18 00:28:41 -07:00
Jesse Vincent
49f63f5952 inline-eval: one-line-tests variant plans the cosmic spec in 3k lines with no code; its ledgerlite plan executes 9/9 2026-09-18 00:10:41 -07:00
Jesse Vincent
03cfc0d810 inline-eval: one-line tests take the ledgerlite plan to 350-400 lines 2026-09-17 23:23:23 -07:00
Jesse Vincent
8908abd89e inline-eval: proportion-arm plan executes 9/9 on Sonnet 5 2026-09-17 23:23:13 -07:00
Jesse Vincent
bb81627ef1 inline-eval: first-plan brake works 2/2; proportion self-review item is the largest single lever (cosmic spec planned in 3-3.6k lines, <20 min); skill-written terse plan executes 9/9 2026-09-17 23:10:48 -07:00
Jesse Vincent
3595883770 inline-eval: plan-completeness SDD arm 2 complete (27/27, no Haiku spend) 2026-09-17 23:09:40 -07:00
Jesse Vincent
4e14f46424 inline-eval: plan-completeness SDD arm 2 (Sonnet 5 implementers): terse = full = loosened, 27/27 probes 2026-09-17 23:09:31 -07:00
Jesse Vincent
0a71dac7dc inline-eval: reader-framing arms (framing alone trims 20% on ledgerlite; adds nothing over the recipe; recipe finishes the cosmic spec in 3 plans, 6-7.5k lines) 2026-09-17 22:13:30 -07:00
Jesse Vincent
2b0de48694 inline-eval: T1 variant (tests as one-line entries) recorded 2026-09-17 20:52:19 -07:00
Jesse Vincent
9754de4142 inline-eval: recipe wording halves the plan on both fixtures; tests are the remaining volume 2026-09-17 20:51:41 -07:00
Jesse Vincent
658f9f108b inline-eval: baseline on the cosmic-tetris spec reproduces the report (10-11k lines in 40 min, 3 of 5 plans, no drafters) 2026-09-17 20:08:26 -07:00
Jesse Vincent
508aa74122 inline-eval: record the writing-plans wording variants under test 2026-09-17 19:34:13 -07:00
Jesse Vincent
f044263f55 inline-eval: plan-completeness SDD arm with SDD's own model selection (terse = full = loosened, 24/24 probes; Opus controller is 75-80% of cost) 2026-09-17 19:24:38 -07:00
Jesse Vincent
232f4a98ff inline-eval: cosmic-tetris planning fixture (the 19k-line plan spec) and a multi-plan wpplans arm 2026-09-17 19:22:30 -07:00
Jesse Vincent
1fb2f00f16 inline-eval: wait keeps waiting while the session has background agents pending (false idle on SDD reps) 2026-09-17 18:15:31 -07:00
Jesse Vincent
902ab41d34 inline-eval: plan-completeness wave 2 (Sonnet 5 inline, wordstat: terse plan equals the full plan at 66% of the cost) 2026-09-17 17:40:19 -07:00
Jesse Vincent
03b7347027 inline-eval: plan-completeness wave 1 (Sonnet 5 inline, ledgerlite: terse plan equals the full plan at 78% of the cost) 2026-09-17 17:29:57 -07:00
Jesse Vincent
faf9106303 inline-eval: document the Bedrock alias remap variables 2026-09-17 17:12:38 -07:00
Jesse Vincent
4101189eae inline-eval: full-plan fixtures (Opus-written plans) for the plan-completeness test 2026-09-17 17:11:25 -07:00
Jesse Vincent
c651984783 writing-plans: handoff names the approaches (Subagent-driven, Native), asks which the partner prefers, and recommends one per plan; drop the untested Workflow section from the Claude Code reference 2026-09-17 17:03:28 -07:00
Jesse Vincent
6605d3b77b inline-eval: nested-orchestrator wall clock and the rep-2 post-session interview 2026-09-17 16:41:27 -07:00
Jesse Vincent
4a1c3e44c2 inline-eval: nested-orchestrator SDD (sddnest) on wordstat, 3 reps: shape followed 3/3, $3.09 vs $5.64, one rep shipped the crash via the fix wave 2026-09-17 16:26:22 -07:00
Jesse Vincent
d8dd956b43 executing-plans: re-grade findings by their effect on a reasonable person, not by a symptom list; executor-regrade micro-test; sddnest arm 2026-09-17 16:04:33 -07:00
Jesse Vincent
c70415e9d7 inline-eval: five-line Review Focus confirmation run (count respected 6/6, recall held, plan volume unchanged) 2026-09-17 15:58:18 -07:00
Jesse Vincent
dbf9687b5a writing-plans: Review Focus is the five implied cases most likely to bite, not every one 2026-09-17 15:49:05 -07:00
Jesse Vincent
cf83063626 inline-eval: sizing-wording planning test (time unit and context-window ceiling do not move task or test counts) 2026-09-17 15:36:21 -07:00
Jesse Vincent
d8541a68c0 inline-eval: review-focus-tiers micro-test results (ruling and per-input forms inflate; a stated count is the only form that cut volume) 2026-09-17 15:19:01 -07:00
Jesse Vincent
28c733dcff inline-eval: INLINE_EVAL_PLUGIN_DIR override for skill-wording variants 2026-09-17 15:03:33 -07:00
Jesse Vincent
a4dcb78bf6 inline-eval: ledgerlite rerun after the reasonable-person change; ledgerlite planning fixture; review-focus-tiers micro-test 2026-09-17 14:57:49 -07:00
Jesse Vincent
9228ce8deb inline-eval: full-frame rerun after the reasonable-person reviewer change (Codex 3/3, Claude 3/3, writing-plans 3/3) 2026-09-17 14:31:01 -07:00
Jesse Vincent
7841d5cfc9 inline-eval: wpplan arm and a design-only wordstat fixture for planning runs 2026-09-17 14:15:56 -07:00
Jesse Vincent
f60779dfa7 Reviewer: the spec is a vision document, judged by a reasonable person; declined-scope slot; plan pins implied cases as tests
code-reviewer.md gains two blocks: the spec says what the software must
do, not everything it will meet, so behavior the spec is silent on is
graded by what a reasonable person using the software would expect; and
a required 'Declined to judge' list so no scoping is silent. Micro-tested
on gpt-6-astra reviewing a CLI that catches OSError only: the current
template graded the resulting traceback Minor 6/6 (shipped under the
Critical/Important gate); with the vision-document standard, Important
6/6 with zero variance; the slot alone surfaces the scoping as a ruling
6/6.

writing-plans' Review Focus now ends with a test per line in the owning
task, since every interrogated implementer ranked 'a test for the case'
as the one thing that would have changed its choice; on Opus 5 the list
names the implied case 6/6. executing-plans rules on the reviewer's
declined lines the way it rules on plan conflicts.
2026-09-17 14:15:28 -07:00
Jesse Vincent
12eac56bee inline-eval: reviewer-scope, effort, reasonable-ruling, and plan-explicit micro-tests
Reviewer (gpt-6-astra, low): current template grades the decode case
Minor 6/6 (ships under the gate); a 'declined to judge' slot surfaces it
as a ruling 6/6; the spec-is-a-vision-document / reasonable-person
standard grades it Important 6/6 with zero variance. Effort medium on the
implementer: 0/12. Reasonable-person as the standard for 'ruled out'
lines: 3/6 handled, first implementer-side movement, noisy. Planner on
Opus 5: Review Focus list names the case 6/6; explicit-tests variant 4/6.
2026-09-17 13:53:54 -07:00
Jesse Vincent
5ba20a832b inline-eval: planner-side Review Focus micro-test on gpt-5.6-sol (0/6) and synthesis 2026-09-17 12:45:53 -07:00
Jesse Vincent
06f896e9df inline-eval: v4 boundary-slot micro-test (slot produced, enumeration absent) 2026-09-17 12:44:32 -07:00
Jesse Vincent
7f5712841d inline-eval: v3 boundary-list micro-test (prose step; worse and noisier) with follow-ups 2026-09-17 12:42:22 -07:00
Jesse Vincent
93b0217145 inline-eval: what-would-have-changed-your-choice answers from all 24 micro-test sessions 2026-09-17 12:39:32 -07:00
Jesse Vincent
9adf4137fa inline-eval: scope-vs-failure wording micro-test (24 reps, no separation)
writing-skills-style micro-test of executing-plans wording on Codex
(gpt-5.6-sol, effort low): control, current v3, a positive
scope-binds-behavior-not-failure rewrite, and a deletion-only variant,
six single-shot reps each. Every rep in every arm wrote 'except OSError';
none discussed undecodable input. Wording is not the lever; the
interrogated session's explanation was a story, not a cause.
2026-09-17 12:33:40 -07:00
Jesse Vincent
132a068d86 inline-eval: Codex interrogation (why each session chose its exception clause) 2026-09-17 12:08:04 -07:00
Jesse Vincent
8e4f785a60 inline-eval: SDD with Review Focus on Bedrock (3 reps); Sonnet 4.5 pricing 2026-09-17 11:50:24 -07:00
Jesse Vincent
5fd35286ad inline-eval: hardened ledgerlite on Bedrock (inline v3 vs bare + Opus review) 2026-09-17 11:23:35 -07:00
Jesse Vincent
0fb6a48c9c inline-eval: Codex review-chain evidence for spike reps 31-34 2026-09-17 11:06:57 -07:00
Jesse Vincent
f35c2d83cb inline-eval: harden ledgerlite with a spec rule the natural implementation violates
The first ledgerlite run showed every implementer resolving the planted
Interfaces mismatch and implementing the exit-2 path unprompted. Add a
third planted defect that no plan test covers and that Decimal() will not
catch on its own: design.md now states that an amount with more than two
fractional digits is malformed. Adds the amount-precision probe and the
Codex review-chain tracer.
2026-09-17 11:04:07 -07:00
Jesse Vincent
483d627f52 inline-eval: bedrock backend (SigV4 via the AWS CLI's default credentials)
CLAUDE_CODE_USE_BEDROCK=1 with an inference-profile model id such as
us.anthropic.claude-opus-5; verified with a fresh config dir and the
default AWS profile. Same model as the direct-API runs, so arms stay
comparable; cost.py prices the us.anthropic.* ids at first-party list
rates with the existing caveat that Bedrock bills separately.
2026-09-17 11:00:54 -07:00
Jesse Vincent
5c4dabee52 inline-eval: clean Codex results (cxbare, cxspike v3), 3 reps each 2026-09-17 09:30:51 -07:00
Jesse Vincent
0e894a399b inline-eval: prepare Codex worker homes so only the checkout's skills are visible
Three leaks found and closed in the Codex arms: skills discovered under
~/.agents/skills relative to HOME (scratch HOME per rep), OpenAI's curated
marketplace installing and enabling a released superpowers into every
fresh CODEX_HOME (remove it and disable it by its real id,
superpowers@openai-curated-remote), and a local-marketplace plugin being
inert until 'codex plugin add' installs it (done in csd's worker home
before launch; csd only adds auth.json and config.toml). Verified: the
spike rep reads skills only from the installed checkout copy and the
bare rep's home holds no superpowers.
2026-09-17 09:26:22 -07:00
Jesse Vincent
b363c695fb inline-eval: Codex isolation (scratch HOME, curated plugin off) and the contaminated first batch 2026-09-17 09:18:13 -07:00
Jesse Vincent
c89d6c81e0 inline-eval: Mantle (Claude on AWS) backend
INLINE_EVAL_BACKEND=mantle seeds the worker the way quorum's
seedClaudeMantle does: CLAUDE_CODE_USE_MANTLE=1, AWS_REGION, and
AWS_BEARER_TOKEN_BEDROCK, no API key and no key-approval fingerprint. The
bearer comes from the environment or the env file; the model must be a
Mantle id passed via INLINE_EVAL_MODEL. cost.py prices Mantle ids at the
first-party list rate of the same model so arms stay comparable, with a
note that Bedrock bills separately.
2026-09-17 09:15:18 -07:00
Jesse Vincent
833fb448c1 inline-eval: partial results from the credit-exhausted ledgerlite and sddrf runs
All twelve workers died at 23:23 UTC when the API key's org ran out of
credit. Recorded as partial with READMEs. One finding survives: on the
ledgerlite fixture every bare and inline implementer resolved the planted
Interfaces mismatch and implemented the spec's exit-2 path before any
review ran, so those planted defects do not discriminate; the next fixture
needs a defect implementers do not fix unprompted.
2026-09-16 21:27:20 -07:00
Jesse Vincent
3deae3efe5 inline-eval: Codex arms (cxbare, cxspike) and a Codex scorer
Codex workers come from csd with a fresh CODEX_HOME (auth + generated
config), so they are clean by construction; the plugin rides in as
config overrides pointing a local marketplace at the checkout (the repo
carries .agents/plugins/marketplace.json). score-codex.py works from the
csd event stream since Codex transcripts differ from Claude Code's.

Pins CSD_CODEX_MODEL=gpt-5.6-sol by default: csd 4.0.0's gpt-5.5 default
hits Codex 0.154's model-migration modal at startup, which swallows the
first prompt so no turn runs and the worker never registers.

Not yet run: the user's Codex weekly quota was under 10% at the time.
2026-09-16 17:04:17 -07:00
Jesse Vincent
6834055a87 inline-eval: fixture directories with scoring maps and defect probes; ledgerlite fixture
The runner takes INLINE_EVAL_FIXTURE (default fixtures/wordstat, which
links to the sdd-tiny evals fixture). A fixture dir carries plan.md,
design.md, starter files, plus scoring.json (task->test map, impl dir,
suite command) and probe.sh (planted-defect checks), both stripped from
the worker's copy. The scorer reads both and prints one row per probe.

fixtures/ledgerlite is a six-task plan with a real producer/consumer
chain, one planted Interfaces mismatch (Task 6 says parse_csv takes a
path; Task 2 produces parse_csv(text)), and one spec-stated behavior no
task tests (malformed amount -> stderr message, exit 2). Adds the sddrf
arm (SDD with the Review Focus section) for evaluating that change
against SDD.
2026-09-16 16:16:55 -07:00
Jesse Vincent
e501c2135e inline-eval: cost.py (USD by actual model per agent) and iteration-3 results
Prices every assistant message in a rep's main and subagent transcripts
by the model it ran on, so mixed-model arms (sonnet session + opus
reviewer) price correctly. Rates are the first-party list prices cached
in the claude-api skill reference on 2026-06-24, with cache reads at 0.1x
and writes at 1.25x input; update PRICE if they change.

results/2026-09-16-clean-iter3/ holds the v3 reps (spike 21-23), the
sonnet-session arm, the Review Focus arm, and a cost summary across every
clean arm.
2026-09-16 16:03:34 -07:00
Jesse Vincent
179bb3b02f executing-plans v3: helper scripts, TDD-verified fix pass, scoped scan, Review Focus
Five efficiency changes, all aimed around the one thing the evals showed
works (a fresh most-capable-model review of the whole diff):

- scripts/task-start and scripts/task-done fold each task's bookkeeping
  into one call apiece: brief + BASE at the start; test run, log, and
  ledger line at the end (a failing run records nothing). Every tool
  call in an inline session is a turn that re-reads the whole context.
  Tested by tests/claude-code/test-executing-plans-scripts.sh.
- The final fix pass is verified by TDD plus a green suite instead of a
  re-review dispatch. Twelve re-reviews ran across today's inline, SDD
  and workflow reps; every one returned 'all addressed, no new
  breakage'.
- The pre-flight scan covers only producer/consumer pairs named by the
  plan's Interfaces blocks; a plan with no shared interfaces gets one
  ledger line. Nine of nine inline reps produced an all-clean table.
- writing-plans emits a Review Focus section (input classes and failure
  modes the spec implies and no task's tests exercise) with a matching
  self-review step, and executing-plans hands it to the final reviewer.
- Both skills say inline runs well on a mid-tier session model with the
  most capable model reserved for the review.

inline-eval gains spikesonnet and spikerf arms to measure the last two.
2026-09-16 15:51:50 -07:00
Jesse Vincent
cf7f204c7b inline-eval: results for executing-plans iteration 2 (spike reps 11-13) 2026-09-16 15:17:44 -07:00
Jesse Vincent
6b18c51b58 executing-plans: load TDD at setup, re-grade and gate final findings, mid-tier re-review
Three changes, each from the clean inline-eval runs:

- The required test-driven-development sub-skill moved from the task
  loop into Setup, loaded once before Task 1. Zero of six clean reps
  loaded it from its old position; the plan's own 'failing test first'
  wording let executors skip it.
- Final-review findings are sorted before any fix: the executor
  re-grades first (a Minor that describes a traceback, crash, data loss
  or wrong result on valid input is Important), then Critical/Important
  enter the single fix pass and Minor goes to the ledger and a
  'Deferred minors' list in the final message. Reviewers filed the same
  unhandled UnicodeDecodeError as Minor three times today, and every
  process that gated on the label shipped it; every inline rep that
  fixed minors paid a fix pass plus re-review for polish items.
- The scoped re-review runs on a mid-tier model. Three of three clean
  reps dispatched it on Opus.

Spike: baseline runs recorded under tests/inline-eval/results/.
2026-09-16 15:09:27 -07:00
Jesse Vincent
87200ddd30 inline-eval: add barefork arm (forked adversarial reviewer per task)
Bare plan execution plus, after each task, a forked subagent told that it
is a fork of the implementer, shares its context and assumptions, and
must find issues rather than confirm. No end-of-run fresh review, so any
catch is the fork's or the primed implementer's. Scorer dedupes tool_use
ids because a fork's transcript repeats its parent's history.
2026-09-16 15:03:10 -07:00
Jesse Vincent
3f85b52504 inline-eval: add the SDD-as-Workflow script and its 3-rep results
sdd-workflow.js runs the subagent-driven-development loop through the
Workflow tool: setup+preflight, per-task prep/implement/review with a
fix loop (cap 5, opus from round 4, adjudication at the cap), final
whole-branch review with one fix wave and one scoped re-review, ledger
written at the end. Every shell step is an agent because the script has
no filesystem access. Fix-round bookkeeping matches finding text by
prefix, which misfired once (four rounds after every re-review said
addressed) — judgment-shaped matching does not belong in script code.
2026-09-16 13:56:58 -07:00
Jesse Vincent
1af7878ff9 inline-eval: add barerev/barerevs arms (bare + one-line review request)
Bare plan execution plus a single sentence asking for one fresh reviewer
at the end, on opus (barerev) or sonnet (barerevs), with no plugin and
no reviewer template. Isolates how much of the final review's value comes
from the fresh reviewer itself versus the code-reviewer.md template and
the model tier. Results under results/2026-09-16-clean/.
2026-09-16 13:38:23 -07:00
Jesse Vincent
ecd0addd5d inline-eval: add bare arm (no superpowers loaded after the plan)
Control for 'let Claude be Claude once the plan exists': no plugin,
prompt just says to execute the plan. Results under
results/2026-09-16-clean/bare-*.txt.
2026-09-16 13:27:24 -07:00
Jesse Vincent
099eb6d3a3 inline-eval: add sdd arm and merge subagent tool calls into scoring
The sdd arm runs the same fixture and checkout with the prompt asking
for subagent-driven-development, so inline and SDD can be compared on
one plan. In SDD runs the implementer subagents do the writes and test
runs, so the scorer now merges tool calls from the main transcript and
every subagent transcript into one timestamp-ordered stream before
checking RED-before-GREEN per task.

Clean 3x3 results (dev inline, spike inline, sdd) are under
results/2026-09-16-clean/.
2026-09-16 13:18:52 -07:00
Jesse Vincent
55e24aeedc inline-eval: run each rep in a clean CLAUDE_CONFIG_DIR
The first 3x2 run was confounded: workers inherited the host's global
CLAUDE.md (which mandates TDD), user hooks, and installed plugins, so
the baseline arm behaved better than a real user's would. Those
scorecards are kept under results/2026-09-16-host-confounded/.

Each rep now gets a fresh CLAUDE_CONFIG_DIR seeded with the three
fields Claude Code's first-run wizard writes (hasCompletedOnboarding,
lastOnboardingVersion, customApiKeyResponses.approved) plus workspace
trust, API-key auth from evals/.env, and a private tmux server per rep
so per-rep env vars reach the worker. The dev checkout is a detached
worktree (a branch already checked out elsewhere cannot be added
again), and its absence now fails loudly instead of launching a worker
with no plugin. wait requires 30s of sustained idle, since csd reports
idle briefly between a turn's stop and a background subagent's first
tool call.

Clean 3x2 results are under results/2026-09-16-clean/.
2026-09-16 12:55:18 -07:00
Jesse Vincent
8a8acb46d3 Add inline-eval: lightweight A/B harness for executing-plans
Drives real Claude Code sessions through claude-session-driver (csd)
against the sdd-tiny fixture with only the superpowers checkout under
test loaded: a --settings override disables every installed plugin and
--plugin-dir loads one copy, so the SessionStart hook, bootstrap, and
Skill tool are the real thing. Two arms: dev (baseline) and this
checkout.

Three environment problems the runner handles: Claude's workspace-trust
dialog for a never-seen directory (pre-recorded in ~/.claude.json), the
macOS keychain being unreachable from an old tmux server (private tmux
socket dir), and the ~104-char unix socket path limit (short socket
path). Workers that dispatch a background Agent end their turn early,
so wait loops until the worker is idle.

score.py reads the session transcript for ordered tool calls (Bash
heredoc writes count as writes, since bypass mode steers workers to
them), checks RED-before-GREEN per task, counts Agent dispatches and
ledger writes, runs the suite, and totals tokens across the main
session and its subagent transcripts.

Results of the first 3x2 run are in results/2026-09-16/.
2026-09-16 12:31:05 -07:00
Jesse Vincent
e26226be87 Add Claude Code platform reference for cheaper plan execution
Two documented Claude Code capabilities let SDD run cheaper without
changing what the skills require: nested subagents (three layers by
default per the sub-agents docs) allow the whole SDD loop to run one
layer down on a mid-tier orchestrator, and the Workflow tool (opt-in
by the user via ultracode) can carry the task loop as a script. Both
are opt-in by the human partner; the workflow mapping is flagged as
untested.

Spike: no baseline or pressure-scenario runs yet.
2026-09-16 12:08:30 -07:00
Jesse Vincent
3ceb303ebb Rebuild executing-plans as a first-class inline execution mode
Subagent-driven development costs a fresh implementer and a fresh
reviewer per task, and users have been vocal about the bill. The
writing-plans handoff already offered inline execution as option 2,
but executing-plans was a 64-line stub framed as a separate-session
mode whose own header told the agent to use SDD instead.

executing-plans is now a real mode: same worktree, plan workspace,
ledger, pre-flight scan, and stopping rules as SDD (so a plan can move
between executors mid-flight), the agent implements every task itself
under TDD with a per-task completion contract, and one fresh-context
whole-branch review at the end replaces the per-task reviewers. Cost
in context loads is roughly 2 versus 2N+2 for SDD.

The handoff prompt now states the cost/thoroughness tradeoff for both
options. SDD stays recommended; its when-to-use graph splits on the
partner's choice rather than on "stay in this session", since both
modes run in-session now. README lines describing executing-plans as
"batch execution with checkpoints" updated to match.

Spike: no baseline or pressure-scenario runs yet.
2026-09-16 12:08:30 -07:00
Drew Ritter
5940bd8d48
Merge pull request #2287 from obra/codex/pri-3127-doctor-evidence-handoff
fix: preserve diagnostic evidence in doctor exports
2026-09-11 15:21:16 -07:00
Drew Ritter
cf1f040986
Clarify scrub audit return contract
Remove the obsolete CLEAN branch beneath the audit prompt's Otherwise return instruction. CLEAN remains governed by the preceding no-misses condition, and MISSED is now the only alternative. Verified with the focused structural test and git diff --check.
2026-09-11 15:20:10 -07:00
Drew Ritter
ac22047ce8
Preserve diagnostic evidence through doctor export
Align scrubber and independent audit prompts around one shared redaction policy while preserving safe command, result, source, session-line, quotation, and linkage structure. Add finished-handoff evidence and reconciliation instructions, provenance labels across case/report/bundle/issue templates, and the structural existence check for the shared reference.\n\nThis patch responds to the retained negative post-report handoff baseline: cited result bodies were removed wholesale, source findings and the positive related-session match were not verifiable, provenance and export statements were stale, and scrub counts disagreed. The behavioral handoff validation remains pending for the follow-up task; this commit records only the focused product guidance and structural RED/GREEN evidence.
2026-09-11 15:20:10 -07:00
Drew Ritter
d4236278fb
Use shared session discovery in diagnosing-superpowers
At Drew's request, apply the evaluated shared-discovery variant to Jesse's
existing PR #2236. Resolve native session sources and record semantics from
available tools, documentation and bounded inspection. Record verified absolute
paths, linkage, extraction queries, human-message distinctions, usage-counter
semantics and uncertainty once in the case for all analysts to consume.
Replace the three per-harness references and update structural checks.

This is exactly the evaluated source tree at
3f0a63e860, applied as one commit on
801badbf71. Fourteen files change;
126 lines added, 285 removed. No private eval fixtures or transcripts ship.

Validation:
- Structural test: 45 passed, 0 failed before and after application.
- Staged tree exactly matches the evaluated candidate; diff check passes.
- Independent read-only review: no actionable blockers.
- Retained before/after full doctor runs: one pair each on native Claude,
  Codex and Pi. All six delivered reports and completed seven dimensions.
  Shared discovered all three native session families without the removed
  references. Both versions had report-quality defects; shared Codex deleted
  its cited case through a fixture symlink. Preserve this negative result.
- Eight fresh Codex follow-ups: original/shared x symlink/ordinary-home x
  two repeats, one retained historical session family. All eight retained
  cases and supported the four core findings. Seven native final deliveries;
  one shared run stopped on provider capacity after writing its report.
  No deletion recurred. One original reused three analysts for seven tasks.
  Recorded follow-up cost $34.8617883, all eight attempts accounted for.

These observations support this scoped simplification, not general equivalence
or a causal claim that reference removal caused or could not cause a failure.
Child assignment/model choices were native behavior; the complete variants
also differ in analyst prompts. Common provenance, citation-verification and
measurement problems remain separate follow-ups. No new paid runs were made
for this publication; evidence and independent audits are retained privately
by Drew. Behavioral evaluation provenance: campaigns
358c7333-c5f0-48bd-a733-61196da992ed and
102d630d-30ef-49a0-97e1-8dd410ed0548.

Prepared with GPT-6 using Codex through Paseo; local codex-cli 0.153.4.
Skills used: superpowers writing-skills, using-git-worktrees,
requesting-code-review, verification-before-completion; primeradiant-ops
linear-ticket-lifecycle. Drew approved publishing this evaluated variant.
Enabled plugins in the publishing checkout's Codex configuration:
- github@openai-curated
- documents@openai-primary-runtime
- spreadsheets@openai-primary-runtime
- presentations@openai-primary-runtime
- primeradiant-ops@primeradiant
- slack@openai-curated
- linear@openai-curated
- codex-security@openai-curated
- pdf@openai-primary-runtime
- template-creator@openai-primary-runtime
- sites@openai-bundled
- visualize@openai-bundled
- computer-use@openai-bundled
- cloud-build@superpowers-cloud-build
- browser@openai-bundled
- superpowers@superpowers-dev
- stream-deck-agents-codex@drew-local
- computer-history@openai-bundled
- codex-app-tools@openai-bundled
- unified-computer-use@openai-bundled
- chrome@openai-bundled
- bits-and-bolts@mcp-extensions-early-access
- visual-probe@visual-probe-local

Tracking: PRI-3127
2026-09-11 15:20:09 -07:00
Jesse Vincent
d3d9d2bec2
diagnosing-superpowers: use gh for issue search and creation
gh handles auth, rate limits, and JSON, and the approval gate on the
exact issue text already covers posting. Keep the public-API and
prefilled-link paths as fallbacks for machines without gh. Note that
GitHub drops labels from reporters without push access, so the template
footer is the durable marker of a skill-filed issue.
2026-09-11 15:20:09 -07:00
Jesse Vincent
be4263611e
diagnosing-superpowers: writing review fixes
Move GitHub search and prefilled-link mechanics to references/github-issues.md.
State the redaction levels neutrally instead of nudging toward more data.
Say that all seven analysts always run and what the quick-reference table
is for. Add a title slot and a bundle slot to the issue template. Drop the
duplicated human-prompts rule from request-conflicts. Prose fixes: active
voice, dangling modifier, vague referents, two lists turned into tables.
2026-09-11 15:20:09 -07:00
Jesse Vincent
06f0ed7bdb
spec: 'agreed to', not 'committed to', in the plan-adherence summary 2026-09-11 15:20:09 -07:00
Jesse Vincent
0a73dd8e0a
diagnosing-superpowers: say 'plan step', not 'commitment'
In a transcript full of git commits, 'commitment' and 'committed to' read
as version control. The plan-adherence and quality-evidence prompts now
say 'agreed plan' and 'plan step'.
2026-09-11 15:20:09 -07:00
Jesse Vincent
3dd5b621a1
diagnosing-superpowers: drop gh; file issues through a prefilled template link
A default gh login carries the repo scope, which is write access to every
repository the user can reach. The skill now searches issues through the
unauthenticated public API and, instead of posting, hands the partner a
prefilled new-issue link. The link uses a new diagnosis_report.md issue
template so the bug and automated-issue-report labels apply regardless of
the reporter's permissions. Addresses arittr's review on #2236.
2026-09-11 15:20:09 -07:00
Jesse Vincent
20c37ac909
diagnosing-superpowers: share the analyst preamble and context-safety rules
The seven analyst prompts opened with an identical 39-line block (role,
inputs, context safety, return format). It now lives once in
prompts/analyst-common.md and each dimension prompt points at it. The
wc -lc / long-line / never-cat rule was restated in nine places; it now
lives in references/context-safety.md and everything else points there.
Addresses arittr's review on #2236.
2026-09-11 15:20:09 -07:00
Jesse Vincent
9682098521
diagnosing-superpowers: build the scrubbed bundle only on request
Never build or push a bundle unprompted. When intake names a bug report as
the goal, say once that a bundle is available on request, then wait. On
handover, state what the bundle contains, point at the scrub log, and say
scrubbing can miss things so every file needs review before sharing.
Raise the SKILL.md word budget to 1000 to fit the added rule.
2026-09-11 15:20:09 -07:00
Jesse Vincent
0c59ebc9c6
docs: add 'When Something Goes Wrong' README section for diagnosing-superpowers 2026-09-11 15:20:09 -07:00
Jesse Vincent
fca1136291
feat: add diagnosing-superpowers skill
Evidence-based diagnosis of superpowers sessions: intake with the human
partner, safe transcript reading for Claude Code and Codex (discovery
procedure for other harnesses), seven analyst subagents, a report with
path:line evidence and a bounded superpowers-involvement line, scrubbed
export bundles, approval-gated GitHub issue search/draft, and
similar-session search. Includes spec, plan, structure test, and README
and docs index lines.

Developed RED-GREEN-REFACTOR per writing-skills: 46 scored scenario runs
across five SKILL.md versions, all twelve scenarios clean against the
final version, micro-tests control 5/5 to skill 0/5 on both
baseline-failing prohibitions, and one end-to-end run. Eval records are
kept by the maintainer outside the repo.

Claude-Session: https://claude.ai/code/session_01DyaGKhTXvHNs2JgPhDktz7
2026-09-11 15:20:09 -07:00
Drew Ritter
3a8bdc11e1
Merge pull request #2250 from obra/fix/platform-support-template-label
Fix platform-support issue template to apply a label that exists
2026-09-08 12:58:52 -07:00
Drew Ritter
929b017964
Merge pull request #2258 from obra/codex/brainstorming-intent-gates
fix: preserve shared intent through design and planning
2026-09-08 11:17:36 -07:00
Jesse Vincent
069edf3ffc fix: review the saved plan before execution
Present the saved, self-reviewed plan for human review before implementation.
Request an execution method when none was supplied; preserve an existing
choice and ask only for plan review when the human already chose a method.

This completes the shared-intent repair without interpreting approval of an
earlier idea or scope as approval of an unseen implementation plan. Four
saved-plan smoke cases covered old/new wording with/without a prior choice;
all passed the narrower handoff checks, including old controls, so this is
not evidence of measured improvement.

Jesse requested consolidation into two commits and removal of the supporting
spec/plan research content from the PR. The skill bytes remain identical to
the reviewed branch; the complete research and original history are retained
in local archives.
2026-09-04 16:03:59 -07:00
Jesse Vincent
3b4f2caf9e fix: establish shared intent before implementation
Discover the intended outcome, audience and success criteria before proposing
features when the request leaves them unclear. Reflect the understanding for
correction and carry it into the selected path's design artifact.

Bind approval to the actual stage presented: new architectural work requires
written-spec review and the planning handoff before implementation. Preserve
the existing lighter spike and bounded paths and clarify the short-design
example accordingly.

Jesse requested this repair after a React todo session advanced from feature
scope approval without establishing purpose. The controlled CLI comparison
observed purpose discovery in 5/5 candidate openings versus 0/5 controls, with
full-chain and holdout outcomes and their limits recorded in the PR. This
commit preserves the independently reviewed skill bytes; research artifacts
and the original development history are archived outside the PR.
2026-09-04 16:03:59 -07:00
Jesse Vincent
573a66b671 Fix platform-support issue template to apply a label that exists
The template auto-applies `platform-support`, but the repo has no such label
(harness requests use `new-harness`). GitHub silently drops labels that don't
exist, so every platform-support request arrives unlabeled — the Amazon Q
request (#2194) is the latest example.

Claude-Session: https://claude.ai/code/session_01UiEfXTZAC5cuH4hgx24mbB
2026-09-03 10:56:32 -07:00
Kattni
fd02874aa5
Update to Prime Radiant Community Code of Conduct. (#2122) 2026-08-12 15:31:42 -07:00
Jesse Vincent
777ceb5e12 Merge main back into dev after the v6.3.0 rebase-merge 2026-08-12 09:56:30 -07:00
Jesse Vincent
41cdb703de Merge main into dev: reconcile We're Hiring removal (44c9b2d) with dev's README rework
# Conflicts:
#	README.md
2026-08-12 05:08:27 +00:00
Jesse Vincent
d4e3c1cb8c chore: bump version to 6.3.0 2026-08-12 05:04:54 +00:00
Jesse Vincent
89d36fe961 docs: release notes for v6.3.0 2026-08-12 04:50:23 +00:00
Drew Ritter
034958f842 docs: keep Hermes in installation navigation
Add Hermes Agent to the installation entries in the table of contents. The removed Quickstart section was the README's only direct link to that existing installation section, so preserving the link avoids a navigation regression.
2026-08-07 17:37:17 -07:00
Drew Ritter
824aabcb21 docs: streamline README getting started navigation
Remove the redundant Quickstart entry and section now that the README has a table of contents. Rename the Installation label in the table of contents to Getting Started while retaining the existing installation anchor and section heading.
2026-08-07 17:37:17 -07:00
Drew Ritter
2d4b675b49
Merge pull request #1995 from caiolopes/add-devin-cli-support
feat: add Devin CLI support
2026-08-07 13:42:37 -07:00
Caio Lopes
d21e171f57 Drop devin-tools.md — not needed for correct operation
Re-ran the clean-session acceptance test with the mapping file and the
SKILL.md Platform Adaptation pointer removed: using-superpowers and
brainstorming still auto-trigger first, and the full workflow chain
(writing-plans, executing-plans, TDD, verification) resolves every action
to Devin's native tools. Devin CLI's own system prompt already documents
its tools (skill invocation, subagent profiles, todo tracking, question
prompts), so the mapping was redundant. Test now validates the manifest only.
2026-08-07 12:47:29 -07:00
Caio Lopes
09a567b6f4 feat: add Devin CLI support
Devin CLI's `devin plugins install obra/superpowers` fails today because the
repo has no `.devin-plugin/plugin.json` manifest. Add the manifest (skills are
auto-discovered from the co-located skills/ directory), a Devin tool mapping
linked from using-superpowers' Platform Adaptation section, a README install
section, version tracking in .version-bump.json, a Codex-sync exclude for the
new dotdir, and a CI-safe test mirroring the kimi/antigravity test style.

Bootstrap rides Devin's native skill surfacing: every installed skill's
name + description is injected into the system prompt at session start with a
standing instruction to invoke matching skills via the native skill tool.
Acceptance test ("Let's make a react todo list") passes in a clean session:
using-superpowers and brainstorming auto-trigger before any code is written.
2026-08-07 12:47:29 -07:00
Drew Ritter
d6a10aba55
Merge pull request #2006 from arimu1/fix/1929-copilot-cli-docs-windows
docs(brainstorming): correct Copilot CLI backgrounding guidance for Windows
2026-08-07 12:39:47 -07:00
Drew Ritter
8f89e512c3
Merge pull request #1919 from boredcity/docs/add-grok-build-cli-to-readme
Docs/add grok build cli to readme
2026-08-07 12:39:06 -07:00
Georgii Perepechko
28125bf284 docs: add Grok Build CLI to README.md 2026-08-07 12:33:35 -07:00
Drew Ritter
c367f804bb
Merge pull request #2063 from obra/fix/t4-brainstorming-three-paths
feat(brainstorming): three-path router — ceremony scales, approval never does
2026-08-06 22:20:58 -07:00
Drew Ritter
5f8f500b1d fix(release): wire Hermes into version bumps
Register the Hermes YAML manifest alongside the existing JSON manifests. Route manifest reads and writes by extension through jq or Mike Farah yq v4, with field names and values passed as data.

Preflight every present manifest before the mutating bump loop so a deterministic YAML read failure cannot leave earlier JSON manifests partially updated. Cover check, audit, bump, registry wiring, and byte-for-byte no-partial-write behavior with one focused fixture test.
2026-08-06 16:21:12 -07:00
Drew Ritter
707b155a38 docs: plan Hermes version-bump wiring
Record Drew's approved reduced design after the second staff review. Limit preflight to the mutating bump path, cover audit's independent read path, and require byte-for-byte proof that deterministic YAML failures cannot partially update earlier JSON manifests.

Provide one TDD implementation task for the Hermes registry entry, jq/yq dispatch, focused preflight, and three behavioral checks. Explicitly defer rollback, audit-status changes, nested YAML, runtime changes, and broader release-tool refactoring.
2026-08-06 16:21:12 -07:00
Drew Ritter
3e1ecde38f docs: reduce Hermes version-bump design
Incorporate the adversarial design review without turning the Hermes wiring follow-up into a general release-script refactor. Keep the existing jq path, add Mike Farah yq v4 only for .yaml, and retain one read-only preflight to prevent deterministic partial bumps.\n\nReduce the test contract to three behavioral cases and explicitly defer .yml support, nested YAML, rollback machinery, audit/status redesign, exhaustive failure matrices, and the separately discovered JSON-expression issue. This follows Drew's direction to avoid ceremony and overengineering.
2026-08-06 16:21:12 -07:00
Drew Ritter
ffe22811bf docs: design Hermes version-bump wiring
Document the agreed follow-up to PR #2025 on a branch based on its merged dev commit. The design registers the Hermes YAML manifest, keeps jq for existing JSON files, and uses Mike Farah yq v4 for a narrow top-level YAML field rather than adding a Bash parser.\n\nDefine focused failure behavior and behavioral tests while explicitly excluding nested YAML, Hermes runtime changes, and unrelated release-script refactors. This captures Drew's request to keep the implementation small and avoid process or abstraction overhead.
2026-08-06 16:21:12 -07:00
Drew Ritter
cfb310c69a
Merge pull request #2089 from obra/fix/x13-illegibility
fix(sdd): reviewers re-read illegible evidence instead of re-running to regenerate it
2026-08-06 16:10:58 -07:00
Drew Ritter
fdd1763d77
Merge pull request #2086 from obra/fix/spec-travels-with-plan
fix(planning): the spec travels with the plan
2026-08-06 15:13:54 -07:00
Drew Ritter
af4bebf762
Merge pull request #2024 from obra/fix/worktree-cleanup-untracked-checkin
fix(finishing): check in with human partner when worktree removal hits untracked files
2026-08-06 12:21:45 -07:00
Drew Ritter
17b42c8128 fix(finishing): name the actual files in the refusal prompt
`git status --porcelain` collapses a wholly-untracked directory to a single
`?? docs/` line. In the shape of the incident this step exists for (#2016 — an
uncommitted plan document under an untracked `docs/` tree), the file list we
show the human partner therefore names no file at all:

    $ git -C "$WORKTREE_PATH" status --porcelain
    ?? docs/
    $ git -C "$WORKTREE_PATH" status --porcelain -uall
    ?? docs/superpowers/plans/2026-08-04-csv-export-rollout.md

Both forms produce identical (empty) output on a clean worktree, so this adds
no over-trigger surface.

Found while running this PR's behavioral micro-tests. Every treatment agent
dug past `?? docs/` unprompted and named the document, so the step did work —
but on the agent's own initiative rather than because the text asked for it.
That initiative is not reliable one tier down: Claude Haiku 4.5 on the control
arm failed for exactly this shape, asking a question that never named the file
and then deciding for the human when they deferred. Nothing in the prior
wording stopped a treatment agent from relaying `?? docs/` verbatim and
satisfying the letter of the instruction.

Re-ran the treatment cells against this amended text — Opus pass (refusal
fired, named the file), Haiku 4.5 pass (named the file) — no regression.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 12:14:55 -07:00
Drew Ritter
1245282b05
Merge pull request #1805 from obra/fix/render-graphs-no-shell
fix(writing-skills): make render-graphs ESM-compatible and shell-free
2026-08-06 11:50:32 -07:00
Drew Ritter
6819b42d97
Merge pull request #2064 from obra/fix/docs-codex-efficiency-campaign
docs: codex-efficiency fix-cycle spec and plan (campaign record)
2026-08-05 23:17:21 -07:00
Drew Ritter
02654f93bf test(writing-skills): cover render-graphs execution 2026-08-05 22:38:40 -07:00
Jesse Vincent
dcd3661b7c fix(writing-skills): run graphviz without a shell in render-graphs.js
The `dot` availability check shelled out to `which dot`, which is not a
command on Windows, so render-graphs.js reported graphviz as missing on
Windows even when it was installed. Replace it with a direct `dot -V`
probe via execFileSync.

Also switch the SVG render call from execSync to execFileSync('dot',
['-Tsvg']). Behavior is identical on macOS/Linux — the diagram source
was already passed via stdin, never interpolated into the command — but
running the binary directly removes the shell entirely.
2026-08-05 22:38:40 -07:00
Drew Ritter
9be44ebf40
Merge pull request #2025 from obra/hermes-harness-rebase
feat(hermes): Hermes Agent harness support — eval-verified pre_llm_call bootstrap
2026-08-05 18:23:46 -07:00
Drew Ritter
695744056e chore(hermes): align plugin version with dev
Update the Hermes plugin manifest from 6.1.1 to 6.2.0 so PR #2025 matches the current release version at the tip of origin/dev.\n\nThis intentionally does not change the version bump tooling. The existing release script supports JSON manifests only; YAML support will be handled separately on its own branch.
2026-08-05 17:57:31 -07:00
Kattni
fb518edf7b Moves Community up, and adds ToC. 2026-08-04 20:16:46 -07:00
Jesse Vincent
80b82abd8d fix(sdd): task reviewers re-read illegible evidence instead of re-running to regenerate it
Interrogation of reviewers who bypassed test-evidence leases showed a
convergent driver: when the report or receipt looked truncated or
couldn't be located, re-running the suite felt cheaper than re-reading —
evidence got regenerated instead of read. This paragraph names that
moment: re-read at the stated path, report a genuine gap to the
controller, and never re-run to regenerate what wasn't read.

Battery: 0/31 reviewer re-runs across 4 treatment reps vs 7/~59
reviewers in 5/8 control reps on the same scenario and classifier.

Claude-Session: https://claude.ai/code/session_0185AJr98gHx5EmwqNeft4Sy
2026-08-04 18:24:10 -07:00
Drew Ritter
05c2393b82
Merge pull request #2078 from obra/fix/x6a-sdd-batch-small-tasks
fix(sdd): batch small same-shape tasks into one dispatch
2026-08-04 14:26:31 -07:00
Drew Ritter
78cc189244 fix(sdd): batch reviews check the diff against the brief's file list
Batching moves N edits under one review, which changes the review's
failure profile: an implementer that silently skips one file of twelve
produces a diff full of correct, uniform edits — nothing conspicuous is
missing, and no seat in the pipeline was assigned to notice. The single
combined review is the only net for a dropped edit, but the reviewer
template never told it to count.

The batch brief already lists every file with its change, so the reviewer
reconciles the diff against that list file by file; a listed file with no
hunk is a Missing finding regardless of how clean the rest of the batch
looks. Conditional on a multi-file brief, so single-task reviews are
unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-04 14:25:40 -07:00
Drew Ritter
be76350536
Merge pull request #2080 from obra/fix/x7a-sdd-evidence-bearing-preflight
fix(sdd): preflight emits its checks as a ledger table and rules on what it surfaces
2026-08-04 14:22:57 -07:00
Drew Ritter
419dec7755 Merge dev into fix/x7a: resolve preflight paragraph with the composed 2077+2080 text
Both PRs rewrote the same preflight paragraph. Resolution is the composed
text published in #2080's description — the configuration the 3/3+3/3
composed eval grades ran.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-04 14:21:58 -07:00
Drew Ritter
2b195749df
Merge pull request #2077 from obra/fix/x9a-sdd-never-stall
fix(sdd): rule and continue — non-catastrophic conflicts get ledgered rulings, not blocking questions
2026-08-04 14:16:16 -07:00
Drew Ritter
7a01a0e83a fix(sdd): one Ruling: token everywhere, exhaustive finish roll-up
The breaker's two ledger formats wrote lowercase 'ruling' (parked
findings, load-bearing adjudications), so the Finish section's
collect-every-`Ruling:`-line step missed exactly the rulings made under
the most pressure. Field evidence from an independent eval rep: a
breaker-cap run adjudicated correctly, wrote everything to the
plan-scoped ledger, deleted the workspace at finish, and left no durable
trace of the adjudication.

Capitalize the two breaker formats to the canonical token, and make the
finish roll-up explicitly exhaustive across preflight, parked, and
breaker rulings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-04 14:13:42 -07:00
Drew Ritter
8acf8e5f24
Merge pull request #2059 from obra/fix/t1-sdd-no-worker-reviewers
fix(sdd): dispatched subagents never dispatch subagents
2026-08-04 14:05:55 -07:00
Drew Ritter
50a924b0c4
Merge pull request #2062 from obra/fix/t5-codex-spawn-routing
fix(codex): explicit model+effort on every spawn, with config backstop
2026-08-04 14:04:31 -07:00
Drew Ritter
2a977c7095
Merge pull request #2061 from obra/fix/t2-codex-event-waits
fix(sdd,codex): event-driven bounded waits — 65-78% wait timeouts to 0%
2026-08-04 14:03:30 -07:00
Drew Ritter
50f787ca5c
Merge pull request #2060 from obra/fix/t3-codex-tools-corrections
fix(codex): correct multi-agent guidance against the Codex source (V2)
2026-08-04 14:00:31 -07:00
Jesse Vincent
538d65120b fix(planning): the spec travels with the plan — Spec: header pointer + SDD reads it at setup
In controlled evals, an identical seeded-incoherence plan yielded 0-1/5
correct conflict resolutions when executed specless (controllers ruled
the conflicts 'internally explained') and 4-5/5 with the spec merely
present and named — even with no other skill-text changes. Cross-task
coherence turns out to be adjudicable only against ground truth above
the plan; this change makes that ground truth travel with the plan.

Claude-Session: https://claude.ai/code/session_0185AJr98gHx5EmwqNeft4Sy
2026-08-04 11:17:05 -07:00
Jesse Vincent
61f669ebc9 fix(sdd): preflight emits its pairwise checks as a ledger table and rules on what it surfaces
The pre-Task-1 conflict scan currently permits 'the scan is clean' with
no evidence the scan happened — mined sessions show controllers skipping
straight to dispatch and plan conflicts surfacing mid-execution as
blocking questions. Requiring the scan to emit one row per task pair
sharing a file/interface and one row per task's self-consistency turns
the claim into an artifact; in controlled evals the table appeared 3/3
with conflicts surfaced pre-dispatch, and the mechanism held 3/3 when
composed with the never-stall ruling change (#2077).

Claude-Session: https://claude.ai/code/session_0185AJr98gHx5EmwqNeft4Sy
2026-08-03 09:03:05 -07:00
Jesse Vincent
e7a4285985 fix(sdd): batch small same-shape tasks into one dispatch
Plans sometimes enumerate many tiny, same-shape edits (one-line fixes,
constant changes, a field added across files) as separate tasks. The
current loop dispatches a fresh implementer plus review per task, so a
12-micro-task plan costs ~24 subagent seats for what one subagent could
do in a single pass. In controlled evals on a micro-task plan, batching
cut cost 73% and dispatches 87% with better completion than control; on
a 5-non-trivial-task plan the rule correctly never batched (dispatch
counts and completion identical to control).

Claude-Session: https://claude.ai/code/session_0185AJr98gHx5EmwqNeft4Sy
2026-08-02 19:36:55 -07:00
Jesse Vincent
39f9602432 fix(sdd): rule and continue — non-catastrophic conflicts get ledgered rulings, not blocking questions
A donated session sat dormant 8h48m waiting for a plan-conflict answer
that cost ~zero tokens to decide. Wrong-ruling rework is bounded;
stalls are not. This encodes the never-stall doctrine: plan conflicts,
ambiguities, and cap exceptions get a controller ruling recorded in
the ledger and work proceeds; only irreversible/destructive actions,
security-sensitive actions, out-of-worktree side effects (merge/push/
publish), and totally-broken plans remain hard stops. Rulings surface
in the Finish report instead of as mid-run questions.

Evals: 3/3 no-stall vs control 3/3 stall-at-preflight on a
seeded-conflict SDD plan; catastrophic guard 5/5 (every rep reaching a
seeded DROP TABLE step refused it); re-validated 3/3 after rebase onto
the current fix-PR text; composes cleanly with the evidence-bearing
preflight treatment.

Claude-Session: https://claude.ai/code/session_0185AJr98gHx5EmwqNeft4Sy
2026-08-02 11:14:01 -07:00
Jesse Vincent
3ff8d15f15 docs: codex-efficiency fix-cycle spec and plan (campaign record) 2026-07-31 10:52:05 -07:00
Jesse Vincent
e9686d5c09 fix(codex): explicit model+effort on every spawn, config backstop
Depth-2 child-issued spawns omitted model 2/2 at CLI 0.146; model
without reasoning_effort resets effort to the model default.
2026-07-31 10:27:51 -07:00
Jesse Vincent
d8189d1587 fix(sdd,codex): bounded wait stretches with reconciliation
Round 2 proved the long-wait mechanism (65.1%->0.0% timeouts) but
20-38 min silent waits starved graders and let 1/51 children vanish;
bounded 5-10 min stretches with a status line and list_agents
reconcile keep the efficiency and restore observability.
2026-07-31 10:27:51 -07:00
Jesse Vincent
db4538fcb8 fix(sdd): controllers wait long or not at all
Docs-only wait guidance in the platform reference changed nothing
(65.1% vs 67.1% baseline wait-timeout rate); the discipline now lives
in the controller loop the session actually re-reads.
2026-07-31 10:27:50 -07:00
Jesse Vincent
9b8b14fe12 fix(codex): event-driven waiting instead of short polls
60-78% of wait_agent calls timed out across every measured corpus;
waits are event subscriptions, so one long wait replaces dozens of
polls at identical wake latency.
2026-07-31 10:27:50 -07:00
Jesse Vincent
4dc71b10b3 fix(brainstorming): bounded means existing code in this repo, not a familiar app genre
Triggering battery: Claude Code classified a brand-new project bounded
3/3 by reading 'existing, understood flow' as genre familiarity — once
while explicitly noting the repo was empty. Gemini routed the same
prompt architectural 3/3.
2026-07-31 09:39:55 -07:00
Jesse Vincent
7c560e048b fix(sdd): reviewers never dispatch subagents either
The first fix-cycle battery moved the depth-2 leak from implementers
(9/9 baseline -> 0/6) to a final reviewer that spawned two
sub-reviewers; the contract now reaches every dispatched role.
2026-07-31 09:39:55 -07:00
Jesse Vincent
75756d2900 fix(codex): correct multi-agent guidance against Codex source
Five claims contradicted by the Codex CLI source (V2 has no
close_agent; followup_task always reaches a child; role files attach
via agent_type; full-history forks accept model/effort; V2 spawn
allowlist). Citations: superpowers-autoresearch
docs/2026-07-29-codex-multiagent-v2-capabilities.md.
2026-07-31 09:39:55 -07:00
Jesse Vincent
b68eaf96bb fix(brainstorming): bounded-path approval is a hard stop
Live ceremony battery: bounded reps produced zero doc ritual (the
measured win) but 2/3 implemented before any approval turn; the
bounded path now states the stop explicitly.
2026-07-31 09:39:55 -07:00
Jesse Vincent
2e7d681591 fix(sdd): implementers never dispatch subagents
Depth-2 worker-spawned reviewers were 9/9 same-task duplicate reviews
across four corpora in the codex-efficiency eval campaign.
2026-07-31 09:39:55 -07:00
Jesse Vincent
6211388f4b feat(brainstorming): three-path router — ceremony scales, approval never does
Spike / bounded / architectural classification said out loud, one-way
upgrade ratchet, approval gate on every path. The measured pathology:
the absolute hard-gate wording forced bounded tasks into the full
two-document ritual 5/5 while a no-guidance control differentiated
paths natively.
2026-07-31 09:39:55 -07:00
Jesse Vincent
bb2a34b2a0 docs: remove the "We're Hiring" section from the README
The community engineer role has a candidate on trial, so the posting no
longer needs to be at the top of the README.
2026-07-27 11:43:14 -07:00
Jesse Vincent
7b4dc4d7fd Merge main back into dev after the v6.2.0 rebase-merge
Converges dev with the rebased release SHAs on main immediately, while
the trees are identical, so the merge is conflict-free and the next
dev -> main release shows only genuinely-new work.
2026-07-23 17:28:00 -07:00
Jesse Vincent
5b73c0f63a Merge main back into dev: converge histories after the 6.1.1 rebase-merge
All 11 conflicts are residue of the v6.1.1 rebase-merge, which replayed
dev's commits onto main as new SHAs: version manifests (dev 6.2.0 vs
main 6.1.1), RELEASE-NOTES.md, the porting-guide table (main predates
the #1969 dead-reference fix), and add/add on the Codex package script
and test (dev carries the later portability fixes). Every conflict
resolves to dev's side; the merged tree is identical to dev's.
2026-07-23 16:18:07 -07:00
Jesse Vincent
d262bc400c
Release v6.2.0: SDD plan-scoped workspace and resume-based fix loop, skills compression sweep, Windows SessionStart fix (#2026)
Release notes for everything on dev since v6.1.1, plus the version bump
to 6.2.0 across all seven declared manifest files (bump-version.sh,
audit clean). Tagging and marketplace publication happen after the
dev -> main merge.
2026-07-23 16:16:32 -07:00
Jesse Vincent
b6613057ae test(hermes): realign suite with the pre_llm_call mechanism; slim docs to the README section
The 20-test suite still exercised the dead on_session_start/inject_message
mechanism (17 failures against the rewritten plugin). Rewritten for the
real contract: pre_llm_call registration + first-turn-only context return,
register_skill receiving pathlib.Path (the conftest mock now raises on str,
mirroring hermes' AttributeError that silently disables a plugin), both
install layouts resolving skills, loud failure when skills are missing,
tool mapping sourced verbatim from hermes-tools.md, and a bootstrap-size
guard against hermes' 10k-char context spill threshold. 19 tests, passing.

Install docs collapse into the README section per maintainer direction:
docs/README.hermes.md and .hermes-plugin/INSTALL.md are gone; the README
carries the two-line install plus the compaction caveat. plugin.yaml
version aligned to 6.1.1.
2026-07-23 16:05:54 -07:00
Jesse Vincent
178528c03e fix(hermes): working bootstrap injection via pre_llm_call + native skill registration
Empirical findings from the quorum eval bring-up (superpowers-evals
docs/experiments/2026-07-23-hermes-target-bringup.md):

- ctx.inject_message exists but returns False when called from
  on_session_start — nothing reaches the model. The documented path,
  a pre_llm_call hook returning {"context": ...} on is_first_turn,
  verifiably delivers (probe model echoed an injected codeword).
- ctx.register_skill requires a pathlib.Path; passing a str raises
  AttributeError inside hermes, which silently disables the entire
  plugin (no log line anywhere). This also means any exception in
  register() is invisible — keep register() failure-proof.
- Registered skills are namespaced by plugin name: models invoke
  skill_view("superpowers:brainstorming") and receive the stock
  SKILL.md — verified live on GLM 5.2, both install layouts.

The plugin now: resolves skills/ for both the git-clone layout
(.hermes-plugin/ and skills/ as siblings) and a flattened install,
raising loudly when neither matches; registers every stock skill with
Hermes' native loader (no per-harness skill copies); injects the
using-superpowers bootstrap via pre_llm_call on the first turn; and
sources the tool mapping from references/hermes-tools.md instead of
duplicating it. Injected context is transient (API-call time only, never
persisted in the session export) — verification of injection must be
behavioral.
2026-07-23 15:17:55 -07:00
Jesse Vincent
7b177613c0 feat(hermes): Hermes Agent harness support, rebased to a Hermes-only diff
Rebase of PR #1922 onto current dev: the ~14 files of v6.1.0-era
codex/release drift are dropped, the porting-guide edits (stale against
the post-prune rewrite, no Hermes content) are dropped, and the Hermes
surface is kept intact: .hermes-plugin/ (on_session_start bootstrap
injection), tests/hermes/ (20 tests, passing), docs/README.hermes.md,
references/hermes-tools.md, the Platform Adaptation row, README section,
and Python ignores.

Known open items from review, unchanged by this rebase: the injection
mechanism uses ctx.inject_message from on_session_start, which the
official plugin guide does not document (pre_llm_call returning
{"context": ...} is the sanctioned path), skills are not registered via
ctx.register_skill, and the acceptance transcript predates the fix.

Co-authored-by: kumarabd <kumarabd@users.noreply.github.com>
2026-07-23 12:18:13 -07:00
Jesse Vincent
1f0e2ab912 fix(finishing): check in with human partner when worktree removal hits untracked files
git worktree remove refuses when the tree holds modified or untracked
files, and the skill gave no guidance for that refusal — the natural
agent response was --force, permanently destroying files that exist
nowhere else (uncommitted plans, notes, scratch work). Reported twice
from real sessions (#2016's plan loss, #1223's dirty-tree ambiguity).

Step 6 now treats the refusal as a stop-and-ask moment: show the
untracked files, offer commit / relocate / delete, and only remove the
worktree after the human partner chooses. Adds a matching rationalization
row so --force-as-cleanup is named as the failure it is.
2026-07-23 11:55:24 -07:00
Jesse Vincent
0146173544 fix(systematic-debugging): find-polluter accepts ./-prefixed patterns and matches top-level tests
Follow-up to #2011 (which fixed the ./-prefix mismatch for the documented
pattern form): strip a leading ./ from the caller's pattern instead of
double-prefixing it into a never-matching ././ form, and also match the
pattern with '**/' collapsed, since find -path cannot match '**/' against
zero directory levels and silently skipped files directly under the base
directory (src/top.test.ts vs src/**/*.test.ts).

Adds a deterministic test suite for the script with a stubbed npm.
2026-07-23 10:54:50 -07:00
dev_Hakaze
54d0efefd7
fix(systematic-debugging): match find -path ./ prefix in find-polluter.sh (#2011)
find . emits ./-prefixed paths, so -path "src/**/*.test.ts" matched
nothing; wc -l on empty stdin then lied as "Found 1". Fixes #2008.

Co-authored-by: arimu1 <19286898+arimu1@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-23 10:53:39 -07:00
Mark Rada
55d28ddf10
docs(using-superpowers): drop dangling subagent-support anchor (#2010)
The prune in e7ddc25e removed the `## Subagent support` section from
antigravity-tools.md but left the inline cross-reference to it in the
dispatch table, so `[Subagent support](#subagent-support)` resolves to
nothing. An agent following the pointer to learn the difference between
the `self` and `research` subagent types lands nowhere.

Drop the dangling parenthetical. The guidance it pointed at survives in
the same table cell -- `self` for full-capability work, `research` for
read-only -- so no content is lost and the row still answers the
question the removed section answered.

gemini-tools.md carries the same cross-reference but retains its
`## Subagent support` heading, so its link is valid and is left alone.
2026-07-23 10:47:44 -07:00
Jesse Vincent
cc690476fc feat(sdd): lifecycle restructure with resume-based fix loop, five-round breaker, and rationalization table 2026-07-19 12:36:33 -07:00
Jesse Vincent
7ce7620d44 feat(sdd): align templates and codex reference with resume-based fix rounds 2026-07-19 12:36:33 -07:00
Jesse Vincent
f428cba185 feat(sdd): add scoped re-review prompt template 2026-07-19 12:36:33 -07:00
Jesse Vincent
eb1ff1f11f docs(plans): SDD fix-loop redesign implementation plan
Eight tasks across two repos: new re-review template, template/reference
alignment, full SKILL.md lifecycle restructure with move map, two
seeded-ledger fixture helpers, three quorum scenarios, and the RED/GREEN/
regression live-run campaign.
2026-07-19 12:36:33 -07:00
Jesse Vincent
bea92dce1a docs(specs): SDD fix-loop redesign design spec
Review-fix loop gets resume-the-implementer semantics, scoped
re-reviews, a five-round circuit breaker, and controller adjudication
at trip. SKILL.md reorganizes by lifecycle; Red Flags converts to a
rationalization table. Brainstormed with Jesse 2026-07-15.
2026-07-19 12:36:33 -07:00
Jesse Vincent
0634449ca6 fix(tests): stop the SDD skill test flaking on timing and prose case
tests/claude-code/test-subagent-driven-development.sh failed
intermittently for two independent reasons:

- Budget mismatch: the file runs 9 prompts with a 90s timeout each
  (810s worst case) inside the runner's 600s per-file ceiling, so slow
  backend days produced spurious timeouts. Raise the runner default to
  900s and fix the help text, which claimed the default was 300.
- Case-sensitive prose matching: the assert helpers grepped free-form
  model output case-sensitively, but models capitalize the skill's own
  headings — observed failures include "Do Not Trust the Report"
  missing pattern "not trust" and a structured answer missing
  "First:.*spec.*compliance". Match case-insensitively in
  assert_contains/assert_not_contains/assert_count/assert_order, widen
  two Test 5 keyword patterns to phrasings observed in real runs, and
  make assert_order dump the output on failure the way assert_contains
  already does, so the next flake is diagnosable.

Observed 3 failures across 4 runs before the change (timeout, two
distinct pattern misses); 3/3 consecutive full runs pass after it.
2026-07-19 12:04:46 -07:00
Jesse Vincent
3fe3cb0530 fix(codex): make package script and its test portable beyond macOS/bsdtar
The packaging pipeline only worked on a Mac with default umask, for
three stacked reasons:

- The deterministic-metadata tar flags (--uid/--gid/--uname/--gname)
  are bsdtar spellings; GNU tar rejects them, so the tar.gz archive
  step died on Linux. Detect the tar flavor and use --owner=:0
  --group=:0 --numeric-owner on GNU tar, which writes byte-identical
  ustar headers (uid/gid 0, empty uname/gname).
- Staged file modes depended on two umasks canceling out: git archive
  masks entry modes with tar.umask (git default 0002 -> 775), and the
  unflagged tar extraction re-masked with the process umask (022 on
  macOS -> 755, but 002 elsewhere -> 775). Pin tar.umask=0022 on the
  archive call and extract with -p so staged modes are canonical
  755/644 on every machine.
- The test's timestamp assertion parsed bsdtar's -tv column layout and
  expected epoch 0 rendered in a US timezone ("Dec 31 1969"); GNU tar
  uses different columns and UTC hosts render "1970-01-01". Assert
  mtime == 0 via python3 tarfile instead, matching how the test
  already checks zip timestamps.

tests/codex/test-package-codex-plugin.sh now passes on Linux/GNU tar;
the bsdtar branch preserves the exact flags that passed on macOS.
2026-07-19 12:04:46 -07:00
Jesse Vincent
fe0b24390e docs(windows): document shell:bash hook dispatch and the PowerShell/CMD fallback hazards 2026-07-19 12:03:59 -07:00
Jesse Vincent
df78c6bfaf fix(hooks): dispatch the SessionStart hook via Git Bash on Windows
The SessionStart command string starts with a quoted path, which breaks
both Windows shells Claude Code may hand it to: PowerShell parses the
leading quoted string as an expression and dies on the next bareword
('Unexpected token session-start', #1751), and cmd.exe's /c quote rule
drops the outer quotes when the path contains a metacharacter, so a
profile dir like C:\Users\Name(External) truncates the command at the
'(' (#1918). Either way the bootstrap silently never loads.

Declare shell: "bash" on the hook. Claude Code >= 2.1.81 then resolves
Git for Windows and runs the polyglot's bash path directly — the same
route it already picks when it detects Git Bash — and when Git Bash is
missing it surfaces an actionable install prompt instead of a parser
error. Older versions ignore the unknown key and behave exactly as
before (verified live on 2.0.77 and 2.1.80).

Verified end-to-end with real claude sessions: Linux (hook fires,
bootstrap injected), Windows 11 + Git Bash under a path containing
'(' and a space (fires, 3276-char context), and Windows 11 without
Git Bash (actionable error replaces the #1751 ParserError, reproduced
verbatim as control).

Fixes #1751
Fixes #1918
2026-07-19 12:03:59 -07:00
Jesse Vincent
30ff376cb6 chore(sdd): consistency sweep for plan-scoped workspace signatures 2026-07-19 12:03:18 -07:00
Jesse Vincent
75f4e9414e eval(sdd): GREEN results — plan-scoped resolution replaces cross-plan forensics 2026-07-19 12:03:18 -07:00
Jesse Vincent
c15e041e03 feat(sdd): plan-scoped durable progress — ledger names its plan, workspace dies at plan end
The start-of-skill ledger check is now scoped to the plan's own
workspace and keyed to the ledger's first line. Baseline eval (25/25
reps) showed controllers already refuse foreign ledgers — at a cost of
6-13 tool calls of cross-plan forensics per resume; plan-scoping makes
the answer structural instead. The workspace is deleted once the final
review is clean — git history is the durable record.
2026-07-19 12:03:18 -07:00
Jesse Vincent
9816a9cee2 feat(sdd): plan-scoped workspace — one .superpowers/sdd/<plan> dir per plan
sdd-workspace now requires the plan file and resolves
.superpowers/sdd/<plan-basename>/; task-brief and review-package write
into their plan's directory (review-package gains PLAN_FILE as its first
argument). Follow-up plans in the same working tree can no longer collide
with a previous plan's briefs, reports, or ledger.
2026-07-19 12:03:18 -07:00
Jesse Vincent
9d9eae52f9 eval(sdd): RED baseline — 25/25 controllers refuse stale ledgers, at a forensic cost 2026-07-19 12:03:18 -07:00
Jesse Vincent
6ddb0bfcd9 docs(specs): record eval re-scope — blind adoption did not reproduce, claims narrowed
25/25 baseline reps refused the stale foreign ledger via git forensics;
the spec's evaluation section now states the honest claims: structural
fix + measured disambiguation-cost delta + same-plan-resume regression
gate, shipping with explicit maintainer sign-off in place of a failing
S1 baseline.
2026-07-19 12:03:18 -07:00
Jesse Vincent
194907435d docs(plans): re-scope eval per maintainer decision — RED compiled, GREEN measures cost
Three RED rounds (25 reps, three framings incl. faithful compaction
resume) never reproduced blind stale-ledger adoption: sonnet controllers
forensically refuse foreign ledgers, spending 6-13 tool calls per resume
doing it. Jesse approved shipping the full change with the eval re-scoped
to what is true: Task 1 compiles the existing RED evidence, Task 4 runs
GREEN on a truthful v3 fixture (real implementations, rotating authors)
with an S2 released-text control, measuring regression safety and the
disambiguation-cost delta instead of an error rate.
2026-07-19 12:03:18 -07:00
Jesse Vincent
c10431b14c docs(plans): fixture v2 — real cited commits, matched task counts
Fixture v1 tripped the Task 1 STOP gate for the right reason: its
ledgers cited fabricated hashes, so RED agents dismissed them via git
forensics (S1 passed for the wrong mechanism, the S2 resume control
failed 5/5). v2 executes plan A's tasks as real commits, gives both
plans five tasks so numbering is ambiguous, adds a symmetric
resume-uncertainty line to the scenario prompt, hard-stops if the S2
control fails twice, and drops rm -rf from cleanup (hook-gated here).
2026-07-19 12:03:18 -07:00
Jesse Vincent
0da87665c8 docs(plans): SDD plan-scoped workspace implementation plan
Five tasks: RED baseline eval (writing-skills Iron Law — before any
skill edit), plan-scoped scripts via TDD, SKILL.md durable-progress
rewrite with mismatch guard and end-of-plan cleanup, GREEN eval with
refinement loop, consistency sweep. Eval = 5 fresh sonnet subagents per
scenario per arm, hand-scored.
2026-07-19 12:03:18 -07:00
Jesse Vincent
20940deae8 docs(specs): SDD plan-scoped workspace design
The .superpowers/sdd workspace has no plan identity and no end-of-life:
follow-up plans in the same worktree read the previous plan's ledger as
their own progress, and artifacts leak into git (observed in serf, three
contamination rounds and ad-hoc progress-p2/p3 workarounds). Structural
fix: per-plan workspace subdirs, ledger names its plan, delete the
workspace when the final review is clean.
2026-07-19 12:03:18 -07:00
arimu1
262ed02103 docs(brainstorming): correct Copilot CLI backgrounding guidance for Windows 2026-07-19 06:15:54 +07:00
Jesse Vincent
fb7b07088e docs: fix dead references to pruned claude-code-tools.md/copilot-tools.md
e7ddc25 deleted claude-code-tools.md and copilot-tools.md but left
writing-skills and the porting guide's reference-integration table
pointing at them. State the current architecture instead: Claude Code's
personal-skills path inline, and "no adapter file needed" for the
harnesses that ride the Claude Code-compatible tool surface.

Reported by @rasibintang (#1969, with a fix proposed in #1970).

Fixes #1969
2026-07-15 19:15:16 +00:00
Gaurav Dubey
7a81eb7177 test(pi): scope mapping assertions to the table, not whole file
The pi tokens (subagent, pi-subagents, Task, TODO.md) also appear in the
surrounding prose, so matching the whole file passed even with the mapping
table deleted — the exact regression this test exists to catch. Filter to
table rows (lines starting with '|') so the assertion fails when the table
is gone and passes on dev.

Reported by @muunkky on #1987 (approach from #1983); verified failing-first
by stripping the table rows from pi-tools.md.
2026-07-15 11:10:55 -07:00
Gaurav Dubey
2b1c06a849 test: realign antigravity + pi mapping assertions with pruned references
Commit e7ddc25 ('Prune per-harness tool-mapping boilerplate') deliberately
removed the skill-loading explainers and generic action->tool tables from
antigravity-tools.md and pi-tools.md, keeping only the harness-specific
notes (subagent dispatch, task tracking). It did not touch tests/, so two
content-assertion tests kept asserting the removed tokens and now fail on
both dev and main:

  - tests/antigravity/test-antigravity-tools.sh: asserted view_file,
    IsSkillFile, run_command, grep_search (all pruned)
  - tests/pi/test-pi-extension.mjs: asserted read/write/edit/bash (pruned)

Update both to assert only the surviving harness-specific mappings. No
reference or skill content is changed; only the stale test assertions.
2026-07-15 11:10:55 -07:00
Jesse Vincent
4562d18dcf refactor(skills): fold TDD Why Order Matters rebuttals into rationalization table
The eval verdict on this cut: deleting Why Order Matters and trusting the
compressed one-line table rows measurably degrades test-first behavior under
the exact pressure the section rebutted ("just write it, tests after") —
control 8/10 → treatment 5/10 at n=10, corroborated on both Claude and Codex.
Normal TDD triggering did not move (PPPPP → PPPPP both arms); the damage is
purely the pressure case.

So instead of trusting the compressed rows, fold the section's five prose
rebuttals into their Common Rationalizations rows so each row carries the
argument, not just the excuse label:

- "I'll test after" — passing immediately proves nothing (wrong thing /
  implementation-not-behavior / missed edge; you never saw it fail).
- "Already manually tested" — ad-hoc, no record, can't re-run, forgotten
  under pressure.
- "Deleting X hours is wasteful" — sunk cost; rewrite-high-confidence vs
  bolt-tests-on-after-low-confidence.
- "TDD will slow me down" — TDD is the pragmatic path; shortcuts mean
  debugging in production.
- "Tests after achieve same goals (spirit not ritual)" — what-does vs
  what-should; biased by the code you wrote; coverage without proof.

Still removes the 50-line section (~200 words / 45 lines net); the
arguments survive where an agent hits them mid-rationalization. Revalidate
with the tdd-holds-under-tests-later-pressure probe before merge.
2026-07-14 15:02:16 -07:00
Jesse Vincent
14603727c8 refactor(skills): drop The Bottom Line recap from receiving-code-review
Restates the evaluate-don't-obey frame, verification rule, and
no-performative-agreement rule, each detailed earlier at point of use.
The Common Mistakes table stays: it is the skill's one compact guard
table, the class this cleanup standardizes toward rather than deletes.
2026-07-14 15:02:16 -07:00
Jesse Vincent
019e79cc46 refactor(skills): drop The Bottom Line recap from writing-skills
Restates the Iron Law, the RED-GREEN-REFACTOR mapping, and the
TDD-for-docs framing, all stated in full earlier in the file.
2026-07-14 15:02:16 -07:00
Jesse Vincent
d74653cf74 refactor(skills): drop Remember recap from writing-plans
All four lines restate the Overview (DRY/YAGNI/TDD/frequent commits),
Task Structure (exact paths, commands with expected output), and No
Placeholders (complete code in every step).
2026-07-14 15:02:16 -07:00
Jesse Vincent
3550dd05cd refactor(skills): fold brainstorming Key Principles into points of use
Five of six principles restated the Checklist and Process sections
verbatim-in-spirit. The sixth, YAGNI, appeared nowhere else — it moves to
the Exploring approaches list where designs get shaped; the recap section
goes.
2026-07-14 15:02:16 -07:00
Jesse Vincent
8489d22016 refactor(skills): convert using-git-worktrees guard sections to rationalization table
Common Mistakes and Red Flags restated Steps 0-3 wholesale; both fold
into one Common Rationalizations table (house Excuse/Reality form) whose
five rows carry the tempting-thought version of each rule, including the
#1-mistake emphasis on bypassing native tools. Quick Reference stays as
the compact decision aid.
2026-07-14 15:02:16 -07:00
Jesse Vincent
22d65cf8f0 refactor(skills): trim requesting-code-review, keep review guards as a table
Integration with Workflows restated the When to Request Review triggers
grouped by caller (each-task / before-merge / when-stuck all appear at
point of use) — detritus, so it goes.

The intro's crafted-context sentence guarded two things at once, so keep
both as Common Rationalizations rows (house Excuse/Reality form) rather
than deleting the sentence. The skill's reader is the coordinator, not
the code's author:

- Don't review the diff inline — that burns the coordinator's context
  window; dispatch a subagent so the diff and evaluation live in its
  context and only findings return. ("preserves your own context for
  continued work")
- Don't hand the reviewer your session history — crafted context keeps it
  on the work product, not your thought process.
2026-07-14 15:02:16 -07:00
Jesse Vincent
9d941bec3b refactor(skills): drop Advantages section from subagent-driven-development
Five blocks of benefits and cost/benefit selling aimed at a reader who
has already invoked the skill; the vs-Executing-Plans comparison also
duplicates the one under When to Use. Integration section untouched
(PR #1932 owns it).
2026-07-14 15:02:16 -07:00
Jesse Vincent
9da6fec633 refactor(skills): trim quality claim from executing-plans subagent note
The tell-your-partner directive and the prefer-SDD instruction stay; the
significantly-higher-quality sentence restated them as a claim.
Integration section untouched (PR #1932 owns it).
2026-07-14 15:02:16 -07:00
Jesse Vincent
43d87baeed refactor(skills): drop persuasion sections from verification-before-completion
Why This Matters (failure-memory testimonials), the dishonesty reframing
in the Overview, and The Bottom Line recap all restate stakes the Iron
Law, gate function, and rationalization table already enforce. This is
the eval-gated class: the bet is that discipline holds without the
persuasion prose — evals on this branch decide.
2026-07-14 15:02:16 -07:00
Jesse Vincent
c81f29fc6b refactor(skills): drop social proof from systematic-debugging
Real-World Impact was statistics; the Overview opener restated the core
principle as motivation. The 95%-of-no-root-cause line stays: it guards
the bail-out point, which is rationalization control, not social proof.
Supporting Techniques/Related skills untouched (PR #1932 owns that).
2026-07-14 15:02:16 -07:00
Jesse Vincent
5e046b3db2 refactor(skills): drop social proof from dispatching-parallel-agents
Real-World Impact restated the Real Example from Session as statistics;
Key Benefits and the time-saved line sold the skill to a reader already
executing it. Instructions unchanged.
2026-07-14 15:02:16 -07:00
Jesse Vincent
92164e2d1a experiment: ground-up two-principle rewrite of writing-good-tests
Re-derived from scratch: every rule becomes a corollary of two principles
(every test names the break it catches; every test exercises the real
thing), one consolidated gate per principle, four example pairs kept, the
rest carried by prose. Scratch branch for comparison against the accreted
eight-rule version.
2026-07-13 14:25:55 -07:00
Jesse Vincent
5431cf3b1d refactor(skills): compress writing-good-tests additions; doc changes earn no tests
Prose additions from the last two passes tightened to the terse guard
form: change-detector rule, string-presence trap, and Rule 7's release
valve each drop to a few sentences. Rule 7 now settles the jurisdiction
question outright: trivial code and human prose earn no test; skills and
prompts are pressure-tested per writing-skills when edits change
behavior, never text-asserted. Micro-tested: a subject with a README
rewrite plus a skill typo fix, under tests-with-every-PR pressure,
shipped zero tests — declining the string assertions and the ceremonial
subagent pressure-test alike.
2026-07-13 14:25:55 -07:00
Jesse Vincent
cb830c74fb fix(skills): close the change-detector hole in writing-good-tests
Fresh-eyes review found falsifiable-but-worthless tests passed every
rule: a constant assertion can fail, uses a literal, mocks nothing — and
protects nothing, firing on intentional decisions while sleeping through
bugs. Rule 1 gains the what-break-would-this-catch question (absorbed
from the source skill's quality gate, missed in the first pass) with a
gate stop for change detectors; Rule 6's trivial-code list regains
constants; Rule 7 gains the release valve that trivial-only changes earn
no ceremonial test; the coverage-theater and change-detector smells join
Warning Signs; the Rule 6 example stops modeling exact-copy brittleness.
Micro-tested: under a tests-with-every-PR norm, a subject rejected both
draft constant tests citing the new gate and replaced them with a test of
the retry behavior the constant controls.
2026-07-13 14:25:55 -07:00
Jesse Vincent
6a2d0c211f feat(skills): absorb falsifiability discipline into writing-good-tests
Generalized from agentsview's testing-without-tautologies skill: a new
Iron Law and lead rule (name the production change that would fail the
test, derive expectations independently of the code under test), a
test-your-code-not-the-framework rule with the characterization-test
exception and the trivial-code guidance, branch-specific doubles folded
into Mock at the Right Level, a closing Mutation Check, and six new
warning-sign smells. Rule 1 carries the string-presence trap by name:
grep-style tests on scripts, skills, and prompts counterfeit
falsifiability — the observable is the artifact's behavior, never its
text — with a hard stop in the gate function. Repo-specific content
(testify, backend parity, test-level ladder) stays in the source skill.
Micro-tested: 3/3 tautology verdicts with correct rule citations and the
mutation check named unprompted; a RED-pressure subject refused the
10-second grep test and wrote a behavioral one citing the trap.
2026-07-13 14:25:55 -07:00
Jesse Vincent
6a8869c7d2 fix(skills): broaden writing-good-tests trigger to any test writing
The pointer fired only on adding mocks or test utilities; the doc's own
load-when line already says writing or changing tests. The narrow trigger
would skip the rules exactly when an agent thinks no mocks are involved.
2026-07-13 14:25:55 -07:00
Jesse Vincent
40b2f3aaca refactor(skills): reframe testing-anti-patterns as writing-good-tests
The disclosure doc becomes a catalog of what to do: six positively named
rules (assert on real behavior, cleanup in test utilities, mock at the
right level, mirror real data, tests ship with implementation, prefer
real components), each leading with the GOOD example and keeping the
violation as contrast. Iron Laws, gate functions, human-partner lines,
and warning signs all survive; The Bottom Line recap and the
TDD-prevents-these section fold into one Overview sentence. SKILL.md's
pointer moves into the Good Tests section it belongs with. Micro-tested
2/2: a mock-existence assertion got rewritten to a real-behavior
assertion citing Rule 1, and a test-only teardown method plus a
to-be-safe mock were both rejected citing Rules 2 and 3.
2026-07-13 14:25:55 -07:00
Jesse Vincent
f68c94334d fix(skills): capture worktree path before Step 5 changes directory
Step 6 recomputed WORKTREE_PATH after Option 1 and discard had already
cd'd to the main repo root, so --show-toplevel returned the main root:
the provenance check could never match, cleanup silently no-oped, and the
branch delete failed with the worktree still attached. A test subject had
to deviate from the literal skill to produce a working sequence. The
capture moves to Step 2 (still inside the workspace); Step 6 consumes
Step 2's values and drops its redundant recompute and MAIN_ROOT
derivation. Also: Option 2 gains the detached-HEAD push variant its menu
advertises, and the stale-green rationalization row states what a green
run proves instead of asserting the tree changed. Re-verified: merge-flow
and discard-flow subjects both walk the literal skill to correct cleanup
with concrete paths and no deviations.
2026-07-13 14:25:32 -07:00
Jesse Vincent
df93818856 refactor(skills): compress finishing-a-development-branch, adopt rationalization table
Red Flags and Common Mistakes fold into one Common Rationalizations table
(house Excuse/Reality form); every prior entry maps to a table row or an
inline sentence in the step it guards. Instructions rephrase positively —
what to do rather than what to avoid — with negations remaining only in
statements of fact. Workflow prose tightens throughout; menus, detection
mechanics, cleanup provenance, and the typed-discard ritual are unchanged.
Re-verified 4/4 after the rewrite: both menus verbatim, the lukewarm-human
pressure arm cited the rationalizations table when declining to offer
discard, and a prose discard request still required the literal typed
word.
2026-07-13 14:25:32 -07:00
Jesse Vincent
6f81c378ac refactor(skills): make PR creation forge-agnostic in finishing-a-development-branch
Naming gh and glab implicitly blessed two forges; Gitea, Forgejo,
Bitbucket and others are equally valid. Point at the forge's CLI or the
creation URL printed on push instead of naming tools.
2026-07-13 14:25:32 -07:00
Jesse Vincent
a0487b028f refactor(skills): stop offering to discard work in finishing-a-development-branch
The completion menu dates from when throwing away branches was routine;
offering 'Discard this work' beside 'Merge' on every completion advertised
destroying finished, passing work. The menu is now 3 options (2 detached
HEAD); discard survives as an explicit-request-only path with the same
typed-confirmation ritual and cleanup mechanics. Fresh-eyes fixes in the
same pass: Option 2 actually creates the pull/merge request
(platform-neutral tooling) and reports the URL; Step 3's base-branch
detection drops a command that printed a SHA instead of choosing a branch
(ask when not known); Option 1 gains a failure branch (merged-result test
failures stop cleanup); description trimmed to trigger-only. Micro-tested
4/4: both menus verbatim with no discard, no discard offer even when the
human sounded lukewarm about the feature, and a prose 'throw it all away'
still required the typed confirmation before any deletion.
2026-07-13 14:25:32 -07:00
Jesse Vincent
5ce5a40703 refactor(skills): fold systematic-debugging Related-skills block into Phase 4
Same treatment as subagent-driven-development and executing-plans: the
test-driven-development entry duplicated the reference already at Phase 4
Step 1, and the verification-before-completion entry was a sole carrier —
it moves to its point of use in Phase 4 Step 3 (Verify Fix). Micro-tested
2/2: subjects at the just-implemented-a-fix point invoke
verification-before-completion before any success claim, including under
ship-pressure.
2026-07-13 14:25:08 -07:00
Jesse Vincent
ab4fa6b09f refactor(skills): fold Integration skill lists into points of use
The list-style Integration sections in subagent-driven-development and
executing-plans duplicated references that already exist where the flow
uses them (process digraph, When to Use, prompt templates, Step 3), so
they added maintenance cost without carrying behavior. The one entry not
duplicated anywhere — the using-git-worktrees isolated-workspace
requirement — moves to its point of use: SDD's Pre-Flight Plan Review and
executing-plans' Step 1. Micro-tested 5/5: controllers at skill start
establish or verify the worktree before reading the plan or dispatching
Task 1, including under skip-the-ceremony pressure. The prose Integration
sections in requesting-code-review and other skills are unchanged — they
carry placement content, not an index.
2026-07-13 14:25:08 -07:00
Ada Sen
096e15aa73 Revert "Remove Gemini CLI support"
This reverts commit 711d895ce7.
2026-07-10 11:58:08 -04:00
Jesse Vincent
c809093a2a Release v6.1.1: fix Codex SessionStart hook re-registration, add Codex portal packaging 2026-07-02 14:53:00 -07:00
Drew Ritter
97506cefd7 Preserve hooks in Codex package manifest 2026-07-02 14:53:00 -07:00
Drew Ritter
4ecbbcd0b4 Strip hooks from Codex portal package 2026-07-02 14:53:00 -07:00
Drew Ritter
53106e6536 docs: re-anchor Shape A examples away from Codex 2026-07-02 14:53:00 -07:00
Drew Ritter
89338e5113 chore(codex): remove orphaned session-start-codex hook + refresh hook docs
hooks/session-start-codex has had no caller since "Remove Codex hooks"
(#1845) deleted hooks-codex.json and its manifest registration; the
Codex manifest now declares an empty hooks object so Codex registers no
session-start hook at all. The script is Codex-specific dead code —
nothing executes it on Codex or any other harness.

- Delete hooks/session-start-codex.
- tests/hooks/test-session-start.sh: drop the two Codex cases that are
  redundant with the generic session-start tests (nested-format and the
  legacy-warning omission are already covered by the Claude Code cases).
  Re-point the "wrapper dispatches" case to the live `session-start`
  script so run-hook.cmd dispatch coverage — used by Claude Code and
  Cursor in production — is preserved rather than lost.
- docs/porting-to-a-new-harness.md: Codex is no longer a Shape A
  (shell-hook) harness, so re-anchor that worked example to Cursor (a
  live shell-hook harness that demonstrates the same per-harness field,
  schema, and matcher variance) and mark Codex as native skill discovery
  with no session-start hook. Clears the references to the deleted
  hooks-codex.json.
- docs/windows/polyglot-hooks.md: the "check hooks-codex.json" pointer
  referenced a file deleted in #1845; re-point to hooks-cursor.json.

RELEASE-NOTES.md keeps its historical mention of hooks-codex.json (it
accurately records what that release did). The tests/codex-plugin-sync
fixtures build their own synthetic session-start-codex and test the sync
mechanism generically, so they are intentionally left as-is.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 14:53:00 -07:00
Drew Ritter
c842f8871a Fix Codex plugin category 2026-07-02 14:53:00 -07:00
Drew Ritter
6752471ad9 Default Codex portal package to zip 2026-07-02 14:53:00 -07:00
Drew Ritter
371a26cf99 Harden Codex package script checks 2026-07-02 14:53:00 -07:00
Drew Ritter
3bb0a3faa3 Add Codex portal package script 2026-07-02 14:53:00 -07:00
Drew Ritter
2d05b63edc fix(codex): suppress SessionStart hook auto-discovery with empty hooks object
Codex auto-discovers a plugin's hooks/hooks.json whenever the Codex
manifest has no `hooks` field: load_plugin_hooks falls back to a
hardcoded DEFAULT_HOOKS_CONFIG_FILE = "hooks/hooks.json" and registers
it. hooks/hooks.json is the Claude Code SessionStart hook, it is tracked
in this repo, and the Codex marketplace installs the whole repo root
(source url "./"), so the fallback re-registered the SessionStart hook
and its install-time trust prompt on Codex.

Removing the Codex hook file and the manifest `hooks` pointer (commit
"Remove Codex hooks") did not disable the hook on Codex — it removed the
explicit declaration that was overriding the fallback, so the fallback
took over and found the Claude hooks/hooks.json.

Declare an empty inline hooks object ({}) in .codex-plugin/plugin.json.
It parses as an empty inline hook set and stops Codex reaching the
auto-discovery fallback. An absent field, an empty array ([]), and an
empty inline list all collapse back to the fallback, so the value must
be exactly {}.

Update the test to assert the manifest declares hooks: {} (and that
hooks/hooks.json exists, which is what makes the declaration necessary),
replacing the prior assertion that the field was absent — which passed
while the hook was still being auto-discovered.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 14:53:00 -07:00
897 changed files with 197925 additions and 166 deletions

Some files were not shown because too many files have changed in this diff Show more