0
Fork 0
mirror of https://github.com/obra/superpowers.git synced 2026-09-25 22:09:05 +00:00
Commit graph

170 commits

Author SHA1 Message Date
Drew Ritter
b5d94eb0ae
Revert movie skill import (#2214) (#2335)
Remove the proving-it-works-with-a-movie import at Drew Ritter’s request because its code quality does not meet the release bar.

Reverts merge dd53fe0b57 and the dependent movie test adjustment in 3979a17bda. Remove the later Muse registration and update the unreleased notes so neither advertises the removed skill. Version changes are left for a separate release step.
2026-09-18 14:27:57 -07:00
Jesse Vincent
3979a17bda
test(movie): expect 8 frames from the slow-capture film test (#2329)
42486db made film() fill every grid slot before the endpoint, but this
test kept its old count of 7. Its fake clock advances 0.01 s per capture
and 0.02 s per sleep, so the prompt is seen at 1.01 s, the 0.4 s hold
ends at 1.41 s, and slots 0.0-1.4 s make 8 frames. The test has failed
on every run since 42486db.
2026-09-18 11:30:01 -07:00
Drew Ritter
8f5a89afe4
Merge pull request #2106 from GoldJohnKing/feat/opencode-v2-support
feat(opencode): support OpenCode 2.0.4+ alongside V1
2026-09-17 21:46:19 -07:00
Drew Ritter
3bb3bd7b39 fix(scripts): exclude root index.js from the Codex plugin sync
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:51:22 -07:00
Drew Ritter
01c004da99 fix(opencode): strip quote pairs after joining multi-line frontmatter values
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:49:55 -07:00
Drew Ritter
d29eaf543c test(opencode): cover bootstrap placement when native compaction retains user messages
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:47:55 -07:00
Drew Ritter
0d7bc1d6b7 chore(opencode): mark test-skill-registration.sh executable
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:46:51 -07:00
Drew Ritter
2c4dcd851a fix(opencode): retain bootstrap after native compaction 2026-09-17 18:24:02 -07:00
Drew Ritter
96ecdeb0fa fix(opencode): retry unsuccessful child-session lookups 2026-09-17 18:21:30 -07:00
Drew Ritter
b1b8d611d4 test(opencode): cover canonical skill paths and registration survival 2026-09-17 18:17:48 -07:00
Jesse Vincent
2b89c4c3f3
Rebuild executing-plans as a first-class inline execution mode
A cheaper execution mode alongside subagent-driven development: the
session implements every task itself under the same workspace, ledger
and stopping rules, with one fresh whole-branch review on the most
capable model at the end. Helper scripts task-start/task-done keep the
ledger and test log honest; the final fix pass re-grades findings and
fixes Critical/Important under TDD. writing-plans' handoff, SDD's
when-to-use text and two README lines change to match.
2026-09-17 17:14:38 -07:00
Jesse Vincent
6692bd289f fix(sdd): ownership markers stop same-basename plans sharing a workspace
sdd-workspace slugged workspaces by basename alone, so docs/alpha/plan.md
and docs/beta/plan.md resolved to one directory and task-brief silently
overwrote the other plan's brief — the single gitignored source of task
requirements, unrecoverable once clobbered.

Each workspace now records its owning plan in a plan-path marker
(repo-relative in-repo, absolute outside). Lookup keeps basename slugs
and existing behavior for the common case: a markerless workspace is
adopted in place (no migration break for in-flight plans), a marker
naming this plan is a match, and a marker naming a different plan
disambiguates with the plan's parent-directory name, then a counter.
Plan paths are normalized (CDPATH-guarded physical cd) so relative,
absolute, and ../ spellings of one plan share one workspace.

task-brief and review-package delegate to sdd-workspace and need no
changes. SKILL.md's workspace bullet no longer promises the exact
<plan-basename> path, since disambiguated workspaces differ.

Reported by @CRGDan; reproduction and test groundwork by @crisnahine
in PR #2120.

Fixes #2045
2026-09-17 15:51:11 -07:00
Drew Ritter
a2b2d1d5ef
Merge pull request #2134 from obra/fix/sdd-helpers-no-exec-bit
fix(sdd): invoke sdd-workspace via bash so helpers survive stripped exec bits
2026-09-17 15:48:31 -07:00
Drew Ritter
b61242d8d0
Merge pull request #2136 from obra/fix/review-package-range-guards
fix(sdd): reject empty or non-descendant BASE..HEAD ranges in review-package
2026-09-17 15:47:55 -07:00
GoldJohnKing
07f9d6f6fc fix(opencode): register skills with Skill.Info 2.0.4 path field; contain per-skill add failures
Reported on PR #2106 (80avin): on OpenCode v2.0.4 the plugin is disabled
at startup with "Plugin disabled after skill.transform failed", losing
both skill registration and bootstrap injection.

Root cause: upstream commit 199aabe9e2 (first released in v2.0.4)
renamed Skill.Info's required file field `location` -> `path` and removed
`slash`. draft.add() decodes payloads with Schema.decodeUnknownSync
against that schema, so our `location` payloads now fail decode with
"Missing key path".

Why the failure was silent: the decode error is thrown during the host's
state rebuild, where the State layer catches it and hard-disables the
whole plugin group asynchronously - the throw never reaches the
try/catch around ctx.skill.transform(), and the session "context" hook
is torn down as collateral.

Fix:
- skill payloads now use `path` (2.0.4 contract); no v2.0.3 compatibility
  retained per review decision
- draft.add() failures are contained per skill inside the transform
  callback, so one rejected payload skips that skill (visible in server
  logs) instead of the host disabling the entire plugin
- new test-skill-registration unit test pins the 2.0.4 payload contract
  (absolute path field, no stale location/slash, hostile-add containment)
  and is registered in run-tests.sh; full suite 3/3 green
2026-09-16 20:35:08 +08:00
Drew Ritter
b6bc191746 Merge origin/dev into import/proving-it-works-skill
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 22:17:24 -07:00
GoldJohnKing
4996eb00bd fix(opencode): align V2 tool mapping with the 2.0.3 catalog; harden plugin internals
- V2 bootstrap mapping now teaches write/edit/websearch (verified against
  a live v2.0.3 /api/plugin tool catalog) instead of routing all file
  mutation through patch
- frontmatter parser tolerates CRLF and YAML block scalars/continuation
  lines
- child-session cache is bounded (512 entries, oldest-quarter eviction)
- INSTALL.md and docs/README.opencode.md tool tables synced to the 2.0.3
  catalog; unit-test needle list extended to cover the new tools
2026-09-14 17:47:03 +08:00
GoldJohnKing
ebeb765e04 Merge branch 'dev' into feat/opencode-v2-support 2026-09-14 14:54:37 +08:00
Drew Ritter
2c3c59ce3f Reject invalid chat transcripts and propagate browser cleanup failures
Address the two final whole-PR findings from the consolidated movie repair
brief. A fresh openai-chat response must provide a speech-bearing string
transcript even when local ASR is off; null must not reuse the no-transcript
sentinel belonging to deterministic engines or cached accepted audio.
Preserve the bounded candidate loop and rejected-byte evidence.

Report failed Windows tree termination as OSError so the recorder owner's
existing per-child handler continues all cleanup and retains failure metadata.
Card rendering checks its acquired browser handle before acting on the PID,
propagates wait failure, and reports locked-profile removal instead of
allowing a pending successful return to hide incomplete cleanup.

Add fake HTTP/process/filesystem boundary regressions for both findings,
including real adapter invocation, accepted chat cache reuse, ASR modes,
normal completed cards, and serve cleanup after a failed leader later exits.
Correct the rejected-chat fixture to use the chat engine. No media or native
browser execution was performed; Drew retains personal video acceptance.

Validation: expected RED failures retained; 40 focused tests and the single
75-test contracts entrypoint pass. Full evidence and self-review are in
.superpowers/sdd/2026-09-11-movie-committee-repairs/final-fix-report.md.
2026-09-11 20:38:31 -07:00
Drew Ritter
b206e0cbe1 docs: make proof-movie recipes preserve failures
Repair the shipped Unix, logging, subtitle, cursor, narration, and recorder guidance against the literal fake-boundary failures recorded for Task 4. The primary pipeline and subtitle recipe now fail fast, measured offsets reach subtitle generation, producer logging preserves the real status under the pipefail owner, and cursor mouseup restores the released state.

Document the accepted narration/cache contract, cooperative recorder cleanup limits, and the safe contracts-suite entrypoint without changing Windows recipes or the existing evidence and human-viewing gates. Include the controller-owned plan bookkeeping and record the executable RED/GREEN results while leaving independent fresh-reader trials pending.

Prompt: implement Task 4 focused executable movie-guide corrections after Tasks 1-3, using writing-skills and only fake producers, text fixtures, and fake DOM execution.

Verification: python3 .superpowers/review/pr2214/committee/recipe-probes.py; uv run --script tests/proving-it-works-with-a-movie/run-tests.py --suite contracts; git diff --check.
2026-09-11 20:16:33 -07:00
Drew Ritter
d340830d6d fix(movie): retain incomplete cleanup and command evidence
Address both independent Task 3 review findings at 42486dbb. When an owned leader exits before tree cleanup, the existing parentage helper cannot confirm orphan descendant cleanup. Treat that state as incomplete, preserve PID metadata, return failure, and never invoke the helper on the exited leader's numeric PID. Continue releasing the other owned resources without adding descendant tracking.

Keep a recording capture failure truthful while preserving available completed command evidence: outcome remains failed, exit status remains 1, and the capture error is retained alongside ok, native exit_code, and cwd. A completed successful command does not make a failed recording successful, and no successful take manifest is published.

TDD regressions reproduced an exited leader incorrectly returning 0 and capture failures dropping completed native status for exits 0 and 7. All 14 covering lifecycle and observation tests pass using fake process handles, fake clocks, mocked capture/log boundaries, and intercepted media writes; git diff --check passes. No real media generation or inspection, process/browser launches, SessionTests, FilmGridTests, or full media suites were executed. Detailed RED/GREEN evidence is appended to .superpowers/sdd/2026-09-11-movie-committee-repairs/task-3-report.md.
2026-09-11 20:07:49 -07:00
Drew Ritter
42486dbb6d fix(movie): own recorder cleanup and bound observation
Implement Task 3 of the consolidated PR 2214 movie repairs. Startup failures could leak an already launched ttyd or browser because acquisitions preceded the cleanup block; close killed historical numeric PIDs and reported success without confirmed cleanup. Register each acquired resource inside serve's try/finally, invalidate readiness on shutdown, retain logs, and atomically retire PID metadata only after confirmed owned-process and profile cleanup. Close now requests stop and waits under one overall 30-second deadline without killing PIDs.

Poll CDP and owner readiness while observing commands without recording, preserving a completed native exit status across simultaneous disconnects and reserving exit 2 for live unfinished commands. Fill only bounded 5 fps slots when a final capture crosses the hard or hold endpoint, correctly strip ST-terminated OSC titles, and emit ASCII-escaped stdout JSON while preserving UTF-8 session files. A serve-only CDP call-boundary check prevents a stop from waiting through multiple consecutive calls.

Add 21 contract tests using fake processes, fake websocket responses, fake clocks, byte-token capture callbacks, and intercepted frame writes. Register real-session test cleanup immediately after Popen and retain live PID snapshots, but do not execute that fixture. RED evidence and implementation report are in .superpowers/sdd/2026-09-11-movie-committee-repairs/task-3-report.md. Validation: 26/26 authorized prompt, serve-argument, and recorder contract tests pass; git diff --check passes. No media generation, inspection, browser/ttyd launch, real session or frame-grid suite was performed. Native movie acceptance remains Drew's review; an abruptly hard-killed owner with a live browser and stale readiness remains the agreed limitation.
2026-09-11 20:01:08 -07:00
Drew Ritter
e48cb57bf1 fix(movie): refine subtitle chunks using allocated durations
Address Task 2 review round 1: initial character budgets count spaces that disappear between chunks, so proportional timing can exceed --max-secs even when a further word-boundary split is feasible. Reallocate after splitting an over-target multiword cue, checking actual integer millisecond intervals each time.

Coalescing runs once before refinement, and refinement stops at the available millisecond count or unsplittable words. This keeps tiny impossible readability targets bounded while preserving every word and the full measured scene interval.

Added a failing six-second unequal-chunk regression and a tiny-target termination/positive-interval guard. The RED emitted 3273 ms for a cue with a 3000 ms target. GREEN: 43 safe contract tests and six covering existing offset/BOM tests pass. All verification remained text/data or mocked media boundaries.
2026-09-11 19:44:06 -07:00
Drew Ritter
8e3414cdb3 fix(movie): keep subtitle timing and track selection faithful
Implement Task 2 of the movie committee repairs. Allocate proportional cue boundaries across each complete rounded scene interval, reserve positive millisecond spans, and coalesce chunks when the available interval cannot represent them separately. Readability limits guide word splitting without dropping text or clipping narration tails; invalid intervals fail before subtitle output is written.

Explicitly select the supplied soft subtitle track with optional source audio, including the hard-burn fallback. Parse only numbered SRT cue timing lines so arrow-bearing captions cannot crash the checker or inflate coverage. Preserve the existing maximum cue end policy, assembly-offset intersection, and partial manual retiming.

Add 17 safe subtitle contracts and register the contracts runner suite. Real-file mocked narrate/assemble/subtitle reruns retain removed narration WAV evidence while omitting stale assembly audio and offsets. Verification: 41 contract tests and 18 authorized existing regressions pass; only text and mocked media boundaries were exercised. Native Windows and human video acceptance remain outside this verification.
2026-09-11 19:39:20 -07:00
Drew Ritter
dd6fe21170 test(movie): cover unsupported ASR and ordinary retries
Close Task 1 review round 2 without production changes. Feed actual unsupported ASR text through auto and strict verification while proving off does not call ASR. Run the second rejected narration invocation without --force, then assert it synthesizes new takes and retains distinct rejected bytes instead of reusing cache evidence.
2026-09-11 19:31:08 -07:00
Drew Ritter
16167e281a fix(movie): retain narration evidence through verification
Address Task 1 review round 1. Treat successful empty ASR output as speech failure rather than an unavailable verifier, reject empty chat claims within the existing bounded retry loop, and extend unsupported segmentation detection to supplementary CJK ideographs.

Withdraw cached acceptance before every revalidation so strict failures and interrupts cannot leave stale publication. Preserve nested manifest WAV paths on cache reacceptance, and measure a unique candidate before promotion so duration failures retain their evidence across reruns. Add structured movie geometry coverage plus both silent and source-audio mapping tests. All tests use uv --no-project with mocked synthesis, ASR, probes, and encoding; no media operation was run.
2026-09-11 19:28:31 -07:00
Drew Ritter
5e02913144 fix(movie): make narration acceptance govern assembly
Repair Task 1 from the 2026-09-11 movie committee plan. Narration now atomically withdraws acceptance before it mutates accepted bytes, stages failed takes as retained evidence, and publishes a manifest entry only after transcript gates and duration measurement. It preflights ffprobe, preserves bounded retries and cache identity, and distinguishes unsupported token comparison from cache identity.

Assembly now validates accepted manifest narration before it starts encoding, ignores generated narration for movie scenes, selects manifest WAV paths, maps source or synthetic movie audio explicitly, fits wide movies within the requested inner rectangle, and escapes only literal directory percent signs in sequence paths. The focused regression suite mocks every media boundary; real-media fixture declarations are updated but not executed. Prompt constraints prohibit real media generation, probing, inspection, synthesis, ASR, browser, checker, or full media suites.
2026-09-11 19:19:47 -07:00
Drew Ritter
b7dd0a51cb Fix inserted narration and movie audio subtitle checks
Address the final three review findings on PR #2214, as approved by Drew. Count both sides of every non-equal transcript span so a short insertion or expanded replacement cannot evade the drift gate on a longer script. Preserve the existing length and run thresholds.

Honor --no-expect-audio for encoded silent tracks. Base subtitle requirements on detected audible speech so opting out of expected audio does not suppress captions for speech that is present.

Extract the first embedded subtitle stream as SRT when no sidecar is present and apply the same cue-end check to either source. Empty cues fail even for short narration, and extraction or malformed timing errors are reported as failures. Preserve the silent end-card allowance.

Validation: the new tests first reproduced ten failing cases across narration insertion, silent-track opt-out, and embedded subtitle handling. All 35 focused narration, checker-policy, and subtitle-text tests now pass. External media commands, audio and picture sampling, contact-sheet creation, synthesis, and ASR were mocked; no actual media inspection or live ASR was performed. Drew retains final video acceptance.
2026-09-11 18:07:49 -07:00
Drew Ritter
ae559110ce Fix narration cache identity and partial subtitle offsets
Address the fresh review on PR #2214 after the #2275 integration. Drew approved fixing the two reproduced bugs and keeping the specified auto/on/off verification semantics.

Cache accepted narration by normalized text plus effective engine, voice, and synthesis model. Resolve voice defaults before rendering, invalidate entries without settings, and retain requested ASR checks on cache hits. Exclude rejected clips as before.

Treat manual subtitle offsets as start-time overrides. Only assembly offsets JSON selects scenes in the cut, including when manual timing overrides are also supplied. Empty narrated cuts write an empty SRT without crashing.

Make the gated Unix example pass --verify on and document the actual local ASR modes. Auto remains permissive if ASR is unavailable; on remains strict.

Validation: observed the new cache and subtitle regressions fail before the fixes; all 23 focused narration and subtitle-text tests now pass. Synthesis, duration probing, and ASR are mocked. No media inspection or live ASR was performed; Drew retains final video acceptance.
2026-09-11 17:40:30 -07:00
Drew Ritter
4d4ede2925 fix(movie): preserve rejected narration and prior takes
Prevent failed narration scenes from entering the cache manifest, while leaving their generated WAV files available as failure evidence. Add filesystem-backed regressions covering repeated rejected chat synthesis, accepted-scene reuse, and strict ASR rejection of cached audio.

Refuse nonempty recording directories both before CLI session side effects and at the direct film boundary. The regression preserves existing numbered frames and a sentinel byte-for-byte across run, key, watch, and direct film refusal.

Make subtitle capability tests independent of the host FFmpeg installation, skip the real pixel test before probing unavailable tools, retain strict skipped-capability rejection, and remove only the unused websockets runner dependency.
2026-09-11 16:42:20 -07:00
Drew Ritter
ac22047ce8
Preserve diagnostic evidence through doctor export
Align scrubber and independent audit prompts around one shared redaction policy while preserving safe command, result, source, session-line, quotation, and linkage structure. Add finished-handoff evidence and reconciliation instructions, provenance labels across case/report/bundle/issue templates, and the structural existence check for the shared reference.\n\nThis patch responds to the retained negative post-report handoff baseline: cited result bodies were removed wholesale, source findings and the positive related-session match were not verifiable, provenance and export statements were stale, and scrub counts disagreed. The behavioral handoff validation remains pending for the follow-up task; this commit records only the focused product guidance and structural RED/GREEN evidence.
2026-09-11 15:20:10 -07:00
Drew Ritter
d4236278fb
Use shared session discovery in diagnosing-superpowers
At Drew's request, apply the evaluated shared-discovery variant to Jesse's
existing PR #2236. Resolve native session sources and record semantics from
available tools, documentation and bounded inspection. Record verified absolute
paths, linkage, extraction queries, human-message distinctions, usage-counter
semantics and uncertainty once in the case for all analysts to consume.
Replace the three per-harness references and update structural checks.

This is exactly the evaluated source tree at
3f0a63e860, applied as one commit on
801badbf71. Fourteen files change;
126 lines added, 285 removed. No private eval fixtures or transcripts ship.

Validation:
- Structural test: 45 passed, 0 failed before and after application.
- Staged tree exactly matches the evaluated candidate; diff check passes.
- Independent read-only review: no actionable blockers.
- Retained before/after full doctor runs: one pair each on native Claude,
  Codex and Pi. All six delivered reports and completed seven dimensions.
  Shared discovered all three native session families without the removed
  references. Both versions had report-quality defects; shared Codex deleted
  its cited case through a fixture symlink. Preserve this negative result.
- Eight fresh Codex follow-ups: original/shared x symlink/ordinary-home x
  two repeats, one retained historical session family. All eight retained
  cases and supported the four core findings. Seven native final deliveries;
  one shared run stopped on provider capacity after writing its report.
  No deletion recurred. One original reused three analysts for seven tasks.
  Recorded follow-up cost $34.8617883, all eight attempts accounted for.

These observations support this scoped simplification, not general equivalence
or a causal claim that reference removal caused or could not cause a failure.
Child assignment/model choices were native behavior; the complete variants
also differ in analyst prompts. Common provenance, citation-verification and
measurement problems remain separate follow-ups. No new paid runs were made
for this publication; evidence and independent audits are retained privately
by Drew. Behavioral evaluation provenance: campaigns
358c7333-c5f0-48bd-a733-61196da992ed and
102d630d-30ef-49a0-97e1-8dd410ed0548.

Prepared with GPT-6 using Codex through Paseo; local codex-cli 0.153.4.
Skills used: superpowers writing-skills, using-git-worktrees,
requesting-code-review, verification-before-completion; primeradiant-ops
linear-ticket-lifecycle. Drew approved publishing this evaluated variant.
Enabled plugins in the publishing checkout's Codex configuration:
- github@openai-curated
- documents@openai-primary-runtime
- spreadsheets@openai-primary-runtime
- presentations@openai-primary-runtime
- primeradiant-ops@primeradiant
- slack@openai-curated
- linear@openai-curated
- codex-security@openai-curated
- pdf@openai-primary-runtime
- template-creator@openai-primary-runtime
- sites@openai-bundled
- visualize@openai-bundled
- computer-use@openai-bundled
- cloud-build@superpowers-cloud-build
- browser@openai-bundled
- superpowers@superpowers-dev
- stream-deck-agents-codex@drew-local
- computer-history@openai-bundled
- codex-app-tools@openai-bundled
- unified-computer-use@openai-bundled
- chrome@openai-bundled
- bits-and-bolts@mcp-extensions-early-access
- visual-probe@visual-probe-local

Tracking: PRI-3127
2026-09-11 15:20:09 -07:00
Jesse Vincent
be4263611e
diagnosing-superpowers: writing review fixes
Move GitHub search and prefilled-link mechanics to references/github-issues.md.
State the redaction levels neutrally instead of nudging toward more data.
Say that all seven analysts always run and what the quick-reference table
is for. Add a title slot and a bundle slot to the issue template. Drop the
duplicated human-prompts rule from request-conflicts. Prose fixes: active
voice, dangling modifier, vague referents, two lists turned into tables.
2026-09-11 15:20:09 -07:00
Jesse Vincent
20c37ac909
diagnosing-superpowers: share the analyst preamble and context-safety rules
The seven analyst prompts opened with an identical 39-line block (role,
inputs, context safety, return format). It now lives once in
prompts/analyst-common.md and each dimension prompt points at it. The
wc -lc / long-line / never-cat rule was restated in nine places; it now
lives in references/context-safety.md and everything else points there.
Addresses arittr's review on #2236.
2026-09-11 15:20:09 -07:00
Jesse Vincent
9682098521
diagnosing-superpowers: build the scrubbed bundle only on request
Never build or push a bundle unprompted. When intake names a bug report as
the goal, say once that a bundle is available on request, then wait. On
handover, state what the bundle contains, point at the scrub log, and say
scrubbing can miss things so every file needs review before sharing.
Raise the SKILL.md word budget to 1000 to fit the added rule.
2026-09-11 15:20:09 -07:00
Jesse Vincent
fca1136291
feat: add diagnosing-superpowers skill
Evidence-based diagnosis of superpowers sessions: intake with the human
partner, safe transcript reading for Claude Code and Codex (discovery
procedure for other harnesses), seven analyst subagents, a report with
path:line evidence and a bounded superpowers-involvement line, scrubbed
export bundles, approval-gated GitHub issue search/draft, and
similar-session search. Includes spec, plan, structure test, and README
and docs index lines.

Developed RED-GREEN-REFACTOR per writing-skills: 46 scored scenario runs
across five SKILL.md versions, all twelve scenarios clean against the
final version, micro-tests control 5/5 to skill 0/5 on both
baseline-failing prohibitions, and one end-to-end run. Eval records are
kept by the maintainer outside the repo.

Claude-Session: https://claude.ai/code/session_01DyaGKhTXvHNs2JgPhDktz7
2026-09-11 15:20:09 -07:00
Drew Ritter
26c44ceee0 feat(movie): simplify the Windows terminal recorder to serve/run/key/watch/close
Windows has no tmux, so the recorder had grown into a 738-line daemon with a
file-based request protocol, request IDs, wait-only result retrieval, and a
Win32 Job Object module. Replace it with the Unix route's shape: serve keeps
ttyd and a headless browser alive and logs the terminal's output; run, key,
watch, and close are one-shot CDP calls against that browser. The installed
prompt reports each command's status through the window title, so run can
print it without any visible marker.

Process cleanup uses taskkill /T (a pgrep walk on Unix) instead of Job
Objects, which also simplifies the card renderer. The session tests run on
macOS too, since nothing in the script is Windows-specific.

Verified: 9 session tests per shell on Windows 11 for PowerShell 5.1,
PowerShell 7, and Git Bash; the browser suite with Chrome and Edge; 44
portable tests on macOS against a real ttyd session.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-09 23:41:25 -07:00
Drew Ritter
4f5304bd31 test(movie): keep tool output out of the test log
The in-process narration and subtitle tests let the tools' stdout and
stderr through to the runner. Capture both and assert the expected
diagnostics.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-09 22:56:28 -07:00
Drew Ritter
ef1e080141 test(movie): keep one portable test suite
The Python suite ported the three shell tests' assertions so they run on
Windows; both copies were kept and the README mapped one to the other. Keep
the portable suite. Drop the one-shot Windows acceptance driver and its
browser fixture, which produced evidence rather than regressions, and the
never-implemented reserved suite names.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-09 22:53:20 -07:00
Drew Ritter
b2f28a41c4 chore(movie): drop the abandoned OS-rollout probe and superseded plans
The first, broader OS-compatibility rollout was stopped and replaced by the
narrower Windows completion. Its feasibility probe, probe cleanup test,
design, review, 12-task plan, results report, and the completion plan and
review record were internal execution artifacts with machine-specific paths.
The one probe-derived test list is inlined into the terminal suite.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-09 22:53:19 -07:00
Drew Ritter
a563302acc docs(movie): correct reserved regression suite names 2026-09-09 21:46:32 -07:00
Drew Ritter
8a31ffc3ff docs(movie): document and verify Windows workflows 2026-09-09 21:36:58 -07:00
Drew Ritter
ce9b3fcf26 fix(movie): verify terminal health before finalizing takes 2026-09-09 20:44:19 -07:00
Drew Ritter
b8a723f474 feat(movie): support native Windows terminal recording 2026-09-09 20:35:45 -07:00
Drew Ritter
5621d1d6aa fix(movie): finish native Windows media tools 2026-09-09 19:57:45 -07:00
Drew Ritter
442a48d990 fix(movie): make frame and concat inputs portable 2026-09-09 18:11:56 -07:00
GoldJohnKing
ca380ba488 fix(opencode): differentiate tool mapping by host flavor and harden child detection
Current opencode2 builds renamed the model-facing tools (bash→shell,
task→subagent with `agent` instead of `subagent_type`, apply_patch→
patch/patchText) and removed todowrite entirely, so the single v1
mapping injected on v2 hosts taught the model stale tool names.

- export V1_MAPPING/V2_MAPPING and inject the flavor-correct one on
  each path (v1 messages.transform → V1; v2 ctx.session.hook("context")
  → V2, incl. no-todo-tool guidance and sessionID continuation)
- child-session detection now keys on parentID presence (primary
  signal on both flavors) with dual-shape unwrapping preserved; v1
  #2160 behavior unchanged
- mirror surfaces updated: INSTALL.md dual mapping tables,
  README.opencode.md host-flavor notes, test-bootstrap-caching.mjs
  asserts both mappings + drives the v2 context hook end-to-end
- skills: add OpenCode to executing-plans' subagent-capable list,
  accurate OpenCode worktree status (git fallback; TUI dialogs are
  user-side only), generalized live-subagent resume guidance
2026-09-10 09:00:50 +08:00
Drew Ritter
2996ad255c test(movie): verify first subtitle cue offset 2026-09-09 17:42:29 -07:00
Drew Ritter
03995d48f1 test(movie): port regression fixtures to Python 2026-09-09 17:18:41 -07:00
Drew Ritter
a2e7743020 fix(movie): release probe resources after evidence failures 2026-09-09 16:45:38 -07:00