0
Fork 0
mirror of https://github.com/obra/superpowers.git synced 2026-09-28 09:42:11 +00:00
obra-superpowers/tests/claude-code
Jesse Vincent 179bb3b02f executing-plans v3: helper scripts, TDD-verified fix pass, scoped scan, Review Focus
Five efficiency changes, all aimed around the one thing the evals showed
works (a fresh most-capable-model review of the whole diff):

- scripts/task-start and scripts/task-done fold each task's bookkeeping
  into one call apiece: brief + BASE at the start; test run, log, and
  ledger line at the end (a failing run records nothing). Every tool
  call in an inline session is a turn that re-reads the whole context.
  Tested by tests/claude-code/test-executing-plans-scripts.sh.
- The final fix pass is verified by TDD plus a green suite instead of a
  re-review dispatch. Twelve re-reviews ran across today's inline, SDD
  and workflow reps; every one returned 'all addressed, no new
  breakage'.
- The pre-flight scan covers only producer/consumer pairs named by the
  plan's Interfaces blocks; a plan with no shared interfaces gets one
  ledger line. Nine of nine inline reps produced an all-clean table.
- writing-plans emits a Review Focus section (input classes and failure
  modes the spec implies and no task's tests exercise) with a matching
  self-review step, and executing-plans hands it to the final reviewer.
- Both skills say inline runs well on a mid-tier session model with the
  most capable model reserved for the review.

inline-eval gains spikesonnet and spikerf arms to measure the last two.
History 2026-09-16 15:51:50 -07:00
..
analyze-token-usage.py fix: replace bare except with except Exception 2026-03-09 17:10:07 -07:00
README.md Address adversarial review findings 2026-06-16 10:09:43 -07:00
run-skill-tests.sh executing-plans v3: helper scripts, TDD-verified fix pass, scoped scan, Review Focus 2026-09-16 15:51:50 -07:00
test-executing-plans-scripts.sh executing-plans v3: helper scripts, TDD-verified fix pass, scoped scan, Review Focus 2026-09-16 15:51:50 -07:00
test-helpers.sh fix(tests): stop the SDD skill test flaking on timing and prose case 2026-07-19 12:04:46 -07:00
test-sdd-workspace.sh feat(sdd): plan-scoped workspace — one .superpowers/sdd/<plan> dir per plan 2026-07-19 12:03:18 -07:00
test-subagent-driven-development-integration.sh Tighten cross-platform tool references 2026-06-16 10:09:43 -07:00
test-subagent-driven-development.sh fix(tests): stop the SDD skill test flaking on timing and prose case 2026-07-19 12:04:46 -07:00
test-worktree-native-preference.sh tests: annotate three kept bash tests with drill coverage notes 2026-06-16 10:09:43 -07:00
test-worktree-path-policy.sh fix: remove global worktree path fallback (#1476) 2026-06-16 10:09:43 -07:00

Claude Code Skills Tests

Automated tests for superpowers skills using Claude Code CLI.

Overview

This test suite verifies that skills are loaded correctly and Claude follows them as expected. Tests invoke Claude Code in headless mode (claude -p) and verify the behavior.

Requirements

  • Claude Code CLI installed and in PATH (claude --version should work)
  • Local superpowers plugin installed (see main README for installation)

Running Tests

./run-skill-tests.sh

Run integration tests (slow, 10-30 minutes):

./run-skill-tests.sh --integration

Run specific test:

./run-skill-tests.sh --test test-subagent-driven-development.sh

Run with verbose output:

./run-skill-tests.sh --verbose

Set custom timeout:

./run-skill-tests.sh --timeout 1800  # 30 minutes for integration tests

Test Structure

test-helpers.sh

Common functions for skills testing:

  • run_claude "prompt" [timeout] - Run Claude with prompt
  • assert_contains output pattern name - Verify pattern exists
  • assert_not_contains output pattern name - Verify pattern absent
  • assert_count output pattern count name - Verify exact count
  • assert_order output pattern_a pattern_b name - Verify order
  • create_test_project - Create temp test directory
  • create_test_plan project_dir - Create sample plan file

Test Files

Each test file:

  1. Sources test-helpers.sh
  2. Runs Claude Code with specific prompts
  3. Verifies expected behavior using assertions
  4. Returns 0 on success, non-zero on failure

Example Test

#!/usr/bin/env bash
set -euo pipefail

SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
source "$SCRIPT_DIR/test-helpers.sh"

echo "=== Test: My Skill ==="

# Ask Claude about the skill
output=$(run_claude "What does the my-skill skill do?" 30)

# Verify response
assert_contains "$output" "expected behavior" "Skill describes behavior"

echo "=== All tests passed ==="

Current Tests

Fast Tests (run by default)

test-subagent-driven-development.sh

Tests skill content and requirements (~2 minutes):

  • Skill loading and accessibility
  • Workflow ordering (spec compliance before code quality)
  • Self-review requirements documented
  • Plan reading efficiency documented
  • Spec compliance reviewer skepticism documented
  • Review loops documented
  • Task context provision documented

Integration Tests (use --integration flag)

test-subagent-driven-development-integration.sh

Full workflow execution test (~10-30 minutes):

  • Creates real test project with Node.js setup
  • Creates implementation plan with 2 tasks
  • Executes plan using subagent-driven-development
  • Verifies actual behaviors:
    • Plan read once at start (not per task)
    • Full task text provided in subagent prompts
    • Subagents perform self-review before reporting
    • Spec compliance review happens before code quality
    • Spec reviewer reads code independently
    • Working implementation is produced
    • Tests pass
    • Proper git commits created

What it tests:

  • The workflow actually works end-to-end
  • Our improvements are actually applied
  • Subagents follow the skill correctly
  • Final code is functional and tested

test-worktree-native-preference.sh

RED-GREEN-REFACTOR validation for the using-git-worktrees skill (~5 minutes):

  • RED: skill without Step 1a — agent should use git worktree add
  • GREEN: skill with Step 1a — agent should use the native EnterWorktree tool
  • PRESSURE: same as GREEN under urgency framing with pre-existing .worktrees/
  • Drill scenario worktree-creation-under-pressure.yaml covers the PRESSURE phase only

Adding New Tests

  1. Create new test file: test-<skill-name>.sh
  2. Source test-helpers.sh
  3. Write tests using run_claude and assertions
  4. Add to test list in run-skill-tests.sh
  5. Make executable: chmod +x test-<skill-name>.sh

Timeout Considerations

  • Default timeout: 5 minutes per test
  • Claude Code may take time to respond
  • Adjust with --timeout if needed
  • Tests should be focused to avoid long runs

Debugging Failed Tests

With --verbose, you'll see full Claude output:

./run-skill-tests.sh --verbose --test test-subagent-driven-development.sh

Without verbose, only failures show output.

CI/CD Integration

To run in CI:

# Run with explicit timeout for CI environments
./run-skill-tests.sh --timeout 900

# Exit code 0 = success, non-zero = failure

Notes

  • Tests verify skill instructions, not full execution
  • Full workflow tests would be very slow
  • Focus on verifying key skill requirements
  • Tests should be deterministic
  • Avoid testing implementation details