Agent-driven development¶
When you need this: an AI coding agent does the building, and you want it operating safely inside the workspace instead of improvising with raw git.
Install the skills¶
mship skill install
This installs the mship skill bundle for Claude Code, including
working-with-mothership — the canonical operating guide an agent loads
when it works in an mship workspace: how to resolve the active task, when to
journal, how phases and gates behave, how to finish.
What the guardrails give you¶
An agent in an mship workspace can't rationalize its way into the classic failure modes, because the boundaries are enforced, not suggested (Concepts):
- It can't edit or commit to
mainwhile a task is active — the pre-commit hook and the editor guard refuse, pointing at the task's worktree instead. - It can't ship undesigned features —
phase devandfinishcheck the work item, the approved spec, and the plan for feature work. - It can't lose the thread —
mship status/mship contextgive it structured state instead of shell archaeology, and the journal records what happened for the next session.
Orchestrator + subagents¶
The pattern that scales: one orchestrator session owns the task — spec,
plan, integration, mship finish — and dispatches fresh subagents to
implement one plan task each:
mship dispatch --task <slug> --plan-task 3
This prints a self-contained prompt: which worktree to cd into, the branch
state, recent journal lines, and exactly one plan task — so the subagent needs
no inherited context. Between subagents, the orchestrator reviews (spec
compliance first, code quality second) before dispatching the next.
Two rules of thumb:
- Parallel subagents are safe to run. Both halves are isolated: each task
gets its own worktree and branch, so code edits never collide, and every write
to
state.yamlgoes through a single read-modify-write under an exclusiveflockwith an atomic replace — so concurrentjournal/test/finishcalls cannot lose each other's updates to it. There are multiprocessing regression tests covering concurrent phase transitions and a same-slug spawn race.
Everything else those commands write is keyed by task, so different tasks
never contend: mship journal appends to logs/<task>.md, and a test run's
output, iteration JSON, and latest.json pointer live under
.mothership/test-runs/<task>/. Only the pass/fail status in state.yaml goes
through the lock; the run artifacts are written outside it.
That leaves one narrow case to avoid: two concurrent runs of the same task.
The test iteration number is chosen by scanning the run directory for its highest
number, with no lock, so two simultaneous mship test --task X calls can pick
the same iteration and overwrite each other's artifacts and latest.json. This
is not what parallel subagents normally do — they work on different tasks — but
do not fan two runs at one task and expect both results to survive.
What actually limits how many agents you can run is elsewhere: one inbox
listener per workspace (the mailbox lease refuses a second, so only one
agent per workspace can hear the phone), one mship serve per workspace,
the review loop — every feature needs an approved spec, and you are the one
approving — and CPU for the test suite when several agents run it at once.
(Ground Control's queue holds and merges items from every workspace; it shows
one card at a time as a presentation choice, so needing more review throughput
is not a reason to add workspaces.) Scale by adding workspaces, or by keeping one
orchestrator per workspace that fans work out to subagents.
- Commit early, journal always. mship journal "<what happened>" after
each meaningful step is what lets any future session — human or agent —
reconstruct the work without replaying it.
Sharing state with another session¶
mship export # bundle the task: journal, plan, spec, diffs
mship export --redacted # same, with opt-in secret redaction
Hand the bundle to a reviewer, a teammate, or another agent — everything needed to evaluate or continue the task, in one artifact.