Pull-API runner: a Claude routine as the unattended-run host¶
Where this fits: this is the pull variant of the unattended cloud runner — instead of scheduling one routine per named spec,
mship item run-nextselects and claims the next eligible item from a backlog. How the worker authenticates is an independent choice (see Prerequisites below): any of the three auth models works with this selection model.
This documents one concrete host for the unattended runner (spec
unattended-runner, AC8): a Claude routine on a cron schedule (or the
cron+claude -p pattern, or a scheduled cloud agent) that drives one tick of
mship's pull API and hands the resulting prompt to an agent turn. The design
is host-agnostic — mship never spawns an agent and doesn't know or care what
runtime calls it — so everything here generalizes to any scheduler that can
run shell commands and invoke an LLM agent. This file is the reference
implementation for that contract, plus a smoke-test checklist for verifying a
deployment actually claims/emits/exits the way this doc says it will.
All commands below were verified against this branch's CLI (mship --help,
mship item run-next --help, mship item bail --help, etc.) rather than
assumed — see the smoke test section for the transcript.
The contract, in one paragraph¶
Split control plane from execution plane. mship is the host-agnostic
control plane: it selects the next eligible WorkItem, claims it (git-backed,
so ephemeral/serverless runs coordinate through nothing but a git remote),
and emits a self-contained dispatch prompt. The agent runtime — Claude
routine, cron, CI, whatever — is the execution plane: it runs that prompt,
calling mship commands as it works, and opens a PR. It never merges.
On an unresolvable fork or unfixable failure it calls mship item bail,
which records the reason, marks the item blocked, and releases the claim so
a later tick (or an attended human) can pick it up. One item per tick (v1 is
one-at-a-time; no parallel runs).
Prerequisites¶
mshipinstalled and onPATHin the routine's execution environment.- A git remote the routine's environment can both fetch from and push to — the run-state ref (see below) is committed and pushed as checkpoints, so a read-only clone is not enough.
- Two environment variables, set on the routine (not mship flags — these are
adapter-level inputs the routine's shell body needs before it can call
mshipat all): - GitHub auth for
bootstrap/finish— pick ONE of the three auth models (full comparison: the runbook's Choosing your setup):- Raw env token — set
GH_TOKEN(orGITHUB_TOKEN;--tokenwins over both,src/mship/core/gh_auth.py). Simplest; only for a trusted execution environment, since the worker holds a real GitHub credential. - The
/gh-tokenbroker — setMSHIP_GH_BROKER_URL+MSHIP_SERVE_TOKENand the worker pulls short-lived repo-scoped tokens frommship serveat the moment of use (cloud-agent-auth.md §1). - Attach-at-relay — the worker holds no GitHub credential at all;
mship bootstrap --relay-url … --run-token …routes git through the credential-attaching egress proxy (cloud-worker-auth-spine.md). Note the PR-open caveat: with only a run token,finishmust run--push-onlyand the PR is opened in an attended step (runbook: Opening the PR).
- Raw env token — set
WORKSPACE_GIT_URL— the git URL of the workspace meta-repo (the one containingmothership.yaml). This is an adapter convention, not an mship flag:mship bootstraphas no positional "clone this workspace" argument — it only clones the member repos a workspace already declares (from each repo'surl:or the workspace'sdefault_remote:inmothership.yaml). Gettingmothership.yamlonto disk in the first place is a plaingit clone, which is the routine's job, not mship's.
Eligibility: what makes an item pick-able¶
mship item run-next selects the oldest-first candidate where all of:
unattendedistrueon the WorkItem (opt-in; set viamship item unattended <id> --onor the Ground Control checkbox on the WorkItem — both flip the same flag),- its phase is
ready(v1:phase_override == "ready"), - it has a linked spec whose status is
approved, - it is not currently claimed on the run-state ref.
If nothing matches, the tick is a clean no-op.
The routine, one tick¶
set -euo pipefail
# --- environment the routine must provide ---
# <auth> ONE of: GH_TOKEN / MSHIP_GH_BROKER_URL+MSHIP_SERVE_TOKEN
# / relay --relay-url+--run-token (see Prerequisites).
# WORKSPACE_GIT_URL git URL of the repo containing mothership.yaml.
WORKDIR="${WORKDIR:-/work/workspace}"
# 1. Materialize the workspace meta-repo. Ephemeral cloud runs assume no
# persistent disk between ticks, so this is a fresh `git clone` every
# time; on a persistent host it's a no-op `git pull` instead. Either way,
# end up at a directory containing mothership.yaml.
if [ ! -f "$WORKDIR/mothership.yaml" ]; then
git clone "$WORKSPACE_GIT_URL" "$WORKDIR"
fi
cd "$WORKDIR"
# 2. Materialize the workspace's member repos. Idempotent: an already-
# present member reports "present" and is left untouched (no-clobber),
# so re-running this every tick on a persistent host is safe too.
mship bootstrap
# 3. Pull: claim + emit the next eligible item's dispatch prompt.
# Non-TTY output is JSON: {"runnable": true, "item_id", "prompt"}
# or {"runnable": false}.
result="$(mship item run-next)"
runnable="$(printf '%s' "$result" | jq -r '.runnable')"
if [ "$runnable" != "true" ]; then
echo "nothing runnable this tick"
exit 0
fi
item_id="$(printf '%s' "$result" | jq -r '.item_id')"
prompt="$(printf '%s' "$result" | jq -r '.prompt')"
# 4. Hand the prompt to a fresh agent turn. This is the execution-plane seam
# — swap this one line for any agent runtime. For the cron+`claude -p`
# pattern:
#
# claude -p "$prompt" --permission-mode acceptEdits
#
# The agent turn (see "What the agent does" below) either finishes cleanly
# (tests green, PR opened) or calls `mship item bail "$item_id" --reason
# "..."` on a fork/failure. Either way this script's job ends when that
# turn returns — one item per tick, no retry loop, no second run-next call.
What the agent does with the prompt¶
The emitted prompt is spec-first: it renders the WorkItem id, the linked
spec's ## Problem section, and its acceptance criteria, then instructs the
agent to "Implement this work item to satisfy its approved spec, then finish
per workspace conventions." If the item already has commits from a prior
(bailed or interrupted) run, mship item run-next prepends a ## RESUMING
prior run preamble naming the branch, commits-ahead count, and the last few
journal lines, and tells the agent not to restart.
Concretely, the agent:
- If the item has no linked task yet (a fresh pickup), spawns one and links
it in the same step:
mship spawn "<title>" --work-item <item_id> --yes(--yesmatters — the routine's shell is never a TTY, and spawn requires it to skip confirmation prompts abovespawn_confirm_threshold). This also links the new task to the WorkItem automatically, which is what lets a later tick's resumable-dispatch wrapper find the branch and journal. - Works the task using the normal mship loop:
mship contextfor a workspace snapshot, edits,mship testuntil green,mship journalto record progress,mship askif it needs a non-blocking decision surfaced to the phone. - On success:
mship finish— creates the PR.mship finishstill enforces its existing gates (approved spec, audit, passing tests); it is never bypassed by the unattended path. - Never merges the PR and never pushes to the base branch directly. A human reviews and merges in the morning.
- On an unresolvable fork (a design decision it can't make unattended) or a
test failure it can't fix:
mship item bail <item_id> --reason "<why>", then stop. It does not retry, does not pick a different item, and does not merge anything.
Rules¶
Never merge. Unattended execution stops at an open PR every time — spec
non-goal, not a bug. mship finish opens PRs; nothing in this loop calls
git merge, gh pr merge, or pushes to the base branch.
Bail, don't block. mship item bail <id> --reason "<reason>"
(checkpoint_bail, src/mship/core/runner.py) does four things in this
order: logs the reason to the item's run-log (durable even if later steps
race), pushes the task branch to origin (best-effort, so the work survives
for a later resume — even if this host is ephemeral), marks the item's linked
task(s) blocked_reason, then releases the claim (authoritatively — see
"Long runs and cross-process claim ownership"). A blocked item is excluded
from run-next selection until the block is cleared, so a bailed item is not
re-picked every tick. The branch is left intact — a bail is a checkpoint, not
a rollback; a later tick (or an attended human) can resume off exactly where
it stopped.
One item per tick. The routine calls mship item run-next exactly once
per invocation. It does not loop internally to drain the backlog — draining
N eligible items takes N scheduled ticks. This is deliberate for v1 (no
parallel runs, see the spec's non-goals); a farm/concurrency model is future
work.
Long runs and cross-process claim ownership¶
The run-state claim's holder token is hostname:pid (_run_holder(),
src/mship/cli/workitem.py), minted fresh by each CLI invocation.
mship item run-next, mship item heartbeat, and mship item bail are
always separate mship process invocations in real use, so a later
invocation's pid never matches the pid that made the original claim. Both
the release and the heartbeat handle this by acting authoritatively —
they read the claim's recorded holder off the run-state ref and act as that
holder rather than as this process's fresh token:
-
bailreleases the claim immediately.checkpoint_bail(src/mship/core/runner.py) logs the reason, pushes the task branch to origin, marks the item blocked, then reads the recorded holder and releases the claim as it. The claim is gone the momentbailreturns — no TTL wait. (Independently, a bailed item carries ablocked_reasonon its task, andrun-next's selector excludes blocked items, so a just-bailed item is not re-offered until a human/decision clears the block — regardless of the claim.) -
Long runs stay claimed. The claim TTL defaults to 4 hours (
RunStateRepo, was 30 minutes), so a typical overnight build does not expire and get double-run mid-flight. For runs that may exceed that, the host callsmship item heartbeat <id>periodically — it reads the recorded holder and advances the heartbeat as that holder, resetting the TTL window. -
Resume survives ephemeral hosts. Because
bailpushes the branch to origin, a later tick — even a fresh clone on a brand-new host with no local worktree — detects the prior work:run-nextreads commits-ahead from the remote branch (_remote_commits_ahead,src/mship/cli/workitem.py) and wraps the prompt with the## RESUMINGpreamble so the agent continues instead of restarting.
Smoke test checklist¶
Run this in a throwaway workspace (not a real one — it creates a scratch
item + spec) before trusting a new deployment of this adapter. It exercises
exactly the two outcomes mship item run-next can produce.
Setup: a workspace with a git origin remote configured (the run-state ref
needs somewhere to push to — even a local bare repo works for this check).
-
No eligible item →
{"runnable": false}. In a workspace with nothing unattended/ready/approved yet:Exit code 0 either way — this is a normal empty tick, not an error.mship --json item run-next # {"runnable": false} -
One approved + unattended item → prints a prompt and records a claim.
Confirm theITEM_ID=$(mship item new "Smoke test item" --kind feature) mship spec new --title "Smoke test item" --id smoke-1 printf '%s' '{"problem":"p","user_story":"u","approach":"a"}' \ | mship spec apply smoke-1 --from-json - mship spec approve smoke-1 --bypass-gate # or a real review pass mship item link-spec "$ITEM_ID" smoke-1 mship item phase "$ITEM_ID" ready mship item unattended "$ITEM_ID" --on mship --json item run-next # {"runnable": true, "item_id": "<ITEM_ID>", "prompt": "# Unattended run: ..."}promptfield mentions the item id and the spec's Problem text. -
The claim actually holds (proves step 2 didn't just print a prompt but also recorded a claim other runs will respect): immediately invoke
run-nextagain, as a separate process (a new shell, not a loop in the same script) — a real second tick would be exactly this.With only one eligible item in the backlog, a second concurrent/immediate pull must come back empty — the first claim is still live. If this instead returns the same item again, the claim isn't being honored and something regressed.mship --json item run-next # {"runnable": false} -
Bail releases immediately and records the reason.
Confirm the item shows as blocked (if it has a linked task,mship item bail "$ITEM_ID" --reason "smoke test" # {"item_id": "<ITEM_ID>", "bailed": true, "reason": "smoke test"}blocked_reasonis set on that task) and the reason string appears somewhere retrievable for the operator. The run-state claim is released the momentbailreturns — even thoughbailruns in a different process from the originalrun-next(the release is authoritative; see "Long runs and cross-process claim ownership" above).run-nextstill won't re-offer this item, but now because it is blocked, not because of a lingering claim — the selector excludes blocked items until the block is cleared. -
Heartbeat keeps a long-running claim alive (optional).
Run against a live claim, it advances the claim'smship item heartbeat "$ITEM_ID" # {"item_id": "<ITEM_ID>", "heartbeat": true}heartbeat_at(as the recorded holder) so the run isn't reclaimed at the 4h TTL.{"heartbeat": false}means there was no live claim to advance.
This exact sequence (minus the throwaway IDs) was run against this branch's CLI while writing this doc and produced the outputs shown above.