Upgrading
mcptask_runner update --self replaces the binary; then run a bare mcptask_runner update in any one project on the host, while no runner is working that checkout — this version changes the installed data pack (the new pr skill, per-host baseline permissions, ci_wait, the test-lock helpers), and the binary does not rewrite those on its own.
Bitbucket Cloud projects can be driven
git_host: bitbucket now selects a real adapter instead of an error. The runner
opens, finds, reads, merges and checks pull requests on Bitbucket Cloud through
its REST API 2.0 — there is no Bitbucket CLI to shell out to — and
mcptask_runner pr works there unchanged, because it goes through the same
interface gh does.
What to do on a Bitbucket host. Put a credential in the environment and
re-run mcptask_runner init, which writes it to
~/.mcptask_env.d/bitbucket_credentials (0600) beside the mcptask token, so it
reaches both your own shell and the scheduled job. One of:
BITBUCKET_ACCESS_TOKEN — a repository or workspace access token, sent as
Bearer. The one to use for a machine.
BITBUCKET_EMAIL + BITBUCKET_API_TOKEN — the Atlassian account's email (not
a Bitbucket username) and an API token, sent as Basic.
App passwords are not accepted. Atlassian stopped serving them on 9 June
2026; BITBUCKET_USERNAME / BITBUCKET_APP_PASSWORD are recognised only so that
init and the adapter can say so, rather than authenticating with something that
fails at merge time, after the work is done.
Setting both schemes at once is refused, naming both, and writes nothing — no
credential is ever picked on your behalf. A re-run of init with nothing set
leaves an existing credential file alone; setting the variables again is how a
rotated token is installed. Values are never printed, by init or by a 401.
What a GitHub host has to do. Nothing. gh still drives it, argv for argv.
One mapping worth knowing. Bitbucket's SUPERSEDED counts as DECLINED, not
MERGED: those commits may well be on the destination branch, but this pull
request was not merged, so an auto-squash run over one ends merge_unverified
and a person looks.
How far this is tested. The conformance suite now drives a Bitbucket project
in both modes against a scripted host, on all three coding CLIs: a manual run
that opens a pull request and puts it on the card, an auto-squash run whose merge
is confirmed by a GET on the pull request rather than by the agent's claim, a run
whose CI fails twice and leaves the pull request open, and the same recording as
that third one with the host answering OPEN — which comes out
merge_unverified. The last two are a pair on purpose: a runner that had stopped
verifying would pass the confirmed one and fail the other.
One environment variable comes with that, and is worth knowing about even though
almost nobody needs it. BITBUCKET_API_BASE_URL points the adapter at an API
root other than api.bitbucket.org — for a host behind a proxy that terminates
Atlassian's API on its own name, and for the conformance suite, which uses it to
reach a fake host on loopback. Unset means Atlassian's own root; a wrong value
fails at the first request, naming the URL it could not reach.
A once_dry run on a project with no assigned work no longer files a bug
once_dry used to have one ending in its prompt: success. So a dry run against
a project whose queue was empty had nothing to answer with — the child fetched
the next piece, was told "No available tasks or stories found", went looking for
one by hand, and then reported status: "error", which is a hard failure. The
runner filed a high-priority piece in its own Errors epic about a project that
was simply idle, and did so again every time that project was dry-run.
An empty queue is now the mode's own answer: the run ends no_more_tasks, the
log says No tasks available, the dry run had nothing to display, and nothing is
filed.
What to do. Nothing on the host — the change is in the prompt the runner
composes, so it takes effect with the new binary. Any pieces already filed under
Runner error: result — <project> for an idle project can be closed.
An API refusal is now named instead of retried
When the coding CLI ends its turn because the API refused it, the runner stops
and says which refusal it was. It no longer reads that as "the child forgot the
result marker", so the three --continue retries thirty seconds apart are gone —
none of them could ever have succeeded — and the ending is no longer filed as
stream_ended.
Four endings, and what each asks of you:
Termination
What happened
What to do
model_unavailable
the configured model is not being served — retired, removed, or spelled differently by this provider
fix
model: in
config/mcptask_runner.yml; the message names the model and quotes the provider
not_authenticated
the CLI has no credentials
sign the CLI in as the user the runner runs as on that host
usage_limit
the account is past a hard cap that clears on a date
nothing, until that date — the message carries it verbatim
api_error
the API refused in words no harness profile recognises
read the quoted text; if it is worth naming, add its phrase to the profile's
usage_limit /
model_unavailable /
not_authenticated list
What changes for an operator. usage_limit files no bug piece — it is a
budget, like the day's quota and the usage window before it, and an account with a
monthly ceiling would otherwise file one piece a month. The day still ends,
because nothing can be spent until the cap clears. The other three do file a
piece, and it names the condition and the remedy rather than saying the stream
ended. The dashboard card and the run log's error_message now say the same
thing, where all three of the reports that prompted this carried null.
What does not change. A 429 still gets its eight patient waits and a 529 its
ten — a refusal the profile could name outranks them, an unnamed one does not — a
context overflow is still an overflow, and a turn that failed for a reason a
resume can fix still takes the marker retry it always had.
Recognition is Claude Code's wire shape for now, that being the only one these
conditions have been recorded on. A Codex or OpenCode host hitting the same wall
still spends the marker retries.
The lock guard no longer refuses commands that merely name a test run
check_test_lock — the PreToolUse hook that stops a second suite starting
beside somebody else's — decided what a test run was by looking for a recognised
invocation anywhere in the command text. A space counted as the start of a
command, and every word inside a quoted string has one in front of it, so a
command that only talked about a test run was refused while a suite was
running:
git commit -m 'fix: bin/ci now takes the lock' BLOCKED, and it runs nothing
grep -q '=== bin/ci exit:' "$LOG" BLOCKED, and it is the line inside /ci-wait
The same anchor was wrong in the other direction, which is the half that
mattered: an invocation had to be preceded by whitespace or nothing at all, so
../other-worktree/bin/ci — how a story branch runs the gate — matched nothing
and a genuine second suite was waved straight through.
The command text is now read the way a shell reads it, into words, and only a
word in command position can be a test run: the first word of the line or of
a new command (after &&, ||, ;, |, &, a newline, (, `, $(),
optionally behind a wrapper that is not the command itself (sudo, env,
time, nohup, an option to one of those, a VAR=value assignment) or an
interpreter that introduces one (bash bin/ci, sh -c 'bin/ci', and
python -m pytest, which was not recognised before). Quotes and # comments
are tracked, and a path-bearing invocation counts. The
exemption for the lock's own helpers (test_lock, run_with_log) is decided
the same way, so a commit message naming one of them no longer exempts a real
test run sharing the line.
Blocking is unchanged for everything that actually runs a suite: bin/ci,
bin/ci --fast, cd worktree && bin/ci, and every command the project declares
in .claude/test-commands.json.
What to do. Run mcptask_runner update on each host — the hook is a helper
in ~/.claude/bin, not a file in any checkout, so a project that only pulls the
new binary keeps the old guard until the helpers are reinstalled. Nothing else
changes, and git commit -m with bin/ci in the message works again.
The machine-wide test lock: reinstall the helpers on every host
~/.claude/bin holds the copies of the helper scripts that actually run, so
none of the fixes below reaches a machine until its helpers are reinstalled —
mcptask_runner update on each host. A repository carrying the fix is not a
fixed machine, and the lock is the one piece of tooling where the difference is
invisible: a host still running the old test_lock keeps serialising its suites
exactly as badly as before, and says nothing about it.
Run the update while no suite is in progress on that host. The scripts are
replaced in place, and a run that is mid-flight through the old one keeps the
file it started with.
test_lock records COMMAND_PID and LOGFILE on Linux as well as macOS
set_command_pid and set_logfile edited the lockfile with sed -i and an
empty backup suffix as a separate argument, which is the BSD spelling: GNU sed
reads that empty string as the script, the real script as a filename, exits 4
and writes nothing. Both call sites discard the failure, so on Linux the lock
never gained either field and nothing said so.
Three consequences, all of them silent, all of them gone now: a lock whose
COMMAND_PID stayed empty was reapable while its suite was still running; ci_wait
could not tell a crashed run from a slow one, because the check for that reads
COMMAND_PID; and a LOCKED verdict could not name the other run's log, so
OTHER_LOG was always unknown on that platform.
Both fields are now written by filtering the key out of the file and renaming a
temporary copy over it — no in-place edit, nothing to spell two ways, and a
reader can no longer catch the file half-built. (task
https://mcptask.online/jchsoft/tasks/12402)
Every run through the helpers is ten seconds shorter
run_with_log started its stall watchdog with a command substitution, and a
backgrounded subshell inherits that substitution's stdout — so the line meant to
start a watchdog and move on blocked until the watchdog's next sleep returned.
Measured against a command that exits instantly: the command was gone at 0.03 s
and the Exit code: footer arrived at 10.8 s. Every run paid it, whatever the
command, because the wait, the tee teardown and the footer all queued behind
that line. The same measurement is now about a second.
/ci-wait and /test-wait answer sooner as a result, and so does anything
watching for the footer. ci_wait's 15-second footer grace is unchanged and now
pure headroom rather than a measured need — nothing pays it on a healthy run,
since it returns the moment the footer appears.
Two more corrections in the same file: the stall watchdog sized the log with
stat -f%z alone, which is BSD-only, so on Linux it read every log as zero
bytes and dumped thread traces into a run that was not stalled; and the watchdog
subshell's own stray output now lands in the log instead of on the launcher's
stderr.
run_with_log is also now maintained in this repository rather than vendored
from the retired Ruby gem, which is what made the fix possible here at all.
(tasks https://mcptask.online/jchsoft/tasks/12401,
https://mcptask.online/jchsoft/tasks/12402)
A lock is now held for as long as the run, and released only by that run
This is the behaviour change to read before updating a shared host. The lock did
not protect what it appeared to protect, in three independent ways, and all three
were silent while they happened.
A lock is stale when nothing it stands for is alive, and nothing else makes it
stale. Age is no longer consulted at all. Before, age >= 900 was tested
first, ahead of any look at the command, so a suite that ran past fifteen
minutes lost its lock while it was working — and a full bin/ci runs to about
that mark. A lock whose command is alive now keeps it at any age; a lock whose
recorded command has died is over immediately.
An empty command pid is no longer a ten-second fuse. acquire used to leave
COMMAND_PID empty and fall through to a pidfile only run_with_log ever
writes, so a lock taken around anything else — a bare bin/ci, an operator
being a good citizen — was reapable from its tenth second, permanently, with no
warning to either side. acquire now records the process that asked for the
lock, and the sentinel's 900 s backs that up, so an unregistered lock lasts as
long as the shell that took it and then as long as the sentinel. The pidfile
branch is gone, and with it a path whose two sides derived the same filename from
different places.
A release now has to prove it is the run that acquired.
release_if_owner compared CALLER and PROJECT — a constant per skill and a
constant per worktree, both reused by a re-run after a rebase — so a run whose
own lock had already been reaped could reach its cleanup and release the NEXT
run's lock, two minutes into a suite that had done nothing wrong. acquire now
issues a token, prints it after ACQUIRED and records it; release_if_owner
<caller> [instance] takes the token, or the command pid, or the holder pid, and
answers not-owner when the instance is not this lock's. run_with_log and
bin/ci pass theirs. A release with no instance still works, for callers that
predate the argument, and now says in as many words that it matched on the name
alone.
test_lock status says which case a lock is in. HELD_BY=command … (does
not age out), HELD_BY=holder … (no command pid recorded yet), HELD_BY=sentinel
… with the time left, or STALE=yes. The first line still starts with LOCKED
or FREE, which is what /wait-unlock reads.
The PreToolUse guard, check_test_lock, was rewritten to the same rules in the
same commit. It had its own copy of the old ones, so on both counts above it
would have waved a test command straight into a running suite while test_lock
was still refusing to hand the lock over.
Two side effects worth knowing. run_with_log now records in the log whether the
lock learned its command pid — the line reads Lock: COMMAND_PID set to …, or
Lock: no lockfile — this run is NOT serialised against other suites on this
machine, which is a legitimate state for a run nobody took a lock for and a
useful thing to be able to check afterwards. And the Windows flake in the
lock-guard tests around the ten-second cliff
(https://mcptask.online/jchsoft/tasks/11897) is gone with the cliff itself.
(task https://mcptask.online/jchsoft/tasks/11852)
Pull requests go through mcptask_runner pr, on whichever host the project is on
The prompts the runner composes used to say gh. Not always out loud, which was
the harder half: the auto-squash spine merged with gh pr merge --squash
--delete-branch and gated its own success on gh pr view --json state, while
the CREATE PULL REQUEST step named no command at all — and a step that names no
command gets gh anyway, because that is what a model reaches for when nothing
says otherwise. Either way a project on Bitbucket Cloud was being told to run a
binary it has not got.
Every one of those is now mcptask_runner pr create|list|view|merge|checks,
which asks the adapter git_host: resolved for that project (task
https://mcptask.online/jchsoft/tasks/12362 put the adapters in). The command
prints one JSON object on stdout, and the workflow tells the child to read
.pull_request.number out of it rather than from its own recollection. A new
bundled skill, pr, carries the five commands and their output shapes.
The runner reads that JSON too. The pull request a run opened used to reach
the dashboard through an event only Claude Code emits, so a Codex or OpenCode run
opened one and the card said nothing; and its NUMBER reached the runner only when
the child put it in the result marker, which about three quarters of real markers
do not. Both now come out of the command's own output, which every harness hands
back the same way. A result with no pr_number takes the one the run was seen to
open — the child's own answer still wins when it gives one, and nothing is
invented when no pull request was opened at all.
Permissions are now per host. Bash(mcptask_runner pr:*) is in the baseline
every project gets. Bash(gh:*), Bash(gh pr checks:*), Bash(gh pr view:*)
and the three github.com WebFetch domains are installed only into a project
whose git host resolves to github; api.bitbucket.org and bitbucket.org only
into a Bitbucket one. A project whose host cannot be worked out gets the shared
list and nothing else — nothing falls back to another host's approvals. Existing
entries are never removed, so a project that already has Bash(gh:*) keeps it.
The pull-request template default is GitHub's alone.
.github/pull_request_template.md is GitHub's own convention, and telling a
Bitbucket project to follow it meant the agent looked, found nothing, and wrote
whatever description it liked, with nothing reporting that the instruction had
missed. A git_host: github project still gets that path and that line, byte for
byte. Any other project gets whatever it declares under pr_template: path: and,
declaring nothing, gets no template line in the prompt at all.
What to do. Run mcptask_runner update on each host: the pr skill is a new
file under the project's skills directory and the permission split is a merge
into .claude/settings.local.json, neither of which arrives with the binary
alone. Run it while no runner is working that checkout — the new files would
otherwise land inside somebody's task commit. Nothing else is required, and a
project already carrying git_host: needs no config change; one that does not
gets its host read off git remote get-url origin at run time, and
mcptask_runner init writes it down for good.
Still naming gh: the two review executors (review and reviews). They
read review threads and enumerate every open pull request on a project, and
internal/prhost models neither — its six questions are create, list-for-task,
view, by-branch, merge and checks. Rewording them to name no host would leave an
agent told what not to use and not told what to use instead, so they are excluded
by name, with the reason in words, in TestSpineNamesNoGitHost. Extending the
interface to cover them is its own piece.
(task https://mcptask.online/jchsoft/tasks/12364)
/ci-runner reads a green run that prints no summary block
A bin/ci that prints one line per step and no aggregation at all — no CI
SUMMARY, no TOTAL RESULTS — matched neither shape ci_wait recognised, so
every green run on such a host ended in NO_SUMMARY_RECOGNISED plus twenty
lines of tail. That tail was not empty, which is what made it hard to spot: on
the run that filed this it happened to carry three of the nine step lines and
lost both of the ones a reader triages by, so the orchestrator reported a ragged
excerpt as if it were the run's summary and could give no step count or timings.
ci_wait now recognises a third shape, the per-step lines themselves:
✅ <step name> passed in 4.88s
❌ <step name> failed in 1m31.30s
They are emitted in order, ANSI-clean, in place of the fallback. A red run
carries both spellings — later steps keep running after one fails — and the
failure path already passed them through its tail filter untouched, so the same
shape now reads on both exits. NO_SUMMARY_RECOGNISED stays for a log that
matches none of the three, and now names all three in its message; it also no
longer spends its twenty lines on the log's own ==== rules, which the two
recognised branches had always dropped and it had not.
/ci-runner's skill body documents both shapes side by side instead of the
older one alone.
What to do. The fix is in ci_wait, which is installed onto the host rather
than compiled into the binary, so a new binary alone does not change it: run
mcptask_runner update in the project to refresh ~/.claude/bin/ci_wait and the
/ci-runner skill. (task https://mcptask.online/jchsoft/tasks/12413)
Changelog
Fixes
- 1b6b96a7af616625a981712fcd6db8537a053382 fix(dry): an empty queue is an answer, not a bug piece (task #12407)
- 1ac4ad2b9e8cadbecc497c93cf811781e469efd6 fix(eventstream): the throttle's clock is shared with the stream's own goroutines, so a test hands it over under a mutex (task #11688)
- b463017bb9094c5304ae4315ff932e9a242a9034 fix(executor): the API refusing a run is named, not retried for a marker it will never give (tasks #12405, #12406, #11850)
- a74dab9409d1133ef38a1a0a4fa2cbb388fd1e58 fix(helpers): a lock is held while its run lives, and given back only by that run (task #11852)
- 584a71dd709ce416d09035aed18e4a88bf10d492 fix(helpers): the lock records its command pid and its logfile on either userland (task #12402)
- 3b202a43d607cbc6957ba3faed2d39c1e3d61094 fix(helpers): the stall watchdog no longer holds the footer for a whole interval (task #12401)
- 50b0446ede6def664d97428f61ce8ac379fa1d66 fix(stress): an orphan may end itself, and the sweep says which ones did (task #12403)
### Other
- 4867ec85fccea0e4354333da18c647dd2b69ae3a Recognise an API refusal on every dialect, not just Claude Code's
- 90353f488c2a233cbe78e862f466798b6e278ca3 Skip the 0600 check for the Bitbucket credential file on Windows
- 0cb0ada0f2c2de1f328769c6bb44210d77f224b3 [#12362] Ask a PRHost, not gh: internal/prhost +
mcptask_runner pr
- e0a591aafc5e606230bddcebdde5d915b68456d3 [#12363] Drive Bitbucket Cloud: a prhost adapter over REST API 2.0
- 6333beb769954a50eb88a6b3d22bab9c2d2fb6f8 [#12364] Say
mcptask_runner pr, not gh: the prompt, a skill, and per-host approvals
- 0d46b797c5cbf4e178324077b0c412c663eaf486 [#12365] Drive a Bitbucket project in conformance, and say so in the README
- 1e4f48e9ef1deb729807cf4f4a679691ef1d3aa0 ci_wait: read a step record that comes with no summary block
- e4418e46f7a757cd60f5264b387b74cecf8cce12 docs(release): v0.3.24 shipped, the next-tag notes start empty
- 2425e84c9189223aafb21084742468d0f723f5f9 fix(check-test-lock): naming a test run is not running one (task #11674)
curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh
Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.