menu_bookRunner documentation

Everything the runner does, in detail

The feature list and the CLI reference come from the catalog the runner publishes with each release, the changelog from its releases on GitHub. Below them, the technical detail: quota, context overflow, concurrency, logs, bug reports, models, self-update and project maintenance.

arrow_backBack to the runner overview
buildKey features

What the runner handles on its own

Each feature as the runner itself describes it in the catalog it publishes with every release.

From an assigned task to a merged pull request

since v0.1.0

The runner takes the tasks assigned to you on mcptask.online one after another. Triage rates each task first and picks one of three model tiers for it; then the coding CLI works it through on its own branch — code, tests, push, pull request, CI. In auto-squash modes the runner merges the pull request itself once CI is green and checks that the merge really happened; in manual modes the pull request waits for a person. Effort and progress are logged back to the task as it goes.

One failing task does not stop the day

since v0.3.22

A task that fails, turns out to be out of scope, or was already started by somebody else is set aside and the loop moves on to the next one. After a while the set-aside task may come back for another attempt, a limited number of times a day, so a passing hiccup does not cost it the day and a genuinely broken task is not paid for over and over.

Surviving a full context window

since v0.1.0

When the coding CLI's context fills up, the session is lost but the work is not — it is a branch with commits on disk. The runner restarts the CLI in a fresh session with its last few actions, and if that runs out of room too, it leaves a short handoff note for the next run: the branch, the commits, the uncommitted files and the unfinished plan, with a reminder to spend context carefully. A task whose pull request is already open counts as delivered, not as an error, and a resumed task runs on the strongest model tier.

A daily budget it reads, never guesses

since v0.1.0

How much work is left for today comes live from your mcptask.online account, and the runner checks it before a task, between tasks and every few minutes during one. When the budget is spent the work stops; the today modes end the day there, while daily sleeps until the next business day and starts again. An optional end of the working day stops new tasks at a set time, and an outage of the quota service is told apart from a spent budget and reported as a bug.

Watchdogs for a stuck or spinning run

since v0.1.0

A watchdog follows the coding CLI's output as it streams. It stops a run that has made no progress for too long or a command that hangs, and it recognises a run spinning in place — the same edit failing again and again, the same failing command, the same steps repeated — while leaving long CI and test waits alone. The task stays in progress and is picked up again later on the strongest model tier.

An unrelated urgent bug becomes its own task

since v0.1.0

When the coding CLI runs into an urgent bug that has nothing to do with its task, the work in progress is committed and pushed, the bug is filed as a new urgent task, and the runner fixes that bug first — even after a restart — before it returns to the original task.

One test suite at a time on a shared machine

since v0.1.0

The test and CI helpers the runner installs share one machine-wide lock, so two agents — or two projects — never run their suites side by side and slow each other into timeouts. A second suite waits for the first, and a lock whose owner is gone is reclaimed on its own. On top of that, one checkout never has two runners working in it at once.

A crash files its own bug task

since v0.1.0

When a run ends in an error, a crash, or a broken guard of the runner's own, the runner files a bug task on mcptask.online with the run log, the tail of the CLI's output and its configuration files with tokens removed attached. Identical failures are recognised and filed once, and a spent quota or a task reassigned to someone else files nothing. The one failure it cannot report is a missing or rejected token, so that one shows in the exit code and the log instead.

See what a run is doing, live and afterwards

since v0.1.0

Every run of the coding CLI gets a JSON run log that is opened the moment the CLI starts and refreshed while it works, so even a hung or crashed run leaves its state on disk; a process-wide log records everything else. While a run is in flight, a live card on mcptask.online shows its state, task, model and progress. Watching never stops a run: a dropped update costs freshness, not the result.

Updates you decide on, adopted without a restart

since v0.2.0

update --self installs the latest release after checking it against the published checksums, and puts the old binary back if the swap fails. The runner never upgrades itself — at the start of a run it only mentions that a newer release exists. Once you have installed a new binary, a long-running loop switches to it between two tasks, and daily switches on its way into the overnight pause, so an unattended machine does not keep running an old version.

Story branches — a whole Story in one merge

since v0.3.35

On unless a host opts out. The tasks of a Story each get their own branch and merge into a shared Story branch instead of the main one; the Story's last task merges the Story branch into main, so the main branch receives finished Stories rather than half of one. The runner names the branches itself and decides where each task starts, so the agent never has to guess. Manual modes never create or merge a Story branch.

A finished Story reaches main on its own

since v0.3.37

A Story whose every subtask is done gets merged into the main branch even when the task that finished it was worked in parallel, on another machine, or on one that does not use Story branches. Right after such a task, and between tasks in the auto-squash modes, the runner looks for finished Stories whose branch is still open and runs the Story's merge on its own — gate, pull request, host checks, merge. It then checks the repository itself, and a Story that still did not make it is reported as a bug, once.

Claude Code, Codex CLI or OpenCode

since v0.3.24

Each project chooses the coding CLI that drives it, once, with init --cli; the choice lives in a per-machine file, so two developers on the same repository can use different CLIs. The work loop, the watchdogs, the dashboard card, the effort log and the bug reports are the same whichever CLI runs. There is no default: a project that never chose is refused by name, and you can add a profile of your own.

Pull requests on GitHub, Bitbucket Cloud and GitLab

since v0.3.25

One command, mcptask_runner pr, opens, reads, checks and merges pull requests on whichever of the three hosts the repository is on, so the same instructions work everywhere. Before an auto-squash merge the runner asks the host for its CI verdict: a failure leaves the pull request open, a pipeline still running is waited for, and a repository with no host checks falls back to the project's own local gate.

Your own models, on a backend of your choice

since v0.1.0

The three model tiers resolve to whatever the harness profile names unless the project pins its own models. Together with a launcher command — local Ollama is the usual example — the coding CLI can run against a backend other than its vendor's. Pin all three tiers or none: a tier left unpinned asks that backend for a model it has never heard of.

A weekday job on macOS, Linux and Windows

since v0.1.0

init generates a scheduled job that starts the runner every weekday morning — a LaunchAgent on macOS, a systemd user timer on Linux, a Task Scheduler job on Windows — at 08:00 unless you pick another time. It does not switch the job on: starting something that spends a quota every morning stays your decision, and init prints the one command that does it.

new_releasesChangelog

What's new in the runner

The release notes of every runner version, as published on GitHub. The newest three are open; the rest are one click away.

v0.3.37

Upgrading

Upgrading starts with one command. mcptask_runner update --self replaces the binary. The bundled .claude/ assets did not change, so nothing on that account asks for a bare update in the projects — the sections below say whether anything else does. A runner idling in a wait keeps the old binary until its next task or a restart (launchctl kill SIGTERM, wait for the job to stop, then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.37 carries this binary; no wrapper changes.


Story branches are on by default

story_branches.enabled no longer has to be written: a host whose
config/mcptask_runner.yml lacks the key now works Stories on story branches
(#13501). A host that must stay
off says enabled: false. Nothing to do on a host that already had true.
Every host of one project should agree — a host that opts out merges its Story
tasks into main beside the others, unless the Story's branch is already live
(see below).

A live story branch binds every host

enabled: false now means "never start a story branch", not "ignore them"
(#13493). When a task's Story
already has a branch on origin, an opted-out host works the task on it like
any other host and logs this host opted out of story branches, but Story #N
has a live branch … — following it
. An opted-out host therefore reads the
task's piece and git ls-remote origin before an auto-squash or manual task,
and a failure there sets the task aside with a story_branch_plan_failed bug
piece, as on an enabled host. Nothing to configure.

A Story task whose verified pull request merged into a branch other than the
planned one now files a story_pr_wrong_base bug piece instead of passing
silently.

A finished Story is merged even when its last task did not know it was last

A Story whose subtasks were finished in parallel, on another host, or on a host
with story branches off used to sit on its live …/story branch until a person
opened the pull request into main
(#13492). The runner now runs a
story-merge run for such a Story: right after the task that finished it, and
at every task boundary of today_auto_squash, queue_auto_squash and
story_auto_squash, where it lists refs/heads/*/story on origin and merges
every Story whose subtasks are all completed or approved — on hosts that opted
out of story branches too. A pull request another host already opened is picked
up, not duplicated. Each Story is tried once per process; a branch still on
origin afterwards is filed as story_merge_failed, and a Story or listing that
cannot be read as story_sweep_failed. The run takes no slot of the daily task
budget and does not start after the end of the work window.

Nothing to configure. Expect, on the first day after upgrading, story-merge runs
for Stories that were finished and left unmerged before — look at origin for
*/story branches first if you would rather merge some of those by hand.


Changelog

Other

  • c693b3f5e78291253db5e242251017b272e149aa [#13492] A finished Story with a live branch is carried into main
  • 64995c943a486fe198ad335f112f8499a0623bb0 [#13493] A live story branch binds every host
  • f0f2876e3ec5a42edb251cd3589d80355a14900b [#13501] story_branches is on unless a host opts out
  • 639a6e78f7be7b78cd4308984877c1b7046c8faf docs(release): v0.3.36 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.36

Upgrading

Upgrading starts with one command. mcptask_runner update --self replaces the binary. The bundled .claude/ assets did not change, so nothing on that account asks for a bare update in the projects — the sections below say whether anything else does. A runner idling in a wait keeps the old binary until its next task or a restart (launchctl kill SIGTERM, wait for the job to stop, then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.36 carries this binary; no wrapper changes.


A feature catalog the website renders (Task #13451)

The public page at mcptask.online/runner was
written by hand and drifted from what the binary does. This repository now
carries catalog.yml: one entry per feature,
subcommand, flag, config key, harness, git host and exit code, each with the
release it first shipped in and a title and summary in English, Czech and
Slovak — the file the page will be rendered from.

internal/catalog loads and validates it strictly. A missing translation, an
unknown kind, a duplicate id, a flag without its subcommand or a since that is
not a release tag is an error that names the entry and the field, never a row
the page quietly drops.

What an operator has to do: nothing. The catalog is documentation shipped
with the source; the binary does not read it and no host behaves differently.

Every release carries catalog.yml as an asset (#13454)

From this release on, docs/catalog.yml — the runner's feature catalog — is attached to the GitHub release on jchsoft/mcptask-releases as catalog.yml. The latest one is at https://github.com/jchsoft/mcptask-releases/releases/latest/download/catalog.yml; mcptask.online reads it from there. bin/release now refuses to start without the file. Nothing for an operator to do.

The catalog and the code may not disagree (Task #13453)

A guard test, TestCatalogMatchesTheCode in internal/catalog/guard_test.go,
compares docs/catalog.yml with the code in both directions: every visible
subcommand and flag in the command tree, every config key the loader reads,
every bundled harness profile, every git host and every exit code has an
entry, and every entry names something that still exists. A readme_anchor
must land on a real README heading. bin/ci runs it, and a red run names the
id to add or remove. To list its exit codes from code rather than from a copy,
cli.Execute's codes are now named constants and the signal handler's list is
read through runner.SignalExitCodes; the codes themselves are unchanged.

What an operator has to do: nothing.


Changelog

Other

  • 547257ed8dec2ed0e229c1b406ce7022ad2e75e6 [#13451] Feature catalog: docs/catalog.yml and internal/catalog
  • bc62ec5003ff4f27129f4ca3c864805af8b25c3f [#13453] Guard test: the catalog and the code may not disagree
  • 8a8b40d54d6330144e28e8b01702eb1e29629080 [#13454] Publish docs/catalog.yml as a release asset on jchsoft/mcptask-releases
  • 3bcf928392c9b3e25c381b3d494103424b3704da docs(release): v0.3.35 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.35

Upgrading

Upgrading starts with one command. mcptask_runner update --self replaces the binary. The bundled .claude/ assets did not change, so nothing on that account asks for a bare update in the projects — the sections below say whether anything else does. A runner idling in a wait keeps the old binary until its next task or a restart (launchctl kill SIGTERM, wait for the job to stop, then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.35 carries this binary; no wrapper changes.


Story branches: a Story reaches main in one merge (opt-in) (Story #13009)

Until now every task merged into main on its own, Story subtasks included, so
main held half-finished Stories and a bug fix released from it shipped them
too. A project can now have the tasks of an unfinished Story merge into a
shared story branch instead, and main gets the whole Story in one merge when
its last task is done. Tasks outside a Story are unchanged.

It is off unless the project turns it on, in config/mcptask_runner.yml:

story_branches:
  enabled: true

A value that is not true or false is not read as off: a Story task then
refuses to start, the task is set aside and a bug piece names the key.

What the runner does with it on:

  • Names. A Story's branch is <story_id>-<slug>/story, its tasks' branches <story_id>-<slug>/<task_id>-<slug> — ASCII slugs of the titles, at most 40 characters. The runner decides both before the child starts and pastes them into the prompt; the child no longer makes up a feature/… name for a Story task.
  • Liveness is the remote. The story branch existing on origin (git ls-remote) is what keeps a Story open, whatever its state on mcptask.online. No branch and an approved Story → the task works against main, as before. No branch and an unapproved Story → this task creates the branch from main; two runners racing to create it end up on the same one. A branch is found by the Story's id, so renaming a Story does not cut a second branch.
  • Every Story task checks out the story branch, merges origin/main into it (so fixes from main reach the Story while it is open), cuts its task branch, opens its PR with --base <story branch> and squash-merges it there. The task link stays the LAST mcptask.online link in the PR body.
  • The last task — every other subtask completed or approved; a blocked one holds the Story — then opens a PR from the story branch to main whose body ends with the Story's link, merges it with a merge commit (mcptask_runner pr merge --squash=false) and deletes the Story's branches. The runner checks the remote afterwards: a story branch still on origin is filed as a story_merge_failed bug piece, and the task's own success stands. A finished Story whose last two tasks ran at once and so left the branch unmerged is filed as story_left_unmerged.
  • Manual modes never create a story branch. With one live, the PR targets it; without one, the task branch is still named under the Story's path and the PR targets main. A person merges.
  • mcptask_runner pr merge --squash=false now asks GitHub for a merge commit by name (gh pr merge --merge); it used to pass no method, which gh refuses without a terminal.

Host caveat. Only GitHub approves a piece when its PR merges (by the last
mcptask.online link in the body, whatever the base branch). On GitLab a merged
task is only set to 100 % ("Hotovo?"), and Bitbucket has no webhook at all —
either way a merged task counts as completed, which is enough for the last
task to be recognised, but nothing approves the Story there automatically. On
GitLab the story → main merge follows the project's merge method; a project set
to fast-forward only refuses the merge commit, and that is reported as
story_merge_failed.

Upgrading: nothing for projects that do not add the key. To turn it on,
add the two lines above to config/mcptask_runner.yml and commit them. Runner
hosts pick the feature up with the binary.


Changelog

Other

  • 1403b0679f1bfe14a19c5a231225e0efe9487d90 [#13010] mcptask client reads a piece with its parent Story and subtask states
  • c6576725026957eb7f002280c7f7ec4607148e64 [#13011] Story branch planner: naming, liveness by the remote, race-safe creation
  • abecf2e79ec5f82a982b3d31f564f903f9e5ebff [#13011] [#13015] Lint: story-branch planner and conformance sandbox
  • fbc3cdbbd5df90b43250a9467fcdddbb1ba68a25 [#13012] Auto-squash prompts take the planned branches: story branch in, main merged first
  • 3742459b7aaee8011d43e203d86bd0c4dea7c2f7 [#13013] The Story's last task is checked against the remote; GitHub merge commits by name
  • 83f4a1d7072706aefdb20c916309471d31c9ed2d [#13014] Manual mode names Story task branches under the story path
  • 3eb3a17b4c955ada3bb66623cc6585bb293d230e [#13015] Conformance scenarios, release notes and README for story branches
  • 4cd3f81e380a070f1c91d4a4fef5b48de75a5542 docs(readme): keep a Mac awake through launcher.command, not the plist
  • 8138647edb6cf2452abf73d22e882e573ea649e6 docs(release): v0.3.34 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

menu_bookOlder releases – the full changelogexpand_more

v0.3.34

Upgrading

Upgrading starts with one command. mcptask_runner update --self replaces the binary. The bundled .claude/ assets did not change, so nothing on that account asks for a bare update in the projects — the sections below say whether anything else does. A runner idling in a wait keeps the old binary until its next task or a restart (launchctl kill SIGTERM, wait for the job to stop, then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.34 carries this binary; no wrapper changes.


A host with Claude Code's read block on refuses to start (#13168)

With permissions.blockReadsOutsideWorkingDirectories turned on, Claude Code
asks the person before any shell command it cannot analyse (a $var in a loop,
a heredoc, ruby -e, python3 -) and before any read outside the project. A
scheduled run has nobody to ask, so those actions are refused even under
bypassPermissions. On the MacBook this ended five runs in two days
(#13149 …
#13167), and the directory grant
from #13137 cannot cover commands.

The runner now reads Claude Code's settings files at startup (user, project,
project-local and managed, in that precedence) and refuses to start when the
effective value is true. The refusal names the file. The runner exits non-zero
and files no bug piece, the same as a missing token.

Upgrading: update --self is enough. On a host that has the setting on,
remove it from the named file, or set it to false, before the next scheduled
run. Until then that host runs nothing and reports the reason in the launcher
log.

A Sonnet card shows the whole plan, not 0/1 (#13424)

The plan block asked for "one item per step of the plan", and each model chose
its own plan. Opus made one item per numbered step, usually seven. Sonnet 5 and
5.5 made one item for the whole task ("Fix + tests + PR + CI + merge"), so every
Sonnet run showed TODO 0/1 on its card from start to finish, on every host.

The block now asks for one item per numbered step under WORKFLOW, and at least
one for every step that ends in a "LOG PROGRESS NOW" line. It also forbids a
single item for the whole task. Claude Code and OpenCode get the new wording.
Codex said "the workflow below", which was wrong in the auto-squash prompts,
where the block sits inside step 2; it now names the section too.

Upgrading: update --self is enough. A run already in progress keeps its
prompt; the next task picks up the new wording.


Changelog

Other

  • b4b962943caa4ab00c98cb8091f6d5d1e2d8ba2a [#13168] A host with Claude Code's read block on refuses to start
  • 688ee72caa82d93cf55622da0a5d82c9c12efb75 [#13424] The plan block asks for one item per WORKFLOW step
  • adfbfd79c7fb882aa6c5728e0ad21fe7ebe5fff3 [#13433] Stress signal tests no longer go red after 17:00
  • c03bed998d70217b24e1266d0ce5c977d8893e01 docs(release): v0.3.33 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.33

Upgrading

Upgrading takes two commands. mcptask_runner update --self replaces the binary; then a bare mcptask_runner update in every project on the host, because the bundled .claude/ assets changed (baseline_permissions.json). Run it while no runner is working that checkout. A runner idling in a wait keeps the old binary until its next task or a restart (launchctl kill SIGTERM, wait for the job to stop, then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.33 carries this binary; wrapper changes in its CHANGELOG.


A run the host refused is filed as a bug, not set aside quietly (#13139)

Sometimes the coding CLI refuses an action because of its own settings: a
permission rule denies it, or a read outside the project is blocked. The child
used to end such a run as failure. The loop then set the task aside for the
day and nothing was filed. On 2026-09-25 that happened to projectoid_ii
#13062 on the MacBook. The same
thing would have stopped every task on that host, and nobody would have been
told.

The runner now remembers each refusal it sees in the stream and logs the first
one as The host refused an action: …. If the run then ends as failure,
error, ci_failed, preexisting_test_errors, stalled_for_genius,
merge_failed or merge_unverified, it is filed as a bug under the termination
host_refused, quoting the first refusal. A run that got past the refusal files
nothing. The prompt also tells the child to end such a run as "status":
"error"
with "reason": "host_refused". That way the refusal gets filed on
Codex and OpenCode too, whose refusal wording nobody has recorded yet, so their
failures.host_refused lists are empty.

For this ending the throttle's fingerprint leaves out the task, so a host that
refuses every task files one piece, not one per task. A focused stress day
filed 76 pieces before this change. The stress day plays the new
host_refused vector and fails if no bug piece names it.

Upgrading: nothing to do. Hosts pick it up with the binary. If you override
a harness profile, add failures.host_refused to your copy (see the bundled
claude.yml).

Test and CI logs stay readable when reads outside the project are blocked (#13137)

The bundled /test-runner and /ci-runner write their logs under
~/.claude/logs/projects/<hash>/. That is outside the project, so a host with
Claude Code's permissions.blockReadsOutsideWorkingDirectories turned on
refused ci_wait's read, and every task stopped at its first test run. On
2026-09-25 this stopped projectoid_ii's
#13062 on the MacBook.

The permission baseline now adds "~/.claude/logs" to
permissions.additionalDirectories in each project's
.claude/settings.local.json. It is merged like allow and deny, so any
directories the project already lists are kept.

Upgrading: run mcptask_runner update in each runner-driven project. Nothing
else to change, and the read block can stay on.

The end of the workday stops the next task, not just an empty queue (#13076)

work_window.end_of_workday_hour (and end_of_workday_minute) now apply in
today_auto_squash, the mode every scheduled job runs. Before this fix it was
checked only when the queue was empty. A runner with work queued went from one
task to the next past the hour. On 2026-09-24 a runner set to stop at 18:00
started new tasks at 18:20 and 18:54. On 2026-09-23 a triage that waited out a
rate limit came back at 21:09, and the task it picked ran until 22:40.

The clock is now checked each time a task would start: once before triage and
again after it, because triage can take hours. This applies in today,
today_auto_squash and daily. A task that is already running still finishes.
The log says Past end of workday (HH:MM) before triage or … after triage,
and the day ends with the new verdict end_of_workday. The queue, story and
single-task modes still ignore the clock, as before.

Nothing to change in configuration. Hosts pick the fix up with the binary.

Auto-squash no longer merges a pull request that no CI has passed (#13089)

Until now an auto-squash run merged its pull request even when the project had
no bin/ci and the git host reported no checks at all (NONE). In that case
no CI had looked at the change. From this release the runner merges only on a
green signal: the host's checks are SUCCESS, or they are NONE and the
project has a bin/ci that the run executed and passed.

The rule is enforced by the runner, not only by the prompt. In every
auto-squash mode the runner sets MCPTASK_MERGE_GATE=auto_squash for the coding
CLI, and mcptask_runner pr merge then checks the host's CI before it merges.
It refuses when the checks are FAILED, still IN_PROGRESS, cannot be read, or
are NONE in a project without bin/ci. The refusal exits non-zero with
merge refused by the auto-squash gate: …, names the status to report, and
merges nothing. This works the same on GitHub, GitLab and Bitbucket.

On GitHub, mcptask_runner pr checks now answers NONE for a pull request
with no checks. Before, it failed with gh's no checks reported on the …
branch
message, so on GitHub a project without CI never reached the NONE
case at all. Running
mcptask_runner pr merge by hand, and every manual mode, is not gated.

A project with neither CI now ends the task with the new status
merge_skipped_no_ci. The pull request stays open, the agent leaves a message
on the piece saying why, and progress stops at 80 %. The loop sets the task
aside for the day and moves on, and the log reads Task #N has a pull request
left open because the project has no CI (no bin/ci, no host checks) and needs a
human to merge it
. The runner card lists the task under set aside with the
reason merge_skipped_no_ci. To get such projects auto-merging again, add a
bin/ci or CI on the git host.

Nothing to change in configuration. Hosts pick the fix up with the binary.

A crash on a customer's runner files a bug piece in the customer's own project (#13090)

Before this change the runner filed a bug piece about its own failure only when
it ran under the jchsoft account. Runners under any other account stopped and
logged locally. Now a hard failure on a customer's runner files a bug piece
in that customer's account, in the project the run was working
(project_relative_id from the host's CLAUDE.md). The piece goes under the
project's bug-bucket Epic (the lowest-numbered Epic marked for errors), or at
the project root when the project has none. It never goes into jchsoft.

A customer's piece carries no raw agent output. The raw runner log is not
attached. Neither is anything else that can quote source code: prompts, tool
inputs and results, the agent's messages and thinking, stderr dumps, or the
values in the CLI's configuration files. The piece carries:

  • the task link, runner version, OS/arch, CLI and its version, model, mode, machine, session, the CLI's exit status and the error message (code-like lines removed);
  • the last tool calls, by name and outcome only;
  • the agent's error lines, sanitized and cut at 200 characters;
  • runner_events.txt, the runner's own log lines with child output removed;
  • a filtered copy of the run record;
  • cli_config_keys.txt, the setting names from the CLI's config files, with no values.

The piece also lists what was left out and gives the local paths of the full
runner log and run record on that machine. Reports under jchsoft are
unchanged.

The run now ends non-zero when it owed a report and could not file one, under
any account, jchsoft included. The log says so at the end:
… runner failure report(s) owed by this run … — ending the run non-zero.
When the server refuses (401/403, no access to the project, an expired trial),
the runner does not try again for the rest of that run. In daily mode a
refusal also stops the loop at the end of the day. A plain outage does not stop
it.

The runner's debug log now writes Tool finished: <name> (error) for a tool
call that failed.

A customer whose token cannot create pieces in its own project gets
REFUSED — this runner may NOT file bug reports and a non-zero exit. The fix is
to give that user project membership. Nothing to change in configuration. Hosts
pick the change up with the binary, so no bare update is needed.

mcptask_runner on PATH for a coding CLI you start yourself (#13123)

A Claude Code session you open by hand in a gem-installed project used to fail
on mcptask_runner pr … with exit 127 — the binary is
~/.mcptask/bin/mcptask_runner and nothing put it on PATH — and then went
hunting for the path in the launchd plist. Runs the runner starts itself were
never affected.

init and update now write ~/.mcptask_env.d/runner_path, which puts the
runner's directory first on PATH (once, however many startup files source the
directory). For the copy bundled inside the Rails wrapper gem that is
~/.mcptask/bin, never the gem directory. A binary not named mcptask_runner
(a worktree build) writes nothing and says so; Windows prints the directory to
add by hand.

To pick it up: bundle exec rake mcptask_runner:update (or
mcptask_runner update), then open a new terminal / restart the session.

A scheduled job set up through npx no longer points into npm's cache (#13095)

npx @mcptask/cli init used to write the scheduled job (LaunchAgent, systemd
timer, Task Scheduler job) with the path of the binary inside npm's cache, for
example ~/.npm/_npx/<hash>/node_modules/@mcptask/cli/vendor/mcptask_runner.
Once npm cleaned that cache or fetched a newer version, the job failed at its
next start on a missing file. The same happened with a project-local or global
npm install.

Now, when the runner is the binary from the @mcptask/cli npm package, init
copies it to ~/.mcptask/bin/mcptask_runner (mcptask_runner.exe on Windows),
the same path the Rails channel uses, and the job runs that copy. The output
says Installed this runner as …, or Replaced … when a different binary was
there before, since other projects' jobs on the same host may run that file.
Keep it current with ~/.mcptask/bin/mcptask_runner update --self.

Jobs written before this change still point into npm's cache. Regenerate them
with npx @mcptask/cli init --schedule --force in each project. A bare
update does not do it.

Elapsed times count the hours the Mac slept (#13132)

On macOS the runner measured every duration with a clock that stops while the
machine sleeps, Maintenance Sleep included. A Mac mini left alone sleeps several
times an hour, so the times the runner reported came out short. One session ran
from 21:09 to 21:42 and was logged as Elapsed time: 0.19 hours.

These now use the wall clock and include sleep:

  • Elapsed time: … hours, Execution finished in …s, and every after N h in the log;
  • the task_worked hours in the result;
  • elapsed_s in the run record, the bug piece and the overflow handoff;
  • elapsed_s of a running tool on the dashboard card.

The 30- and 60-minute "ask the queue again" waits now end at the clock time
they announce. Before, a wait that crossed a sleep ran later than that time.

The watchdog still ignores sleep. Its inactivity and hung-tool deadlines, and
the waiting for: Bash since Ns part of the heartbeat line, measure only time
the machine was awake. A child cannot work while the machine sleeps, so a run
that was healthy before a sleep is not killed as stalled afterwards.

The same change affects Linux hosts after a suspend. Nothing to configure.
Hosts pick the change up with the binary, so no bare update is needed.

A piece name with braces no longer stops the day (#13121)

The runner lost a well-formed TASKRUNNER_RESULT when a string in it held
unbalanced braces. The usual case is a task_name copied from a piece name such
as a raw i18n hash, {one: "až %{count_max} uživatel…. It found the end of the
JSON by counting braces, strings included, so the object never closed. The
triage was logged as Missing TASKRUNNER_RESULT marker, retried, then failed
with no task_id, and the loop stopped with A task failed, stopping. The next
run picked the same piece and stopped again. On 2026-09-25 this stopped a
runner after four tasks. The only workaround was renaming the piece.

The marker is now decoded by the JSON decoder itself, which skips string
contents. A marker the runner finds in the raw stream is also unescaped
properly, so one spread over several lines is read too.

When a marker is present but its JSON cannot be used, the log now says so
instead of calling it missing:

  • result marker found but its JSON never closed: {"TASKRUNNER_RESULT": … (or did not parse (<decoder error>)), with the first 200 characters;
  • TASKRUNNER_RESULT marker present but unusable (…) - will retry with marker-only instruction;
  • after the retries: Unusable TASKRUNNER_RESULT after retries exhausted.

A marker that is really missing is still logged as Missing TASKRUNNER_RESULT
marker
. The retries are unchanged. Nothing to change in configuration. Hosts
pick the change up with the binary, so no bare update is needed.

Adopting the bundle's runner now updates the binary the scheduled job runs (#13133)

On a Rails host, a runner whose project Gemfile.lock names a newer
mcptask-rails-runner adopts that bundled binary, at the start of a run and
between tasks. Until this fix it installed the binary into the project's own
bin/mcptask_runner instead of ~/.mcptask/bin/mcptask_runner, and then ran
it from there. The day's run was on the new version, but the LaunchAgent
(systemd timer, Task Scheduler job) went on starting the old binary in
~/.mcptask/bin the next morning, so every scheduled run had to adopt again.

Now the adoption installs into ~/.mcptask/bin/mcptask_runner
(mcptask_runner.exe on Windows), the file the scheduled job runs. init,
the scheduler and the runner-on-PATH export compute that path with the same
function. A home directory the runner cannot find is an error that says so,
never a path relative to the working directory.

Copies the old code left behind are not deleted. At the start of run, in a
project that bundles the wrapper gem, the log says
[Adopt] <project>/bin/mcptask_runner looks like a runner binary adopted into
the project by mistake
. If that file is not the project's own, delete it. It
is about 10 MB. Check that it was never committed, because not every project
gitignores bin/mcptask_runner.

Nothing to change in configuration, and no bare update needed. Hosts pick
the fix up with the binary.


Changelog

Other

  • 21e8e9360eecd096c8b92827dc2b6776eeabe09d [#13076] No task starts after the end of the workday, and the stress day proves it
  • 6e96bc636e2938067a4fa825091ebf29e378f8a1 [#13088] The Bitbucket install tests no longer read the host's credentials
  • e7f96ff03899d44f7d8a9ef8f7bef2ed27452671 [#13089] Auto-squash merges only on a green CI signal, and the runner enforces it
  • 642d79ce7b4ee805e6359d46830ad4815c33c1f2 [#13089] GitHub checks with none reported answer NONE instead of failing
  • 4cb249304a5e7604346d9ad91f784ec175da0888 [#13090] A customer's runner files its crash into its own project, without raw agent output
  • ff34c2a772936bdebef01b25d8b6bcd505afcc47 [#13095] A job scheduled through npx runs a stable copy, not npm's cache
  • 5d754ec579a491ff7a2cfd3fc5e4b32bc9a43393 [#13121] A marker whose strings hold braces is decoded, not lost to a brace count
  • c55e20219b28d57cc57a67e88ac1e661e4133144 [#13123] init and update put mcptask_runner on PATH for sessions started by hand
  • 68be0fe333a3e29a48c21896c06cf8143b0f9267 [#13128] The updater resolves the mcptask token through an injected environment
  • 926062e754ed01351a51670424868f0a011175d9 [#13132] Reported durations count host sleep; the watchdog still does not
  • 44d445dfc19f779cd26a8a6a2217aec8d2f18ab7 [#13133] Adoption installs into ~/.mcptask/bin, the binary the scheduled job runs
  • fc43f98a756ab29a243bd1dce1d19d86af01f7a7 [#13133] [#13123] The Windows test run no longer assumes unix paths or an env directory
  • eae7ae45b239b386e86ca62f7bfe9c80400909d6 [#13137] The installer grants ~/.claude/logs so a read block cannot stop test and CI verification
  • 42cfa4cc388f4093557c46bf7afc7d1b890a32a5 [#13139] A run the host refused is filed as a bug, not set aside quietly
  • da351b1452fa33e03c2901533aaefe0e6a9143ef docs(release): v0.3.32 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.32

Upgrading

Upgrading starts with one command. mcptask_runner update --self replaces the binary. The bundled .claude/ assets did not change, so nothing on that account asks for a bare update in the projects — the sections below say whether anything else does. A runner idling in a wait keeps the old binary until its next task or a restart (launchctl kill SIGTERM, wait for the job to stop, then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.32 carries this binary; no wrapper changes.


A second bundle update in one day is taken up too (#12905)

Since 0.3.31 an iterating runner takes up a new runner from the project's bundle
at the next task boundary. It did so once per process lineage: the exec that
swapped it set MCPTASK_ADOPTED=1, every later process inherited it, and a
process carrying it never looked at the bundle again — so the second bundle
update
of the day waited for a manual restart. Now the marker names the version
that was installed (MCPTASK_ADOPTED=0.3.32, say), the new process takes it off
its own environment at startup, and it refuses only a re-adoption of that same
version. What an operator sees: each bundle update produces its own
"Updating from X to Y" on the card at the next boundary.

The marker no longer reaches the coding CLI's environment either.

One new line can appear in the log, and it means a broken release rather than a
broken host: [Adopt] This process was installed as the bundle's X but reports
itself as Y
— the gem's binary does not know its own version, and the runner
stays on it rather than reinstalling it in a loop.

No update is needed: nothing in a project's .claude/ changes. A runner still
on 0.3.31 sets the old flag form when it hands over; the new binary treats 1
as no version, so the handover onto this release is unaffected.

A refused push ends the task; it is no longer worked around (#12903)

In a smoke run a child's git push was refused because its SSH key had dropped
out of the agent. The child got past it by rewriting the project's origin to
another transport, finished the task, and left the clone changed — so the next
run failed on the rewritten URL, with an error that no longer pointed at the
missing key.

From this tag every PUSH step in every workflow (auto-squash, honest, manual,
story, review) tells the child that a refused push, fetch or pull —
authentication, permission, host key, protected branch — is not permission to
change the remote, switch transport, or touch credentials, SSH keys or git
config. It may re-run the same command once; after that it stops with status
failure and quotes git's refusal word for word, so the task note says what
the host said.

The runner also reads git remote get-url origin before and after every child.
If it changed, the run log gets an ERROR line naming the old and the new URL and
the git remote set-url origin <old> that restores it, and an anomaly piece
(origin_rewritten) is filed in the runner's Errors epic. The runner does NOT
restore the URL itself and does not change the task's result: whether the new
URL is wrong is for a person to decide. A project with no readable origin is
not checked (a DEBUG line says so).

No update is needed: the prompt and the check are both in the binary.

An expired trial stops the runner with one line instead of spinning (#12953, #12967)

mcptask.online now refuses an account whose trial has ended with no payment
method on file — HTTP 402 on the REST API, trial_expired: true on every MCP
tool (#12948). Until now the
runner read that as a passing hiccup: the queue check fell through to triage,
triage was refused too, a bug report was attempted and refused, and the wait
cycle started it all again.

Now the run ends at the first refusal it meets — the profile read at startup,
the queue check, or the quota poll — with one line on stdout and in the run log
naming the account, the server's own message and the payment link it contains,
and a non-zero exit:

❌ [WorkLoop] mcptask.online refuses account "kamr" — the trial has expired: "The trial has ended, please order a subscription: https://mcptask.online/kamr/standalone_payment" (payment page: https://mcptask.online/kamr/standalone_payment) — stopping the run: no triage, no retry, and no bug report, because the report would be refused the same way

No bug piece is filed for it, on purpose: the reporter uses the same token and
would be refused the same way. A scheduled host therefore shows a failed job in
its launcher log each weekday morning until the account is paid for; the next
start after payment runs normally. Nothing to do on upgrade — a bare update is
not needed.

A fan-out of subagents over one file is no longer killed as a loop (#12966)

The watchdog's stall detector counted every tool call on the stream into one
window, including the calls of subagents the session had forked. A skill that
sends six persona subagents over one shared file — each reading it once, by
design — produced six identical Reads, and the fourth ended the run:

Stall detected: reason=loop_signature signature=Read:/tmp/zuboklik-persona-target-wizard.md:: count=4

The run was filed as stalled_for_genius, the piece was re-run on genius,
stalled again and was set aside for the day. Zuboklik
#12842 spent four days that way.

From this tag the detector keeps its window, its failed-edit streak and its
Bash failure records per agent, reading who made each call off Claude Code's
parent_tool_use_id. One agent repeating itself four times is still a loop,
exactly as before; six agents each doing a thing once are not. A successful
edit still counts as progress for the session that forked the subagent, so a
root that delegates a fix and re-runs its failing tests is not killed either —
but never for a sibling subagent, whose loop is its own.

What you see in the run log. A verdict inside a subagent names it:
Stall detected: reason=loop_signature signature=Read:/x.md:: count=4
detail=agent=toolu_… — terminating for genius escalation
, and a Bash loop
reads detail=exit=1 agent=toolu_…. A verdict in the session you launched
reads exactly as it did.

The context-cost report had the same fault. After an overflow the fresh
session was told you read X in full 4 times about Reads its subagents had
made in their own context. It now counts only the session that ran out of
room, and the [context_cost] summary line says how much it left out when
there was anything: …, 0 failed edits; 6 subagent calls not counted (their
context is their own)
. A run that forked no subagent logs the line it always
did.

Codex and OpenCode are unchanged. Codex's exec --json items name no
thread and OpenCode prints no child session's calls at all, so on those two
every call the stream shows is counted as the root's, as before.

No update is needed: the change is in the binary.

ollama launch works again, and a CLI that cannot start says so (#12965)

Hosts whose launcher.command is [ollama, launch, claude] (or codex,
opencode) stopped working when Ollama 0.32 began taking a CLI's arguments only
after --: every launch died in a tenth of a second with Error: unknown
shorthand flag: 'p' in -p
, and the runner reported it as the model forgetting
its result marker — two --continue retries, then a bug piece titled
stream_ended.

  • The runner now composes ollama launch <cli> --model <model> -y -- <the CLI's arguments>. Nothing to change in the config: [ollama, launch, claude] is still the form to write. A launcher.command that itself carries --model, -- or another Ollama flag is refused before the run starts, with the reason.
  • Any CLI that exits non-zero before its first stream line now ends the run at once as launch_failed, with the exit status and the child's own stderr in the log, on the card and in the bug piece. No marker retry, no --continue.

Upgrading: no update is needed for this — nothing the installer writes
changed; the fix is the binary. Nor any config change: [ollama, launch, claude]
stays as it is. On an Ollama-launched host that has been failing since Ollama
0.32, install the new binary and restart the runner's LaunchAgent (an idle
runner keeps the binary it started with).


Changelog

Other

  • 877c4581f67ad907d299cdcaaf56224dbf04d356 [#12903] A refused push ends the task instead of being repaired by rewriting origin
  • f91a67c253643cf41d988cb575fc6b1eab856512 [#12905] The adoption marker names the version it installed, so a second bundle update in one day is taken up
  • 5623aaaaaa14668fc4ae8bea34abb89bc1ea0ae8 [#12953] An expired trial ends the run with one line, not a retried failure
  • b2e10d0883b56de3f5253ef1332b7b8c7131b5fd [#12965] A CLI that cannot start is a launch failure, and ollama launch gets the command line Ollama 0.32 accepts
  • 22435d77c05f3a8d5780d7c7d7362f2b13713a34 [#12966] The context-cost report counts only the session that ran out of room
  • c115499c802a2c2f5d924c29a56d8199541a4be4 [#12966] The stall detector judges each subagent on its own calls
  • 41f2b23b5b60b92ab0a40c4ae9d784dd86581b74 bin/release: the upgrading paragraph reports the asset diff, it no longer concludes a host has nothing else to do [skip ci]
  • b08156576375b0edbd75d758237b23d5d7cbf678 docs(release): v0.3.31 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.31

Upgrading

Upgrading is two commands this once. mcptask_runner update --self replaces the binary. Then a bare mcptask_runner update once in every project on the host (it also stamps the machine's helper home) — not because the bundled .claude/ assets changed (they did not), but because this tag starts recording which version installed them, and until that record exists every run start prints a WARN and leaves the files alone; see "The installed skills and helpers follow the binary by themselves" below. From then on installing a new binary is enough. A runner idling in a wait keeps the old binary until its next task or a restart (launchctl kill SIGTERM, wait for the job to stop, then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.31 carries this binary; no wrapper changes.


The installed skills and helpers follow the binary by themselves

Until now the skills in a project's .claude/skills, the helper scripts in
~/.claude/bin, the permission baseline and the test-commands file were
written by mcptask_runner init and refreshed by mcptask_runner update, and
by nothing else. A fix that lives in one of those files rather than in Go code
reached no host until somebody typed update there — task
#11669 corrected three helper
scripts on 2026-08-21 and every host kept the broken ones
(#11673).

From this tag every run starts by comparing the version that installed the
files with the version running, and refreshes them from the running binary
when the binary is the newer of the two. It is a string comparison against a
record on disk; nothing is fetched. The refresh is update without --force:
a skill or helper you edited by hand is left exactly as it is and named in the
log, never overwritten. The process that takes over between tasks after a
binary swap passes through the same start, so a binary installed at 10:00 has
its assets in place before its first task too.

What you see in the run log. One line per record, each naming the project:

  • [Reconcile] <project>: the skills in <project> were installed from 0.3.10 and this process is 0.3.11 — refreshed from it: 2 updated (ci-runner, test-runner), 1 left as edited by hand (discover), 9 already up to date — on stdout and in the file; the updater's per-file table is in the file at DEBUG, and a skill it left alone gets the usual conflict — … locally modified, skipped line at WARN. The helpers get the same line for ~/.claude/bin.
  • Nothing above DEBUG when the record already names the running version, which is every start after the first.
  • no skills manifest at … — this project was not set up by mcptask_runner init at INFO on a scratch clone run by hand; nothing is written there.
  • … were installed from 0.3.12, which is NEWER than the 0.3.10 this process is running — left exactly as they are at WARN. A host can legitimately be ahead of its binary — assets installed from a checkout, a hand-updated host where an older per-machine binary still runs — and the runner refuses to move them backwards. The binary is what is behind; install the newer one.
  • … records no runner version, so whether the skills … are older or newer than this 0.3.11 cannot be told — left as they are; run mcptask_runner update in this project once at WARN. This is every existing host at its first start on this tag: the record that makes the check possible did not exist before it, and an absent version is not read as "older than everything" — that guess is exactly what would put a corrected helper back to its broken version on a host whose binary is behind its files.
  • A refresh that could not be carried out (a directory that cannot be written) is an ERROR line, a bug piece under assets_reconcile_failed, and the run goes on with the files as they are.

What you have to do. Once per project, and once for the machine's helper
home: mcptask_runner update, from the binary you want the files to follow.
It records its version in .claude/skills/.mcptask_runner_manifest.json and
~/.claude/bin/.mcptask_runner_helpers.json, and from then on installing a
new binary is enough — update is no longer needed for an asset fix to land.
The order still matters for the typed command: release the binary, then
update. update from an older binary over newer files now says so before it
runs (… the files are going BACKWARDS), and runs anyway, because you asked.

Run logs and bug-report attachments no longer carry escape codes

The stream formatter used to wrap the task announcement in bold whenever the
emoji icons were on, whatever it was writing to — so the per-run log file got
\x1b[1m in it, and so did the copy bug-report attaches to a piece, which is
the one artefact meant to be read by somebody who was not there. Thirteen of
sixteen logs on the host that found this had escape bytes in them.

Colour is now decided per destination, by the same two variables as everywhere
else in the runner: MCPTASK_RUNNER_COLOR first (it wins both ways), then
NO_COLOR, then whether the writer is a terminal. The log file is never asked
and is never painted; stdout is painted only when it is a terminal that renders
it.

MCPTASK_RUNNER_ASCII is now about the icon set and nothing else. It used to
strip the bold as a side effect, which meant an operator who wanted clean log
files had to give up the emoji to get them — that is what the colour variables
are for now, and MCPTASK_RUNNER_ASCII=1 at a terminal keeps the emphasis it
used to lose.

Nothing to do on upgrade. Logs written before it still contain the escapes;
bug-report attachments from this version on do not.

The release is one script

bin/release X.Y.Z in this repository does the two-repository release —
gem CHANGELOG and version bump, both tags, the Go and gem Release workflows
watched, this file's contents put above what goreleaser wrote, this file
emptied, the six platform gems verified on rubygems — and refuses, by name,
before anything irreversible: a dirty or off-main checkout, a main that
differs from origin/main in either repository, a tag that already exists, a
version that does not go up, an empty notes file, a wrapper change with no
CHANGELOG line, or gh not signed in. bin/release X.Y.Z --dry-run runs
every check and prints every step it would take. Nothing changes for a host
that runs the binary; this is for whoever cuts the tags (task #12464).


A child that forks subagents in the foreground is no longer killed at five minutes

Zuboklik task #12842 asks for a persona review, which the project skill runs
as six subagents in one message with the parent silent until they answer. Four
runs on 2026-09-21/22 each read the inputs, logged 20 %, started the review
and died 5 minutes later with no commit and nothing on the piece; sibling
tasks went through fine, so to the owner it looked like the runners kept
skipping that one (bug #12900).

The profile knew the subagent tool as Task, the name Claude Code used to give
it, and gave that kind the long-running ceiling. Claude Code 2.1.x calls the
tool Agent, which no profile line named, so the watchdog filed it under
"other" and applied the quick ceiling: a kill at 300 s once the stream had
been quiet for 180 s, a failure verdict, the day's set-aside list, and a
resume that died the same way. Agent is now a subagent in the bundled claude
profile, beside Task, which stays for the CLIs that still use it.

One command on a host: mcptask_runner update --self. A host that overrides
the claude profile with a file of its own has to add the line itself —
Agent: {kind: subagent, name: Agent} under tools:.


A transient 500 from the API no longer ends the day

Bug #12896: on 2026-09-22 two
runners on the claude harness hit an HTTP 500 during triage. The CLI ended its
turn on the terminal event that a refusal arrives on, with a sentence that says
the condition is temporary; the runner filed it as a refusal it had no name for,
returned status error, and the quota Decider stopped the whole working day on
both machines over one 918ms request.

What an operator sees now:

  • API Error: 500, 502, 503 and 504 on Claude Code's terminal event are read as an overload and ride the same budget a 529 does — ten waits, doubling from 60 s to 600 s — instead of ending the run. The 500 was recorded from the incident; the other three are the same CLI template and were not.
  • The log says so in the CLI's own words. At the moment the line arrives: The API answered with a server-side error — the CLI said: "API Error: 500 …"; not a refusal: if this attempt ends without an answer it is retried on the overload budget (a WARN, not an ERROR). At the retry: API server-side error - retry 1/10, waiting 60s before next attempt... — the CLI said: "…", and the dashboard card says the same while it waits. A bare 529 keeps its old line, API overloaded (529) - retry 1/10 …, byte for byte.
  • On codex and opencode every overload already arrives as a sentence, so their retry lines now quote it too (… — the CLI said: "Selected model is at capacity…") rather than naming a 529 the CLI never sent.
  • If the ten waits run out, the run ends as before with status error, and the message carries the CLI's sentence rather than "API overloaded (529)".
  • A refusal the profile has no words for (a 403, for instance) still stops the run and the day, with no named remedy in the log: that is the cue to record the words in the profile, under overload if the CLI calls them temporary.

Nothing to do on upgrade. The profile change ships inside the binary. A host
that overrides the claude profile with its own copy — config/harnesses/claude.yml
in the project, or ~/.mcptask/harnesses/claude.yml — keeps running its copy
and has to add the four literals under failures.overload itself.

A running runner takes up a newly installed binary between tasks

Until now a binary installed while the runner was working was not picked up
until the process ended: in daily at the day's end, in today_auto_squash
at tomorrow's scheduled start. From this tag every mode that works more than
one task — today, today_auto_squash, daily, the queue and story modes —
checks at the boundary between two tasks whether the installed binary
changed, and if it did, swaps itself onto it there
(#11867). The check is two
stat calls and a read of Gemfile.lock; nothing is fetched from the
network, ever.

The channel does not matter: homebrew, mcptask_runner update --self,
curl | sh, another project's run rewriting ~/.mcptask/bin/mcptask_runner,
or bundle update mcptask-rails-runner — the bundle is adopted into
~/.mcptask/bin at the same boundary, which is what makes the process stale.

What you see. The runner's card on the dashboard blinks: it closes with
the message Updating from <old> to <new> and a new card opens under a fresh
session id a moment later. The run log says the same, with the number of
results carried over and where they were written. The new process continues
the same day — the tasks worked so far, the moment the day began — so the
quota check and the stop rules see the whole day, not a second one starting
at 11:00. The skip list and the urgent pin were already on disk and survive as
before. The day-end swap in daily stays as it was, in addition.

What you have to do. Nothing beyond installing the new version. Do not
restart the runner by hand; the running process swaps itself at the next task
boundary, never while a child is alive. Before handing over it runs
<new binary> version once; a binary that cannot say its version is refused
with a line in the log, and the runner stays on the one it has. A handover
whose exec fails is likewise logged and left for the day's end to try again.

State file. The day's state travels through
tmp/mcptask_runner/handover.json under the project, written a moment before
the exec and consumed once by the process that takes over. A file left there
by a run that did not continue is reported in the log and removed at the next
start, never resumed.

The story loop asks the server which subtask is next

Triage used to be handed a whole Story and a rule in prose — first subtask not
finished, not blocked, not on the skip list — and re-derived a database query
from a list that carried no priority, no task type and no blocker. On
jchsoft/Zuboklik on 2026-09-15 it walked that list in creation order, looped on
a medium task while an urgent bug in the same Story waited, and handed over a
blocked subtask that cost a whole session of the strongest model to discover
(task #12568).

The runner now asks mcptask.online's story-scoped queue — GET
/api/:account/pieces/:story/next
, task #12567 — before any triage session
runs, with today's skip list passed as exclude_relative_ids. The server
orders the Story the way it orders the queue (priority tier, bug or complaint
first, oldest first), never offers a blocked or finished subtask, and triage
receives one concrete subtask and chooses only the model. The subtask scan is
gone from every prompt that met a Story: the story-locked triage, the Story
branch of the discovery triage, and the dry run.

The three empty answers are three exits. Nothing left ends the story loop as a
finished Story. Every remaining subtask blocked ends it with no_more_tasks
naming the blocked subtasks — in the verdict, in the log and in the card's
message — rather than spinning on a Story it can take nothing from. Nothing
beyond the skip list ends it the way the skip list always has: the loop waits,
as on an empty queue.

Requires mcptask.online with the story-scoped @next (task #12567, deployed
2026-09-15).
Against an older server the REST call answers 404 and the loop
falls back to the same query through MCP: triage is handed the Story alone and
its prompt reads mcptask://pieces/{account}/{story}/@next?exclude_relative_ids=…
itself — on a server without that resource, a story loop cannot pick a subtask
at all and answers no_more_tasks with the server's refusal in its message.
Hosts with a .claude/ installed by an earlier release need no update: the
change is entirely in the binary and its prompts.

GetNextStorySubtaskTool now counts as fetch evidence in every bundled
harness profile, so a triage child that reached the Story's queue through the
tool rather than the resource is not discarded as an unverified pick.


Changelog

Other

  • bd3b65e6f6eef0f68903a903041510348672b0cf [#11673] The installed skills and helpers are reconciled with the running binary at every run's start, by version, never backwards
  • e56cf6ed50f6b870023a951d1cb9fb6f5abbc382 [#11780] the stream formatter paints a terminal and writes the log file plain
  • b0b6cc3fb415e10812d6dac48253707881a69a13 [#11867] The in-process stop cuts a session start and the closing-frame redial short
  • c57cd0a0b2ceaa592d335c4c0ec0d4ea40f20ed4 [#11867] The runner takes up a newly installed binary between tasks, not only at day end
  • ce2aceecdb4d290f9b1f096055fc536ec86698a0 [#12464] bin/release: the two-repository release as one script that refuses by name
  • 0d4c1f885553e4bf2004aa8d5449a468f2671ed2 [#12466] the conformance driver is built as mcptask_runner, and every scenario checks the child can find it
  • 9c20400e407555e830195ef674b42e5129041c13 [#12568] The story loop asks the server which subtask is next
  • 45953b16f287d6bd524b59786c53bc1b3880a300 [#12896] A transient 500 on the terminal event rides the overload budget instead of ending the day
  • 7d93df8475e9aff0dca905ba5ee10f1e13bf045f [#12896] Scenario 25 asserts the retry line in each dialect's own words
  • adf337334a4b44152054787ba20af6b203c81836 [#12900] the claude profile names Agent as the subagent tool, beside Task
  • eacefab166e2bd6f7300f133a406268a86c6ebda docs(release): v0.3.30 shipped, the next-tag notes start empty
  • 9e09a9be67ffe3fdfd83f920ab6dab75fd793a1c test(eventstream): the stop tests group their imports the way goimports wants

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.30

Upgrading takes two commands

mcptask_runner update --self replaces the binary. Then a bare
mcptask_runner update in every project on the host, because the bundled pr
skill changed — it now names the shape of the task link a pull request must
carry and tells the agent to copy the one the CREATE PULL REQUEST step prints,
which is the other half of the pr create check below. Run it while no runner
is working that checkout. A runner idling in a wait keeps the old binary until
its next task or a restart (launchctl kill SIGTERM, wait for the job to stop,
then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.30 carries this binary; no wrapper
changes.


init can be told "no scheduled job"

The scheduled-job prompt at the end of init offered 1) today_auto_squash
and 2) today and nothing else. There was no way to say "I do not want the
runner started automatically": any other answer was invalid mode, and because
the schedule is the last step, that error failed an install whose every earlier
step had already succeeded (task #12790).

The prompt now has 3) none — no scheduled job, start the runner by hand. It
writes no LaunchAgent, systemd timer or Task Scheduler job, prints
Scheduling: skipped (none chosen …) and the install ends green. A
provisioning script says the same with --mode none. Nothing changes for a
host that already answered 1 or 2; pick it up with `mcptask_runner update

--self`.

The end of init can be read without reading every word

The tail of an install was forty [Installer] lines in one voice: paths
written, commands to run, and the sentences explaining both, all at the same
weight — and an operator who had just answered three prompts could not tell
which lines to do something about. On a terminal each line is now painted by
what it is. A path, a command and a value somebody set keep the terminal's own
colour; an explanation (why the hour matters, what the log retention is, what
stopping a run costs) is dimmed whole; and the one step to take now — "To
activate, run:" — is bold. The config summary follows the same rule: section
names bold, meanings and values still at their default dimmed, so what you
configured is the only thing on the page that is neither.

The words are unchanged, and so are the bytes a pipe, a log file or CI
receives: colour reaches a terminal and nowhere else, as before. NO_COLOR
and MCPTASK_RUNNER_COLOR still decide it outright.


A pull request without a task link is refused where no template would catch it

PR jchsoft/mcp4mail#48 went out with `Task: https://mcptask.online (jchsoft

12761)` in its description. mcptask.online reads a pull request back to its

task through that link, and its parser wants the whole path —
https://mcptask.online/<account>/tasks/<relative_id> — so it found no task,
answered the webhook with "Not body supported format!" and task #12761 was
never approved. The prompt said "mcptask.online task link" and left the shape
to the agent, and mcp4mail has no pull-request template, which is exactly where
the shape got invented (task #12865).

Two changes, one at each end. The CREATE PULL REQUEST step now prints the
task's link finished, in that shape, with nothing left to invent; when the run
picks its own task the account is filled in and the id is named as the one
from the fetch step. And on a project with no pull-request template,
mcptask_runner pr create refuses a description that carries no such link,
before the host is asked, naming the shape — the same way it has refused a
description that ignores the template since #12599. A project with a template
of its own is unchanged: the template says where the link goes.

Two commands on a host. mcptask_runner update --self puts the check in the
binary and the wording in the prompt; then a bare mcptask_runner update in
every project, because the bundled pr skill changed too — it now names the
link's shape and tells the agent to copy the one the CREATE PULL REQUEST step
prints. Run it while no runner is working that checkout. A project that was
opening pull requests without any task link will now see them refused — add
the link; that is the pull request mcptask.online could not follow anyway.


Changelog

Features

  • 94346af8cdc3284ec1f349489c76988645f5d8f2 feat(install): the end of init weighs each line by what it is ### Other
  • 589b7ce2b1f56c6f2698b299571f582b93347241 [#12790] Offer 3) none in the scheduled-job prompt, and --mode none on the CLI
  • daabf1e7e835b8e176e69a530da3c7c398bf42b8 [#12865] pr create refuses a description with no task link where no template would catch it
  • 9071ff135fa92d1cca4310a04b60b2a492ebf830 docs(release): the #12865 note names the bare update the changed pr skill needs
  • 8adb4fefdd6ddbf79e73e62100bf014de6d705cf docs(release): v0.3.29 shipped, the next-tag notes start empty
  • 1aa5fe3e30b0853f698b1f2c70fd469665034e62 test(install): the painted-fact check uses strings.Cut, not an Index slice

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.29

Upgrading takes two commands

mcptask_runner update --self replaces the binary. Then a bare
mcptask_runner update in every project on the host, because the bundled pr
skill changed — it now tells the agent to write the pull-request body under
the project template's own headings, which is the other half of the pr create
guard below. Run it while no runner is working that checkout. A runner idling
in a wait keeps the old binary until its next task or a restart
(launchctl kill SIGTERM + kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.29 carries this binary; no wrapper
changes.


A stalled task really does come back on the strongest model

A stall — the same call over and over, the same edit refused three times, the
same command failing the same way — kills the attempt and leaves the piece in
progress with the verdict stalled_for_genius. The name was a promise the daily
loop did not keep: triage in discovery is told never to call a pick a resume, so
@next handed the same piece back, triage graded it "smart" again, and one host
launched the same model on the same task several times in a row on 2026-09-16
(task #12613).

The loop now remembers what stalled today and on which tier, and decides the
escalation itself. A piece that stalled on a weaker tier runs on genius the next
time triage picks it, whatever triage recommends — the log says so in one line.
A piece that stalls on genius goes on the skip list for the rest of the day,
like a failed one, and gets its fresh look tomorrow. Nothing to configure; pick
it up with mcptask_runner update --self.


A blocked subtask no longer costs a session to discover

Triage walked a Story's subtask list excluding the two Czech labels for finished
work and nothing for blocked, so a subtask sitting at "Blokováno" passed it. The
loop then started the strongest model on the piece, which fetched it, read
is_blocked: true and stopped — a whole session to learn something the Story
payload triage had already read contained, while the work that was actually
available waited (task #12568, seen on jchsoft/Zuboklik on 2026-09-15).

Blocked subtasks are now excluded where the selection is made, in every prompt
that walks a Story: the triage prompts and the dry run. Where the server sends
the machine fields — is_blocked, task_state_code, priority_code,
task_type_code, blocked_by — the child is told to use them and to take the
list in the order the server gives it, which is what makes an urgent bug filed
after a blocked task reachable at all. A server that does not send them behaves
exactly as before: the state label and the progress are the whole test, and
nothing is inferred.

A Story whose remaining subtasks are all blocked now ends the story loop with
no_more_tasks and a message naming the pieces and their blockers, instead of
handing a blocked subtask over again. No new status, and nothing to configure —
pick it up with mcptask_runner update --self.


pr create refuses a description that ignores the project's template

About one pull request in four opened on projectoid_ii carried Claude Code's
own body — ## Summary, ## Test plan, the robot footer — instead of the
project's .github/pull_request_template.md: 9 of the last 40, on both
accounts and both models (task #12599, seen on
https://github.com/jchsoft/projectoid_ii/pull/2036). The CREATE PULL REQUEST
step names the template in one bullet, the agent never opens the file, and
mcptask_runner pr create forwarded whatever it was given.

The prompt has not grown a line. Instead pr create reads the template the
prompt names — GitHub's or GitLab's convention, or pr_template: path: —
and refuses a description missing any heading the template does not mark
(optional), before the host is asked. The refusal names the missing sections
and prints the template's skeleton, so the agent's next turn is the rewrite
with nothing else to look up. A project with no template on disk is checked
against nothing, and a declared path with no file behind it is said on stderr
rather than guessed around. Pick it up with mcptask_runner update --self.

Changelog

Other

  • 7989e4cc35ae390e09c18668e96d0543e90c3f8f [#12568] Skip a blocked subtask in triage, before it costs a session (#12)
  • d4e8fbc6d5427826ae23eb0fc080789a23fa7c57 [#12599] Refuse a pull-request description that ignores the project's template
  • 5d1baa4031bebbbf7809aa1e883c9b9603b8ac42 [#12613] Escalate a stalled task to genius runner-side, and set it aside when genius stalls
  • 7f952fb60a3a831633a0b9620d3b821ef7fcb300 docs(release): v0.3.28 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.28

Jobs with ignore_quota: true now wake up on an assignment

A runner started with ignore_quota never read its user profile, so it never
learned its own user id and ignored every assignment broadcast — the wait ran
its full length and the log said Assignment event ignored: this runner does
not know its own user id yet
on each one, while the card on mcptask.online
named the user the whole time (bug #12473). The profile is now read on every
start; ignore_quota only decides whether the working-hours flag is applied.

Nothing to configure. A running job picks the fix up at its next start: after
mcptask_runner update --self, restart an idle runner (launchctl kill SIGTERM
+ kickstart), or wait for the next scheduled run. The startup log of an
ignore_quota job now carries a [QuotaGuard] user id N read from the user
profile at startup
line; a job whose profile read fails three times goes ahead
without its id and says what that costs, instead of refusing to start over a
flag it would not have applied.

Changelog

Other

  • ee730bd334f95d1f5592e2f9abd98360d4b99382 [#12473] Read the user profile under ignore_quota too, so the stream knows whose assignments to act on
  • 6926a1de6b8ef860d4801211d0d0cb6acab146a7 docs(release): v0.3.27 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.27

Upgrading

Two commands: mcptask_runner update --self replaces the binary; then a bare
mcptask_runner update in any one project on the host, because the bundled
pr skill and the permission baseline both changed. Run it while no runner is
working that checkout. A GitLab project also needs glab installed and
glab auth login done on the host.

mcptask_runner is now on the child's PATH on every project

Binary only — no helper, skill or launcher file changes.

The prompt and the bundled pr skill both tell the coding CLI to run
mcptask_runner pr …, and nothing ever put that name where the child could
find it. Only Ruby projects got away with it: the mcptask-rails-runner gem
installs an executable called mcptask_runner into the bundle's bin directory,
which Bundler has already put on PATH. On a Go, Node, Python or plain
repository — and on any launchd run, whose generated launcher execs the binary
by absolute path and adds nothing to PATH — the first mcptask_runner pr
create
came back

Exit code 127: command not found: mcptask_runner

and the child, having no other way to open a merge request, typed glab mr
create
instead: the one thing the skill tells it never to reach for.

The runner now prepends the directory of its own running binary to the child's
PATH, so the mcptask_runner the prompt names is the binary driving the run
rather than an older copy installed somewhere else. Nothing changes on a Ruby
project: the bundle's shim and this binary are the same version, because the
gem's version is coupled to it.

If you run a binary you built yourself, name it mcptask_runner. The
directory goes on the child's PATH only when a file of that name is in it —
reaching for some other mcptask_runner elsewhere on the machine would be a
different build wearing this run's name — and a build named after the module
(mcptask_go_runner) gets one line in the run log saying exactly that:

[TaskAutoSquash] WARNING: mcptask_runner is NOT on the child's PATH. This runner is …

Releases and mcptask_runner init both install the binary under that name, so
this affects development builds only.

Adoption now works on an rbenv, asdf or mise host

Binary only — no helper, skill or launcher file changes.

Until now the runner looked for the gem its project bundles in GEM_HOME and
GEM_PATH, and nowhere else. rvm exports both; rbenv, asdf and mise export
neither — they put shims on PATH and initialise from ~/.zshrc, which the
bash -l launcher never reads. So on those hosts the search had no roots at
all, every task boundary printed

[Adopt] Gemfile.lock names 0.3.26 but its gem is not installed here — still running 0.3.24 (run `bundle install`)

and both halves of that were wrong: the gem was unpacked under
~/.rbenv/versions/<ruby>/lib/ruby/gems/<x.y.0>, and bundle install had
already run. The host stayed on its old binary indefinitely.

The search now keeps the environment roots first, then derives the project's
Ruby from .ruby-version (walking up) or the Gemfile's ruby line and probes
the documented layouts of rbenv, asdf, mise, rvm, chruby, frum and Homebrew,
~/.gem/ruby/<x.y.0> for user installs, and finally asks bundle show
mcptask-rails-runner
through that Ruby's bin — the only source that can name
a BUNDLE_PATH or vendor/bundle install. Nothing is taken on trust: a
candidate counts only once its interpreter reports the pinned RUBY_VERSION and
names its own Gem.dir, or — where no interpreter can be run — once the
directory is found to actually hold a gems/. Every adoption now logs which
source answered, and a failure lists the roots it searched. The two failures are
told apart as well: a root that exists without the version in it is a missing
bundle install and says so, while finding no gem directory at all names the
pinned Ruby and every candidate it turned down instead.

An affected host cannot adopt its way onto this fix, because the binary that
would do the adopting is the broken one. Once, in the project directory, in the
foreground:

bundle exec rake mcptask_runner:update     # or: mcptask_runner update --self

then a bare mcptask_runner update for the helpers. After that the host adopts
by itself again.


GitLab is a third git host

git_host: gitlab now drives merge requests, and mcptask_runner init writes
the key off a gitlab.com origin by itself. A self-managed instance is on the
customer's own hostname, so nothing in its remote says "gitlab" and nothing
guesses from one: such a project is installed with

mcptask_runner init --git-host gitlab

There is no credential to place. GitLab is driven through glab, which holds
its own login the way gh does, so init writes no token file and no
environment variable — it names the one command that makes the host work, and
glab auth login on that host is the whole of the setup. Nothing about Bitbucket
or GitHub changed: a project on either gets exactly the install it got before.

Two things follow that an operator sees. A GitLab project's
.claude/settings.local.json picks up Bash(glab:*) and
WebFetch(domain:gitlab.com) on the next mcptask_runner update — approvals no
other host is handed, the same way gh belongs to a github project alone. And
the pull-request template a GitLab project is told to follow is
.gitlab/merge_request_templates/Default.md, which is what GitLab's own web UI
puts into a new merge request; pr_template: path: still overrides it, and a
Bitbucket project still gets no template line at all, because Bitbucket Cloud has
no such convention to name.

A self-managed instance needs no address here. glab auth login --hostname
gitlab.example.com
once per machine, git_host: gitlab in the project, and that
is the setup: every question the runner asks is a glab run in the project's own
checkout, and glab resolves the host from that checkout's remote against the
credentials it stored. The runner keeps no URL of its own — a second place for the
address to be written is the first one to go stale. Hostname derivation stays
exact, so gitlab.com derives the key by itself and a customer's own domain does
not: a substring match on "gitlab" would point glab at somebody else's service.

mcptask_runner pr checks prints ONE row on GitLab, and that row is the whole
verdict rather than a summary of others. The host has no per-check object: a merge
request has one pipeline, and the pipeline's single status is what GitLab
publishes and gates merging on, so the report carries the head pipeline as one
check named after it. success is the only pass; failed, canceled and
canceling are red; everything else — skipped and manual included — is
IN_PROGRESS and is waited on, because a pipeline that never ran is not a green
one. A project with no .gitlab-ci.yml answers NONE, which is not a pass either.

Four conformance scenarios pin the behaviour (21–24: a manual merge request, a
green pipeline merged and verified, a red pipeline twice, and a merge the host
would not confirm). They drive a fake glab first on the PATH and assert its argv
word for word — including the --auto-merge=false that stops glab mr merge from
scheduling itself, exiting zero, and letting GitLab merge the branch ten minutes
later with nobody watching.

A failing gh or glab now says what it said

Binary only — no helper, skill or launcher file changes.

Until now, a refusal from either CLI reached the operator and the run log as its
exit status and nothing else:

Error: glab mr create --title x --description y --draft --yes: exit status 1

What glab actually said was a 409 — this branch already has a merge request —
and it was never lost on the wire: exec.Cmd.Output() captured the child's
stderr into ExitError.Stderr, and the wrap that named the command kept only
ExitError.Error(), which is the string "exit status 1". Every reason a merge
can be refused on those two hosts arrives that way: a protected branch, an
approval rule not satisfied, a merge train, an expired login. The runner reported
all of them identically.

It now repeats what the CLI said, in the form the Bitbucket adapter already used
for the host's own error envelope:

Error: glab mr create --title x --description y --draft --yes: exit status 1: ERROR Post https://gitlab.com/api/v4/projects/1/merge_requests: 409 {message: [Another open merge request already exists for this source branch: !1]}

One line: a two-line ERROR block is joined rather than allowed to wrap the run
log around it. stderr is quoted when there is any, and the first line of stdout
when there is not, because glab prints some refusals there. A failure with no
child behind it — a timeout, a glab that is not installed on this host — gets
nothing appended, since it has nothing to quote and its own text already names
what happened. Nothing about exit statuses or control flow changed: gh pr
checks
still exits non-zero on a red CI and that is still read as a verdict.


A machine's own pull-request skill no longer outranks the runner

Binary and skills — a host that installs this release should re-run
mcptask_runner update so the bundled pr skill and the permission baseline
land beside the new binary.

The first real GitLab run found the gap the mcptask_runner pr work of the
previous release had left open. The step said "CREATE PULL REQUEST", named the
command, and forbade "a host-specific CLI" — and the child used neither. It
reached for a skill installed machine-wide called pull-request-creator, whose
description offers to create a pull request "following WorkVector conventions",
which reads as the house style rather than as somebody else's host. That skill
drives GitHub's CLI, so on a GitLab project it produced

gh pr create …
→ none of the git remotes configured for this repository point to a known GitHub host

before improvising its way to glab mr create. The merge request was opened by
luck: on a project whose host has no CLI at all, the same route produces two
failures and nothing else.

Nothing was violated in the child's reading of the prompt, which is the point.
The step's own title is word for word what that skill advertises, and the step
never said the WORDS belonged to the command. It does now: "create a pull
request", "create a PR" and "open a merge request" are declared to mean
mcptask_runner pr create, and anything else offering to answer them — another
skill, another script, a command you know — is declared wrong, because each is
written for one git host and this project may not be on it. The bundled pr
skill leads with the same phrases and says outright that it outranks any other
pull-request skill on the machine. Neither names a host or a CLI; the neutrality
guards would refuse it, and the runner cannot know what is installed anyway.

One standing approval is gone from the permission baseline:
Skill(pull-request-creator). The installer was handing every project it
touched an approval for the very skill the prompts forbid. Removing it is all
this file can honestly do — the merge is a union, so a project that already has
the entry keeps it, and an implement run is launched in a permission mode that
skips these checks. What decides it is the prompt and the skill description.

A published pull request now reaches the run record even when the command was
not used.
The same run recorded {"status": "success"} with no pr_number,
while the runner had printed the merge request's number on the card and in the
log: only the mcptask_runner pr envelope ever stored one, and the child had
never run it. The CLI's own published-change event is a second source now,
below the envelope — which still wins, whichever arrives first — and strictly
read: an identifier that is not a positive integer is dropped rather than turned
into a number. Downstream of that are the merge verification, the effort line
and the card's closing message.

Changelog

Other

  • a10dab6b5fcf0b38a0b420c00b406b90f332a419 [#12418] Give the dry prompt the harness's fetch block
  • 9fba8128f0c481d15f67200de2e48d012827d85c [#12419] A close code is never destroyed, only misfiled, and the chaos test cannot prove it either way
  • 791cce8861066bb03f70ec4efc968f732130c99b [#12435] Drop a GEM_HOME that names no directory, and say so
  • 9aafbe1dd50684dc1e04f225f4b1aa67ccc28832 [#12435] Find the bundled gem on an rbenv, asdf or mise host
  • f95cc32f54aa4edc41a997118551525a00f54fc4 [#12435] Keep the gem search inside the roots the caller named, and off Windows's ruby.exe mismatch
  • fc4affd7ee63ec88e93d7bbdea32ca79b2b10e53 [#12437] Keep the host's test lock out of reach of tests and the stress day
  • 3cf39989ab8addfc1bbe3291c57d7b7412070819 [#12437] Put the conformance sandboxes outside every git repository
  • 54b60f2e96ea3710ccafb3f702f4571f62e8064b [#12449] A gitlab adapter that drives glab, and the places glab is not gh
  • 2abf0832e6337cf3ca95ef07ac71ea07f655241d [#12450] Everything that names a host by name learns the third one
  • d5999c68335e64dd9bf47ee21bb3ca49b2bbaf44 [#12451] GitLab proved: four scenarios against a fake glab, and two runs against the real one
  • 3e5c44a3adb7872fc6204d82d5ef81889e9053d3 [#12454] A refused gh or glab command repeats what the CLI said
  • 7095704f11f52dcda65147c623fead2d76004b3a [#12455] Claim the words "create a pull request", and read the number off the stream
  • 137eda5f6711e269680576f9dc03d7f6a956e536 [#12456] Put the runner's own binary on the child's PATH, under the name the prompt uses
  • c2a68e166540891078eedbb0323db2020b0c3a7f docs(release): v0.3.26 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.26

Upgrading

mcptask_runner update --self replaces the binary; then run a bare mcptask_runner update in any one project on the host, while no runner is working that checkout — the bundled pr skill gained the review commands and the binary does not rewrite installed skills on its own.

A close the server explains is logged with its explanation

When mcptask.online ends the dashboard socket with a close code — going away
during a deploy, a policy violation, an internal error — the runner's log now
says so, instead of sometimes reporting only use of closed network connection.

A dying socket is noticed twice: by the read loop, which has the close code, and
by whatever snapshot write was in flight, which has only the symptom. The
reconnect notice could be composed before the read loop was scheduled, and the
close code — the half that says who ended the connection — was then dropped on
the floor. It is now printed on a line of its own when it arrives that late:

[EventStream] WebSocket closed (the snapshot write failed: use of closed network connection), attempting reconnect...
[EventStream] the read loop's account of that close landed after the notice: server closed it, code 1001 going away

Nothing branches on the close code, so no behaviour changes: the runner redialled
before and redials now. What changes is that a run log attached to a bug report
says whether the server let the connection go on purpose — which decides whether
there is anything to go and look at on the server side at all.

What to do: mcptask_runner update --self. Binary only — no config change,
no reinstall of the .claude/ assets.


No prompt the runner composes names a git host any more

The review and reviews executors were the last two that did. Everything else
moved onto mcptask_runner pr in v0.3.24; these two could not follow, because
prhost.Host had no question for "what did the reviewers say on this pull
request" or "what is open on this repository", and a prompt cannot be pointed at
a command that cannot answer it. So they went on typing gh pr view, two
gh api repos/{owner}/{repo}/pulls/{n}/… reads and a gh pr list --state open —
four instructions that do nothing on a project hosted anywhere but GitHub.

The interface has both questions now, and two new subcommands expose them:

mcptask_runner pr list --open      # every open pull request, no search
mcptask_runner pr reviews 1234     # the verdicts, then the comments

reviews is one flat list across what each host splits differently — GitHub's
submitted reviews and inline comments, Bitbucket's participant verdicts and its
comments. A note carrying a state is a verdict (APPROVED,
CHANGES_REQUESTED, COMMENTED); one carrying a path and a line is inline
on the diff.

Two things it deliberately does not claim. bot: true is the host marking an
app, and Bitbucket Cloud marks none, so the field's absence is not a promise
that a person wrote the note — the pr skill says to read author there and to
say which you did. And there is no resolved flag on either host: whether a
review has been addressed is a judgement made by reading it, which is what these
prompts did when they called the REST endpoints directly, and a field only one
adapter could ever fill would be an invitation to trust a false the other host
never said.

The guard list in TestSpineNamesNoGitHost is now empty, which is the point of
keeping it: it was the record of what was left, and an empty one is a claim every
prompt is held to.

What to do: mcptask_runner update --self for the binary, then a bare
mcptask_runner update in each project — the bundled pr skill gained the two
commands, and a project still carrying the old copy will not know they exist.

Auto-squash asks the git host whether its checks passed, before merging

The auto-squash CI step used to have one verdict in it: the local bin/ci. A
green run merged, a red one left the pull request open, and the host — which is
the thing actually running the pipeline the project configured — was never
asked. A run on a project with no bin/ci merged on no verdict at all: it
skipped the local gate, merged immediately, and the host's pipeline reported
SUCCESS a few minutes later. It happened to be green. Nothing in the runner
would have noticed if it had not been.

The step now has two verdicts and the merge waits for both. Between the local
gate and the merge the child runs mcptask_runner pr checks <pr_number> and
reads .checks.state:

  • SUCCESS — merge.
  • FAILED — status ci_failed, nothing merged, the pull request stays open with the red check named.
  • IN_PROGRESS — wait two minutes and ask again, at most ten times. Still running after the last ask is ci_failed; an unfinished check is not a passed one.
  • NONE — merge. A check nobody reported is not a pass, but a project whose host runs no CI must still be able to auto-squash, and the local gate is what carries that case.

Which of the four decided the run is now a required field of the result marker,
status_detail, so a ci_failed in the effort trail says whether it was the
local suite, a red check on the host, or a build that never finished.

One text for every host, as the rest of the spine is: the command is
mcptask_runner pr checks on GitHub and on Bitbucket alike, and the adapter
behind it answers in the same four words.

What to do: mcptask_runner update --self. Binary only — no config change,
no reinstall of the .claude/ assets.

init --home-dir re-roots the Bitbucket credential, and a re-run says what it did with it

Two fixes to the step that places a Bitbucket credential during init
(#12424).

--home-dir re-roots everything an install writes outside the project: the
launcher script, the token directory, the shell rc file, the scheduled job. The
Bitbucket credential was the one file that escaped — it resolved the account's
own home and wrote there, so an install aimed at a scratch directory put a secret
into ~/.mcptask_env.d/bitbucket_credentials instead, where a second machine's
credential would then sit in a directory nothing sources. It now goes under the
same home the mcptask token does, out of one shared resolution rather than two.

The second half is what a re-run says. The credential step used to hang off the
path that newly wrote git_host:, so the two shapes of re-run that change
nothing — the key already in the file, or --git-host bitbucket repeating what
the file already says — printed git_host: bitbucket already set … and stopped.
That is exactly the run an operator types after making a token. It now runs
whenever the resolved host is bitbucket, and says which file it left alone:

[Installer] git_host: bitbucket already set in …/config/mcptask_runner.yml
[Installer] [Bitbucket] credentials already in ~/.mcptask_env.d/bitbucket_credentials — left alone. To replace them (a rotated token, or a different account), re-run init with BITBUCKET_ACCESS_TOKEN, or BITBUCKET_EMAIL and BITBUCKET_API_TOKEN, set.

The no-clobber rule is unchanged: an install handed nothing still writes nothing.
It just no longer does it silently.

What to do: mcptask_runner update --self. Binary only — no config change,
no reinstall of the .claude/ assets. A host installed with --home-dir that
has a Bitbucket credential is worth one look: the file may be in the account's
home rather than under that directory.

Run records say how the attempt ended

log/runs/run_*.json used to end like this on a run that had gone perfectly:

"status": "processing",
"termination": "result",
"final_status": "processing",
"result": { "status": "success", "pr_number": 3, … }

Two of those three lines were about a run that was still going. The record is
stamped on the way out of the streaming loop, one line before the session is
moved off processing, and it is never written again — so the field a reader
reaches for to tell how a run ended reported the state the run had been in while
it was working, and every clean run looked identical to a run that hung.

Both now answer the question they are named for. final_status carries the
child's own verdict where there is one — success, ci_failed,
merge_unverified, no_more_tasks — and the session's terminal status where the
child never answered, which is error on every kill path, exactly as it already
read there. status carries the terminal frame: finished on a clean ending,
error on a kill.

The record's shape is unchanged: same members, same order, no new field, nothing
for the dashboard to learn. Records written by older binaries keep their old
values; this only changes what is written from here on.

What to do: mcptask_runner update --self. Binary only — no config change,
no reinstall of the .claude/ assets.


Changelog

Fixes

  • 2e7ddea2857effacff589672145c8e02c6fdd236 fix(eventstream): a close code that arrives after the notice gets its own line (task #12421) ### Other
  • f3a9d1dbb050938efe31d3eb8f16dc47de2fad62 Ask mcptask_runner pr for reviews and for the open queue (task #12420)
  • 6ca5eb9ae3a8f8af8d35ae2baa90d637d7f53855 Ask the git host's checks before the auto-squash merge (task #12425)
  • b7595569ae0a929c8f48cb6fbd83d0d47980524d [#12424] Re-root the Bitbucket credential, and speak about it on a re-run
  • 47792ccc135245fb7b555b3a95b58d890543aeb1 [#12426] Make the run record say how the attempt ended
  • 6daeb941aeaa518cf1b819b35f918d7c2ab59ff7 docs(release): v0.3.25 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.25

Upgrading

mcptask_runner update --self replaces the binary; then run a bare mcptask_runner update in any one project on the host, while no runner is working that checkout — this version changes the installed data pack (the new pr skill, per-host baseline permissions, ci_wait, the test-lock helpers), and the binary does not rewrite those on its own.

Bitbucket Cloud projects can be driven

git_host: bitbucket now selects a real adapter instead of an error. The runner
opens, finds, reads, merges and checks pull requests on Bitbucket Cloud through
its REST API 2.0 — there is no Bitbucket CLI to shell out to — and
mcptask_runner pr works there unchanged, because it goes through the same
interface gh does.

What to do on a Bitbucket host. Put a credential in the environment and
re-run mcptask_runner init, which writes it to
~/.mcptask_env.d/bitbucket_credentials (0600) beside the mcptask token, so it
reaches both your own shell and the scheduled job. One of:

  • BITBUCKET_ACCESS_TOKEN — a repository or workspace access token, sent as Bearer. The one to use for a machine.
  • BITBUCKET_EMAIL + BITBUCKET_API_TOKEN — the Atlassian account's email (not a Bitbucket username) and an API token, sent as Basic.

App passwords are not accepted. Atlassian stopped serving them on 9 June
2026; BITBUCKET_USERNAME / BITBUCKET_APP_PASSWORD are recognised only so that
init and the adapter can say so, rather than authenticating with something that
fails at merge time, after the work is done.

Setting both schemes at once is refused, naming both, and writes nothing — no
credential is ever picked on your behalf. A re-run of init with nothing set
leaves an existing credential file alone; setting the variables again is how a
rotated token is installed. Values are never printed, by init or by a 401.

What a GitHub host has to do. Nothing. gh still drives it, argv for argv.

One mapping worth knowing. Bitbucket's SUPERSEDED counts as DECLINED, not
MERGED: those commits may well be on the destination branch, but this pull
request was not merged, so an auto-squash run over one ends merge_unverified
and a person looks.

How far this is tested. The conformance suite now drives a Bitbucket project
in both modes against a scripted host, on all three coding CLIs: a manual run
that opens a pull request and puts it on the card, an auto-squash run whose merge
is confirmed by a GET on the pull request rather than by the agent's claim, a run
whose CI fails twice and leaves the pull request open, and the same recording as
that third one with the host answering OPEN — which comes out
merge_unverified. The last two are a pair on purpose: a runner that had stopped
verifying would pass the confirmed one and fail the other.

One environment variable comes with that, and is worth knowing about even though
almost nobody needs it. BITBUCKET_API_BASE_URL points the adapter at an API
root other than api.bitbucket.org — for a host behind a proxy that terminates
Atlassian's API on its own name, and for the conformance suite, which uses it to
reach a fake host on loopback. Unset means Atlassian's own root; a wrong value
fails at the first request, naming the URL it could not reach.

A once_dry run on a project with no assigned work no longer files a bug

once_dry used to have one ending in its prompt: success. So a dry run against
a project whose queue was empty had nothing to answer with — the child fetched
the next piece, was told "No available tasks or stories found", went looking for
one by hand, and then reported status: "error", which is a hard failure. The
runner filed a high-priority piece in its own Errors epic about a project that
was simply idle, and did so again every time that project was dry-run.

An empty queue is now the mode's own answer: the run ends no_more_tasks, the
log says No tasks available, the dry run had nothing to display, and nothing is
filed.

What to do. Nothing on the host — the change is in the prompt the runner
composes, so it takes effect with the new binary. Any pieces already filed under
Runner error: result — <project> for an idle project can be closed.

An API refusal is now named instead of retried

When the coding CLI ends its turn because the API refused it, the runner stops
and says which refusal it was. It no longer reads that as "the child forgot the
result marker", so the three --continue retries thirty seconds apart are gone —
none of them could ever have succeeded — and the ending is no longer filed as
stream_ended.

Four endings, and what each asks of you:

Termination What happened What to do model_unavailable the configured model is not being served — retired, removed, or spelled differently by this provider fix model: in config/mcptask_runner.yml; the message names the model and quotes the provider not_authenticated the CLI has no credentials sign the CLI in as the user the runner runs as on that host usage_limit the account is past a hard cap that clears on a date nothing, until that date — the message carries it verbatim api_error the API refused in words no harness profile recognises read the quoted text; if it is worth naming, add its phrase to the profile's usage_limit / model_unavailable / not_authenticated list

What changes for an operator. usage_limit files no bug piece — it is a
budget, like the day's quota and the usage window before it, and an account with a
monthly ceiling would otherwise file one piece a month. The day still ends,
because nothing can be spent until the cap clears. The other three do file a
piece, and it names the condition and the remedy rather than saying the stream
ended. The dashboard card and the run log's error_message now say the same
thing, where all three of the reports that prompted this carried null.

What does not change. A 429 still gets its eight patient waits and a 529 its
ten — a refusal the profile could name outranks them, an unnamed one does not — a
context overflow is still an overflow, and a turn that failed for a reason a
resume can fix still takes the marker retry it always had.

Recognition is Claude Code's wire shape for now, that being the only one these
conditions have been recorded on. A Codex or OpenCode host hitting the same wall
still spends the marker retries.

The lock guard no longer refuses commands that merely name a test run

check_test_lock — the PreToolUse hook that stops a second suite starting
beside somebody else's — decided what a test run was by looking for a recognised
invocation anywhere in the command text. A space counted as the start of a
command, and every word inside a quoted string has one in front of it, so a
command that only talked about a test run was refused while a suite was
running:

git commit -m 'fix: bin/ci now takes the lock'    BLOCKED, and it runs nothing
grep -q '=== bin/ci exit:' "$LOG"                 BLOCKED, and it is the line inside /ci-wait

The same anchor was wrong in the other direction, which is the half that
mattered: an invocation had to be preceded by whitespace or nothing at all, so
../other-worktree/bin/ci — how a story branch runs the gate — matched nothing
and a genuine second suite was waved straight through.

The command text is now read the way a shell reads it, into words, and only a
word in command position can be a test run: the first word of the line or of
a new command (after &&, ||, ;, |, &, a newline, (, `, $(),
optionally behind a wrapper that is not the command itself (sudo, env,
time, nohup, an option to one of those, a VAR=value assignment) or an
interpreter that introduces one (bash bin/ci, sh -c 'bin/ci', and
python -m pytest, which was not recognised before). Quotes and # comments
are tracked, and a path-bearing invocation counts. The
exemption for the lock's own helpers (test_lock, run_with_log) is decided
the same way, so a commit message naming one of them no longer exempts a real
test run sharing the line.

Blocking is unchanged for everything that actually runs a suite: bin/ci,
bin/ci --fast, cd worktree && bin/ci, and every command the project declares
in .claude/test-commands.json.

What to do. Run mcptask_runner update on each host — the hook is a helper
in ~/.claude/bin, not a file in any checkout, so a project that only pulls the
new binary keeps the old guard until the helpers are reinstalled. Nothing else
changes, and git commit -m with bin/ci in the message works again.

The machine-wide test lock: reinstall the helpers on every host

~/.claude/bin holds the copies of the helper scripts that actually run, so
none of the fixes below reaches a machine until its helpers are reinstalled —
mcptask_runner update on each host. A repository carrying the fix is not a
fixed machine, and the lock is the one piece of tooling where the difference is
invisible: a host still running the old test_lock keeps serialising its suites
exactly as badly as before, and says nothing about it.

Run the update while no suite is in progress on that host. The scripts are
replaced in place, and a run that is mid-flight through the old one keeps the
file it started with.

test_lock records COMMAND_PID and LOGFILE on Linux as well as macOS

set_command_pid and set_logfile edited the lockfile with sed -i and an
empty backup suffix as a separate argument, which is the BSD spelling: GNU sed
reads that empty string as the script, the real script as a filename, exits 4
and writes nothing. Both call sites discard the failure, so on Linux the lock
never gained either field and nothing said so.

Three consequences, all of them silent, all of them gone now: a lock whose
COMMAND_PID stayed empty was reapable while its suite was still running; ci_wait
could not tell a crashed run from a slow one, because the check for that reads
COMMAND_PID; and a LOCKED verdict could not name the other run's log, so
OTHER_LOG was always unknown on that platform.

Both fields are now written by filtering the key out of the file and renaming a
temporary copy over it — no in-place edit, nothing to spell two ways, and a
reader can no longer catch the file half-built. (task
https://mcptask.online/jchsoft/tasks/12402)

Every run through the helpers is ten seconds shorter

run_with_log started its stall watchdog with a command substitution, and a
backgrounded subshell inherits that substitution's stdout — so the line meant to
start a watchdog and move on blocked until the watchdog's next sleep returned.
Measured against a command that exits instantly: the command was gone at 0.03 s
and the Exit code: footer arrived at 10.8 s. Every run paid it, whatever the
command, because the wait, the tee teardown and the footer all queued behind
that line. The same measurement is now about a second.

/ci-wait and /test-wait answer sooner as a result, and so does anything
watching for the footer. ci_wait's 15-second footer grace is unchanged and now
pure headroom rather than a measured need — nothing pays it on a healthy run,
since it returns the moment the footer appears.

Two more corrections in the same file: the stall watchdog sized the log with
stat -f%z alone, which is BSD-only, so on Linux it read every log as zero
bytes and dumped thread traces into a run that was not stalled; and the watchdog
subshell's own stray output now lands in the log instead of on the launcher's
stderr.

run_with_log is also now maintained in this repository rather than vendored
from the retired Ruby gem, which is what made the fix possible here at all.
(tasks https://mcptask.online/jchsoft/tasks/12401,
https://mcptask.online/jchsoft/tasks/12402)

A lock is now held for as long as the run, and released only by that run

This is the behaviour change to read before updating a shared host. The lock did
not protect what it appeared to protect, in three independent ways, and all three
were silent while they happened.

A lock is stale when nothing it stands for is alive, and nothing else makes it
stale.
Age is no longer consulted at all. Before, age >= 900 was tested
first, ahead of any look at the command, so a suite that ran past fifteen
minutes lost its lock while it was working — and a full bin/ci runs to about
that mark. A lock whose command is alive now keeps it at any age; a lock whose
recorded command has died is over immediately.

An empty command pid is no longer a ten-second fuse. acquire used to leave
COMMAND_PID empty and fall through to a pidfile only run_with_log ever
writes, so a lock taken around anything else — a bare bin/ci, an operator
being a good citizen — was reapable from its tenth second, permanently, with no
warning to either side. acquire now records the process that asked for the
lock, and the sentinel's 900 s backs that up, so an unregistered lock lasts as
long as the shell that took it and then as long as the sentinel. The pidfile
branch is gone, and with it a path whose two sides derived the same filename from
different places.

A release now has to prove it is the run that acquired.
release_if_owner compared CALLER and PROJECT — a constant per skill and a
constant per worktree, both reused by a re-run after a rebase — so a run whose
own lock had already been reaped could reach its cleanup and release the NEXT
run's lock, two minutes into a suite that had done nothing wrong. acquire now
issues a token, prints it after ACQUIRED and records it; release_if_owner
<caller> [instance]
takes the token, or the command pid, or the holder pid, and
answers not-owner when the instance is not this lock's. run_with_log and
bin/ci pass theirs. A release with no instance still works, for callers that
predate the argument, and now says in as many words that it matched on the name
alone.

test_lock status says which case a lock is in. HELD_BY=command … (does
not age out), HELD_BY=holder … (no command pid recorded yet), HELD_BY=sentinel
…
with the time left, or STALE=yes. The first line still starts with LOCKED
or FREE, which is what /wait-unlock reads.

The PreToolUse guard, check_test_lock, was rewritten to the same rules in the
same commit. It had its own copy of the old ones, so on both counts above it
would have waved a test command straight into a running suite while test_lock
was still refusing to hand the lock over.

Two side effects worth knowing. run_with_log now records in the log whether the
lock learned its command pid — the line reads Lock: COMMAND_PID set to …, or
Lock: no lockfile — this run is NOT serialised against other suites on this
machine
, which is a legitimate state for a run nobody took a lock for and a
useful thing to be able to check afterwards. And the Windows flake in the
lock-guard tests around the ten-second cliff
(https://mcptask.online/jchsoft/tasks/11897) is gone with the cliff itself.
(task https://mcptask.online/jchsoft/tasks/11852)

Pull requests go through mcptask_runner pr, on whichever host the project is on

The prompts the runner composes used to say gh. Not always out loud, which was
the harder half: the auto-squash spine merged with gh pr merge --squash
--delete-branch
and gated its own success on gh pr view --json state, while
the CREATE PULL REQUEST step named no command at all — and a step that names no
command gets gh anyway, because that is what a model reaches for when nothing
says otherwise. Either way a project on Bitbucket Cloud was being told to run a
binary it has not got.

Every one of those is now mcptask_runner pr create|list|view|merge|checks,
which asks the adapter git_host: resolved for that project (task
https://mcptask.online/jchsoft/tasks/12362 put the adapters in). The command
prints one JSON object on stdout, and the workflow tells the child to read
.pull_request.number out of it rather than from its own recollection. A new
bundled skill, pr, carries the five commands and their output shapes.

The runner reads that JSON too. The pull request a run opened used to reach
the dashboard through an event only Claude Code emits, so a Codex or OpenCode run
opened one and the card said nothing; and its NUMBER reached the runner only when
the child put it in the result marker, which about three quarters of real markers
do not. Both now come out of the command's own output, which every harness hands
back the same way. A result with no pr_number takes the one the run was seen to
open — the child's own answer still wins when it gives one, and nothing is
invented when no pull request was opened at all.

Permissions are now per host. Bash(mcptask_runner pr:*) is in the baseline
every project gets. Bash(gh:*), Bash(gh pr checks:*), Bash(gh pr view:*)
and the three github.com WebFetch domains are installed only into a project
whose git host resolves to github; api.bitbucket.org and bitbucket.org only
into a Bitbucket one. A project whose host cannot be worked out gets the shared
list and nothing else — nothing falls back to another host's approvals. Existing
entries are never removed, so a project that already has Bash(gh:*) keeps it.

The pull-request template default is GitHub's alone.
.github/pull_request_template.md is GitHub's own convention, and telling a
Bitbucket project to follow it meant the agent looked, found nothing, and wrote
whatever description it liked, with nothing reporting that the instruction had
missed. A git_host: github project still gets that path and that line, byte for
byte. Any other project gets whatever it declares under pr_template: path: and,
declaring nothing, gets no template line in the prompt at all.

What to do. Run mcptask_runner update on each host: the pr skill is a new
file under the project's skills directory and the permission split is a merge
into .claude/settings.local.json, neither of which arrives with the binary
alone. Run it while no runner is working that checkout — the new files would
otherwise land inside somebody's task commit. Nothing else is required, and a
project already carrying git_host: needs no config change; one that does not
gets its host read off git remote get-url origin at run time, and
mcptask_runner init writes it down for good.

Still naming gh: the two review executors (review and reviews). They
read review threads and enumerate every open pull request on a project, and
internal/prhost models neither — its six questions are create, list-for-task,
view, by-branch, merge and checks. Rewording them to name no host would leave an
agent told what not to use and not told what to use instead, so they are excluded
by name, with the reason in words, in TestSpineNamesNoGitHost. Extending the
interface to cover them is its own piece.
(task https://mcptask.online/jchsoft/tasks/12364)

/ci-runner reads a green run that prints no summary block

A bin/ci that prints one line per step and no aggregation at all — no CI
SUMMARY
, no TOTAL RESULTS — matched neither shape ci_wait recognised, so
every green run on such a host ended in NO_SUMMARY_RECOGNISED plus twenty
lines of tail. That tail was not empty, which is what made it hard to spot: on
the run that filed this it happened to carry three of the nine step lines and
lost both of the ones a reader triages by, so the orchestrator reported a ragged
excerpt as if it were the run's summary and could give no step count or timings.

ci_wait now recognises a third shape, the per-step lines themselves:

✅ <step name> passed in 4.88s
❌ <step name> failed in 1m31.30s

They are emitted in order, ANSI-clean, in place of the fallback. A red run
carries both spellings — later steps keep running after one fails — and the
failure path already passed them through its tail filter untouched, so the same
shape now reads on both exits. NO_SUMMARY_RECOGNISED stays for a log that
matches none of the three, and now names all three in its message; it also no
longer spends its twenty lines on the log's own ==== rules, which the two
recognised branches had always dropped and it had not.

/ci-runner's skill body documents both shapes side by side instead of the
older one alone.

What to do. The fix is in ci_wait, which is installed onto the host rather
than compiled into the binary, so a new binary alone does not change it: run
mcptask_runner update in the project to refresh ~/.claude/bin/ci_wait and the
/ci-runner skill. (task https://mcptask.online/jchsoft/tasks/12413)


Changelog

Fixes

  • 1b6b96a7af616625a981712fcd6db8537a053382 fix(dry): an empty queue is an answer, not a bug piece (task #12407)
  • 1ac4ad2b9e8cadbecc497c93cf811781e469efd6 fix(eventstream): the throttle's clock is shared with the stream's own goroutines, so a test hands it over under a mutex (task #11688)
  • b463017bb9094c5304ae4315ff932e9a242a9034 fix(executor): the API refusing a run is named, not retried for a marker it will never give (tasks #12405, #12406, #11850)
  • a74dab9409d1133ef38a1a0a4fa2cbb388fd1e58 fix(helpers): a lock is held while its run lives, and given back only by that run (task #11852)
  • 584a71dd709ce416d09035aed18e4a88bf10d492 fix(helpers): the lock records its command pid and its logfile on either userland (task #12402)
  • 3b202a43d607cbc6957ba3faed2d39c1e3d61094 fix(helpers): the stall watchdog no longer holds the footer for a whole interval (task #12401)
  • 50b0446ede6def664d97428f61ce8ac379fa1d66 fix(stress): an orphan may end itself, and the sweep says which ones did (task #12403) ### Other
  • 4867ec85fccea0e4354333da18c647dd2b69ae3a Recognise an API refusal on every dialect, not just Claude Code's
  • 90353f488c2a233cbe78e862f466798b6e278ca3 Skip the 0600 check for the Bitbucket credential file on Windows
  • 0cb0ada0f2c2de1f328769c6bb44210d77f224b3 [#12362] Ask a PRHost, not gh: internal/prhost + mcptask_runner pr
  • e0a591aafc5e606230bddcebdde5d915b68456d3 [#12363] Drive Bitbucket Cloud: a prhost adapter over REST API 2.0
  • 6333beb769954a50eb88a6b3d22bab9c2d2fb6f8 [#12364] Say mcptask_runner pr, not gh: the prompt, a skill, and per-host approvals
  • 0d46b797c5cbf4e178324077b0c412c663eaf486 [#12365] Drive a Bitbucket project in conformance, and say so in the README
  • 1e4f48e9ef1deb729807cf4f4a679691ef1d3aa0 ci_wait: read a step record that comes with no summary block
  • e4418e46f7a757cd60f5264b387b74cecf8cce12 docs(release): v0.3.24 shipped, the next-tag notes start empty
  • 2425e84c9189223aafb21084742468d0f723f5f9 fix(check-test-lock): naming a test run is not running one (task #11674)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.24

Upgrading

Two commands this time: mcptask_runner update --self for the binary, then a bare mcptask_runner update in each host project — it writes the now-required harness: key and rewrites the helper scripts in ~/.claude/bin. Details below.

harness: is now required in config/mcptask_runner.yml

The runner drives a coding CLI named by a profile, and there is no default. A
project whose config does not name one refuses to start:

harness: is not set in config/mcptask_runner.yml, run `mcptask_runner init --cli <name>`; known profiles: claude, codex, opencode

What to do. Run mcptask_runner update on each host. A project this runner
already installed into — one with a .mcp.json and a
.claude/skills/.mcptask_runner_manifest.json — has harness: claude written
for it once, and the update says so in its output. That is the whole migration
and it changes nothing else about how the project runs.

A project update does not recognise (and any new project) needs
mcptask_runner init --cli <name> instead. init without --cli on a project
that has never named one is refused rather than assuming claude: assuming it is
how a host that installed a different CLI gets a Claude command line built for
it, fails at the first launch, and is told about a binary rather than about the
key nobody set.

A host that already sets launcher.command has said which CLI it starts, so
init and update now derive the key from it when exactly one profile claims
that name — printed, written into the file, never inferred at run time — and a
run whose launcher.command contradicts its harness: key is refused before it
starts, while a command whose program cannot be read (a shell, a wrapper script)
is logged and allowed; point launcher.command at a mock and you write
launcher.cli_check: false beside it, and a CLI that really runs under another
name goes in the profile's binary.aliases.

Where the harness shows up: the runner card on mcptask.online carries a harness
label beside the model (an additive snapshot member the server accepts since
2026-09-08; older runners keep working without it); mcptask_runner version and
the scheduled job's per-run log open with a Harness: <name> -> <binary> line
after the version line; a runner bug piece names the harness and attaches the
profile's own config files instead of the Claude Code settings only.

init --cli also adds config/mcptask_runner.yml to .gitignore if it is not
ignored already. The file is per-machine — harness, schedule times, model pins —
and two developers sharing one committed copy overwrite each other's answers.

Codex CLI support

mcptask_runner init --cli codex sets a project up for the Codex CLI instead of
Claude Code. The work loop, the watchdog, the dashboard card and the mcptask.online
protocol are unchanged; what differs is the command line, the words the prompt
uses for tools and skills, and what init writes:

  • .codex/config.toml in the project, with the mcptask.online MCP server marked required = true. The entry names the environment variable holding the bearer token (bearer_token_env_var) and never the token itself, so the file is safe to commit.
  • [projects."<path>"] trust_level = "trusted" in ~/.codex/config.toml (or under $CODEX_HOME), because Codex reads a project-level config only for a project it trusts.
  • Eight skills into .agents/skills, invoked as $test-runner, $ci-runner and so on. Helper scripts stay in ~/.claude/bin, shared with any claude project on the same machine.
  • AGENTS.md, pointing at the project's existing CLAUDE.md rather than copying it.

Both config writes append and never rewrite: an entry that already says what the
runner would have written is left alone, and one that says something else is
reported as a conflict and left as it is.

One thing to do by hand: codex login. The runner never logs a CLI in. It
checks, before a run starts, that this one is logged in already — codex login
status
has to exit 0 — and refuses with the probe's own output when it does not.
An unauthenticated headless Codex produces a failed turn with nothing done, for
as long as the loop runs.

Read the sandbox note before scheduling an unattended Codex run. The
implementing executors launch codex exec with
--dangerously-bypass-approvals-and-sandbox: full access, no approval prompts,
which is the parity with Claude Code's bypassPermissions this runner has always
used. Dry runs, triage and review get --sandbox read-only, plus one -c
override that approves the mcptask-online server's MCP tools and nothing else.
That second token is not optional: codex exec is non-interactive, which means
its approval policy is never, and never REFUSES an MCP tool call rather than
waving it through — without the override, every read-only-class run is cut off
from mcptask.online and cannot fetch the task it was started for. A write from a
read-only class is still refused by the sandbox. The checkout the runner is
pointed at is the only boundary there is.

How far Codex support has been proven. The conformance suite runs on
--dialect codex as its own CI job (twelve of fourteen scenarios; 06 and 07
describe a state Codex cannot reach and say so in the file), and the chaos stress
day runs on it too. Both drive a mock CLI.

Since 2026-09-09 the real codex binary has been driven as well, on codex-cli
0.153.4: a dry run, and a task worked end to end — branch, edit, go test,
bin/ci, pull request — with five progress logs left on the piece along the way.
Two faults that only a real CLI could show came out of that and are fixed here.
The approval-policy one is described above. The other is that continuations
were refused outright
: codex exec resume is a sub-subcommand with a smaller
option set than codex exec, taking neither -C nor --sandbox, so every retry
and every marker retry died with exit 2 and an argument-parser usage message. A
Codex host running an earlier build has never had a working retry. The flags now
precede resume, which is the order codex exec [OPTIONS] <COMMAND> documents.

Still unexercised: the rate-limit and context-overflow literals are source-derived
and no real run has reached either.

Claude Code hosts are unaffected by any of this beyond the harness: key.

Codex asks for its plan tool by name. codex-cli 0.153.4 does not offer
update_plan unless it is enabled, so every Codex command line now carries
-c tools.update_plan.enabled=true. Without it the child could log its progress
but never post a plan, and the runner card showed a run with no steps for its
whole duration. The prompt now also describes the tool the CLI actually has:
pending, in_progress and completed, exactly one step in progress at a
time, the whole list on every call.

Codex runs the test helpers directly, instead of a skill it cannot invoke.
A Codex run used to reach "UNIT TESTS", read $test-runner — codex exec does
load .agents/skills without being asked, so the skill really was there — and
stop. That skill is a state machine which drives itself through a sub-skill tool
Codex has not got, and its own rules forbid running the test command directly
because that would skip the machine-wide test lock, so the child was left with no
legal move. It refused, correctly, and every task_manual run on this harness
died at the same step. The prompt now names the lock-taking helper scripts the
installer already puts in ~/.claude/bin: test_start or ci_start to take the
lock and launch the run detached, then ci_wait polled for the exit code. Same
lock, one wrapper layer down — nothing about how runs are serialised on a shared
machine changes. Claude Code hosts keep invoking /test-runner and /ci-runner
and are unaffected.

A detached run now starts in a session of its own, and a fast one is no longer
called dead.
Two more defects, found by the first Codex run that actually
reached the helpers, both in the helpers themselves and both fixed for every
harness. test_start and ci_start used to detach with nohup and disown,
which survives the launching shell exiting but not its process group being
signalled — and a Codex child runs each command in a group of its own and
kills that group when the command returns, so the run died within seconds,
leaving no Exit code: footer and the lock held by a dead pid. Both scripts now
start the run in a new session: setsid where there is one, otherwise Ruby,
which this helper family already requires, and if neither is present they refuse
and say so rather than launching something that will be reaped. Separately,
ci_wait allowed two seconds between the command's process disappearing and the
footer arriving, where the real gap is about ten — the footer is written after
run_with_log has torn down its stall watchdog — so a sub-second command was
reported FAILED_EXTERNAL on a suite that then exited 0. That window is now 15
seconds, polled twice a second, and such a run is reported by its actual exit
code. The Codex vocabulary also spells the invocation as
test_start 600 cmd arg arg rather than test_start 600 <command>, and
test_start now refuses a whole command quoted into one argument by name
instead of guessing where to split it.

What to do. The same two-step upgrade the next section describes: these are
helper scripts in ~/.claude/bin rather than code in the binary, so
mcptask_runner update --self and then a bare mcptask_runner update in any one
project on the host.

bin/ci takes the test lock itself — reinstall the helpers

A bin/ci started by hand, or by a CLI child that found the script on disk,
used to run beside whoever was mid-suite; only the /ci-runner and
/test-runner skills took the machine-wide lock. The script now takes it
itself, and refuses out loud (exit 2, naming the holder) when somebody else has
it. The skills keep working through a marker: ci_start and test_start take
the lock first, as they always did, and export MCPTASK_TEST_LOCK_HELD so the
script knows the lock is its own.

What to do. The helpers live in ~/.claude/bin, not in the binary, so this
version needs the two-step upgrade: mcptask_runner update --self for the
binary, then a bare mcptask_runner update in any one project on the host,
which rewrites the helper scripts from the new binary. Until that second step
runs, a suite started through /ci-runner finds the lock held by the old
ci_start — which does not export the marker — and refuses itself as if
another agent held it. A hand-run bin/ci is fine either way.

Changelog

Features

  • c4eccf4f65a2cadec7e33fcff2896a27b63dec35 feat(bugreport): the bug piece names the harness and attaches the profile's config files (task #12300)
  • 9362d47e1beb2748c46af30e3c87d3f236d6dc86 feat(capture): capture mode records the child's raw stream for building new dialect fixtures (task #12287)
  • 037ee1d0bed0f907879ce6a9919586abeffaf28f feat(config): the version banner and the launcher log name the harness and its resolved binary (task #12301)
  • 48084db83a10dd77dfd9704861ab26d6f020940e feat(conformance): scenarios in normalized events, mock-cli --dialect with a renderer, chaos on the renderer (task #12286)
  • e804acbfb869f8c34e52a670166d04e7d640171a feat(conformance): the suite, mock-cli, chaos and the stress day run on the codex dialect (task #12291)
  • d23be25ec8018da875a9387ac27970667c2047d0 feat(harness): bundled codex.yml profile; command.order for subcommand CLIs; codex argv goldens (task #12288)
  • 044eea0e434dbc711436eeae289860aea6fb3e73 feat(harness): bundled opencode.yml profile; positional prompt with no token; readonly enforced via config (task #12295)
  • 81ad327cc9c069ee70d34aeb8b1f09f0e229147a feat(harness): codex failure literals from codex-rs 16ff14c, each cited by file and line (task #12293, part 1)
  • 20464a1d8ba009e00e04c6bc06b418bcd10d5197 feat(harness): codex prompt vocabulary: $test-runner/$ci-runner, MCP-tool piece fetch, update_plan block (task #12290)
  • 7fab6b61730ab89a47c14b363d3c7fba067df432 feat(harness): codex_jsonl decoder and renderer; resume by thread id from the captured session (task #12289)
  • ab5e95d01d01983505ecbcd435a7c22808483b89 feat(harness): every progress anchor in the spine spells the log tool through vocabulary.log_tool (task #12370)
  • 991f448a2e463139e69704d7e6d502a83c617c8d feat(harness): launcher.command is cross-checked against harness: at run time and derives the key at install time, never inferred at run time (task #12357)
  • 59bc636ee6aae849f258084538e03ffae4463b50 feat(harness): log lines name the harness display name; the engine delegates binary lookup to the profile (task #12285)
  • 7fcce96d85d09f100f68e1e1cc0b88af47b3047b feat(harness): opencode failure names from the OpenCode source, model tiers from the real CLI, real-run findings (task #12298)
  • 14a937f96d8b3b8dc1b9ad891ed0dd4ad4d98fdc feat(harness): opencode_jsonl decoder and renderer; result from marker and exit, never from step_finish (task #12296)
  • 438efa65eddc437eb5861090cbea78fc7a2689c8 feat(harness): profile schema, loader and the harness: config key with no default (task #12280)
  • 77962344c5bce58bdb178ae91d2958f45ae23a90 feat(harness): the claude_stream_json decoder emits normalized events; every consumer reads events and ToolKind (task #12282)
  • 8c883865a82d75827ac5ce69f437f10acf464419 feat(harness): the current-user fetch comes from vocabulary.current_user, not a literal resource line in the spine (task #12331)
  • 6af102c1f138dc30fc1cb5e44d01ace65c47262a feat(harness): the engine builds the child command from the claude profile (task #12281)
  • 6a989e2dd9db0c6b02d4808d53a76f8e50d56c5f feat(harness): the prompt spine reads its harness fragments from the profile vocabulary (task #12283)
  • 9c94208a607e941e7db261149cda8d11e2f7a19b feat(install): codex install section: .codex/config.toml MCP server, project trust, .agents/skills, AGENTS.md pointer, login preflight (task #12292)
  • 3182be07fffc1e83a7ddd102d4ab2900c7f3bd1b feat(install): opencode install section: opencode.json remote MCP writer, readonly overlay via OPENCODE_CONFIG_CONTENT, real vocabulary (task #12297)
  • 29452358747c03f1348feb98265499694706ce1a feat(install): the helper home is the profile's install.helpers_dir, rendered into skills and helpers at install time (task #12303)
  • 81ec85ca672e6e679e57e7b906ae27d7d30ddc9c feat(install): the installer reads its layout from the harness profile; init --cli, loud update migration, binary preflight (task #12284)
  • 35bb8f0110288dd68d07eee79c8bd36509086cb1 feat(snapshot): the additive harness member on the wire, in the run log and in the handoff note (task #12299) ### Fixes
  • 7d4bffd66f8eb1fe6b5d1efdaa9737e3a126dba7 fix(banner): the Models line names the profile whose ids it prints (task #12381)
  • 55df871aae8eff9c3bd5cdfbdfef45420125fd09 fix(ci): bin/ci takes the machine-wide test lock itself when not started through ci_start, and refuses when it is held (task #12369)
  • 298dc2022dcebf501c96f5a9d1b2fe003820eed1 fix(codex): ask for the plan tool, and describe the one this CLI has (task #12383)
  • b0f4dfa8113e57847d2f58d6fddccd4f0107a75c fix(codex): real runs on the ChatGPT plan — read-only class keeps its MCP server, continuations parse, refusals read from the wire (task #12293)
  • 7e42be2c3800bd4216fe7ac33d64d8912f9943b9 fix(codex): the prompt drives the lock's helper scripts, not a skill this CLI cannot invoke (task #12395)
  • 26a8adfc848505219155fcb529598dcdd78756dd fix(eventstream): a failed dial no longer prints the token it dialled with (task #12379)
  • 1b3a269360341a021ef2b557c95865e54bc96ce5 fix(executor): a run that answered with the marker is labelled result on every dialect, not stream_ended (task #12368)
  • 5815b2545a0afa62b046aef520afb23e56a48f56 fix(executor): the stderr label names the CLI that wrote the line (task #12380)
  • fc4178d720ed8b9483b063995268457491235523 fix(harness): codex exec takes --dangerously-bypass-approvals-and-sandbox or --sandbox read-only, never -a (task #12332)
  • f39750da54afbfb2935f773947d08a69c61a6780 fix(harness): opencode progress anchors name the callable log tool on every line (task #12367)
  • cb5a352cfb9f3ec3096e1baef06e1a90435b4071 fix(harness): the opencode readonly class keeps the shell and refuses writes by pattern (task #12353)
  • 77767eef92910fa8981498d2c5f13151ff21c496 fix(helpers): a detached run gets a session of its own, and a fast one is no longer called dead (tasks #12399, #12397)
  • 7c1d5abae8159d505e245705432c849dd68160ce fix(init): a scripted install is asked of the OS, not of the file mode (task #12382)
  • 5f4360b7d9c46d069093365af33c7be1895c558d fix(opencode): the plan comment stops claiming update_plan has no in_progress (task #12391)
  • 068f188f5dc5ff2adc3fa7bf0e5645dd9fd7390e fix(smoke): the banner checks ask a directory with no config (task #12384)
  • 381260a399407a1a859cf691eed5c8e0181d723b fix(tests): the init refusal test expects the host path, not the git spelling (story #12276) ### Other
  • d34e6c905f24482cf5c5f36c0e5184947023b4cc docs(ci): the codex conformance job comment counts its exclusions correctly (task #12294)
  • c3d811c3b2b7fe14b7c44f33dc9fa2aa5bcfbe45 docs(codex): the plan comment stops asking for a correction opencode.yml already carries (task #12391)
  • b679dd3947a85012cbc7064ab38528f6f62d9a74 docs(harness): README harness section, Codex host specifics, requirements, npm and CLI texts, release notes draft (task #12294)
  • c0583d6f795df1c4f39bc43b8815d490d86aa5ce docs(harness): the opencode skills comment reflects the self-locking local gate from task #12369 (story #12278)
  • 39ae784fb4cc661975f43cebceee59036c3f3c06 docs(release): the plan-tool flag, and the two-step upgrade the self-locking gate needs (story #12277)
  • 8ce7005b0b01d1df78928c5750d8cacf4d8c2e77 docs(release): where the harness shows up: card, version banner, launcher log, bug piece (story #12279)
  • 6dfba41f87808b2d5ec372b487b28cb74739676f merge: story #12276 — harness profile layer: Claude Code behind a data profile, no behaviour change
  • 0aa6545ba0faca27e80c14ccde4042d37f3b50fd merge: story #12277 follow-ups — token-free dial errors, per-harness stderr label, profile-true banner, scripted init, codex plan tool, smoke in a clean dir (tasks #12379-#12384)
  • b058a7d06faf042aaaed16c8d1d2d13ff52ef396 merge: story #12277 round 3 — plan-state comments agree, first real codex plan captured, release notes for the plan flag and the helper upgrade (tasks #12391, #12392)
  • bf250676dd324a0c6698be4bc00f10f0138a413b merge: story #12277 round 4 — codex drives the lock helpers, helpers survive their launcher, ci_wait waits for the footer (tasks #12395, #12399, #12397)
  • d5d8a86187669a6b0a2ab505b15be2173eae1c0b merge: story #12277 — Codex CLI harness (profile, decoder, vocabulary, install, conformance, docs; real-CLI runs pending login)
  • e2ee3141f3f18e6e9400388f27df42a934ee1c80 merge: story #12278 — OpenCode harness (profile, decoder, install, readonly by pattern, real runs verified; self-locking local gate)
  • 59e7ebd7a4066966e9425816201ff6add0225dae merge: story #12279 — harness visibility (snapshot member, bug piece, version banner, helpers_dir, launcher.command cross-check, log_tool vocabulary)
  • 650a002d65c1d9d3dbcc9fa223ef8ffd2bf01ad4 merge: task #12293 part 2 — codex driven for real (read-only MCP approval, resume argv, wire-shaped refusals)
  • 333b34da96f1ef8662e13bae51db95b91cfc0ced test(codex): the first real plan of a runner run, from the capture (task #12392)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.23

Upgrading

The usual one command replaces the binary:

mcptask_runner update --self

The bundled skills-and-helpers pack did not change in this release, so a
follow-up bare update is not required. Two behaviour changes worth knowing
before an unattended host picks this up:

  • --ignore-quota now means ignore, everywhere. The daily mode's day-level scheduler no longer queries mcptask.online about the budget when the flag is set, so a REST outage cannot park an unattended daily host until the next business day. The end-of-workday clock still applies — it is a time rule, not a quota rule (#11891).
  • The work window keeps its minutes. init --until 18:30 stores the new additive work_window.end_of_workday_minute key and the loop stops picking up new work at 18:30, not 18:00. Configs without the key mean :00, so existing hosts change nothing; a value that cannot be stored is refused loudly, never rounded (#11892, #11896).

And one to expect on your next bare update: locally modified helper
scripts
now get the same five-outcome manifest protection as skills — a
hand-edited helper reports conflict-skipped instead of being silently
overwritten, and FORCE=1 leaves a .bak beside anything it replaces
(#11894). Also in this release: the streamable-HTTP .mcp.json entry for
fresh installs (#11873, #11875), the token reaching interactive terminals
(#11874), and the ruby conformance baseline moving to nightly-only CI
(#11893).

Changelog

Features

  • 335d101837b7738278e105c3c46d93ca41802bcc feat(config): the work window keeps its minutes (task #11896)
  • 937d21249d4b1f6af9f16998672ddd731275d19a feat(update): carry a legacy SSE mcptask-online entry forward (task #11875) ### Fixes
  • e3ac5976b98e2c189534e3174803fea81935fc88 fix(install): --until refuses minutes it cannot store (task #11892)
  • 525daab5394e4d244bb5f52434a9195a24eec7fa fix(install): a fresh .mcp.json gets the streamable-HTTP entry (task #11873)
  • a8277924b6ea1dcdb274d9ec4cce211443650067 fix(install): helpers get the same five-outcome protection as skills (task #11894)
  • 22525d047a05382f6842127ab1d5953b6f99a958 fix(install): the token reaches a terminal, not only the scheduled job (task #11874)
  • 232399064826689e4dded85488bb5eed1b90ef80 fix(runner): --ignore-quota covers the daily scheduler too (task #11891)
  • 4c870dcc01c07fdf47f3eea5445e3ca817c434d7 fix(tests): the lock-guard test resets its clock before every subtest (task #11897)
  • 9e7abcc1af46b18293a54eb9fb5d2ebf3cbc2ce5 fix(tests): the shell-file table joins its paths instead of spelling them (task #11880) ### Other
  • aadfcf8d051948d4d9db6a04af39a1e4bcd96cb4 ci: the ruby baseline runs nightly only, and the docs stop contradicting it (task #11893)
  • be020fa81612407e1cfddbdfdb03054d7e1b94f0 docs(comments): the outage window is named by its constants, and childEnv stops teaching the pre-#11670 model story (task #11895)
  • 495223cb63bf2f8e1bee3ddbd3079163a51cc6bd merge: runner-audit fixes — tasks #11891-#11896

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.22

Upgrading takes two commands this time, not one

mcptask_runner update --self     # replaces the binary
mcptask_runner update            # in each host project

The first is the usual one and covers every loop change below. It is not
enough for the skills and helpers: those are a data pack the binary carries, and
only the second command writes them into a project.

Skipping it leaves the old /test-runner and /ci-runner in place — and on a
project that is not Rails those name commands it cannot run, which is the very
thing this version fixes. A host that has locally edited a skill keeps its edit;
FORCE=1 takes the new one and leaves a .bak beside it.

Worth knowing why this is easy to get wrong: the data pack lives inside the
binary, so an update run through the old binary reinstalls the old skills.
The order above is the whole trick.

One runner, every framework

The bundled CI and test skills stop naming a framework. /test-runner and
/ci-runner used to spell out bin/rails test, test:system and RuboCop, and
check_test_lock recognised a running suite by six regexes, five of which were
Ruby. On a Django, Node or Swift host the commands did not run, and the guard
against a second concurrent suite waved every real test command through — blind
on exactly the projects a universal runner exists for.

Both now read what the project declares in .claude/test-commands.json, which
mcptask_runner init writes and update backfills. With no such file the
toolchain is detected from go.mod, Gemfile, package.json, manage.py,
pyproject.toml or Package.swift — and when nothing is recognised the runner
says so rather than inventing a command that will fail later and elsewhere.
(#11853)

Two more places the prompt spoke Rails to everyone. The auto-squash workflow
told every project to run bin/rails assets:precompile RAILS_ENV=test between
its unit and system suites, and the manual screenshot step named a Rails class.
Neither is measured by what it breaks — the first costs a turn or invites a
plausible substitute, the second is silently ignored on a project that has no
such class — but both contradict the claim the product makes.
(#11858,
#11864)

A test now guards it: every executor prompt, both configs, is checked against
two lists — commands only one ecosystem can run, and identifiers that exist in
only one.

The working day

A failed task waits an hour, not the rest of the day. A task that fails is
set aside and the runner carries on with the next one.
(#11846)

The dashboard can take a piece off the skip list. A piece set aside earlier
can be released from the card, and the runner picks it up again without a
restart. (#11847)

The dashboard connection

A closed socket is described by whoever actually saw it close. Two paths
reported a dropped connection — the read loop, which knows the close code the
server sent, and a failed write, which knows only that the socket was gone.
Whichever finished first wrote the notice, so a deliberate close with a code
could be reported as a bare transport error.

The reader's account now wins. That is what makes "the server closed it, code
1001" worth trusting when you are deciding whether the server needs looking at,
and it is the difference between a server letting a connection go and a
connection dying underneath both ends.
(#11860)

Under the hood

The loop's test logger is safe to write from goroutines — a data race the race
detector caught on CI while forty local runs stayed green
(#11859) — and a stress day that
overruns its wall clock no longer leaves a child writing into a temporary
directory that has already been removed
(#11865). Neither changes runner
behaviour.


Verified before tagging: the full local bin/ci green on all ten steps with no
skips, including the stress day and both conformance suites, and GitHub CI green
across three operating systems and the Go floor from go.mod.

v0.3.21

The one that was costing whole days

A deploy of mcptask.online now costs one attempt instead of the working day (task #11832).

While the site is being deployed its MCP server is unreachable for minutes and then comes back on its own. The runner spent its entire recovery budget inside about forty seconds — a child dies in ten, five seconds between restarts, two restarts — so all three attempts lost the same race before the server could possibly be back, and the run ended with error. That status is what ends the day: one deploy could stop a day that had already finished eighteen tasks.

Restarting faster cannot fix a wait. The two situations that reach that path are now told apart where they are detected, rather than by reading the message afterwards:

  • A deferred tool that never loaded is the child's own miss. It keeps its two fast restarts and still ends in an error, because it is not going to fix itself.
  • An outage gets six restarts at twelve times the wait — about six minutes, which spans a deploy — and when that runs out it ends the way a stall does: a verdict the loop carries on from, the task left in_progress for the next triage, and no bug piece, because a deploy is not a defect anybody can fix. The wait keeps the heartbeat going, so a card watched through a deploy says the runner is waiting rather than falling silent.

This was the largest cluster in the Errors epic — five of twelve real failures, across three machines and two host projects.

The one you will see

Today's skip list is visible (task #11835). A task the runner has set aside — out of scope, failed, already taken by somebody else, a merge it cannot finish — is skipped by every later triage that day, and until now the only trace was one log line at the moment it was decided. It now shows on the startup banner, in every wait, and on the dashboard card. A quiet runner can be read rather than guessed at.

Also

main is unprotected on purpose, and the two habits that stand in for a branch rule are now written down (task #11831); a new push into a pull request stops the CI run it supersedes (task #11826).

Upgrading

Run mcptask_runner update --self on each host, or move a project's gem pin to 0.3.21 and run rake mcptask_runner:update. The binary is per machine and the newest one any project offers is the one that stays, so one project is enough.

Until a host actually runs this binary, a deploy of mcptask.online goes on costing it the day.

v0.3.20

Upgrading: run mcptask_runner update --self on each host, and do it for
this one rather than at leisure. The fix below can only reach a runner that is
running the new binary, and until it does, a task moved between two agents keeps
being worked by both. A daily process takes the new copy at its night swap — but
only once that copy is on disk, so the update still has to happen. Nothing to
migrate; no configuration changes.

What changed

A piece assigned to somebody else while this runner is working it now stops
the child.
The dashboard socket has always carried both halves of an
assignment — "given to you" cut an idle wait short, "given to somebody else"
reached nothing — so a task moved between two agents was picked up by its new
owner while its old one worked on to the end. Two agents on one piece, the same
branch name pushed twice, two sets of efforts logged. The runner now abandons
the piece, says so in the log, and goes back to the queue for one that is still
its own. It is not a reason to stop the day, and the piece is deliberately not
set aside: the queue has already stopped offering it, and it stays workable if
it is handed back. Neither quota ending loses its priority — a run killed by a
crossed budget and a reassignment in the same breath still stops the day, since
that is the claim about the whole day and this is a claim about one piece.

The run log is coloured as it is read, and the file itself stays plain. The
two lines that matter in a wall of uniform text — the ones Warn and Error
already mark with a glyph — are painted red and yellow, a child's stderr is
yellow, and the [Component] tags, cost lines and phase separators are dimmed
so the shape of a run is visible in a scroll. Only what the runner itself marked
gets coloured; guessing a severity from the wording would be wrong on the day it
mattered. Nothing is written to the file, which is what bug reports attach and
what people grep. runner-log --paint is the same filter over stdin, so
tail -f any.log | runner-log --paint works for a log read some other way.

The stress harness gained a day that proves the first of these end to end: two
otherwise identical days differ only in whether the broadcast names the piece in
flight, and they come out opposite — one abandons, the other leaves the child
alone. That second half is the one worth having, because killing healthy work
over a stranger's assignment would be worse than the bug this fixes.


Changelog

Features

  • 1fb910ebee499b6f89efc1935806a84ad4c7ce14 feat(stress): a day proves the runner lets go of a piece taken from it, and keeps working (task #11790) ### Fixes
  • 82c2ddddd93abee15fe8c3ed560f436c110128a3 fix(eventstream): a piece given to somebody else stops the runner that was working it (task #11788) ### Other
  • e3de44be2da659b116dcb8d824769cb78211a13a feat(runner-log): the run log is coloured as it is read, and the file stays plain (task #11786)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.19

The one you will notice

Every step that does work says so on the piece before it moves on (task #11784). A run on 2026-08-25 opened its pull request having called LogWorkProgressTool zero times — while its plan moved on the dashboard card the whole time. The card said "working", the piece said nothing, and the hours landed nowhere.

Task #11616 had already established the shape of the fix: the cadence has to be written into a numbered step, because a block sitting past the last step never gets read. It put three anchors on the auto-squash spine — and left the seven steps between CREATE BRANCH and the CI gate saying nothing about logging, while the three manual workflows got no anchor at all.

Now every substantive step carries its own LOG PROGRESS NOW line, and each one names the TaskUpdate that closes the same step's todo item:

4. IMPLEMENT TASK (incremental commits, clear messages)
   - LOG PROGRESS NOW — TaskUpdate this step's item "completed", then
     LogWorkProgressTool at ~45%: implementation written, tests not run yet

That pairing is the whole trick. The child was calling TaskUpdate reliably all along — that is why the card moved — so hanging the log on the call it already makes is what ties the effort trail to the plan. Auto-squash now walks 20 / 45 / 60 / 70 / 75 / 80 and reaches 100% only after gh pr view says MERGED; the manual workflows top out at 90%, because the PR is left for a human; a review run logs at 25 / 70 / 100.

It binds the model rather than forcing it — the runner posts no effort of its own, and has no way to speak to a child mid-run. What changed is that skipping the log is now skipping a step.

The one that was quietly stopping the 08:00 job

The child searches the project, not the whole disk (task #11782). A child that cannot place a file falls back to find / or find ~ — two days of one project's logs hold five such calls, one of them find /Users/<you> -maxdepth 3.

On macOS that walk enters ~/Documents and ~/Desktop, which are gated behind TCC, so the OS raises a permission dialog. The dialog names mcptask_runner — macOS asks on behalf of the responsible process, the launchd-spawned root of the tree, not the find underneath it — which is why the binary gets blamed for folders it never opens. And the job fires at 08:02 with nobody at the keyboard: the dialog blocks the call until it times out, and then denies it.

The prompt now bounds the search to the project directory, and — the part that makes a prohibition stick — names where to look instead: bundle show <gem> / gem which <file> for Ruby, go env GOMODCACHE for Go, then search under the path it prints.

Upgrading

Move your project's gem pin to 0.3.19 and run rake mcptask_runner:update. The binary is per machine and the newest one any project offers is the one that stays, so one project is enough to update the host.

Both changes are prompt text sent to the child, so they take effect on the next run the updated binary starts. Nothing to migrate, nothing to re-install beyond the binary.

v0.3.18

The one you will notice

The install and update output is coloured (task #11779). The update summary was a wall of grey in which the one row that needed attention read exactly like the thirty that did not. Now: green for up-to-date and added, yellow for updated and force-updated, red for conflict-skipped, dimmed section tags, warnings in red on stderr.

It is decided per stream and only when the stream is really a terminal — a pipe, a redirect and a launchd log all get exactly the bytes they got before. MCPTASK_RUNNER_COLOR (1/true/always/force, or 0/false/never) overrides, and NO_COLOR is honoured. Your run logs stay greppable.

The one that was quietly costing you processes

A child's process group is remembered, so what the child left behind is still reachable (task #11773).

A process group outlives its leader: while any member is alive the group id stays allocated and a signal to -group still reaches them. But the kernel can no longer name that group once the leader has been reaped — and killTree looked it up at kill time, got nothing, and fell through to signalling a bare pid that no longer resolved.

So anything a child started and then exited on went on running. Every claude session that launches a dev server, backgrounds a command, or leaves a watcher behind and then finishes its turn is exactly that shape. Nothing raises, so nothing ever reported it — the processes just accumulate on an unattended machine.

This had shipped in every version of the runner there has ever been. It is fixed by reading the group when the child starts, while it is certainly alive, instead of asking about it when it is already gone.

How it was found, which is the part worth repeating

Not by the test suite. bin/ci runs the chaos stress harness against one fixed seed, and the invariant that catches this — the orphan sweep — has existed since the harness was built. They never met, because the fixed seed draws a different misbehaviour at the moment the signal lands.

It took eight random seeds. Two failed on the current code, and a third failed all the way back at v0.3.16. Scheduling that sweep is now task #11778.

Also in here

The stress harness grew to cover the dashboard-card behaviour that landed after it was built (story #11765, seven tasks) — none of which ships in this binary, but it is what turned up both of the above.

Upgrading

Move your project's gem pin to 0.3.18 and run rake mcptask_runner:update. The binary is per machine and the newest one any project offers is the one that stays, so one project is enough to update the host.

v0.3.17

What this is

One story: the chaos stress harness grown to cover the behaviours that landed after it was built (story #11765, seven tasks).

For an operator, nothing changes. No new behaviour, no changed wording, no new configuration. The entire shipping delta against v0.3.16 is 29 lines, and the only one that touches a code path is a refactor that leaves production behaviour identical.

The one real fix

Stream.SetOnAssignment (new), wired unconditionally by the work loop.

The assignment wake-up added in v0.3.16 was wired inside the branch that builds a stream when the caller supplied none. Production always takes that branch, so the feature worked — but every caller that supplies its own Stream got one with no way to fill the wake channel, which meant no test could ever see it fire. The harness proved it: 299 assignments delivered to the runner's own user, not one wait cut short. The wiring now happens on whatever stream the loop ends up with.

If you are on v0.3.16 this changes nothing you can observe. It is here so the next change to that path cannot break it silently.

Everything else is the harness

internal/stress and internal/chaos do not ship in this binary. What grew there:

  • the stub's dashboard socket pushes as well as records, and can end a connection either with a close frame or with none at all
  • a stop names which of the three ended the day, and a reconnect names who closed the socket
  • the card's message must stand, a clean run must say how it went, and a retry must say what it is retrying for
  • an assignment wakes the short waits and must not touch the overnight sleep or the daily restart's anchor
  • a forked task keeps the name that buys it the watchdog's generous ceiling, while a tool that never closes is still killed
  • a plan reaches the card, and every child is proved to have been launched with the tools that make one

The suite now costs about four minutes rather than ninety seconds; bin/ci says so.

Upgrading

Nothing to do. The gem ships at the same version as always — move a project's pin to 0.3.17 and run rake mcptask_runner:update when convenient, or wait for the next release that actually changes something.

v0.3.16

Upgrading

Run mcptask_runner update --self on each host. Everything here is binary
behaviour — no new skills or helpers ship with this version — so plain
mcptask_runner update has nothing to do, and neither does bundle update:
the binary the 08:00 job runs lives in ~/.mcptask/bin and only --self
replaces it.

The wrapper gem mcptask-rails-runner is published at the same version, on all
six platforms.

What this release is about

The runner card on the dashboard is filled in rather than sparse. It names
the binary drawing it, says how a finished run went, counts out the retries the
CLI is riding out, links a pull request the run has just opened, and says what
an MCP tool call is acting on instead of showing a bare tool name. A task forked
inside a subagent reaches the card too, and the notification that ends it is no
longer thrown away. A card whose stream goes quiet keeps saying the last thing
it knew instead of emptying out, and the TODO list fills in on Opus and Sonnet
as well.

An idle wait ends the moment mcptask.online says this runner has been given a
piece
, rather than sitting out the rest of the poll interval.

A stop on the daily quota says which quota, and how many hours it means —
worked_today=8.2h of per_day=8h. Those hours are the user's day on
mcptask.online, not the run's: another session spends the same budget, so a
runner that idled all morning could end its day on hours it never worked, and
the old one-line message gave a reader no way to tell that from a bug. The same
line now separates a spent budget from an endpoint that never answered, and a
stop caused by a failed task no longer claims to be a quota stop.


Changelog

Features

  • e921a308b8fbcf2a25170d045efa21a28661bc59 feat(executor): a finished run tells the card how it went (task #11758)
  • ee0af55aaf5b9c2413245d5f9e5e13d8f53b6869 feat(executor): the card counts out the retries the CLI is riding out (task #11754)
  • 7ab052bb17845a0b24b2f353a6b96f19119f72f4 feat(executor): the pull request a run just opened shows up on its card (task #11753)
  • cfa61d63b0f4d26311cbf1253a5ad4d5bb4dc8e1 feat(runner): an idle wait ends when mcptask.online says this runner was given a piece (task #11741)
  • 55872bda6abec508541898b59a2ccafbb6f152b9 feat(snapshot): the runner card names the binary that is drawing it (task #11747) ### Fixes
  • 82ebb34127201c463e683996d27d755e999b331f fix(eventstream): the reconnect notice names what closed the socket
  • 94c3df906f741e11924f1f9a81238964d1d763e9 fix(executor): a task forked inside a subagent reaches the card, and the notification that ends it is not thrown away (task #11751)
  • 29978c739ec623b14188fc6708e5447ba3120881 fix(executor): an MCP tool call says what it is acting on, instead of arriving as a bare name (task #11752)
  • 1f8c1cf4cd755524d7ff565cda919344dd37bc5b fix(executor): the dashboard TODO card fills in on Opus and Sonnet too (task #11755)
  • cae366f659a7abd6287c0337c1aa387f566a8762 fix(runner): a daily-quota stop names the quota and the hours behind it (#5)
  • 87cbd6820e787192ab590e5d04f44648dafbca42 fix(snapshot): the card keeps saying the last thing it knew (task #11757) ### Other
  • 3a73198713dd540b1fefaf46dec7ec1f3c9b7eb8 docs(ci): a piped bin/ci reports tail's exit code, not the suite's
  • a2555366c751739e2513be0580c793e79c930d30 docs(snapshot): the two session views, and why neither is a superset (task #11759)
  • 3a1056cee599be8b20ba258926e78a8814618f03 test(executor): the branches the two card stories left uncovered (tasks #11750, #11756)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.15

What this release is for

Two things an unattended host feels directly, and one the dashboard does.

A 429 rate limit is recognised. The CLI retries a spent usage window ten
times internally and then stops mid-answer with no result marker. The runner had
no branch for that, so it called the ending a missing marker, spent its three
thirty-second retries against a window that had not reset, failed the task, filed
a bug piece about the marker — and, because the quota Decider stops the loop on
any error, ended the day's remaining work. The window that prompted this came
back about a minute after the last retry. A 429 now has its own flag, its own
budget of eight, and a back-off measured in minutes (5 → 10 → 20 → 30, capped),
classified ahead of the missing-marker path because a 429 always also looks like
one. A spent usage window is no longer filed as a bug: nobody can fix one.

A daily runner adopts a new binary by itself. Everything that notices a new
version — the bundle adopt, the update hint — runs once, at startup, and a
daily process never ends. It kept whichever binary it started with, however
many releases went by, on exactly the host nobody is watching. On its way into
the overnight sleep it now compares its own executable with what is installed and,
if they differ, retires the card, releases the instance lock and replaces itself
— at the configured end of the workday, or at midnight on a host with no window.

A failed task now shows as failed on the card. A run is marked finished the
moment its stream ends, before the ending is classified, and finished → error
was not a legal transition — so a failed task filed a bug piece while the
dashboard went from finished straight back to triage. Anyone watching saw a
healthy runner through every failure of the day.

Upgrading

Nothing to do beyond the usual install. One exception, and it is the reason to
bother: a daily host will not pick this up by itself. The self-restart
above only exists from this version on, so the process now running has to be
stopped once by hand — after that it keeps itself current.

launchctl bootout gui/$(id -u)/online.mcptask.runner-<slug>
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/online.mcptask.runner-<slug>.plist

today hosts need nothing: the next morning's job starts the current binary.
Check which one a machine is running with ~/.mcptask/bin/mcptask_runner version.

Changelog

Features

  • be9e6ef541de78e0cb923fbc32f4c4de13791f37 feat(chaos): a CLI stand-in that starts well, decays through five phases, and records what it did to the runner (task #11715)
  • 9b8b4cc4a50ceca2cf32614e40e554b8a7f11383 feat(chaos): a stub mcptask.online that freezes, lies, drops the dashboard socket, and keeps an honest ledger (task #11716)
  • 184cc296585c53afc787ed3bdb89769a1843ca5f feat(executor): a 429 rate limit is waited out on its own budget instead of misread as a missing marker (task #11731)
  • c5e50fa889774937a76d2fe8a70a2f9253cf977a feat(runner): a daily runner replaces itself with the installed binary at the day's end (task #11729)
  • dcef1a9ef3d1421d5221b85111a76fc927b9af60 feat(stress): a real signal to a real binary, and everything it has to take with it (task #11719)
  • e0eb4f012390d344327db1f970d12b0722212ea4 feat(stress): an in-process driver that compresses a workday into ninety seconds and states how the loop may end (task #11717)
  • e38e3638755ce63536aa8bdcbd3fc5c00a2e57be feat(stress): one fixed-seed day in bin/ci, bin/stress for the rest, and the harness written down (task #11720)
  • 36a22107e88c62e2bf5f3613025a18c08a374550 feat(stress): the card has to keep telling the truth while everything burns (task #11718) ### Fixes
  • f799bfe05a5d6490c2cdf691fc58ff499044df93 fix(eventstream): a first subscription confirmation has nothing to resend, so a fresh session stops duplicating its opening frame (task #11725)
  • ea6fc1839a472f7bbaeb4b7e3507d92c9652215a fix(executor): a marker whose JSON does not parse is not the child's answer — keep looking, and take the retry (task #11724)
  • 74dc7ad64df2f477e6d8306632d25327bbf98add fix(executor): a server that has not finished connecting is recognised as the MCP route failing, not left to become an empty bug report (task #11712)
  • 693d609f24dff2143e40924c879ed26907bf2e11 fix(executor): the card says what it is retrying for, because the runner already knew (task #11732)
  • c8a70925de6f362b0fb156740529ba395cb5faf9 fix(loop): a day that ended on the clock sleeps until the next one instead of spinning until midnight (task #11726)
  • c9644cf59a609d39643aac69134fb839b88a874a fix(runner): a failed task shows as failed on the card, and the check that says so cannot be fooled (task #11734)
  • 8678677b6c093f7e312956b35d762983bb8c95d9 fix(snapshot): the dashboard TODO card follows TaskCreate/TaskUpdate, the plan tools that replaced TodoWrite (task #11710) ### Other
  • 2658215af771f52833587916fc08b23defc13792 docs(config): every wording of work_window says which mode exits and which sleeps (task #11728)
  • bfc76427946075e8fffac21a29db48c7fe01dbd1 merge main into story-11714 before the story merge

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.14

Upgrading

Run mcptask_runner update --self on each host. This release changes loop
behaviour inside the binary, and nothing replaces the binary for you: the copy
the scheduled job runs lives in ~/.mcptask/bin, and neither bundle update
nor a gem bump touches it. mcptask_runner version says what a host actually
has.

No new skills or helpers ship here, so plain mcptask_runner update has
nothing to do in this one.

What changed

An empty queue is asked about every half hour, not every five minutes.
The REST next-task check added in v0.3.13 made an idle round cheap, and cheap
rounds ran twelve times an hour: two GETs each, and two rows in the dashboard's
activity feed each, all saying nothing had happened.

waiting_strategy.short_wait_minutes still governs the first rounds — it is
the right answer in the minutes just after a task finishes, when work is
expected any moment. What is new is that a run of empty rounds counts as
evidence: once those short waits have added up to one long_wait_minutes, the
long wait takes over and keeps it until there is work again. Nothing waits
longer than long_wait_minutes, the ceiling daily mode always had, and the
switch is announced in the log in both directions.

On the installer's 5 and 30 that is 21 idle rounds across an eight-hour
afternoon instead of 96. A configuration whose long wait is no longer than its
short one has nothing to back off to, and is left alone.

Two smaller log fixes in the same path. The quota check made after an empty
round used to end the day on a bare break with no line at any level; it now
says why the runner stopped. And daily mode no longer claims "will wait 1 hour
before retry" while waiting long_wait_minutes.

Changelog

Features

  • adad985d17350cb40bd716dbe08d02fd92b13816 feat(loop): a queue that has been empty for half an hour is asked about every half hour, not every five minutes (task #11699) ### Fixes
  • a65a5fda79fdaf748efb8a6204ef7393f553fb21 fix(loop): the backoff line names its waits the way the waiting strategy does (task #11699)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.13

Upgrading: mcptask_runner update --self (or bundle update mcptask-rails-runner to 0.3.13) upgrades the binary; then run mcptask_runner update in each project — this release ships a new helper (runner-log) and rewritten config sections, and nothing installs those for you. The scheduled job runs the per-machine binary in ~/.mcptask/bin, so check mcptask_runner version there afterwards.

What changes in behaviour:

  • Before paying for a triage session the loop asks GET /api/:account/pieces/next; an empty queue answers no_more_tasks with no model involved. Until mcptask.online ships that endpoint (task #11691) it answers 404 and the log shows next-task REST check failed (HTTP 404 …); falling through to triage every cycle — expected, triage behaves exactly as before.
  • The configured end of workday is the one hard stop, checked after a task finishes. There is no default any more: a config whose work_window.end_of_workday_hour is unset or null has no end of day (the 18:00 default is gone; the quota and the day gate still apply), and a running task is never interrupted by it.
  • The skip list is persisted, so a restart does not re-pick what today already declined; failure, task_already_started, merge_failed, merge_unverified and out_of_scope set a task aside for the day instead of stopping the run.
  • respect_working_hours is read from the user profile once at startup — a change to the checkbox is heard at the next start.

Changelog

Features

  • a1798c3cb6bc2f707d31e01c4d5d5e5f4d89aef0 feat(bugreport): a guard of the runner's own that did not hold files a piece instead of a log line (task #11684)
  • a14790c3995e036d6368c7b5742768246f617fcc feat(install): ship runner-log, the helper an operator types to tail the current run
  • 3fb7637cc491069964281b2f729ecf3c2f949491 feat(install): the install ends with a summary of every config section the runner reads, and the sections it writes carry a comment (task #11690)
  • 8f78435eb604330ea5267a1cd1085c521da4ff08 feat(loop): ask the REST next-task endpoint before triage — an empty queue is no_more_tasks without a model (task #11692) ### Fixes
  • dde40a14bf9a2fefe495970f3accf73a5e5dc9ba fix(loop): a task declined as out of scope goes on today's skip list, and the loop never stops over it
  • dd3a529977087feb39b27a0626c0e6dc63c35596 fix(loop): failure, task_already_started, merge_failed and merge_unverified set the task aside for the day instead of stopping (task #11680)
  • f727a15afec1b66ff369fc56daaa397f763d9040 fix(loop): the configured end of workday is the one hard stop, checked after a task finishes — and there is no default any more (task #11689)
  • 0b18385f30d5a1f429a68be46714ce8763350b6d fix(loop): the skip list outlives the process, so a restart does not re-pick what today already declined (task #11683)
  • 32c00200b5849084d91bb1e0beaa2249fe4293bd fix(quota): read respect_working_hours from the user profile once at startup, not from time_status (task #11686)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.12

This release is mostly about a guard that had stopped guarding, and a setting that could not be right for everyone reading it.

The lock that let tests through. check_test_lock refuses a direct bin/ci while another agent's suite is running. It looked for pidfiles in a flat directory nothing has written in a long time, so every lock older than ten seconds read as stale and the command went through — the guard worked for the first ten seconds of a lock's life and never again. The path is now rebuilt from the lock's own project. Two smaller failures in the same lock went with it: the directory that serialises the moment of acquiring could not go stale, so an acquire killed midway wedged every later run behind a lock reported as unknown; and a lockfile missing its PID= line killed acquire outright before the staleness check could reap it, leaving nothing but a blank error to act on. (#11669, #11671, #11672)

Bundled skills no longer declare a model:, and the install-time rewrite that resolved it is gone.

A model name written into .claude/skills/ can only be right for one backend, and that directory is read by two: the runner's child process and whoever opens the same project in a session of their own. Whichever way the name was written, it was wrong for one of them — and wrong silently, because the CLI returns its model error as the skill's own result and the agent above reads that as the answer. Each fork now runs on its session's own model. context: fork is untouched, and that is what actually keeps a fork's output out of the parent's context. (#11670)

Upgrading

Run mcptask_runner update in each project after upgrading the binary. Nothing does it for you — upgrading the binary does not refresh the skills and helpers already on disk, so without this step every fix above stays installed nowhere. Helpers are replaced in ~/.claude/bin; skills are refreshed per project. Do it in every project you use, not just one. (Making the upgrade reconcile them by itself is tracked as #11673.)

No --force is needed. Hosts whose skills carry a rewritten model name are replaced normally by mcptask_runner update, because the installer recorded the manifest hash from the same content it wrote — so the rewritten form is the baseline, and the updater reads those files as untouched. Verified against three hosts carrying rewritten copies (gemma4:31b-cloud, minimax-m3:cloud and haiku): all eight bundled skills hash equal to their manifest entry on each. A skill somebody edited by hand is still reported as a conflict and skipped, exactly as before.

One known rough edge, not fixed here: now that the lock guard blocks correctly, it also blocks commands that merely mention bin/ci — grepping a CI log, for instance — because the pattern matches the substring rather than the command position. If a refusal looks wrong, check ~/.claude/bin/test_lock status to see whether a run is genuinely in progress. Tracked as #11674.

Changelog

Fixes

  • cccbcab0c3ec7d60f581c594405a91121f9945b6 fix(shutdown): a SIGTERM is not a failed attempt, so it announces no retry (task #11655)
  • 481f3a986071f48906736ef6dc791eb0d4a8329f fix(skills): a pinned model can only be right for one of the two backends reading the same directory (task #11670) ### Other
  • f6d5a9815968c7fa3f2c7443e150928740e698ed fix(ci-lock): the acquiring guard could not go stale, and a nameless lock killed acquire outright (tasks #11671, #11672)
  • bb0b389ebab93a7153d7f439e3d8600367a4a13a fix(ci-lock): the guard against a direct test run looked for pidfiles nobody writes (task #11669)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.11

Changelog

Fixes

  • 230496500ec4a402e8d4bdd4bcff1918147f74df fix(quota): a role that ignores working hours is not stopped by its hour goal (task #11643)
  • 886a3355e74423c1d01e1a367a3d39a613a4731c fix(triage): a disabled fetch tool is a detour, not the end of the day (task #11642)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.10

Changelog

Fixes

  • 5fef8b8808e0f121d83d093370ddc01964e07cb1 fix(helpers): resolve Ruby the same way on rbenv as on rvm (task #11641)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.9

Changelog

Fixes

  • bc156c306eb0ce4f115e07d004fdcd6802851175 fix(ci): bump pinned golangci-lint to v2.13.0 for Go 1.27 support
  • 642002843a54df7469862fd332bae3adb55e1c74 fix(install): print which account a scheduled job will authenticate as (task #11494)
  • b782fbecb65c01aa7072cf72b78b5ea5db07b958 fix(launcher): one log file per run instead of one that grows forever (task #11617)
  • a40f46624f890ca770b34e0181d5d204d35ab7ce fix(loop): a declined task skips, it does not end the day (task #11622)
  • e38b8d6025f8662f11a9a6c2cdb8b69652dfda15 fix(triage): discard a pick that belongs to a different project (task #11619) ### Other
  • 55f52d52fb2f96a49b26aa4c7c2d1361246eabb9 docs(skills): follow the MCP parameter renames ahead of the server change (task #11609)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.8

Changelog

Features

  • bb18d74c555e82f11ac1324d9b3b92ac9b3e4d22 feat(config): make end-of-workday hour configurable at install time (task #11495) ### Fixes
  • 7264ca3901696ef4c861fada80b586d35793394a fix(prompt): anchor the progress-logging milestones inside the auto-squash steps (task #11616)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.7

Changelog

Features

  • b66ced087bcd9d6afae57c80e57cf5ea1ce4283a feat(bugreport): remove the manual bug-report CLI command (task #11481) ### Fixes
  • 66bab4931cd4555bab6cfc0655994d9ae61bd275 fix(bugreport): retarget automatic error reports to this project's own Epic (task #11482) ### Other
  • 0ae13df7ff594e7e7db44aa841892ed4a0ee6df4 chore(claude): each story gets its own worktree to protect against concurrent session conflicts
  • c1af57c30c09d8a8f16ecd8618aaf109d0443325 chore(claude): ignore /.claude/settings.local.json repo-wide
  • 30d47d8edc50821df2fb312e2f9eeccfff6b5075 chore(claude): install discover/mcptask-read/mcptask-write/memory-search skills

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

list_altCLI reference

Five public subcommands, two persistent flags, three exit codes

The runner is a single binary on a cobra root: run, init, update, version, pr. A sixth subcommand, internal, is hidden from --help because generated launcher scripts call it, not operators: internal launcher-log-name is how the Windows launcher names its own log. Hidden is not provisional — the launcher depends on it at run time. Two persistent flags --verbose and --ignore-quota are read case-insensitively from the matching env var (verbose=true, ignore_quota=true). Per-subcommand flags are listed below each command. Exit code is 0 on success, 3 for an unimplemented mode, 1 for anything else, and 128+N on a signal. 3 is a designed sentinel rather than an ending you can reach today: every mode is implemented, so nothing returns ErrModeNotImplemented.

menu_bookEvery subcommand, flag and exit code – referenceexpand_more

Subcommands

run — the work loopsince v0.1.0

Starts the work loop in one of fifteen modes: a single task, everything assigned for today, a whole queue, the subtasks of one Story, one named task, pending code reviews, or an unattended loop that runs every business day. Auto-squash modes take a task all the way to a merged pull request; manual modes leave the pull request open for a person to review.

  • --story-idsince v0.1.0

    The Story whose subtasks story_manual and story_auto_squash work through.

  • --task-idsince v0.1.0

    The one task that task_manual and task_auto_squash work on.

init — set up a projectsince v0.1.0

Prepares a project for the runner in one command: it records which coding CLI drives the project and which git host its pull requests live on, installs the bundled skills and helper scripts, writes the MCP configuration and the approvals the CLI needs, stores the access token, and generates a weekday scheduled job that you switch on yourself.

  • --clisince v0.3.24

    Names the coding CLI that drives the project — claude, codex or opencode. Required the first time; there is no default, and later runs reuse the answer.

  • --git-hostsince v0.3.25

    Names the git host — github, bitbucket or gitlab — when the origin remote does not say it, for example behind an SSH alias, on a mirror or on a self-managed GitLab.

  • --epic-idsince v0.1.0

    The Epic on mcptask.online that collects the bug tasks the runner files about itself; 0 keeps them at the project root. Answers the question without a prompt.

  • --epic-namesince v0.1.0

    A display name recorded next to --epic-id.

  • --modesince v0.1.0

    The work-loop mode the scheduled job runs, such as today_auto_squash, or none for no scheduled job at all.

  • --atsince v0.3.3

    The time of day the scheduled job starts, as HH:MM; 08:00 when not given. A value that is not a time of day is refused, never rounded.

  • --untilsince v0.3.8

    The time of day after which the runner starts no further task, as HH:MM. A task already running still finishes. Left out, the working day has no end and only the daily quota stops the runner.

  • --schedulesince v0.1.0

    Regenerates the scheduled job and touches nothing else.

  • --forcesince v0.1.0

    Overwrites skills, helpers, config sections and the scheduled job that already exist. Also set by FORCE=1.

  • --helper-bin-dirsince v0.1.0

    Where the CI and test helper scripts are installed, and what the skills are written to call. Defaults to the directory the harness profile names.

  • --home-dirsince v0.1.0

    Puts the token directory, the shell start-up file and the scheduled job under another home directory than the current user's.

update — refresh skills, or the runner itselfsince v0.1.0

Brings an installed project up to date with the runner it now runs on. Every bundled skill and helper is compared with what was installed and reported as added, up to date, updated, skipped because you edited it, or overwritten on request with a backup. With --self it replaces the runner binary from the latest release instead.

  • --selfsince v0.2.0

    Replaces the runner binary with the latest release, verified against the release's checksums; a failed or mismatched download leaves the installed binary untouched.

  • --checksince v0.2.0

    Reports whether a newer runner release exists, and changes nothing.

  • --forcesince v0.1.0

    Overwrites skills and helpers you edited locally, keeping a .bak copy of each one next to it. Also set by FORCE=1.

  • --helper-bin-dirsince v0.1.0

    Where the CI and test helper scripts are refreshed, and what the skills are written to call. Defaults to the directory the harness profile names.

  • --home-dirsince v0.1.0

    Re-roots whatever update writes outside the project under another home directory.

version — what is running, with which settingssince v0.1.0

Prints the runner version together with the configuration it resolved for this project: the coding CLI, the models for each tier, the launcher command, the git host and the other settings, each with where its value came from.

pr — pull requests on any supported git hostsince v0.3.25

Creates, lists, reads, merges and checks pull requests on whichever git host the project is on — GitHub, Bitbucket Cloud or GitLab — and prints one JSON object per call. The bundled skills and the runner's prompts use only this command, so the same instructions work on every host.

  • pr createsince v0.3.25

    Opens a pull request for the current branch — a merge request on GitLab — using the project's pull-request template when it has one.

    • --titlesince v0.3.25

      The pull request title. Required.

    • --bodysince v0.3.25

      The pull request description, given inline.

    • --body-filesince v0.3.25

      Reads the pull request description from a file instead.

    • --basesince v0.3.25

      The branch to merge into; the host's default branch when not given.

    • --headsince v0.3.25

      The branch the work is on; the current branch when not given.

    • --draftsince v0.3.25

      Opens the pull request as a draft.

  • pr listsince v0.3.25

    Finds pull requests: the ones that mention an mcptask.online task, the one opened from a given branch, or every open pull request of the repository.

    • --tasksince v0.3.25

      Searches for pull requests that mention an mcptask.online task. The results are candidates — a search for one number also finds longer ones that contain it.

    • --branchsince v0.3.25

      Finds the pull request opened from this branch, in any state.

    • --statesince v0.3.25

      Narrows a --task search to open, merged, declined or all pull requests.

    • --opensince v0.3.26

      Lists every open pull request of the repository — the whole review queue.

  • pr viewsince v0.3.25

    Reads one pull request — its state, branches and link — as JSON.

  • pr reviewssince v0.3.26

    Reads everything reviewers left on a pull request in one list: the verdicts first, then the comments, inline ones with their file and line, bots included and marked as such.

  • pr mergesince v0.3.25

    Merges a pull request and reads it back, so the output says what really happened rather than what was asked for. Inside an auto-squash run the merge is refused unless the host's CI said yes, or the host runs no checks and the project has its own local gate.

    • --squashsince v0.3.25

      Squashes the branch into one commit. On by default.

    • --delete-branchsince v0.3.25

      Deletes the head branch after the merge. On by default.

  • pr checkssince v0.3.25

    Reports the git host's CI verdict on a pull request: success, failed, still in progress, or none — and none is never read as success, because a repository whose CI never ran has not passed anything.

  • --git-hostsince v0.3.25

    Overrides the git host for one pr command, whatever the configuration file or the origin remote say.

Global flags

--verbosesince v0.1.0

Prints every line the coding CLI streams, not just the summary. Also set by the environment variable verbose=true.

--ignore-quotasince v0.1.0

Skips the daily quota checks: the runner neither asks mcptask.online for the remaining budget nor refuses a task over it, and the operator takes responsibility for the spend. Also set by ignore_quota=true.

Configuration keys

harnesssince v0.3.24

Which coding CLI drives the project, by profile name: claude, codex, opencode, or a profile of your own. The one required key — there is no default, and a project that names none is refused rather than assumed to be anything. init --cli writes it.

git_hostsince v0.3.25

Which git host the project's pull requests live on: github, bitbucket or gitlab. Optional — left out, the runner reads it off the origin remote, and a remote it cannot recognise is refused by name rather than guessed at.

models.geniussince v0.1.0

The model behind the strongest tier, used for heavy coding and for a task that has to be finished after an earlier attempt fell short. Without it the harness profile's own model is used.

models.smartsince v0.1.0

The model behind the middle tier, used for triage and reviews. Without it the harness profile's own model is used.

models.primitivesince v0.1.0

The model behind the fastest tier, used for simple, read-mostly work. Without it the harness profile's own model is used. On a backend that is not the CLI's own, pin all three tiers or none.

launcher.commandsince v0.1.0

A command that starts the coding CLI in place of its plain binary — for example through ollama launch, or wrapped in caffeinate so a Mac does not fall asleep mid-task. It is checked against harness, so it cannot quietly start a different CLI than the one the project declared.

launcher.flagssince v0.1.0

Overrides individual command-line flags the harness profile passes to the coding CLI; a flag set to null is left out entirely.

launcher.cli_checksince v0.3.24

On by default. Set to false only when launcher.command points at a test double instead of a real coding CLI; it switches off the check that the command starts the CLI the project declared.

waiting_strategy.short_wait_minutessince v0.1.0

How long today_auto_squash waits before asking again when no task is ready. 30 minutes by default.

waiting_strategy.long_wait_minutessince v0.1.0

How long daily waits before asking again when no task is ready. 60 minutes by default.

skip_list.revive_after_minutessince v0.3.22

How long a task the loop had to set aside — one that failed, or that somebody else had already started — stays out of triage before it may be picked again. 60 minutes by default; 0 keeps it out for the rest of the day.

skip_list.max_revivalssince v0.3.22

How many times in one day a set-aside task may come back for another attempt. 2 by default; 0 means never.

bug_destination.epic_relative_idsince v0.1.0

The Epic on mcptask.online that receives the bug tasks the runner files about its own failures. Without it they land at the project root.

bug_destination.epic_namesince v0.1.0

A label for that Epic, shown in the start-up summary only.

pr_template.pathsince v0.1.0

The pull-request template the agent is told to follow, relative to the project root. GitHub and GitLab projects get their host's usual path without it; on Bitbucket this key is the only way to name one.

work_window.end_of_workday_hoursince v0.3.8

The hour after which today, today_auto_squash and daily start no further task; a running task still finishes. Left out, the working day has no end. init --until writes it.

work_window.end_of_workday_minutesince v0.3.23

The minute within that hour, so a day can end at 18:30. Optional and 0 when absent; an unreadable value leaves the whole window unset and says so at start-up.

story_branches.enabledsince v0.3.35

Story branches for this host: the tasks of a Story merge into a shared Story branch, and the Story's last task brings the whole Story into the main branch in one merge. On unless set to false, and even a host set to false follows a Story branch that already exists; a value that is not true or false is refused rather than guessed.

debug.capturesince v0.3.24

Records the coding CLI's raw output into the project's log directory, with secrets redacted, for diagnosing a run. It changes nothing about how the run behaves, and an environment variable overrides it either way.

Coding CLIs

Claude Codesince v0.1.0

Anthropic's coding CLI. It gets the full skill set, including the three skills that hand a lookup to a separate sub-session so its raw output never fills the main context, and it can run against any Anthropic-compatible backend through launcher.command and pinned models.

Codex CLIsince v0.3.24

OpenAI's coding CLI. The runner checks that it is logged in before a run starts and refuses the run otherwise, rather than spending the day's quota on a CLI that would fail at its first call.

OpenCodesince v0.3.24

The open-source coding CLI that takes models from many providers. It has no read-only switch, so the runner supplies a permission map that gives the same guarantee by another route.

Git hosts

GitHubsince v0.3.25

Driven through the gh CLI and its own login. The pull-request template is read from the place GitHub expects it. A GitHub Enterprise instance is a different host and is refused by name rather than treated as github.com.

Bitbucket Cloudsince v0.3.25

Driven through Bitbucket Cloud's REST API, with either an access token for a machine or an e-mail and API token for a person — never both, and never an app password. init stores the credential in a file only you can read and never prints it.

GitLabsince v0.3.27

Driven through the glab CLI, which keeps its own login, on gitlab.com and on self-managed instances alike; a self-managed project names gitlab in its configuration because its address does not say so. What the runner opens there is a merge request.

Exit codes

0 — the run completedsince v0.1.0

The command ran to its end. For run this means the loop finished its work, not that every task succeeded — a task that failed is reported on mcptask.online and in the log, and is not a reason for the process to fail.

1 — the command could not do its jobsince v0.1.0

Something the command needs was missing or refused — no harness named, no token, a coding CLI that is not logged in, an unknown mode, a failed download — and the error printed on stderr names it. The runner refuses to start rather than spend an hour exiting with success over nothing.

3 — mode not implementedsince v0.1.0

Reserved for a work-loop mode the binary recognises but does not carry out. Every mode is implemented today, so a current runner does not return it; a scheduled job may still treat it as "upgrade the runner".

130 — stopped with Ctrl-Csince v0.3.0

The runner was interrupted (SIGINT). It stops the coding CLI it started, marks the session as ended on the dashboard, and exits with the shell's usual 128 plus the signal number. The task stays in progress and the next run picks it up again.

143 — stopped by the systemsince v0.3.0

The runner was asked to terminate (SIGTERM), typically by the scheduler or at shutdown. It stops cleanly the same way as after Ctrl-C, so a stopped run is never recorded as a successful one.

monetization_onQuota gates

One number, three places to check it

The runner never trusts its own guess. Every quota decision reads a live number from mcptask.online over REST, then enforces it before, between, and during tasks.

sourceWhere the number comes from

The daily budget is fetched live from mcptask.online over REST (GET /api/{account}/users/current/time_status). The runner never asks the agent to estimate its own budget and never reads a cached value — the account is the single source of truth.

gpp_goodThree gates, in this order

WhenWhat gets checked
PRE-RUN (after triage)Re-polls the live REST endpoint after triage. Triage itself takes minutes; a fresh pull catches anything the user spent while the runner was classifying the task.
BETWEEN-TASK (decider)Stops the loop on a mid-task quota kill, a spent daily budget, or a task that ended with status error. A failed, out-of-scope, already-started or merge-failed task goes on the skip list instead, and the loop moves on to the next one.
MID-TASK (every 360 s)Polls every DefaultQuotaPollInterval. A crossing kills the child and ends the loop with quota_exceeded_mid_task — no retry. Only the daily mode comes back on its own, the next day, starting again from triage. The kill streak (DefaultQuotaFailureKillStreak = 3) absorbs about an 18-minute REST outage before the runner gives up.

tuneA deliberate fail-closed / fail-open split

The two questions are not answered the same way on purpose. "Is the quota exceeded?" fails CLOSED; "Can I work today?" fails OPEN. One error blocks a run; the other lets it start.

blockFail closed

When the question is "have we already spent today's budget?" and the REST call cannot answer it, the answer is YES — the budget is treated as spent and no new work starts. Better to skip a task than to overspend.

check_circleFail open

When the question is "is there budget left for me to work today?" and the REST call cannot answer it, the answer is YES. Better to start a run than to waste a working day on a transient outage.

reportOutage vs spent budget

A REST outage that killed healthy work is a different bug from an exhausted daily budget. The runner tells them apart on purpose (ErrQuotaPollOutage); a spent budget is nobody's fault, while a REST outage is an 18-minute incident worth a bug task.

ScenarioWhat the runner does
Daily budget spentBetween-task decider stops the loop. No bug filed.
REST poll outage (under kill streak)Retry on the next interval. Loop continues.
REST poll outage (over kill streak, ~18 min)Ends the loop fail-closed with status error and its own termination quota_poll_outage — not quota_exceeded_mid_task. That termination is not on the soft list, so a bug task is filed unconditionally and the 18-minute incident gets tracked.

--ignore-quota skips all three gates. The runner never asks for the time status and never refuses a task — the one REST call left is the startup check that identifies the user (/users/current); the operator takes responsibility for the spend.

memoryContext overflow

What survives when the context overflows

The session is unrecoverable. The work is not — because the work is a branch and commits on disk. Two independent mechanisms carry the rest forward, and a third verdict tells the harness to stop filing noise.

menu_bookHow the runner survives a context overflow – technical detailsexpand_more
check_circle

Survives

What the runner hands to the next attempt

  • account_tree

    The git branch and its commits

    The work itself is a branch with commits on disk. Even if every byte of the session context is gone, the diff is still reviewable, the PR is still openable, and the changes are still recoverable.

  • history

    Last 3 actions (in-process restart)

    When the budget is fresh and the process restarts in place, the restart preamble carries the last three actions (RecentActionsCap = 3) plus the measured context-cost findings, so the new attempt resumes with momentum instead of cold.

  • rule

    ContextBudget rulebook

    Every restart inherits the shared handoff.ContextBudget() rules — what counts toward the limit, what does not, and how the runner decides the next attempt is safe to start.

block

Lost

What no mechanism can save

  • chat_bubble_outline

    The full conversation context

    --continue would reload the same oversized context, so the SESSION is unrecoverable. Only the last three actions ride along in the restart preamble; everything before that is gone.

layersTwo independent mechanisms

MechanismScopeWhat carries forward
In-process fresh restartInside the same runner process (retry.go, handleContextOverflow)Budget = 1 restart per process (maxOverflowRestarts = 1, deliberate — kept at 1 after a review). The new attempt reads the last 3 actions, the cost findings, and the budget rulebook.
Cross-process TaskHandoffinternal/handoff/ — one note per task (engine.go, recordTaskHandoff)Writes log/handoffs/task_<id>.json on BOTH the restart and the terminal branch (terminal ends THIS process, not the task). The next NEW process reads it as the prompt preamble for the first attempt — never on a --continue retry. overflow_count accumulates ACROSS runner processes. Notes retire the moment a run ends any other way, and are pruned after 30 days.

gavelThe third verdict: overflow_pr_open

If the restart budget is spent AND an open PR for the task exists, the status is overflow_pr_open, NOT an error. The PR is the deliverable; an open PR with green checks means the work shipped. Filing an error here costs the whole day's loop on a task that is already done. A real task lost a finished commit, an open PR and the rest of the day over exactly this misclassification, which on top of that auto-filed a bug task it had not earned.

overflow_pr_open (not an error)error (what it used to be)
restart budget spent + PR open → overflow_pr_openrestart budget spent + no PR → error
memoryFor developers

Why this isn't magic

Four concerns run in parallel: a child in its own process group, a stream reader, a watchdog, and the event stream. The state machine underneath has ten named states and an explicit allow-list - a rejected transition is a warning, never a reason to stop.

menu_bookWhat runs at the same time – process group, watchdog and state machineexpand_more
lan

Child in its own process group

The runner spawns the coding CLI in a new process group (syscall.SysProcAttr{Setpgid: true}, internal/executor/process_unix.go, configureProcessGroup), so SIGTERM and SIGKILL from the runner hit the whole tree and never leak to the parent. Reader joins are bounded at 30 s (engine.go, stderrJoinTimeout + defaultStdoutJoinTimeout) so a wedged child cannot block shutdown.

stream

Stream reader

Reads the CLI's stream-json line by line as it is produced, in whichever dialect the harness profile declares. Parses events, hunts for the TASKRUNNER_RESULT completion contract, and feeds the StallDetector. No waiting for the process to end — it reacts incrementally to every line that lands.

timer

Watchdog

Independent guard thread on a 30-second heartbeat. Idle kill at 20 minutes, frozen soft-warn at 3 minutes, per-tool hang ceilings (quick / long), an absolute backstop, and a live REST quota poll every 6 minutes. The only outside supervisor over the child.

wifi

Event stream

Persistent WebSocket to mcptask.online over ActionCable (RunnerSessionChannel). Snapshots throttled to ~0.5 s; FSM state changes force-flush. Async reconnect with a 30-second throttle on drop, and a 0.5-second grace window on the final 'closed' frame so the live card never freezes on a stale state. MCPTASK_RUNNER_DISABLE is not a stream-only opt-out — it is the global kill switch for all mcptask.online traffic: the event stream, the quota polls and the bug reports alike (internal/mcptask/endpoint.go, DisableEnv).

hubThe state machine underneath

schema version 3

Ten named states. Every transition is on the allow-list; anything else logs a warning and keeps working. Frozen, pending, stalled and closed are always-allowed transition targets from any state; they are forced by the watchdog (frozen, pending), the stall detector (stalled) and the loop's session close (closed).

States

startingtriageprocessingwaitingfinishedstalledfrozenpendingerrorclosed

Allow-list (a sample)

A few of the transitions the loop uses every minute — and the forced ones the watchdogs fire on their own.

fromtobynote
startingtriageloopAfter spawn, before the first task is picked
triageprocessingloopTask chosen, context delivered to the CLI
processingwaitingloopBacking off between attempts of the same task
processingstalledstall detectorRepeated Edit failures, repeated identical Bash failures, or the same tool signature within the sliding window
anyfrozenwatchdogIdle past FROZEN_WARN_THRESHOLD, no active tools
anypendingwatchdogSingle tool past its warn ceiling
anyclosedloopend_session — the final frame
visibilityObservability

What you can see while a run is in flight

Three independent surfaces — the on-disk run log, the process-wide operational log, and the live ActionCable card — plus the explicit opt-out variables. They are best-effort by construction: watching a run — snapshots and logs — never aborts it.

menu_bookWhat the runner writes to its logs and what the live card shows – technical detailsexpand_more
data_object

Run log — one JSON file per child execution

The most useful surface when the question is "why has the runner been stuck for 90 minutes". The file is opened IMMEDIATELY when the child is spawned, so even an instant crash leaves something to read.

folderWhere
log/runs/run_*.json (one file per child execution)
  • lock_open

    Opened

    At spawn — before the first line of stream-json is even read

  • favorite

    Heartbeat

    Refreshes stream_events, inactive_s and stream_quiet_s every tick, so a hung run leaves its in-flight state on disk

  • task_alt

    Finalize

    Stamps why the attempt ended

Turn "why has the runner been stuck for 90 minutes" from a grep over a 268 000-line stream log into a single file read.

description

Process-wide operational log — broader than the run log

Not a human-formatted mirror of the JSON run log. It is a chronological text log of the whole process: every subsystem writes its Debug/Info/Warn/Error calls into it, so it holds things no run log ever shows — and it does not carry the run log's fields such as session_id or stream_events.

folderWhere
log/mcptask_runner_YYYYMMDD_HHMMSS.log
short_textLine shape
[timestamp] SEVERITY - message (internal/observe/logger.go)
wifi

Live card — one ActionCable WebSocket

One persistent WebSocket carries whole snapshots, never individual events. A dropped frame costs freshness, not correctness, and a reconnect needs no replay.

ChannelPayloadThrottleTimeouts
RunnerSessionChannel (internal/eventstream/eventstream.go, channelIdentifier)Whole snapshots only — no event replay500 ms snapshot, 30 s reconnect throttle, 500 ms final-frame grace10 s subscribe + dial

Finished card: Lingers 60 s after the loop ends (snapshotCloseTTL)

No token: A runner started without a token says so once, loudly, at startup — never silently misses the first event

shield

The boundary — and the opt-outs

Every one of these is best-effort by construction. Outbound observation never aborts a run; a nil *Log is a working no-op. Two environment variables turn the stores off cleanly.

Run logTask handoff
MCPTASK_RUN_LOG=0 – Turns the JSON run log offMCPTASK_TASK_HANDOFF=0 – Turns the cross-process task-handoff note off

The one exception — reassignment: RunnerSessionChannel is duplex: alongside the outbound snapshots it receives the inbound control events piece_assigned and set_aside_cleared (internal/eventstream/eventstream.go). A reassignment is wired through abandonReassigned (internal/runner/loop.go) into the engine's AbandonReassigned — it kills the in-flight child and reclassifies its death as a reassignment. That is the channel's one deliberate control role.

What this does NOT mean: None of the observability surfaces — the run log, the operational log, the outbound snapshots — can stop a run, restart a run, or classify a run. The best-effort property is the precise boundary of the project's no-fallbacks premise: the one place it does degrade, and says so out loud.

Runner fleet: Every runner instance of the account is also listed on its Runner fleet page (user menu): one row per machine and project with the CLI, model, state, last reported task, today's hours against the daily limit and the runner version, updated live. Account managers and company owners see every runner; project members see the runners of their own projects.

bug_reportBug reporting

When a hard failure happens, a bug task files itself

There is no manual command to run. When a hard failure happens, the runner opens a bug task for it automatically. A real crash becomes a task somebody can pick up — not a silent gap nobody knows to go looking for.

menu_bookWhen the runner files a bug task by itself, what it attaches and how it avoids duplicatesexpand_more
rule

When a bug task gets filed

Only status error, status crash or status anomaly counts as a hard failure. Those are the three values in hardStatuses (internal/bugreport/runner_error.go). Every other outcome — success, stalled_for_genius, urgent_bug_pending, or a graceful quota bail — is the runner working as designed and files nothing.

check_circle

Filed

status error, status crash, status anomaly

What an anomaly is

a guard of the runner's own that did not hold in a run which then carried on — nothing failed, the day's work continued, and from the outside there is nothing to see. An error announces itself by ending a run and a crash by killing a process; an anomaly announces itself to nobody, which is exactly why it gets a piece.

horizontal_rule

NOT filed

success, stalled_for_genius, urgent_bug_pending, graceful quota bail — these are working-as-designed verdicts

priority_high

Soft exception

a mid-task quota kill, a spent rate-limit window, a spent CLI usage limit, or a piece taken away from the runner all arrive wearing status error and are never filed (softTerminations = {"quota", "rate_limited", "usage_limit", "reassigned"}, internal/bugreport/runner_error.go)

attach_file

What the bug task carries with it

Three artefacts are attached to the filed piece (internal/bugreport/runner_error.go, attachArtifacts), so the next person can read the failure without re-running it.

description

Failing attempt's run log

the JSON the on-disk run log finalized with

terminal

Stream log tail

the newest stream log, capped at 512 KB (internal/bugreport/runner_error.go, streamTailBytes)

visibility_off

Redacted config snapshots

depend on the harness — .mcp.json, .claude/settings.json and .claude/settings.local.json on Claude Code, .codex/config.toml on Codex CLI, opencode.json on OpenCode — tokens stripped (mcptask/piece.go, ConfigFiles + Redact)

fingerprint

One incident is one piece

Identical failures are fingerprinted, throttled, and reported exactly once. Without that, a tight loop would file the same bug 200 times before somebody noticed.

FingerprintNormalizationThrottle windowStamp claim before file
sha256 over termination + normalized message + task id + project, truncated to 16 hex chars (internal/bugreport/runner_error.go, fingerprint)strips paths, hex strings, and decimal numbers — what looks like the same crash with a different timestamp counts as the same crashsix hours (internal/bugreport/runner_error.go, throttleWindow + claim) — the second identical fingerprint within the window is droppedthe fingerprint stamp is claimed BEFORE the piece is created and RELEASED if the piece never gets created. The failure most likely to break CreatePiece is the one a retry would otherwise silently suppress
warning

Panics — re-raised, with a stack trace

A panic is a hard failure with a forensic trail. The piece gets the trace; the loop keeps the panic (loop.go, reportPanic). Otherwise the trace would be eaten by the panic recovery and lost to whoever picks up the task next.

shield

Fail-safe throughout

Every reporting path swallows its own errors, including its own panic (internal/bugreport/runner_error.go, MaybeReport defer/recover). A reporting failure can never end a run. The whole premise is that the worst-case reporter is the one that costs the least — a swallowed log line, not a missed crash.

visibility_off

The blind spot, stated openly

The reporter authenticates with the same token as everything else, so it CANNOT file the bug that says the token is missing. A missing or unconfigured token skips the reporter silently — OutcomeSkippedDisabled writes no error line, so only the runner's own exit code and log carry it. A token the server refuses is the loud path instead: OutcomeRefused logs an error naming the token (internal/bugreport/runner_error.go). The alternative was to hide the gap; the choice here is to name it.

swap_horizFor developers

Model-agnostic by configuration

Three tiers map to any provider and to any of the three coding CLIs. Claude Code points at any Anthropic-API-compatible endpoint (local Ollama is the documented example) by changing the launcher and the models section. OpenCode takes provider/model IDs from any provider, Codex CLI ships its own defaults, and `ollama launch <cli>` works for all three.

menu_bookModel configuration YAML examples – technical detailsexpand_more

The models: section is optional. Without it each harness uses its own generic aliases, resolved by that CLI at runtime: opus / sonnet / haiku on Claude Code, gpt-6-astra / gpt-5.6-terra / gpt-5.6-luna on Codex CLI, and provider/model ids on OpenCode. Pin versioned IDs only when you want deterministic retries, or when routing through a non-Anthropic backend — and once you do, set all three of genius / smart / primitive. The runner does not refuse a partial models: map, but the background work the CLI does on its own then fails with "model may not exist".

terminalPer-harness defaults
# models: is optional. Without it each harness falls back to the generic
# aliases its own CLI resolves at run time.

# Claude Code
models:
  genius:    opus
  smart:     sonnet
  primitive: haiku

# Codex CLI
models:
  genius:    gpt-6-astra
  smart:     gpt-5.6-terra
  primitive: gpt-5.6-luna

# OpenCode — every id is provider/model
models:
  genius:    <provider>/<strongest>
  smart:     <provider>/<middle>
  primitive: <provider>/<fastest>
terminalExample: non-Anthropic backend
models:
  genius:    minimax-m3:cloud
  smart:     kimi-k2.7-code:cloud
  primitive: deepseek-v4-flash:cloud

launcher:
  # Local ollama
  command: [env, "ANTHROPIC_BASE_URL=http://localhost:11434", "ANTHROPIC_AUTH_TOKEN=ollama", "claude"]

The IDs below are one example, not a recommendation — they age fastest of anything on this page. The shape is what matters.

launcher.command flags (launcher.go, Flag* constants): a null flag is OMITTED from the spawned argv, a value flag with no token is passed POSITIONALLY.

auto_fix_highThe loader never fails

A missing, empty, or malformed config/mcptask_runner.yml is treated as "no config" (config.go, LoadFile) and the built-in defaults apply — the loader itself never fails. The run preflight is a separate check, and it does refuse to start: when harness: is unset, when the token is missing, or when the project has no CLAUDE.md.

check_circleNo pin

No shipped skill declares `model:` — ForkModelEnv therefore reaches every forked subagent, including ones running on a non-Anthropic launcher.

info

The configuration lives in config/mcptask_runner.yml, resolved relative to the working directory of the launcher (config.go, FileName; README.md) — the same file the Ruby gem reads, with the same semantics. You can also set additional options there, such as the destination Epic for auto-found bugs or the waiting strategy between runs.

GEM — exact file path, exact semantics

system_update_altSelf-update

update --self, and the deliberate non-feature

The runner is a static binary installed by download, not a gem that bumps on every bundle. Two commands move it: update --self fetches a newer release and swaps it in, update --check reports whether one exists and does nothing else. The third decision — whether the runner should ever move itself before a scheduled run — is the one we deliberately do not automate.

menu_bookHow the runner swaps itself for a new version – technical detailsexpand_more
swap_horiz

How update --self swaps the binary

The new binary lands in the same directory as the old one and a rename dance moves it into place (selfupdate.go). That sequence matters more than it looks: a rename is only atomic within ONE filesystem, and /tmp usually is not the same one, so the new file is staged alongside, not copied over from somewhere else. A failed or mismatched download never reaches the swap — the binary on disk is untouched and the command exits non-zero.

file_download

1. Fetch from the public releases mirror

the binary for this GOOS/GOARCH is downloaded from jchsoft/mcptask-releases (selfupdate.go, DefaultRepo). The source repo is private, so the public mirror is the only host the download can come from without a token

verified

2. Verify against the release's checksums.txt

the asset's sha256 is compared to the one the release published (selfupdate.go, verify). A mismatch aborts before any write touches the install directory

file_copy

3. Stage alongside, rename the running binary aside

the new binary is written next to the running one, then the running one is renamed to a sibling name. That is also what makes Windows work — overwriting a file that is mapped for execution fails on Windows, and the rename-aside step sidesteps that

published_with_changes

4. Rename the staged binary in, atomically

a final os.Rename on the same filesystem moves the staged binary into the install path. Atomic within that filesystem. The command exits 0 only after the swap

visibility

update --check — a tag lookup, no download

update --check resolves the latest release tag, compares it to the running version, reports the answer and exits. It downloads no asset and verifies no checksum — none of the four steps above run. The binary on disk is not touched. It is the right command to wire into a nightly job that wants to alert without installing.

shield

The deliberate non-feature

The runner does NOT update itself before a scheduled run. It prints a one-line hint when a newer release exists and does nothing else, because one bad release must not reach every unattended host at 08:05 simultaneously. Upgrading stays the operator's decision, even when the binary runs unattended.

notifications

What the hint does

selfupdate.Hint prints one line when a newer release exists (hint.go, hintTimeout + hintInterval). Capped at 2 s, cached for 24 h, ignoring every error — a network blip or a rate-limit never reaches the run

block

How to turn it off

set MCPTASK_NO_UPDATE_CHECK=1 in the host environment. That is what an air-gapped or metered host wants, and the check is silenced before the cache file is even read

inventory_2

One of three paths that restart the runner — bundle adoption

BUNDLE ADOPTION: on run, if the project's own Gemfile.lock names a newer wrapper-gem version than the running binary, the run re-execs on that binary. It FETCHES NOTHING — the operator already made that decision by merging the lockfile. The 08:05 host reads the lock it pulled from main and runs what the lock says to run. The third path sits between tasks: an iterating mode takes up a newly installed binary before it starts the next task — including one update --self put into ~/.mcptask/bin — and checks the lockfile for adoption at that same boundary.

repeat

MCPTASK_ADOPTED

set in the environment of the re-exec'd child, this stops adoption from becoming an exec loop (adopt.go, EnvMarker). A binary that reports older than the lock would otherwise adopt, exec, find itself outdated again, and spin

push_pin

MCPTASK_NO_ADOPT

set to 1 to pin the installed binary no matter what projects carry. A host that wants one runner version across every project it runs sets this once

bedtime

Another path — daily hands over to a newer binary before its overnight pause

On the way into its overnight sleep, a daily run checks whether the binary on disk is still the one it started from. If not, it releases the instance lock and execs the installed binary — restartOnCurrentBinary (restart.go). That handover DOWNLOADS NOTHING either — it only adopts what update --self or an operator already put on disk. The ordering is load-bearing: the day's work is finished, so no child is alive and no working tree is half-written, and the process that wakes up tomorrow is the one that started this evening.

help_outline

Why it exists

everything that notices a new version happens once, at startup. A daily process starts once and never ends, so an unattended host would otherwise keep whichever binary it happened to start with however many releases go by — exactly the host nobody is watching. Making daily exit instead was considered and rejected: a runner that goes away is one a missing or unloaded scheduled job never brings back

lock_open

ErrLockLost — the third ending

restart.go declares ErrLockLost, a handover that let go of the instance lock and then could not take it back. That ending DOES stop the loop, deliberately and loudly — carrying on unlocked would let the next scheduled start put a second runner in the same checkout

autorenewMaintenance

Skill Updates Without Losing Your Edits

An installed project drifts as the runner gains skills. One command re-syncs it: mcptask_runner update. It compares every skill and helper against the install manifest, reports what it did to each one, and refuses to overwrite anything you edited yourself. It is also how an installed project stops calling `gh`: the bundled skills migrated to `mcptask_runner pr`, and `update` is what brings that migration to a project installed before it.

The command

mcptask_runner update
menu_bookWhat update does to your project – the five outcomes and --forceexpand_more
fact_check

Five outcomes, one per file

The updater holds a content hash for every file it installed. Each skill and helper is classified against that manifest and the result is printed — five outcomes, not a silent copy. A project installed before the PR work reads its skills as `updated` — they are the shipped version, untouched by you — so the `gh` calls inside them are replaced without asking.

add_circle_outlineadded

The file is not on disk. The updater writes it. New runner capabilities land this way without you asking for them by name.

check_circle_outlineup-to-date

The file on disk hashes to the manifest entry. Nothing is written. A gem-installed host reads as up-to-date to this binary — checked directly against the file, not assumed from the install channel.

syncupdated

The file matches an older manifest entry, so it is the shipped version and you never touched it. Safe to replace, and it is replaced.

blockconflict-skipped

The file hashes to neither the current nor a known shipped version — you edited it. The updater leaves it exactly as it is and says so. This is the default, and it is why the command is safe to run on a project you have customised.

restore_pageforce-updated

The same conflict, resolved the other way because you asked. Only reachable with --force or FORCE=1.

restore_page

Overwriting a conflict, with a way back

mcptask_runner update --force

--force (or FORCE=1) overwrites locally modified files instead of skipping them — and keeps a .bak next to each one it overwrites. Your edits are not deleted; they are moved aside where you can diff them.

shield

What plain update deliberately leaves alone

update touches skills and helpers — and one outdated .mcp.json entry, described below. Everything else the installer once wrote stays exactly where it is. That restraint is the selling point, not an omission: re-syncing skills must not silently re-activate a schedule or rewrite a config file you tuned by hand. `git_host:` and `harness:` are config, so `update` never touches them either — moving a project to another PR host or another coding CLI stays an `init` decision.

schedule

The scheduled job

The LaunchAgent, systemd user timer or Task Scheduler job is not regenerated and not re-enabled. Use init --schedule when you actually mean to regenerate it.

data_object

.mcp.json

Your MCP client configuration is not regenerated. The token variable you declared survives the update untouched; the one change update makes is moving a legacy SSE mcptask-online entry to HTTP, the transport the server speaks today.

tune

The config sections

config/mcptask_runner.yml keeps the values you set — the work window, the launcher command, the end-of-workday hour. update has no opinion about them.

system_update_alt

Updating the binary is a separate decision

update re-syncs the project. It does not move the binary that did the syncing. Two other commands do that, and neither of them ever runs on its own: mcptask_runner update --self swaps in a newer release, and mcptask_runner update --check reports whether one exists and changes nothing.

person

The runner never updates itself before a run

At the start of a run the runner may print one line saying a newer release exists. That is the whole feature — capped at 2 s, cached for 24 h, every error ignored, and silenced entirely by MCPTASK_NO_UPDATE_CHECK=1. Upgrading stays the operator's decision, because a runner that moved itself before a scheduled run would be changing the thing doing the work without being asked. That matters if you are leaving a binary running unattended.

verified

A bad download never reaches the swap

update --self downloads from the public mirror jchsoft/mcptask-releases and verifies the asset against the release's checksums.txt. A failed or mismatched download aborts before any write touches the install directory: the binary on disk is untouched and the command exits non-zero.

verifiedProof

Conformance scenarios and the CI matrix

The evidence behind every reliability claim on this page: recovery, watchdog, stall detection and quota each have a recorded scenario of their own.

playlist_play

Fourteen conformance scenarios

Each scenario is a recorded conversation with the claude CLI plus the behaviour a runner is required to show while replaying it. The scenarios are implementation-neutral: plain YAML plus the contract in conformance/README.md, with no Go or Ruby in them. Driven with `go run ./cmd/conformance {list,run --impl go}`.

Every failure mode the runner claims to survive has its own recorded scenario in conformance/scenarios/. The two implementations run the same files; the Ruby gem is the retired, frozen reference implementation the suite can be compared against, and the gate is the Go suite being green.

menu_bookAll 14 conformance scenarios the runner has to passexpand_more
Scenarios 
01 — successthe happy path: a clean run ends with a green verdict, no retries, and a branch ready to push
02 — missing_marker_retrythe run log never finalises a marker; the runner retries until the budget runs out and then bails with a verdict instead of a hang
03 — context_overflow_fresh_restartcontext overflow while the engine is still running — the in-process restart is what carries the work forward and keeps the run alive
04 — context_overflow_terminalcontext overflow after the engine has stopped — the carry-forward machinery hands the file to a new process and that is the only way the run gets out alive
05 — api_overload_529the upstream returns 529 — the runner backs off, the run stays on the right task, and the marker does not advance on a transient
06 — tool_not_enabled_fresh_restarta tool the run needs is missing on the first attempt; `--continue` re-runs it as a fresh restart and the run recovers
07 — tool_not_enabled_terminalthe same tool is still missing after a fresh restart — the run bails with a verdict, not a hang
08 — stall_edit_failuresthe engine edits a file the harness watches; the edits diverge and the watchdog catches it before the marker advances on a lie
09 — inactivity_kill_then_recoverthe child process stops talking — the watchdog kills it, the debug dump is left on disk, and the next attempt picks up the same task
10 — recoverable_retries_exhaustedevery recoverable retry has been spent and the run is still not done — that is a verdict, and the verdict is the only safe outcome
11 — quota_mid_taskthe budget runs out while the engine is mid-edit — the runner kills the child, the attempt ends as quota_exceeded_mid_task with no retry and the loop ends; only the daily mode starts again from triage the next day
12 — stream_closedthe upstream closes the stream mid-turn — the runner detects the close, treats the run as incomplete, and the next attempt re-derives state from disk rather than the gone stream
13 — hung_tool_killa tool the engine called has stopped answering — the watchdog kills the child with the kill-on-hang signal and the run is left re-runnable
14 — triage_unverified_pickthe next-task picker returns a task that has not been verified — the runner refuses to start it, files nothing, and waits for a verified candidate rather than picking something blindly
science

The binary carries no test hook of its own

The runner is pointed at a mock CLI through the ordinary `launcher.command` override (README.md) — the same per-host configuration a developer uses to drive a real one. The suites run in all three harness dialects (Claude Code, Codex CLI, OpenCode), and a conformance project on Bitbucket Cloud and one on gitlab.com cover the other two PR hosts, so a green result holds for every dialect and every host. There is no test-only code path in the binary for the conformance suite to lean on, so a green run here is the same code a user installs.

grid_view

The CI matrix

Every commit runs a 3-OS x 2-Go-version matrix (ubuntu / macos / windows x 1.25.x / stable), go vet, race-detector tests on Unix + stable, bin/smoke on the BUILT binary, gofmt, golangci-lint, a goreleaser config check plus a six-target cross-compile, and the whole conformance suite (.github/workflows/ci.yml). The matrix exists because the runner is installed on developer machines — every desktop platform has to build and pass.

fact_check

What a local run does not cover

A green LOCAL run prints what it did NOT cover — "This machine only: <goos>/<goarch>" — because one OS, one architecture and one Go version is not what CI runs (bin/ci). Green locally is necessary, not sufficient. The Ruby comparison run is a reference check, not a gate: it is skipped, not failed, when the retired gem is not checked out — a machine that cannot reach the reference should still be able to check its own work.

visibility_off

The boundary, named

The conformance harness runs one runner process per scenario, so it cannot express the cross-process half of overflow survival; that is covered by tests running two engines in sequence sharing only the file (README.md). The Ruby comparison run is skipped, not failed, when the retired gem is not checked out, and it holds up no merge either way.

rocket_launch

Turn the machine on in the morning

One static binary, one install line, and the runner starts taking tasks from your queue.

verified_userFree 30-day trial. No credit card. Cancel any time.