menu_bookDokumentace runneru

Všechno, co runner dělá, do detailu

Seznam funkcí a referenci příkazové řádky bereme z katalogu, který runner zveřejňuje s každou verzí, přehled změn z jeho vydání na GitHubu. Pod nimi technické podrobnosti: kvóta, přetečení kontextu, souběh, logy, hlášení chyb, modely, samoaktualizace a údržba projektu.

arrow_backZpět na přehled runneru
buildKlíčové vlastnosti

Co runner zvládne sám

Každá funkce tak, jak ji popisuje sám runner v katalogu, který zveřejňuje s každou verzí.

Od přiděleného úkolu ke sloučenému pull requestu

od v0.1.0

Runner bere úkoly, které máte na mcptask.online přidělené, jeden po druhém. Nejdřív každý úkol roztřídí a vybere pro něj jednu ze tří úrovní modelu; pak ho kódovací nástroj zpracuje na vlastní větvi — kód, testy, push, pull request, CI. V režimech auto-squash runner pull request po zeleném CI sám sloučí a ověří, že se sloučení opravdu stalo; v manuálních režimech pull request počká na člověka. Odpracovaný čas a postup se průběžně zapisují zpět k úkolu.

Jeden nepovedený úkol nezastaví celý den

od v0.3.22

Úkol, který selže, ukáže se jako mimo zadání nebo ho už začal někdo jiný, se odloží a smyčka pokračuje dalším. Po čase se odložený úkol může vrátit k dalšímu pokusu, jen omezeněkrát za den — přechodný výpadek ho tak nestojí celý den a za skutečně rozbitý úkol se neplatí pořád dokola.

Přežití zaplněného kontextu

od v0.1.0

Když se kontext kódovacího nástroje zaplní, sezení je ztracené, ale práce ne — je to větev s commity na disku. Runner nástroj spustí znovu v čistém sezení a předá mu jeho posledních pár kroků, a pokud dojde místo i tam, nechá dalšímu běhu krátkou předávací poznámku: větev, commity, necommitnuté soubory a nedokončený plán, s upozorněním, ať s kontextem šetří. Úkol, jehož pull request už je otevřený, se počítá jako odvedený, ne jako chyba, a navázaný úkol běží na nejsilnější úrovni modelu.

Denní rozpočet, který čte, nikdy neodhaduje

od v0.1.0

Kolik práce na dnešek zbývá, zjišťuje runner živě z vašeho účtu na mcptask.online, a to před úkolem, mezi úkoly i každých pár minut během něj. Když je rozpočet vyčerpaný, práce se zastaví; režimy today tím den končí, režim daily přespí do dalšího pracovního dne a začne znovu. Volitelný konec pracovního dne zastaví nové úkoly v nastavený čas a výpadek služby s kvótou runner rozliší od vyčerpaného rozpočtu a nahlásí ho jako chybu.

Hlídání zaseknutého nebo zacykleného běhu

od v0.1.0

Hlídač sleduje výstup kódovacího nástroje průběžně, jak přichází. Zastaví běh, který se příliš dlouho nikam neposunul, nebo příkaz, který visí, a pozná běh, který se točí na místě — tatáž úprava selhává znovu a znovu, tentýž příkaz padá, opakují se stejné kroky — přičemž dlouhé čekání na CI a testy nechá být. Úkol zůstane rozpracovaný a později se na něj naváže na nejsilnější úrovni modelu.

Nesouvisející urgentní chyba se stane samostatným úkolem

od v0.1.0

Když kódovací nástroj narazí na urgentní chybu, která s jeho úkolem nesouvisí, rozpracovaná práce se commitne a pushne, chyba se založí jako nový urgentní úkol a runner ji opraví jako první — i po restartu — a teprve pak se vrátí k původnímu úkolu.

Na sdíleném stroji jedna testovací sada naráz

od v0.1.0

Pomocné skripty pro testy a CI, které runner instaluje, sdílejí jeden zámek pro celý stroj, takže dva agenti — ani dva projekty — nikdy nepouštějí své testy současně a nezpomalí se navzájem až k vypršení časových limitů. Druhá sada počká na první a zámek, jehož vlastník už neexistuje, se uvolní sám. Navíc v jednom pracovním adresáři nikdy nepracují dva runnery zároveň.

Pád si sám založí chybový úkol

od v0.1.0

Když běh skončí chybou, pádem nebo porušením vlastní pojistky runneru, runner založí na mcptask.online chybový úkol a přiloží k němu log běhu, konec výstupu nástroje a jeho konfigurační soubory s odstraněnými tokeny. Shodná selhání pozná a založí jen jednou; vyčerpaná kvóta nebo úkol přeřazený někomu jinému nezaloží nic. Jediné selhání, které nahlásit nemůže, je chybějící nebo odmítnutý token — to se proto ukáže v návratovém kódu a v logu.

Co běh dělá — živě i zpětně

od v0.1.0

Každé spuštění kódovacího nástroje dostane JSON log běhu, který se otevře hned při startu nástroje a průběžně se obnovuje, takže i zaseknutý nebo spadlý běh po sobě nechá na disku svůj stav; vše ostatní zachytí log celého procesu. Během běhu ukazuje živá karta na mcptask.online jeho stav, úkol, model a postup. Sledování běh nikdy nezastaví: vypadlá aktualizace stojí čerstvost údajů, ne výsledek.

Aktualizace, o kterých rozhodujete vy, převzaté bez restartu

od v0.2.0

update --self nainstaluje poslední vydání po ověření proti zveřejněným kontrolním součtům, a když se výměna nepovede, vrátí starý binární soubor. Sám se runner nikdy neaktualizuje — na začátku běhu jen zmíní, že existuje novější vydání. Jakmile nový binární soubor nainstalujete, dlouho běžící smyčka na něj přejde mezi dvěma úkoly a režim daily při přechodu do noční pauzy, takže stroj bez dozoru nezůstane na staré verzi.

Větve Story — celá Story jedním sloučením

od v0.3.35

Zapnuté, dokud je hostitel nevypne. Úkoly jedné Story dostanou každý vlastní větev a slučují se do společné větve Story místo do hlavní; poslední úkol Story pak sloučí větev Story do hlavní větve, takže ta dostává hotové Story, ne jejich poloviny. Runner větve pojmenuje sám a sám rozhodne, odkud který úkol začíná, takže agent nic nehádá. Manuální režimy větev Story nikdy nezakládají ani neslučují.

Hotová Story se do hlavní větve dostane sama

od v0.3.37

Story, jejíž podúkoly jsou všechny hotové, se do hlavní větve sloučí i tehdy, když úkol, který ji dokončil, běžel souběžně, na jiném stroji nebo na stroji, který větve Story nepoužívá. Hned po takovém úkolu a mezi úkoly v režimech auto-squash runner hledá hotové Story, jejichž větev je pořád otevřená, a sloučení Story spustí sám — kontrola, pull request, kontroly hostitele, sloučení. Pak sám ověří repozitář, a Story, která se tam stejně nedostala, nahlásí jako chybu, jednou.

Claude Code, Codex CLI nebo OpenCode

od v0.3.24

Každý projekt si jednou, příkazem init --cli, zvolí kódovací nástroj, který ho pohání; volba je uložená v souboru konkrétního stroje, takže dva vývojáři na stejném repozitáři mohou používat různé nástroje. Pracovní smyčka, hlídání, karta na nástěnce, záznam odpracovaného času i hlášení chyb jsou stejné, ať běží kterýkoli nástroj. Výchozí volba neexistuje: projekt, který si nevybral, runner jmenovitě odmítne, a můžete přidat i vlastní profil.

Pull requesty na GitHubu, Bitbucket Cloudu a GitLabu

od v0.3.25

Jediný příkaz, mcptask_runner pr, zakládá, čte, kontroluje a slučuje pull requesty na tom ze tří hostingů, kde repozitář je, takže stejné pokyny fungují všude. Před sloučením v režimu auto-squash se runner zeptá hostingu na verdikt CI: selhání nechá pull request otevřený, na běžící pipeline počká a repozitář bez kontrol na hostingu se spolehne na vlastní lokální bránu projektu.

Vlastní modely na backendu podle vaší volby

od v0.1.0

Tři úrovně modelů se překládají na to, co určuje profil kódovacího nástroje, dokud si projekt nenastaví vlastní modely. Spolu se spouštěcím příkazem — typickým příkladem je lokální Ollama — pak kódovací nástroj může běžet proti jinému backendu, než je backend jeho výrobce. Nastavte všechny tři úrovně, nebo žádnou: nenastavená úroveň se takového backendu zeptá na model, o kterém nikdy neslyšel.

Úloha na pracovní dny na macOS, Linuxu i Windows

od v0.1.0

init vygeneruje naplánovanou úlohu, která runner spustí každé ráno pracovního dne — na macOS jako LaunchAgent, na Linuxu jako uživatelský časovač systemd, na Windows jako úlohu Plánovače úloh — v 08:00, pokud nezvolíte jiný čas. Sama ji nezapne: spustit něco, co každé ráno čerpá kvótu, zůstává vaším rozhodnutím, a init vypíše jediný příkaz, který to udělá.

new_releasesChangelog

Co je v runneru nového

Poznámky k vydání každé verze runneru tak, jak vyšly na GitHubu. Nejnovější tři jsou otevřené, ostatní jsou o jedno kliknutí dál.

Poznámky k vydání jsou v angličtině.

v0.3.37

Upgrading

Upgrading starts with one command. mcptask_runner update --self replaces the binary. The bundled .claude/ assets did not change, so nothing on that account asks for a bare update in the projects — the sections below say whether anything else does. A runner idling in a wait keeps the old binary until its next task or a restart (launchctl kill SIGTERM, wait for the job to stop, then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.37 carries this binary; no wrapper changes.


Story branches are on by default

story_branches.enabled no longer has to be written: a host whose
config/mcptask_runner.yml lacks the key now works Stories on story branches
(#13501). A host that must stay
off says enabled: false. Nothing to do on a host that already had true.
Every host of one project should agree — a host that opts out merges its Story
tasks into main beside the others, unless the Story's branch is already live
(see below).

A live story branch binds every host

enabled: false now means "never start a story branch", not "ignore them"
(#13493). When a task's Story
already has a branch on origin, an opted-out host works the task on it like
any other host and logs this host opted out of story branches, but Story #N
has a live branch … — following it
. An opted-out host therefore reads the
task's piece and git ls-remote origin before an auto-squash or manual task,
and a failure there sets the task aside with a story_branch_plan_failed bug
piece, as on an enabled host. Nothing to configure.

A Story task whose verified pull request merged into a branch other than the
planned one now files a story_pr_wrong_base bug piece instead of passing
silently.

A finished Story is merged even when its last task did not know it was last

A Story whose subtasks were finished in parallel, on another host, or on a host
with story branches off used to sit on its live …/story branch until a person
opened the pull request into main
(#13492). The runner now runs a
story-merge run for such a Story: right after the task that finished it, and
at every task boundary of today_auto_squash, queue_auto_squash and
story_auto_squash, where it lists refs/heads/*/story on origin and merges
every Story whose subtasks are all completed or approved — on hosts that opted
out of story branches too. A pull request another host already opened is picked
up, not duplicated. Each Story is tried once per process; a branch still on
origin afterwards is filed as story_merge_failed, and a Story or listing that
cannot be read as story_sweep_failed. The run takes no slot of the daily task
budget and does not start after the end of the work window.

Nothing to configure. Expect, on the first day after upgrading, story-merge runs
for Stories that were finished and left unmerged before — look at origin for
*/story branches first if you would rather merge some of those by hand.


Changelog

Other

  • c693b3f5e78291253db5e242251017b272e149aa [#13492] A finished Story with a live branch is carried into main
  • 64995c943a486fe198ad335f112f8499a0623bb0 [#13493] A live story branch binds every host
  • f0f2876e3ec5a42edb251cd3589d80355a14900b [#13501] story_branches is on unless a host opts out
  • 639a6e78f7be7b78cd4308984877c1b7046c8faf docs(release): v0.3.36 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.36

Upgrading

Upgrading starts with one command. mcptask_runner update --self replaces the binary. The bundled .claude/ assets did not change, so nothing on that account asks for a bare update in the projects — the sections below say whether anything else does. A runner idling in a wait keeps the old binary until its next task or a restart (launchctl kill SIGTERM, wait for the job to stop, then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.36 carries this binary; no wrapper changes.


A feature catalog the website renders (Task #13451)

The public page at mcptask.online/runner was
written by hand and drifted from what the binary does. This repository now
carries catalog.yml: one entry per feature,
subcommand, flag, config key, harness, git host and exit code, each with the
release it first shipped in and a title and summary in English, Czech and
Slovak — the file the page will be rendered from.

internal/catalog loads and validates it strictly. A missing translation, an
unknown kind, a duplicate id, a flag without its subcommand or a since that is
not a release tag is an error that names the entry and the field, never a row
the page quietly drops.

What an operator has to do: nothing. The catalog is documentation shipped
with the source; the binary does not read it and no host behaves differently.

Every release carries catalog.yml as an asset (#13454)

From this release on, docs/catalog.yml — the runner's feature catalog — is attached to the GitHub release on jchsoft/mcptask-releases as catalog.yml. The latest one is at https://github.com/jchsoft/mcptask-releases/releases/latest/download/catalog.yml; mcptask.online reads it from there. bin/release now refuses to start without the file. Nothing for an operator to do.

The catalog and the code may not disagree (Task #13453)

A guard test, TestCatalogMatchesTheCode in internal/catalog/guard_test.go,
compares docs/catalog.yml with the code in both directions: every visible
subcommand and flag in the command tree, every config key the loader reads,
every bundled harness profile, every git host and every exit code has an
entry, and every entry names something that still exists. A readme_anchor
must land on a real README heading. bin/ci runs it, and a red run names the
id to add or remove. To list its exit codes from code rather than from a copy,
cli.Execute's codes are now named constants and the signal handler's list is
read through runner.SignalExitCodes; the codes themselves are unchanged.

What an operator has to do: nothing.


Changelog

Other

  • 547257ed8dec2ed0e229c1b406ce7022ad2e75e6 [#13451] Feature catalog: docs/catalog.yml and internal/catalog
  • bc62ec5003ff4f27129f4ca3c864805af8b25c3f [#13453] Guard test: the catalog and the code may not disagree
  • 8a8b40d54d6330144e28e8b01702eb1e29629080 [#13454] Publish docs/catalog.yml as a release asset on jchsoft/mcptask-releases
  • 3bcf928392c9b3e25c381b3d494103424b3704da docs(release): v0.3.35 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.35

Upgrading

Upgrading starts with one command. mcptask_runner update --self replaces the binary. The bundled .claude/ assets did not change, so nothing on that account asks for a bare update in the projects — the sections below say whether anything else does. A runner idling in a wait keeps the old binary until its next task or a restart (launchctl kill SIGTERM, wait for the job to stop, then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.35 carries this binary; no wrapper changes.


Story branches: a Story reaches main in one merge (opt-in) (Story #13009)

Until now every task merged into main on its own, Story subtasks included, so
main held half-finished Stories and a bug fix released from it shipped them
too. A project can now have the tasks of an unfinished Story merge into a
shared story branch instead, and main gets the whole Story in one merge when
its last task is done. Tasks outside a Story are unchanged.

It is off unless the project turns it on, in config/mcptask_runner.yml:

story_branches:
  enabled: true

A value that is not true or false is not read as off: a Story task then
refuses to start, the task is set aside and a bug piece names the key.

What the runner does with it on:

  • Names. A Story's branch is <story_id>-<slug>/story, its tasks' branches <story_id>-<slug>/<task_id>-<slug> — ASCII slugs of the titles, at most 40 characters. The runner decides both before the child starts and pastes them into the prompt; the child no longer makes up a feature/… name for a Story task.
  • Liveness is the remote. The story branch existing on origin (git ls-remote) is what keeps a Story open, whatever its state on mcptask.online. No branch and an approved Story → the task works against main, as before. No branch and an unapproved Story → this task creates the branch from main; two runners racing to create it end up on the same one. A branch is found by the Story's id, so renaming a Story does not cut a second branch.
  • Every Story task checks out the story branch, merges origin/main into it (so fixes from main reach the Story while it is open), cuts its task branch, opens its PR with --base <story branch> and squash-merges it there. The task link stays the LAST mcptask.online link in the PR body.
  • The last task — every other subtask completed or approved; a blocked one holds the Story — then opens a PR from the story branch to main whose body ends with the Story's link, merges it with a merge commit (mcptask_runner pr merge --squash=false) and deletes the Story's branches. The runner checks the remote afterwards: a story branch still on origin is filed as a story_merge_failed bug piece, and the task's own success stands. A finished Story whose last two tasks ran at once and so left the branch unmerged is filed as story_left_unmerged.
  • Manual modes never create a story branch. With one live, the PR targets it; without one, the task branch is still named under the Story's path and the PR targets main. A person merges.
  • mcptask_runner pr merge --squash=false now asks GitHub for a merge commit by name (gh pr merge --merge); it used to pass no method, which gh refuses without a terminal.

Host caveat. Only GitHub approves a piece when its PR merges (by the last
mcptask.online link in the body, whatever the base branch). On GitLab a merged
task is only set to 100 % ("Hotovo?"), and Bitbucket has no webhook at all —
either way a merged task counts as completed, which is enough for the last
task to be recognised, but nothing approves the Story there automatically. On
GitLab the story → main merge follows the project's merge method; a project set
to fast-forward only refuses the merge commit, and that is reported as
story_merge_failed.

Upgrading: nothing for projects that do not add the key. To turn it on,
add the two lines above to config/mcptask_runner.yml and commit them. Runner
hosts pick the feature up with the binary.


Changelog

Other

  • 1403b0679f1bfe14a19c5a231225e0efe9487d90 [#13010] mcptask client reads a piece with its parent Story and subtask states
  • c6576725026957eb7f002280c7f7ec4607148e64 [#13011] Story branch planner: naming, liveness by the remote, race-safe creation
  • abecf2e79ec5f82a982b3d31f564f903f9e5ebff [#13011] [#13015] Lint: story-branch planner and conformance sandbox
  • fbc3cdbbd5df90b43250a9467fcdddbb1ba68a25 [#13012] Auto-squash prompts take the planned branches: story branch in, main merged first
  • 3742459b7aaee8011d43e203d86bd0c4dea7c2f7 [#13013] The Story's last task is checked against the remote; GitHub merge commits by name
  • 83f4a1d7072706aefdb20c916309471d31c9ed2d [#13014] Manual mode names Story task branches under the story path
  • 3eb3a17b4c955ada3bb66623cc6585bb293d230e [#13015] Conformance scenarios, release notes and README for story branches
  • 4cd3f81e380a070f1c91d4a4fef5b48de75a5542 docs(readme): keep a Mac awake through launcher.command, not the plist
  • 8138647edb6cf2452abf73d22e882e573ea649e6 docs(release): v0.3.34 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

menu_bookStarší verze – celý changelogexpand_more

v0.3.34

Upgrading

Upgrading starts with one command. mcptask_runner update --self replaces the binary. The bundled .claude/ assets did not change, so nothing on that account asks for a bare update in the projects — the sections below say whether anything else does. A runner idling in a wait keeps the old binary until its next task or a restart (launchctl kill SIGTERM, wait for the job to stop, then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.34 carries this binary; no wrapper changes.


A host with Claude Code's read block on refuses to start (#13168)

With permissions.blockReadsOutsideWorkingDirectories turned on, Claude Code
asks the person before any shell command it cannot analyse (a $var in a loop,
a heredoc, ruby -e, python3 -) and before any read outside the project. A
scheduled run has nobody to ask, so those actions are refused even under
bypassPermissions. On the MacBook this ended five runs in two days
(#13149 …
#13167), and the directory grant
from #13137 cannot cover commands.

The runner now reads Claude Code's settings files at startup (user, project,
project-local and managed, in that precedence) and refuses to start when the
effective value is true. The refusal names the file. The runner exits non-zero
and files no bug piece, the same as a missing token.

Upgrading: update --self is enough. On a host that has the setting on,
remove it from the named file, or set it to false, before the next scheduled
run. Until then that host runs nothing and reports the reason in the launcher
log.

A Sonnet card shows the whole plan, not 0/1 (#13424)

The plan block asked for "one item per step of the plan", and each model chose
its own plan. Opus made one item per numbered step, usually seven. Sonnet 5 and
5.5 made one item for the whole task ("Fix + tests + PR + CI + merge"), so every
Sonnet run showed TODO 0/1 on its card from start to finish, on every host.

The block now asks for one item per numbered step under WORKFLOW, and at least
one for every step that ends in a "LOG PROGRESS NOW" line. It also forbids a
single item for the whole task. Claude Code and OpenCode get the new wording.
Codex said "the workflow below", which was wrong in the auto-squash prompts,
where the block sits inside step 2; it now names the section too.

Upgrading: update --self is enough. A run already in progress keeps its
prompt; the next task picks up the new wording.


Changelog

Other

  • b4b962943caa4ab00c98cb8091f6d5d1e2d8ba2a [#13168] A host with Claude Code's read block on refuses to start
  • 688ee72caa82d93cf55622da0a5d82c9c12efb75 [#13424] The plan block asks for one item per WORKFLOW step
  • adfbfd79c7fb882aa6c5728e0ad21fe7ebe5fff3 [#13433] Stress signal tests no longer go red after 17:00
  • c03bed998d70217b24e1266d0ce5c977d8893e01 docs(release): v0.3.33 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.33

Upgrading

Upgrading takes two commands. mcptask_runner update --self replaces the binary; then a bare mcptask_runner update in every project on the host, because the bundled .claude/ assets changed (baseline_permissions.json). Run it while no runner is working that checkout. A runner idling in a wait keeps the old binary until its next task or a restart (launchctl kill SIGTERM, wait for the job to stop, then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.33 carries this binary; wrapper changes in its CHANGELOG.


A run the host refused is filed as a bug, not set aside quietly (#13139)

Sometimes the coding CLI refuses an action because of its own settings: a
permission rule denies it, or a read outside the project is blocked. The child
used to end such a run as failure. The loop then set the task aside for the
day and nothing was filed. On 2026-09-25 that happened to projectoid_ii
#13062 on the MacBook. The same
thing would have stopped every task on that host, and nobody would have been
told.

The runner now remembers each refusal it sees in the stream and logs the first
one as The host refused an action: …. If the run then ends as failure,
error, ci_failed, preexisting_test_errors, stalled_for_genius,
merge_failed or merge_unverified, it is filed as a bug under the termination
host_refused, quoting the first refusal. A run that got past the refusal files
nothing. The prompt also tells the child to end such a run as "status":
"error"
with "reason": "host_refused". That way the refusal gets filed on
Codex and OpenCode too, whose refusal wording nobody has recorded yet, so their
failures.host_refused lists are empty.

For this ending the throttle's fingerprint leaves out the task, so a host that
refuses every task files one piece, not one per task. A focused stress day
filed 76 pieces before this change. The stress day plays the new
host_refused vector and fails if no bug piece names it.

Upgrading: nothing to do. Hosts pick it up with the binary. If you override
a harness profile, add failures.host_refused to your copy (see the bundled
claude.yml).

Test and CI logs stay readable when reads outside the project are blocked (#13137)

The bundled /test-runner and /ci-runner write their logs under
~/.claude/logs/projects/<hash>/. That is outside the project, so a host with
Claude Code's permissions.blockReadsOutsideWorkingDirectories turned on
refused ci_wait's read, and every task stopped at its first test run. On
2026-09-25 this stopped projectoid_ii's
#13062 on the MacBook.

The permission baseline now adds "~/.claude/logs" to
permissions.additionalDirectories in each project's
.claude/settings.local.json. It is merged like allow and deny, so any
directories the project already lists are kept.

Upgrading: run mcptask_runner update in each runner-driven project. Nothing
else to change, and the read block can stay on.

The end of the workday stops the next task, not just an empty queue (#13076)

work_window.end_of_workday_hour (and end_of_workday_minute) now apply in
today_auto_squash, the mode every scheduled job runs. Before this fix it was
checked only when the queue was empty. A runner with work queued went from one
task to the next past the hour. On 2026-09-24 a runner set to stop at 18:00
started new tasks at 18:20 and 18:54. On 2026-09-23 a triage that waited out a
rate limit came back at 21:09, and the task it picked ran until 22:40.

The clock is now checked each time a task would start: once before triage and
again after it, because triage can take hours. This applies in today,
today_auto_squash and daily. A task that is already running still finishes.
The log says Past end of workday (HH:MM) before triage or … after triage,
and the day ends with the new verdict end_of_workday. The queue, story and
single-task modes still ignore the clock, as before.

Nothing to change in configuration. Hosts pick the fix up with the binary.

Auto-squash no longer merges a pull request that no CI has passed (#13089)

Until now an auto-squash run merged its pull request even when the project had
no bin/ci and the git host reported no checks at all (NONE). In that case
no CI had looked at the change. From this release the runner merges only on a
green signal: the host's checks are SUCCESS, or they are NONE and the
project has a bin/ci that the run executed and passed.

The rule is enforced by the runner, not only by the prompt. In every
auto-squash mode the runner sets MCPTASK_MERGE_GATE=auto_squash for the coding
CLI, and mcptask_runner pr merge then checks the host's CI before it merges.
It refuses when the checks are FAILED, still IN_PROGRESS, cannot be read, or
are NONE in a project without bin/ci. The refusal exits non-zero with
merge refused by the auto-squash gate: …, names the status to report, and
merges nothing. This works the same on GitHub, GitLab and Bitbucket.

On GitHub, mcptask_runner pr checks now answers NONE for a pull request
with no checks. Before, it failed with gh's no checks reported on the …
branch
message, so on GitHub a project without CI never reached the NONE
case at all. Running
mcptask_runner pr merge by hand, and every manual mode, is not gated.

A project with neither CI now ends the task with the new status
merge_skipped_no_ci. The pull request stays open, the agent leaves a message
on the piece saying why, and progress stops at 80 %. The loop sets the task
aside for the day and moves on, and the log reads Task #N has a pull request
left open because the project has no CI (no bin/ci, no host checks) and needs a
human to merge it
. The runner card lists the task under set aside with the
reason merge_skipped_no_ci. To get such projects auto-merging again, add a
bin/ci or CI on the git host.

Nothing to change in configuration. Hosts pick the fix up with the binary.

A crash on a customer's runner files a bug piece in the customer's own project (#13090)

Before this change the runner filed a bug piece about its own failure only when
it ran under the jchsoft account. Runners under any other account stopped and
logged locally. Now a hard failure on a customer's runner files a bug piece
in that customer's account, in the project the run was working
(project_relative_id from the host's CLAUDE.md). The piece goes under the
project's bug-bucket Epic (the lowest-numbered Epic marked for errors), or at
the project root when the project has none. It never goes into jchsoft.

A customer's piece carries no raw agent output. The raw runner log is not
attached. Neither is anything else that can quote source code: prompts, tool
inputs and results, the agent's messages and thinking, stderr dumps, or the
values in the CLI's configuration files. The piece carries:

  • the task link, runner version, OS/arch, CLI and its version, model, mode, machine, session, the CLI's exit status and the error message (code-like lines removed);
  • the last tool calls, by name and outcome only;
  • the agent's error lines, sanitized and cut at 200 characters;
  • runner_events.txt, the runner's own log lines with child output removed;
  • a filtered copy of the run record;
  • cli_config_keys.txt, the setting names from the CLI's config files, with no values.

The piece also lists what was left out and gives the local paths of the full
runner log and run record on that machine. Reports under jchsoft are
unchanged.

The run now ends non-zero when it owed a report and could not file one, under
any account, jchsoft included. The log says so at the end:
… runner failure report(s) owed by this run … — ending the run non-zero.
When the server refuses (401/403, no access to the project, an expired trial),
the runner does not try again for the rest of that run. In daily mode a
refusal also stops the loop at the end of the day. A plain outage does not stop
it.

The runner's debug log now writes Tool finished: <name> (error) for a tool
call that failed.

A customer whose token cannot create pieces in its own project gets
REFUSED — this runner may NOT file bug reports and a non-zero exit. The fix is
to give that user project membership. Nothing to change in configuration. Hosts
pick the change up with the binary, so no bare update is needed.

mcptask_runner on PATH for a coding CLI you start yourself (#13123)

A Claude Code session you open by hand in a gem-installed project used to fail
on mcptask_runner pr … with exit 127 — the binary is
~/.mcptask/bin/mcptask_runner and nothing put it on PATH — and then went
hunting for the path in the launchd plist. Runs the runner starts itself were
never affected.

init and update now write ~/.mcptask_env.d/runner_path, which puts the
runner's directory first on PATH (once, however many startup files source the
directory). For the copy bundled inside the Rails wrapper gem that is
~/.mcptask/bin, never the gem directory. A binary not named mcptask_runner
(a worktree build) writes nothing and says so; Windows prints the directory to
add by hand.

To pick it up: bundle exec rake mcptask_runner:update (or
mcptask_runner update), then open a new terminal / restart the session.

A scheduled job set up through npx no longer points into npm's cache (#13095)

npx @mcptask/cli init used to write the scheduled job (LaunchAgent, systemd
timer, Task Scheduler job) with the path of the binary inside npm's cache, for
example ~/.npm/_npx/<hash>/node_modules/@mcptask/cli/vendor/mcptask_runner.
Once npm cleaned that cache or fetched a newer version, the job failed at its
next start on a missing file. The same happened with a project-local or global
npm install.

Now, when the runner is the binary from the @mcptask/cli npm package, init
copies it to ~/.mcptask/bin/mcptask_runner (mcptask_runner.exe on Windows),
the same path the Rails channel uses, and the job runs that copy. The output
says Installed this runner as …, or Replaced … when a different binary was
there before, since other projects' jobs on the same host may run that file.
Keep it current with ~/.mcptask/bin/mcptask_runner update --self.

Jobs written before this change still point into npm's cache. Regenerate them
with npx @mcptask/cli init --schedule --force in each project. A bare
update does not do it.

Elapsed times count the hours the Mac slept (#13132)

On macOS the runner measured every duration with a clock that stops while the
machine sleeps, Maintenance Sleep included. A Mac mini left alone sleeps several
times an hour, so the times the runner reported came out short. One session ran
from 21:09 to 21:42 and was logged as Elapsed time: 0.19 hours.

These now use the wall clock and include sleep:

  • Elapsed time: … hours, Execution finished in …s, and every after N h in the log;
  • the task_worked hours in the result;
  • elapsed_s in the run record, the bug piece and the overflow handoff;
  • elapsed_s of a running tool on the dashboard card.

The 30- and 60-minute "ask the queue again" waits now end at the clock time
they announce. Before, a wait that crossed a sleep ran later than that time.

The watchdog still ignores sleep. Its inactivity and hung-tool deadlines, and
the waiting for: Bash since Ns part of the heartbeat line, measure only time
the machine was awake. A child cannot work while the machine sleeps, so a run
that was healthy before a sleep is not killed as stalled afterwards.

The same change affects Linux hosts after a suspend. Nothing to configure.
Hosts pick the change up with the binary, so no bare update is needed.

A piece name with braces no longer stops the day (#13121)

The runner lost a well-formed TASKRUNNER_RESULT when a string in it held
unbalanced braces. The usual case is a task_name copied from a piece name such
as a raw i18n hash, {one: "až %{count_max} uživatel…. It found the end of the
JSON by counting braces, strings included, so the object never closed. The
triage was logged as Missing TASKRUNNER_RESULT marker, retried, then failed
with no task_id, and the loop stopped with A task failed, stopping. The next
run picked the same piece and stopped again. On 2026-09-25 this stopped a
runner after four tasks. The only workaround was renaming the piece.

The marker is now decoded by the JSON decoder itself, which skips string
contents. A marker the runner finds in the raw stream is also unescaped
properly, so one spread over several lines is read too.

When a marker is present but its JSON cannot be used, the log now says so
instead of calling it missing:

  • result marker found but its JSON never closed: {"TASKRUNNER_RESULT": … (or did not parse (<decoder error>)), with the first 200 characters;
  • TASKRUNNER_RESULT marker present but unusable (…) - will retry with marker-only instruction;
  • after the retries: Unusable TASKRUNNER_RESULT after retries exhausted.

A marker that is really missing is still logged as Missing TASKRUNNER_RESULT
marker
. The retries are unchanged. Nothing to change in configuration. Hosts
pick the change up with the binary, so no bare update is needed.

Adopting the bundle's runner now updates the binary the scheduled job runs (#13133)

On a Rails host, a runner whose project Gemfile.lock names a newer
mcptask-rails-runner adopts that bundled binary, at the start of a run and
between tasks. Until this fix it installed the binary into the project's own
bin/mcptask_runner instead of ~/.mcptask/bin/mcptask_runner, and then ran
it from there. The day's run was on the new version, but the LaunchAgent
(systemd timer, Task Scheduler job) went on starting the old binary in
~/.mcptask/bin the next morning, so every scheduled run had to adopt again.

Now the adoption installs into ~/.mcptask/bin/mcptask_runner
(mcptask_runner.exe on Windows), the file the scheduled job runs. init,
the scheduler and the runner-on-PATH export compute that path with the same
function. A home directory the runner cannot find is an error that says so,
never a path relative to the working directory.

Copies the old code left behind are not deleted. At the start of run, in a
project that bundles the wrapper gem, the log says
[Adopt] <project>/bin/mcptask_runner looks like a runner binary adopted into
the project by mistake
. If that file is not the project's own, delete it. It
is about 10 MB. Check that it was never committed, because not every project
gitignores bin/mcptask_runner.

Nothing to change in configuration, and no bare update needed. Hosts pick
the fix up with the binary.


Changelog

Other

  • 21e8e9360eecd096c8b92827dc2b6776eeabe09d [#13076] No task starts after the end of the workday, and the stress day proves it
  • 6e96bc636e2938067a4fa825091ebf29e378f8a1 [#13088] The Bitbucket install tests no longer read the host's credentials
  • e7f96ff03899d44f7d8a9ef8f7bef2ed27452671 [#13089] Auto-squash merges only on a green CI signal, and the runner enforces it
  • 642d79ce7b4ee805e6359d46830ad4815c33c1f2 [#13089] GitHub checks with none reported answer NONE instead of failing
  • 4cb249304a5e7604346d9ad91f784ec175da0888 [#13090] A customer's runner files its crash into its own project, without raw agent output
  • ff34c2a772936bdebef01b25d8b6bcd505afcc47 [#13095] A job scheduled through npx runs a stable copy, not npm's cache
  • 5d754ec579a491ff7a2cfd3fc5e4b32bc9a43393 [#13121] A marker whose strings hold braces is decoded, not lost to a brace count
  • c55e20219b28d57cc57a67e88ac1e661e4133144 [#13123] init and update put mcptask_runner on PATH for sessions started by hand
  • 68be0fe333a3e29a48c21896c06cf8143b0f9267 [#13128] The updater resolves the mcptask token through an injected environment
  • 926062e754ed01351a51670424868f0a011175d9 [#13132] Reported durations count host sleep; the watchdog still does not
  • 44d445dfc19f779cd26a8a6a2217aec8d2f18ab7 [#13133] Adoption installs into ~/.mcptask/bin, the binary the scheduled job runs
  • fc43f98a756ab29a243bd1dce1d19d86af01f7a7 [#13133] [#13123] The Windows test run no longer assumes unix paths or an env directory
  • eae7ae45b239b386e86ca62f7bfe9c80400909d6 [#13137] The installer grants ~/.claude/logs so a read block cannot stop test and CI verification
  • 42cfa4cc388f4093557c46bf7afc7d1b890a32a5 [#13139] A run the host refused is filed as a bug, not set aside quietly
  • da351b1452fa33e03c2901533aaefe0e6a9143ef docs(release): v0.3.32 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.32

Upgrading

Upgrading starts with one command. mcptask_runner update --self replaces the binary. The bundled .claude/ assets did not change, so nothing on that account asks for a bare update in the projects — the sections below say whether anything else does. A runner idling in a wait keeps the old binary until its next task or a restart (launchctl kill SIGTERM, wait for the job to stop, then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.32 carries this binary; no wrapper changes.


A second bundle update in one day is taken up too (#12905)

Since 0.3.31 an iterating runner takes up a new runner from the project's bundle
at the next task boundary. It did so once per process lineage: the exec that
swapped it set MCPTASK_ADOPTED=1, every later process inherited it, and a
process carrying it never looked at the bundle again — so the second bundle
update
of the day waited for a manual restart. Now the marker names the version
that was installed (MCPTASK_ADOPTED=0.3.32, say), the new process takes it off
its own environment at startup, and it refuses only a re-adoption of that same
version. What an operator sees: each bundle update produces its own
"Updating from X to Y" on the card at the next boundary.

The marker no longer reaches the coding CLI's environment either.

One new line can appear in the log, and it means a broken release rather than a
broken host: [Adopt] This process was installed as the bundle's X but reports
itself as Y
— the gem's binary does not know its own version, and the runner
stays on it rather than reinstalling it in a loop.

No update is needed: nothing in a project's .claude/ changes. A runner still
on 0.3.31 sets the old flag form when it hands over; the new binary treats 1
as no version, so the handover onto this release is unaffected.

A refused push ends the task; it is no longer worked around (#12903)

In a smoke run a child's git push was refused because its SSH key had dropped
out of the agent. The child got past it by rewriting the project's origin to
another transport, finished the task, and left the clone changed — so the next
run failed on the rewritten URL, with an error that no longer pointed at the
missing key.

From this tag every PUSH step in every workflow (auto-squash, honest, manual,
story, review) tells the child that a refused push, fetch or pull —
authentication, permission, host key, protected branch — is not permission to
change the remote, switch transport, or touch credentials, SSH keys or git
config. It may re-run the same command once; after that it stops with status
failure and quotes git's refusal word for word, so the task note says what
the host said.

The runner also reads git remote get-url origin before and after every child.
If it changed, the run log gets an ERROR line naming the old and the new URL and
the git remote set-url origin <old> that restores it, and an anomaly piece
(origin_rewritten) is filed in the runner's Errors epic. The runner does NOT
restore the URL itself and does not change the task's result: whether the new
URL is wrong is for a person to decide. A project with no readable origin is
not checked (a DEBUG line says so).

No update is needed: the prompt and the check are both in the binary.

An expired trial stops the runner with one line instead of spinning (#12953, #12967)

mcptask.online now refuses an account whose trial has ended with no payment
method on file — HTTP 402 on the REST API, trial_expired: true on every MCP
tool (#12948). Until now the
runner read that as a passing hiccup: the queue check fell through to triage,
triage was refused too, a bug report was attempted and refused, and the wait
cycle started it all again.

Now the run ends at the first refusal it meets — the profile read at startup,
the queue check, or the quota poll — with one line on stdout and in the run log
naming the account, the server's own message and the payment link it contains,
and a non-zero exit:

❌ [WorkLoop] mcptask.online refuses account "kamr" — the trial has expired: "The trial has ended, please order a subscription: https://mcptask.online/kamr/standalone_payment" (payment page: https://mcptask.online/kamr/standalone_payment) — stopping the run: no triage, no retry, and no bug report, because the report would be refused the same way

No bug piece is filed for it, on purpose: the reporter uses the same token and
would be refused the same way. A scheduled host therefore shows a failed job in
its launcher log each weekday morning until the account is paid for; the next
start after payment runs normally. Nothing to do on upgrade — a bare update is
not needed.

A fan-out of subagents over one file is no longer killed as a loop (#12966)

The watchdog's stall detector counted every tool call on the stream into one
window, including the calls of subagents the session had forked. A skill that
sends six persona subagents over one shared file — each reading it once, by
design — produced six identical Reads, and the fourth ended the run:

Stall detected: reason=loop_signature signature=Read:/tmp/zuboklik-persona-target-wizard.md:: count=4

The run was filed as stalled_for_genius, the piece was re-run on genius,
stalled again and was set aside for the day. Zuboklik
#12842 spent four days that way.

From this tag the detector keeps its window, its failed-edit streak and its
Bash failure records per agent, reading who made each call off Claude Code's
parent_tool_use_id. One agent repeating itself four times is still a loop,
exactly as before; six agents each doing a thing once are not. A successful
edit still counts as progress for the session that forked the subagent, so a
root that delegates a fix and re-runs its failing tests is not killed either —
but never for a sibling subagent, whose loop is its own.

What you see in the run log. A verdict inside a subagent names it:
Stall detected: reason=loop_signature signature=Read:/x.md:: count=4
detail=agent=toolu_… — terminating for genius escalation
, and a Bash loop
reads detail=exit=1 agent=toolu_…. A verdict in the session you launched
reads exactly as it did.

The context-cost report had the same fault. After an overflow the fresh
session was told you read X in full 4 times about Reads its subagents had
made in their own context. It now counts only the session that ran out of
room, and the [context_cost] summary line says how much it left out when
there was anything: …, 0 failed edits; 6 subagent calls not counted (their
context is their own)
. A run that forked no subagent logs the line it always
did.

Codex and OpenCode are unchanged. Codex's exec --json items name no
thread and OpenCode prints no child session's calls at all, so on those two
every call the stream shows is counted as the root's, as before.

No update is needed: the change is in the binary.

ollama launch works again, and a CLI that cannot start says so (#12965)

Hosts whose launcher.command is [ollama, launch, claude] (or codex,
opencode) stopped working when Ollama 0.32 began taking a CLI's arguments only
after --: every launch died in a tenth of a second with Error: unknown
shorthand flag: 'p' in -p
, and the runner reported it as the model forgetting
its result marker — two --continue retries, then a bug piece titled
stream_ended.

  • The runner now composes ollama launch <cli> --model <model> -y -- <the CLI's arguments>. Nothing to change in the config: [ollama, launch, claude] is still the form to write. A launcher.command that itself carries --model, -- or another Ollama flag is refused before the run starts, with the reason.
  • Any CLI that exits non-zero before its first stream line now ends the run at once as launch_failed, with the exit status and the child's own stderr in the log, on the card and in the bug piece. No marker retry, no --continue.

Upgrading: no update is needed for this — nothing the installer writes
changed; the fix is the binary. Nor any config change: [ollama, launch, claude]
stays as it is. On an Ollama-launched host that has been failing since Ollama
0.32, install the new binary and restart the runner's LaunchAgent (an idle
runner keeps the binary it started with).


Changelog

Other

  • 877c4581f67ad907d299cdcaaf56224dbf04d356 [#12903] A refused push ends the task instead of being repaired by rewriting origin
  • f91a67c253643cf41d988cb575fc6b1eab856512 [#12905] The adoption marker names the version it installed, so a second bundle update in one day is taken up
  • 5623aaaaaa14668fc4ae8bea34abb89bc1ea0ae8 [#12953] An expired trial ends the run with one line, not a retried failure
  • b2e10d0883b56de3f5253ef1332b7b8c7131b5fd [#12965] A CLI that cannot start is a launch failure, and ollama launch gets the command line Ollama 0.32 accepts
  • 22435d77c05f3a8d5780d7c7d7362f2b13713a34 [#12966] The context-cost report counts only the session that ran out of room
  • c115499c802a2c2f5d924c29a56d8199541a4be4 [#12966] The stall detector judges each subagent on its own calls
  • 41f2b23b5b60b92ab0a40c4ae9d784dd86581b74 bin/release: the upgrading paragraph reports the asset diff, it no longer concludes a host has nothing else to do [skip ci]
  • b08156576375b0edbd75d758237b23d5d7cbf678 docs(release): v0.3.31 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.31

Upgrading

Upgrading is two commands this once. mcptask_runner update --self replaces the binary. Then a bare mcptask_runner update once in every project on the host (it also stamps the machine's helper home) — not because the bundled .claude/ assets changed (they did not), but because this tag starts recording which version installed them, and until that record exists every run start prints a WARN and leaves the files alone; see "The installed skills and helpers follow the binary by themselves" below. From then on installing a new binary is enough. A runner idling in a wait keeps the old binary until its next task or a restart (launchctl kill SIGTERM, wait for the job to stop, then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.31 carries this binary; no wrapper changes.


The installed skills and helpers follow the binary by themselves

Until now the skills in a project's .claude/skills, the helper scripts in
~/.claude/bin, the permission baseline and the test-commands file were
written by mcptask_runner init and refreshed by mcptask_runner update, and
by nothing else. A fix that lives in one of those files rather than in Go code
reached no host until somebody typed update there — task
#11669 corrected three helper
scripts on 2026-08-21 and every host kept the broken ones
(#11673).

From this tag every run starts by comparing the version that installed the
files with the version running, and refreshes them from the running binary
when the binary is the newer of the two. It is a string comparison against a
record on disk; nothing is fetched. The refresh is update without --force:
a skill or helper you edited by hand is left exactly as it is and named in the
log, never overwritten. The process that takes over between tasks after a
binary swap passes through the same start, so a binary installed at 10:00 has
its assets in place before its first task too.

What you see in the run log. One line per record, each naming the project:

  • [Reconcile] <project>: the skills in <project> were installed from 0.3.10 and this process is 0.3.11 — refreshed from it: 2 updated (ci-runner, test-runner), 1 left as edited by hand (discover), 9 already up to date — on stdout and in the file; the updater's per-file table is in the file at DEBUG, and a skill it left alone gets the usual conflict — … locally modified, skipped line at WARN. The helpers get the same line for ~/.claude/bin.
  • Nothing above DEBUG when the record already names the running version, which is every start after the first.
  • no skills manifest at … — this project was not set up by mcptask_runner init at INFO on a scratch clone run by hand; nothing is written there.
  • … were installed from 0.3.12, which is NEWER than the 0.3.10 this process is running — left exactly as they are at WARN. A host can legitimately be ahead of its binary — assets installed from a checkout, a hand-updated host where an older per-machine binary still runs — and the runner refuses to move them backwards. The binary is what is behind; install the newer one.
  • … records no runner version, so whether the skills … are older or newer than this 0.3.11 cannot be told — left as they are; run mcptask_runner update in this project once at WARN. This is every existing host at its first start on this tag: the record that makes the check possible did not exist before it, and an absent version is not read as "older than everything" — that guess is exactly what would put a corrected helper back to its broken version on a host whose binary is behind its files.
  • A refresh that could not be carried out (a directory that cannot be written) is an ERROR line, a bug piece under assets_reconcile_failed, and the run goes on with the files as they are.

What you have to do. Once per project, and once for the machine's helper
home: mcptask_runner update, from the binary you want the files to follow.
It records its version in .claude/skills/.mcptask_runner_manifest.json and
~/.claude/bin/.mcptask_runner_helpers.json, and from then on installing a
new binary is enough — update is no longer needed for an asset fix to land.
The order still matters for the typed command: release the binary, then
update. update from an older binary over newer files now says so before it
runs (… the files are going BACKWARDS), and runs anyway, because you asked.

Run logs and bug-report attachments no longer carry escape codes

The stream formatter used to wrap the task announcement in bold whenever the
emoji icons were on, whatever it was writing to — so the per-run log file got
\x1b[1m in it, and so did the copy bug-report attaches to a piece, which is
the one artefact meant to be read by somebody who was not there. Thirteen of
sixteen logs on the host that found this had escape bytes in them.

Colour is now decided per destination, by the same two variables as everywhere
else in the runner: MCPTASK_RUNNER_COLOR first (it wins both ways), then
NO_COLOR, then whether the writer is a terminal. The log file is never asked
and is never painted; stdout is painted only when it is a terminal that renders
it.

MCPTASK_RUNNER_ASCII is now about the icon set and nothing else. It used to
strip the bold as a side effect, which meant an operator who wanted clean log
files had to give up the emoji to get them — that is what the colour variables
are for now, and MCPTASK_RUNNER_ASCII=1 at a terminal keeps the emphasis it
used to lose.

Nothing to do on upgrade. Logs written before it still contain the escapes;
bug-report attachments from this version on do not.

The release is one script

bin/release X.Y.Z in this repository does the two-repository release —
gem CHANGELOG and version bump, both tags, the Go and gem Release workflows
watched, this file's contents put above what goreleaser wrote, this file
emptied, the six platform gems verified on rubygems — and refuses, by name,
before anything irreversible: a dirty or off-main checkout, a main that
differs from origin/main in either repository, a tag that already exists, a
version that does not go up, an empty notes file, a wrapper change with no
CHANGELOG line, or gh not signed in. bin/release X.Y.Z --dry-run runs
every check and prints every step it would take. Nothing changes for a host
that runs the binary; this is for whoever cuts the tags (task #12464).


A child that forks subagents in the foreground is no longer killed at five minutes

Zuboklik task #12842 asks for a persona review, which the project skill runs
as six subagents in one message with the parent silent until they answer. Four
runs on 2026-09-21/22 each read the inputs, logged 20 %, started the review
and died 5 minutes later with no commit and nothing on the piece; sibling
tasks went through fine, so to the owner it looked like the runners kept
skipping that one (bug #12900).

The profile knew the subagent tool as Task, the name Claude Code used to give
it, and gave that kind the long-running ceiling. Claude Code 2.1.x calls the
tool Agent, which no profile line named, so the watchdog filed it under
"other" and applied the quick ceiling: a kill at 300 s once the stream had
been quiet for 180 s, a failure verdict, the day's set-aside list, and a
resume that died the same way. Agent is now a subagent in the bundled claude
profile, beside Task, which stays for the CLIs that still use it.

One command on a host: mcptask_runner update --self. A host that overrides
the claude profile with a file of its own has to add the line itself —
Agent: {kind: subagent, name: Agent} under tools:.


A transient 500 from the API no longer ends the day

Bug #12896: on 2026-09-22 two
runners on the claude harness hit an HTTP 500 during triage. The CLI ended its
turn on the terminal event that a refusal arrives on, with a sentence that says
the condition is temporary; the runner filed it as a refusal it had no name for,
returned status error, and the quota Decider stopped the whole working day on
both machines over one 918ms request.

What an operator sees now:

  • API Error: 500, 502, 503 and 504 on Claude Code's terminal event are read as an overload and ride the same budget a 529 does — ten waits, doubling from 60 s to 600 s — instead of ending the run. The 500 was recorded from the incident; the other three are the same CLI template and were not.
  • The log says so in the CLI's own words. At the moment the line arrives: The API answered with a server-side error — the CLI said: "API Error: 500 …"; not a refusal: if this attempt ends without an answer it is retried on the overload budget (a WARN, not an ERROR). At the retry: API server-side error - retry 1/10, waiting 60s before next attempt... — the CLI said: "…", and the dashboard card says the same while it waits. A bare 529 keeps its old line, API overloaded (529) - retry 1/10 …, byte for byte.
  • On codex and opencode every overload already arrives as a sentence, so their retry lines now quote it too (… — the CLI said: "Selected model is at capacity…") rather than naming a 529 the CLI never sent.
  • If the ten waits run out, the run ends as before with status error, and the message carries the CLI's sentence rather than "API overloaded (529)".
  • A refusal the profile has no words for (a 403, for instance) still stops the run and the day, with no named remedy in the log: that is the cue to record the words in the profile, under overload if the CLI calls them temporary.

Nothing to do on upgrade. The profile change ships inside the binary. A host
that overrides the claude profile with its own copy — config/harnesses/claude.yml
in the project, or ~/.mcptask/harnesses/claude.yml — keeps running its copy
and has to add the four literals under failures.overload itself.

A running runner takes up a newly installed binary between tasks

Until now a binary installed while the runner was working was not picked up
until the process ended: in daily at the day's end, in today_auto_squash
at tomorrow's scheduled start. From this tag every mode that works more than
one task — today, today_auto_squash, daily, the queue and story modes —
checks at the boundary between two tasks whether the installed binary
changed, and if it did, swaps itself onto it there
(#11867). The check is two
stat calls and a read of Gemfile.lock; nothing is fetched from the
network, ever.

The channel does not matter: homebrew, mcptask_runner update --self,
curl | sh, another project's run rewriting ~/.mcptask/bin/mcptask_runner,
or bundle update mcptask-rails-runner — the bundle is adopted into
~/.mcptask/bin at the same boundary, which is what makes the process stale.

What you see. The runner's card on the dashboard blinks: it closes with
the message Updating from <old> to <new> and a new card opens under a fresh
session id a moment later. The run log says the same, with the number of
results carried over and where they were written. The new process continues
the same day — the tasks worked so far, the moment the day began — so the
quota check and the stop rules see the whole day, not a second one starting
at 11:00. The skip list and the urgent pin were already on disk and survive as
before. The day-end swap in daily stays as it was, in addition.

What you have to do. Nothing beyond installing the new version. Do not
restart the runner by hand; the running process swaps itself at the next task
boundary, never while a child is alive. Before handing over it runs
<new binary> version once; a binary that cannot say its version is refused
with a line in the log, and the runner stays on the one it has. A handover
whose exec fails is likewise logged and left for the day's end to try again.

State file. The day's state travels through
tmp/mcptask_runner/handover.json under the project, written a moment before
the exec and consumed once by the process that takes over. A file left there
by a run that did not continue is reported in the log and removed at the next
start, never resumed.

The story loop asks the server which subtask is next

Triage used to be handed a whole Story and a rule in prose — first subtask not
finished, not blocked, not on the skip list — and re-derived a database query
from a list that carried no priority, no task type and no blocker. On
jchsoft/Zuboklik on 2026-09-15 it walked that list in creation order, looped on
a medium task while an urgent bug in the same Story waited, and handed over a
blocked subtask that cost a whole session of the strongest model to discover
(task #12568).

The runner now asks mcptask.online's story-scoped queue — GET
/api/:account/pieces/:story/next
, task #12567 — before any triage session
runs, with today's skip list passed as exclude_relative_ids. The server
orders the Story the way it orders the queue (priority tier, bug or complaint
first, oldest first), never offers a blocked or finished subtask, and triage
receives one concrete subtask and chooses only the model. The subtask scan is
gone from every prompt that met a Story: the story-locked triage, the Story
branch of the discovery triage, and the dry run.

The three empty answers are three exits. Nothing left ends the story loop as a
finished Story. Every remaining subtask blocked ends it with no_more_tasks
naming the blocked subtasks — in the verdict, in the log and in the card's
message — rather than spinning on a Story it can take nothing from. Nothing
beyond the skip list ends it the way the skip list always has: the loop waits,
as on an empty queue.

Requires mcptask.online with the story-scoped @next (task #12567, deployed
2026-09-15).
Against an older server the REST call answers 404 and the loop
falls back to the same query through MCP: triage is handed the Story alone and
its prompt reads mcptask://pieces/{account}/{story}/@next?exclude_relative_ids=…
itself — on a server without that resource, a story loop cannot pick a subtask
at all and answers no_more_tasks with the server's refusal in its message.
Hosts with a .claude/ installed by an earlier release need no update: the
change is entirely in the binary and its prompts.

GetNextStorySubtaskTool now counts as fetch evidence in every bundled
harness profile, so a triage child that reached the Story's queue through the
tool rather than the resource is not discarded as an unverified pick.


Changelog

Other

  • bd3b65e6f6eef0f68903a903041510348672b0cf [#11673] The installed skills and helpers are reconciled with the running binary at every run's start, by version, never backwards
  • e56cf6ed50f6b870023a951d1cb9fb6f5abbc382 [#11780] the stream formatter paints a terminal and writes the log file plain
  • b0b6cc3fb415e10812d6dac48253707881a69a13 [#11867] The in-process stop cuts a session start and the closing-frame redial short
  • c57cd0a0b2ceaa592d335c4c0ec0d4ea40f20ed4 [#11867] The runner takes up a newly installed binary between tasks, not only at day end
  • ce2aceecdb4d290f9b1f096055fc536ec86698a0 [#12464] bin/release: the two-repository release as one script that refuses by name
  • 0d4c1f885553e4bf2004aa8d5449a468f2671ed2 [#12466] the conformance driver is built as mcptask_runner, and every scenario checks the child can find it
  • 9c20400e407555e830195ef674b42e5129041c13 [#12568] The story loop asks the server which subtask is next
  • 45953b16f287d6bd524b59786c53bc1b3880a300 [#12896] A transient 500 on the terminal event rides the overload budget instead of ending the day
  • 7d93df8475e9aff0dca905ba5ee10f1e13bf045f [#12896] Scenario 25 asserts the retry line in each dialect's own words
  • adf337334a4b44152054787ba20af6b203c81836 [#12900] the claude profile names Agent as the subagent tool, beside Task
  • eacefab166e2bd6f7300f133a406268a86c6ebda docs(release): v0.3.30 shipped, the next-tag notes start empty
  • 9e09a9be67ffe3fdfd83f920ab6dab75fd793a1c test(eventstream): the stop tests group their imports the way goimports wants

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.30

Upgrading takes two commands

mcptask_runner update --self replaces the binary. Then a bare
mcptask_runner update in every project on the host, because the bundled pr
skill changed — it now names the shape of the task link a pull request must
carry and tells the agent to copy the one the CREATE PULL REQUEST step prints,
which is the other half of the pr create check below. Run it while no runner
is working that checkout. A runner idling in a wait keeps the old binary until
its next task or a restart (launchctl kill SIGTERM, wait for the job to stop,
then kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.30 carries this binary; no wrapper
changes.


init can be told "no scheduled job"

The scheduled-job prompt at the end of init offered 1) today_auto_squash
and 2) today and nothing else. There was no way to say "I do not want the
runner started automatically": any other answer was invalid mode, and because
the schedule is the last step, that error failed an install whose every earlier
step had already succeeded (task #12790).

The prompt now has 3) none — no scheduled job, start the runner by hand. It
writes no LaunchAgent, systemd timer or Task Scheduler job, prints
Scheduling: skipped (none chosen …) and the install ends green. A
provisioning script says the same with --mode none. Nothing changes for a
host that already answered 1 or 2; pick it up with `mcptask_runner update

--self`.

The end of init can be read without reading every word

The tail of an install was forty [Installer] lines in one voice: paths
written, commands to run, and the sentences explaining both, all at the same
weight — and an operator who had just answered three prompts could not tell
which lines to do something about. On a terminal each line is now painted by
what it is. A path, a command and a value somebody set keep the terminal's own
colour; an explanation (why the hour matters, what the log retention is, what
stopping a run costs) is dimmed whole; and the one step to take now — "To
activate, run:" — is bold. The config summary follows the same rule: section
names bold, meanings and values still at their default dimmed, so what you
configured is the only thing on the page that is neither.

The words are unchanged, and so are the bytes a pipe, a log file or CI
receives: colour reaches a terminal and nowhere else, as before. NO_COLOR
and MCPTASK_RUNNER_COLOR still decide it outright.


A pull request without a task link is refused where no template would catch it

PR jchsoft/mcp4mail#48 went out with `Task: https://mcptask.online (jchsoft

12761)` in its description. mcptask.online reads a pull request back to its

task through that link, and its parser wants the whole path —
https://mcptask.online/<account>/tasks/<relative_id> — so it found no task,
answered the webhook with "Not body supported format!" and task #12761 was
never approved. The prompt said "mcptask.online task link" and left the shape
to the agent, and mcp4mail has no pull-request template, which is exactly where
the shape got invented (task #12865).

Two changes, one at each end. The CREATE PULL REQUEST step now prints the
task's link finished, in that shape, with nothing left to invent; when the run
picks its own task the account is filled in and the id is named as the one
from the fetch step. And on a project with no pull-request template,
mcptask_runner pr create refuses a description that carries no such link,
before the host is asked, naming the shape — the same way it has refused a
description that ignores the template since #12599. A project with a template
of its own is unchanged: the template says where the link goes.

Two commands on a host. mcptask_runner update --self puts the check in the
binary and the wording in the prompt; then a bare mcptask_runner update in
every project, because the bundled pr skill changed too — it now names the
link's shape and tells the agent to copy the one the CREATE PULL REQUEST step
prints. Run it while no runner is working that checkout. A project that was
opening pull requests without any task link will now see them refused — add
the link; that is the pull request mcptask.online could not follow anyway.


Changelog

Features

  • 94346af8cdc3284ec1f349489c76988645f5d8f2 feat(install): the end of init weighs each line by what it is ### Other
  • 589b7ce2b1f56c6f2698b299571f582b93347241 [#12790] Offer 3) none in the scheduled-job prompt, and --mode none on the CLI
  • daabf1e7e835b8e176e69a530da3c7c398bf42b8 [#12865] pr create refuses a description with no task link where no template would catch it
  • 9071ff135fa92d1cca4310a04b60b2a492ebf830 docs(release): the #12865 note names the bare update the changed pr skill needs
  • 8adb4fefdd6ddbf79e73e62100bf014de6d705cf docs(release): v0.3.29 shipped, the next-tag notes start empty
  • 1aa5fe3e30b0853f698b1f2c70fd469665034e62 test(install): the painted-fact check uses strings.Cut, not an Index slice

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.29

Upgrading takes two commands

mcptask_runner update --self replaces the binary. Then a bare
mcptask_runner update in every project on the host, because the bundled pr
skill changed — it now tells the agent to write the pull-request body under
the project template's own headings, which is the other half of the pr create
guard below. Run it while no runner is working that checkout. A runner idling
in a wait keeps the old binary until its next task or a restart
(launchctl kill SIGTERM + kickstart on macOS), so restart it.

The wrapper gem mcptask-rails-runner 0.3.29 carries this binary; no wrapper
changes.


A stalled task really does come back on the strongest model

A stall — the same call over and over, the same edit refused three times, the
same command failing the same way — kills the attempt and leaves the piece in
progress with the verdict stalled_for_genius. The name was a promise the daily
loop did not keep: triage in discovery is told never to call a pick a resume, so
@next handed the same piece back, triage graded it "smart" again, and one host
launched the same model on the same task several times in a row on 2026-09-16
(task #12613).

The loop now remembers what stalled today and on which tier, and decides the
escalation itself. A piece that stalled on a weaker tier runs on genius the next
time triage picks it, whatever triage recommends — the log says so in one line.
A piece that stalls on genius goes on the skip list for the rest of the day,
like a failed one, and gets its fresh look tomorrow. Nothing to configure; pick
it up with mcptask_runner update --self.


A blocked subtask no longer costs a session to discover

Triage walked a Story's subtask list excluding the two Czech labels for finished
work and nothing for blocked, so a subtask sitting at "Blokováno" passed it. The
loop then started the strongest model on the piece, which fetched it, read
is_blocked: true and stopped — a whole session to learn something the Story
payload triage had already read contained, while the work that was actually
available waited (task #12568, seen on jchsoft/Zuboklik on 2026-09-15).

Blocked subtasks are now excluded where the selection is made, in every prompt
that walks a Story: the triage prompts and the dry run. Where the server sends
the machine fields — is_blocked, task_state_code, priority_code,
task_type_code, blocked_by — the child is told to use them and to take the
list in the order the server gives it, which is what makes an urgent bug filed
after a blocked task reachable at all. A server that does not send them behaves
exactly as before: the state label and the progress are the whole test, and
nothing is inferred.

A Story whose remaining subtasks are all blocked now ends the story loop with
no_more_tasks and a message naming the pieces and their blockers, instead of
handing a blocked subtask over again. No new status, and nothing to configure —
pick it up with mcptask_runner update --self.


pr create refuses a description that ignores the project's template

About one pull request in four opened on projectoid_ii carried Claude Code's
own body — ## Summary, ## Test plan, the robot footer — instead of the
project's .github/pull_request_template.md: 9 of the last 40, on both
accounts and both models (task #12599, seen on
https://github.com/jchsoft/projectoid_ii/pull/2036). The CREATE PULL REQUEST
step names the template in one bullet, the agent never opens the file, and
mcptask_runner pr create forwarded whatever it was given.

The prompt has not grown a line. Instead pr create reads the template the
prompt names — GitHub's or GitLab's convention, or pr_template: path: —
and refuses a description missing any heading the template does not mark
(optional), before the host is asked. The refusal names the missing sections
and prints the template's skeleton, so the agent's next turn is the rewrite
with nothing else to look up. A project with no template on disk is checked
against nothing, and a declared path with no file behind it is said on stderr
rather than guessed around. Pick it up with mcptask_runner update --self.

Changelog

Other

  • 7989e4cc35ae390e09c18668e96d0543e90c3f8f [#12568] Skip a blocked subtask in triage, before it costs a session (#12)
  • d4e8fbc6d5427826ae23eb0fc080789a23fa7c57 [#12599] Refuse a pull-request description that ignores the project's template
  • 5d1baa4031bebbbf7809aa1e883c9b9603b8ac42 [#12613] Escalate a stalled task to genius runner-side, and set it aside when genius stalls
  • 7f952fb60a3a831633a0b9620d3b821ef7fcb300 docs(release): v0.3.28 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.28

Jobs with ignore_quota: true now wake up on an assignment

A runner started with ignore_quota never read its user profile, so it never
learned its own user id and ignored every assignment broadcast — the wait ran
its full length and the log said Assignment event ignored: this runner does
not know its own user id yet
on each one, while the card on mcptask.online
named the user the whole time (bug #12473). The profile is now read on every
start; ignore_quota only decides whether the working-hours flag is applied.

Nothing to configure. A running job picks the fix up at its next start: after
mcptask_runner update --self, restart an idle runner (launchctl kill SIGTERM
+ kickstart), or wait for the next scheduled run. The startup log of an
ignore_quota job now carries a [QuotaGuard] user id N read from the user
profile at startup
line; a job whose profile read fails three times goes ahead
without its id and says what that costs, instead of refusing to start over a
flag it would not have applied.

Changelog

Other

  • ee730bd334f95d1f5592e2f9abd98360d4b99382 [#12473] Read the user profile under ignore_quota too, so the stream knows whose assignments to act on
  • 6926a1de6b8ef860d4801211d0d0cb6acab146a7 docs(release): v0.3.27 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.27

Upgrading

Two commands: mcptask_runner update --self replaces the binary; then a bare
mcptask_runner update in any one project on the host, because the bundled
pr skill and the permission baseline both changed. Run it while no runner is
working that checkout. A GitLab project also needs glab installed and
glab auth login done on the host.

mcptask_runner is now on the child's PATH on every project

Binary only — no helper, skill or launcher file changes.

The prompt and the bundled pr skill both tell the coding CLI to run
mcptask_runner pr …, and nothing ever put that name where the child could
find it. Only Ruby projects got away with it: the mcptask-rails-runner gem
installs an executable called mcptask_runner into the bundle's bin directory,
which Bundler has already put on PATH. On a Go, Node, Python or plain
repository — and on any launchd run, whose generated launcher execs the binary
by absolute path and adds nothing to PATH — the first mcptask_runner pr
create
came back

Exit code 127: command not found: mcptask_runner

and the child, having no other way to open a merge request, typed glab mr
create
instead: the one thing the skill tells it never to reach for.

The runner now prepends the directory of its own running binary to the child's
PATH, so the mcptask_runner the prompt names is the binary driving the run
rather than an older copy installed somewhere else. Nothing changes on a Ruby
project: the bundle's shim and this binary are the same version, because the
gem's version is coupled to it.

If you run a binary you built yourself, name it mcptask_runner. The
directory goes on the child's PATH only when a file of that name is in it —
reaching for some other mcptask_runner elsewhere on the machine would be a
different build wearing this run's name — and a build named after the module
(mcptask_go_runner) gets one line in the run log saying exactly that:

[TaskAutoSquash] WARNING: mcptask_runner is NOT on the child's PATH. This runner is …

Releases and mcptask_runner init both install the binary under that name, so
this affects development builds only.

Adoption now works on an rbenv, asdf or mise host

Binary only — no helper, skill or launcher file changes.

Until now the runner looked for the gem its project bundles in GEM_HOME and
GEM_PATH, and nowhere else. rvm exports both; rbenv, asdf and mise export
neither — they put shims on PATH and initialise from ~/.zshrc, which the
bash -l launcher never reads. So on those hosts the search had no roots at
all, every task boundary printed

[Adopt] Gemfile.lock names 0.3.26 but its gem is not installed here — still running 0.3.24 (run `bundle install`)

and both halves of that were wrong: the gem was unpacked under
~/.rbenv/versions/<ruby>/lib/ruby/gems/<x.y.0>, and bundle install had
already run. The host stayed on its old binary indefinitely.

The search now keeps the environment roots first, then derives the project's
Ruby from .ruby-version (walking up) or the Gemfile's ruby line and probes
the documented layouts of rbenv, asdf, mise, rvm, chruby, frum and Homebrew,
~/.gem/ruby/<x.y.0> for user installs, and finally asks bundle show
mcptask-rails-runner
through that Ruby's bin — the only source that can name
a BUNDLE_PATH or vendor/bundle install. Nothing is taken on trust: a
candidate counts only once its interpreter reports the pinned RUBY_VERSION and
names its own Gem.dir, or — where no interpreter can be run — once the
directory is found to actually hold a gems/. Every adoption now logs which
source answered, and a failure lists the roots it searched. The two failures are
told apart as well: a root that exists without the version in it is a missing
bundle install and says so, while finding no gem directory at all names the
pinned Ruby and every candidate it turned down instead.

An affected host cannot adopt its way onto this fix, because the binary that
would do the adopting is the broken one. Once, in the project directory, in the
foreground:

bundle exec rake mcptask_runner:update     # or: mcptask_runner update --self

then a bare mcptask_runner update for the helpers. After that the host adopts
by itself again.


GitLab is a third git host

git_host: gitlab now drives merge requests, and mcptask_runner init writes
the key off a gitlab.com origin by itself. A self-managed instance is on the
customer's own hostname, so nothing in its remote says "gitlab" and nothing
guesses from one: such a project is installed with

mcptask_runner init --git-host gitlab

There is no credential to place. GitLab is driven through glab, which holds
its own login the way gh does, so init writes no token file and no
environment variable — it names the one command that makes the host work, and
glab auth login on that host is the whole of the setup. Nothing about Bitbucket
or GitHub changed: a project on either gets exactly the install it got before.

Two things follow that an operator sees. A GitLab project's
.claude/settings.local.json picks up Bash(glab:*) and
WebFetch(domain:gitlab.com) on the next mcptask_runner update — approvals no
other host is handed, the same way gh belongs to a github project alone. And
the pull-request template a GitLab project is told to follow is
.gitlab/merge_request_templates/Default.md, which is what GitLab's own web UI
puts into a new merge request; pr_template: path: still overrides it, and a
Bitbucket project still gets no template line at all, because Bitbucket Cloud has
no such convention to name.

A self-managed instance needs no address here. glab auth login --hostname
gitlab.example.com
once per machine, git_host: gitlab in the project, and that
is the setup: every question the runner asks is a glab run in the project's own
checkout, and glab resolves the host from that checkout's remote against the
credentials it stored. The runner keeps no URL of its own — a second place for the
address to be written is the first one to go stale. Hostname derivation stays
exact, so gitlab.com derives the key by itself and a customer's own domain does
not: a substring match on "gitlab" would point glab at somebody else's service.

mcptask_runner pr checks prints ONE row on GitLab, and that row is the whole
verdict rather than a summary of others. The host has no per-check object: a merge
request has one pipeline, and the pipeline's single status is what GitLab
publishes and gates merging on, so the report carries the head pipeline as one
check named after it. success is the only pass; failed, canceled and
canceling are red; everything else — skipped and manual included — is
IN_PROGRESS and is waited on, because a pipeline that never ran is not a green
one. A project with no .gitlab-ci.yml answers NONE, which is not a pass either.

Four conformance scenarios pin the behaviour (21–24: a manual merge request, a
green pipeline merged and verified, a red pipeline twice, and a merge the host
would not confirm). They drive a fake glab first on the PATH and assert its argv
word for word — including the --auto-merge=false that stops glab mr merge from
scheduling itself, exiting zero, and letting GitLab merge the branch ten minutes
later with nobody watching.

A failing gh or glab now says what it said

Binary only — no helper, skill or launcher file changes.

Until now, a refusal from either CLI reached the operator and the run log as its
exit status and nothing else:

Error: glab mr create --title x --description y --draft --yes: exit status 1

What glab actually said was a 409 — this branch already has a merge request —
and it was never lost on the wire: exec.Cmd.Output() captured the child's
stderr into ExitError.Stderr, and the wrap that named the command kept only
ExitError.Error(), which is the string "exit status 1". Every reason a merge
can be refused on those two hosts arrives that way: a protected branch, an
approval rule not satisfied, a merge train, an expired login. The runner reported
all of them identically.

It now repeats what the CLI said, in the form the Bitbucket adapter already used
for the host's own error envelope:

Error: glab mr create --title x --description y --draft --yes: exit status 1: ERROR Post https://gitlab.com/api/v4/projects/1/merge_requests: 409 {message: [Another open merge request already exists for this source branch: !1]}

One line: a two-line ERROR block is joined rather than allowed to wrap the run
log around it. stderr is quoted when there is any, and the first line of stdout
when there is not, because glab prints some refusals there. A failure with no
child behind it — a timeout, a glab that is not installed on this host — gets
nothing appended, since it has nothing to quote and its own text already names
what happened. Nothing about exit statuses or control flow changed: gh pr
checks
still exits non-zero on a red CI and that is still read as a verdict.


A machine's own pull-request skill no longer outranks the runner

Binary and skills — a host that installs this release should re-run
mcptask_runner update so the bundled pr skill and the permission baseline
land beside the new binary.

The first real GitLab run found the gap the mcptask_runner pr work of the
previous release had left open. The step said "CREATE PULL REQUEST", named the
command, and forbade "a host-specific CLI" — and the child used neither. It
reached for a skill installed machine-wide called pull-request-creator, whose
description offers to create a pull request "following WorkVector conventions",
which reads as the house style rather than as somebody else's host. That skill
drives GitHub's CLI, so on a GitLab project it produced

gh pr create …
→ none of the git remotes configured for this repository point to a known GitHub host

before improvising its way to glab mr create. The merge request was opened by
luck: on a project whose host has no CLI at all, the same route produces two
failures and nothing else.

Nothing was violated in the child's reading of the prompt, which is the point.
The step's own title is word for word what that skill advertises, and the step
never said the WORDS belonged to the command. It does now: "create a pull
request", "create a PR" and "open a merge request" are declared to mean
mcptask_runner pr create, and anything else offering to answer them — another
skill, another script, a command you know — is declared wrong, because each is
written for one git host and this project may not be on it. The bundled pr
skill leads with the same phrases and says outright that it outranks any other
pull-request skill on the machine. Neither names a host or a CLI; the neutrality
guards would refuse it, and the runner cannot know what is installed anyway.

One standing approval is gone from the permission baseline:
Skill(pull-request-creator). The installer was handing every project it
touched an approval for the very skill the prompts forbid. Removing it is all
this file can honestly do — the merge is a union, so a project that already has
the entry keeps it, and an implement run is launched in a permission mode that
skips these checks. What decides it is the prompt and the skill description.

A published pull request now reaches the run record even when the command was
not used.
The same run recorded {"status": "success"} with no pr_number,
while the runner had printed the merge request's number on the card and in the
log: only the mcptask_runner pr envelope ever stored one, and the child had
never run it. The CLI's own published-change event is a second source now,
below the envelope — which still wins, whichever arrives first — and strictly
read: an identifier that is not a positive integer is dropped rather than turned
into a number. Downstream of that are the merge verification, the effort line
and the card's closing message.

Changelog

Other

  • a10dab6b5fcf0b38a0b420c00b406b90f332a419 [#12418] Give the dry prompt the harness's fetch block
  • 9fba8128f0c481d15f67200de2e48d012827d85c [#12419] A close code is never destroyed, only misfiled, and the chaos test cannot prove it either way
  • 791cce8861066bb03f70ec4efc968f732130c99b [#12435] Drop a GEM_HOME that names no directory, and say so
  • 9aafbe1dd50684dc1e04f225f4b1aa67ccc28832 [#12435] Find the bundled gem on an rbenv, asdf or mise host
  • f95cc32f54aa4edc41a997118551525a00f54fc4 [#12435] Keep the gem search inside the roots the caller named, and off Windows's ruby.exe mismatch
  • fc4affd7ee63ec88e93d7bbdea32ca79b2b10e53 [#12437] Keep the host's test lock out of reach of tests and the stress day
  • 3cf39989ab8addfc1bbe3291c57d7b7412070819 [#12437] Put the conformance sandboxes outside every git repository
  • 54b60f2e96ea3710ccafb3f702f4571f62e8064b [#12449] A gitlab adapter that drives glab, and the places glab is not gh
  • 2abf0832e6337cf3ca95ef07ac71ea07f655241d [#12450] Everything that names a host by name learns the third one
  • d5999c68335e64dd9bf47ee21bb3ca49b2bbaf44 [#12451] GitLab proved: four scenarios against a fake glab, and two runs against the real one
  • 3e5c44a3adb7872fc6204d82d5ef81889e9053d3 [#12454] A refused gh or glab command repeats what the CLI said
  • 7095704f11f52dcda65147c623fead2d76004b3a [#12455] Claim the words "create a pull request", and read the number off the stream
  • 137eda5f6711e269680576f9dc03d7f6a956e536 [#12456] Put the runner's own binary on the child's PATH, under the name the prompt uses
  • c2a68e166540891078eedbb0323db2020b0c3a7f docs(release): v0.3.26 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.26

Upgrading

mcptask_runner update --self replaces the binary; then run a bare mcptask_runner update in any one project on the host, while no runner is working that checkout — the bundled pr skill gained the review commands and the binary does not rewrite installed skills on its own.

A close the server explains is logged with its explanation

When mcptask.online ends the dashboard socket with a close code — going away
during a deploy, a policy violation, an internal error — the runner's log now
says so, instead of sometimes reporting only use of closed network connection.

A dying socket is noticed twice: by the read loop, which has the close code, and
by whatever snapshot write was in flight, which has only the symptom. The
reconnect notice could be composed before the read loop was scheduled, and the
close code — the half that says who ended the connection — was then dropped on
the floor. It is now printed on a line of its own when it arrives that late:

[EventStream] WebSocket closed (the snapshot write failed: use of closed network connection), attempting reconnect...
[EventStream] the read loop's account of that close landed after the notice: server closed it, code 1001 going away

Nothing branches on the close code, so no behaviour changes: the runner redialled
before and redials now. What changes is that a run log attached to a bug report
says whether the server let the connection go on purpose — which decides whether
there is anything to go and look at on the server side at all.

What to do: mcptask_runner update --self. Binary only — no config change,
no reinstall of the .claude/ assets.


No prompt the runner composes names a git host any more

The review and reviews executors were the last two that did. Everything else
moved onto mcptask_runner pr in v0.3.24; these two could not follow, because
prhost.Host had no question for "what did the reviewers say on this pull
request" or "what is open on this repository", and a prompt cannot be pointed at
a command that cannot answer it. So they went on typing gh pr view, two
gh api repos/{owner}/{repo}/pulls/{n}/… reads and a gh pr list --state open —
four instructions that do nothing on a project hosted anywhere but GitHub.

The interface has both questions now, and two new subcommands expose them:

mcptask_runner pr list --open      # every open pull request, no search
mcptask_runner pr reviews 1234     # the verdicts, then the comments

reviews is one flat list across what each host splits differently — GitHub's
submitted reviews and inline comments, Bitbucket's participant verdicts and its
comments. A note carrying a state is a verdict (APPROVED,
CHANGES_REQUESTED, COMMENTED); one carrying a path and a line is inline
on the diff.

Two things it deliberately does not claim. bot: true is the host marking an
app, and Bitbucket Cloud marks none, so the field's absence is not a promise
that a person wrote the note — the pr skill says to read author there and to
say which you did. And there is no resolved flag on either host: whether a
review has been addressed is a judgement made by reading it, which is what these
prompts did when they called the REST endpoints directly, and a field only one
adapter could ever fill would be an invitation to trust a false the other host
never said.

The guard list in TestSpineNamesNoGitHost is now empty, which is the point of
keeping it: it was the record of what was left, and an empty one is a claim every
prompt is held to.

What to do: mcptask_runner update --self for the binary, then a bare
mcptask_runner update in each project — the bundled pr skill gained the two
commands, and a project still carrying the old copy will not know they exist.

Auto-squash asks the git host whether its checks passed, before merging

The auto-squash CI step used to have one verdict in it: the local bin/ci. A
green run merged, a red one left the pull request open, and the host — which is
the thing actually running the pipeline the project configured — was never
asked. A run on a project with no bin/ci merged on no verdict at all: it
skipped the local gate, merged immediately, and the host's pipeline reported
SUCCESS a few minutes later. It happened to be green. Nothing in the runner
would have noticed if it had not been.

The step now has two verdicts and the merge waits for both. Between the local
gate and the merge the child runs mcptask_runner pr checks <pr_number> and
reads .checks.state:

  • SUCCESS — merge.
  • FAILED — status ci_failed, nothing merged, the pull request stays open with the red check named.
  • IN_PROGRESS — wait two minutes and ask again, at most ten times. Still running after the last ask is ci_failed; an unfinished check is not a passed one.
  • NONE — merge. A check nobody reported is not a pass, but a project whose host runs no CI must still be able to auto-squash, and the local gate is what carries that case.

Which of the four decided the run is now a required field of the result marker,
status_detail, so a ci_failed in the effort trail says whether it was the
local suite, a red check on the host, or a build that never finished.

One text for every host, as the rest of the spine is: the command is
mcptask_runner pr checks on GitHub and on Bitbucket alike, and the adapter
behind it answers in the same four words.

What to do: mcptask_runner update --self. Binary only — no config change,
no reinstall of the .claude/ assets.

init --home-dir re-roots the Bitbucket credential, and a re-run says what it did with it

Two fixes to the step that places a Bitbucket credential during init
(#12424).

--home-dir re-roots everything an install writes outside the project: the
launcher script, the token directory, the shell rc file, the scheduled job. The
Bitbucket credential was the one file that escaped — it resolved the account's
own home and wrote there, so an install aimed at a scratch directory put a secret
into ~/.mcptask_env.d/bitbucket_credentials instead, where a second machine's
credential would then sit in a directory nothing sources. It now goes under the
same home the mcptask token does, out of one shared resolution rather than two.

The second half is what a re-run says. The credential step used to hang off the
path that newly wrote git_host:, so the two shapes of re-run that change
nothing — the key already in the file, or --git-host bitbucket repeating what
the file already says — printed git_host: bitbucket already set … and stopped.
That is exactly the run an operator types after making a token. It now runs
whenever the resolved host is bitbucket, and says which file it left alone:

[Installer] git_host: bitbucket already set in …/config/mcptask_runner.yml
[Installer] [Bitbucket] credentials already in ~/.mcptask_env.d/bitbucket_credentials — left alone. To replace them (a rotated token, or a different account), re-run init with BITBUCKET_ACCESS_TOKEN, or BITBUCKET_EMAIL and BITBUCKET_API_TOKEN, set.

The no-clobber rule is unchanged: an install handed nothing still writes nothing.
It just no longer does it silently.

What to do: mcptask_runner update --self. Binary only — no config change,
no reinstall of the .claude/ assets. A host installed with --home-dir that
has a Bitbucket credential is worth one look: the file may be in the account's
home rather than under that directory.

Run records say how the attempt ended

log/runs/run_*.json used to end like this on a run that had gone perfectly:

"status": "processing",
"termination": "result",
"final_status": "processing",
"result": { "status": "success", "pr_number": 3, … }

Two of those three lines were about a run that was still going. The record is
stamped on the way out of the streaming loop, one line before the session is
moved off processing, and it is never written again — so the field a reader
reaches for to tell how a run ended reported the state the run had been in while
it was working, and every clean run looked identical to a run that hung.

Both now answer the question they are named for. final_status carries the
child's own verdict where there is one — success, ci_failed,
merge_unverified, no_more_tasks — and the session's terminal status where the
child never answered, which is error on every kill path, exactly as it already
read there. status carries the terminal frame: finished on a clean ending,
error on a kill.

The record's shape is unchanged: same members, same order, no new field, nothing
for the dashboard to learn. Records written by older binaries keep their old
values; this only changes what is written from here on.

What to do: mcptask_runner update --self. Binary only — no config change,
no reinstall of the .claude/ assets.


Changelog

Fixes

  • 2e7ddea2857effacff589672145c8e02c6fdd236 fix(eventstream): a close code that arrives after the notice gets its own line (task #12421) ### Other
  • f3a9d1dbb050938efe31d3eb8f16dc47de2fad62 Ask mcptask_runner pr for reviews and for the open queue (task #12420)
  • 6ca5eb9ae3a8f8af8d35ae2baa90d637d7f53855 Ask the git host's checks before the auto-squash merge (task #12425)
  • b7595569ae0a929c8f48cb6fbd83d0d47980524d [#12424] Re-root the Bitbucket credential, and speak about it on a re-run
  • 47792ccc135245fb7b555b3a95b58d890543aeb1 [#12426] Make the run record say how the attempt ended
  • 6daeb941aeaa518cf1b819b35f918d7c2ab59ff7 docs(release): v0.3.25 shipped, the next-tag notes start empty

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.25

Upgrading

mcptask_runner update --self replaces the binary; then run a bare mcptask_runner update in any one project on the host, while no runner is working that checkout — this version changes the installed data pack (the new pr skill, per-host baseline permissions, ci_wait, the test-lock helpers), and the binary does not rewrite those on its own.

Bitbucket Cloud projects can be driven

git_host: bitbucket now selects a real adapter instead of an error. The runner
opens, finds, reads, merges and checks pull requests on Bitbucket Cloud through
its REST API 2.0 — there is no Bitbucket CLI to shell out to — and
mcptask_runner pr works there unchanged, because it goes through the same
interface gh does.

What to do on a Bitbucket host. Put a credential in the environment and
re-run mcptask_runner init, which writes it to
~/.mcptask_env.d/bitbucket_credentials (0600) beside the mcptask token, so it
reaches both your own shell and the scheduled job. One of:

  • BITBUCKET_ACCESS_TOKEN — a repository or workspace access token, sent as Bearer. The one to use for a machine.
  • BITBUCKET_EMAIL + BITBUCKET_API_TOKEN — the Atlassian account's email (not a Bitbucket username) and an API token, sent as Basic.

App passwords are not accepted. Atlassian stopped serving them on 9 June
2026; BITBUCKET_USERNAME / BITBUCKET_APP_PASSWORD are recognised only so that
init and the adapter can say so, rather than authenticating with something that
fails at merge time, after the work is done.

Setting both schemes at once is refused, naming both, and writes nothing — no
credential is ever picked on your behalf. A re-run of init with nothing set
leaves an existing credential file alone; setting the variables again is how a
rotated token is installed. Values are never printed, by init or by a 401.

What a GitHub host has to do. Nothing. gh still drives it, argv for argv.

One mapping worth knowing. Bitbucket's SUPERSEDED counts as DECLINED, not
MERGED: those commits may well be on the destination branch, but this pull
request was not merged, so an auto-squash run over one ends merge_unverified
and a person looks.

How far this is tested. The conformance suite now drives a Bitbucket project
in both modes against a scripted host, on all three coding CLIs: a manual run
that opens a pull request and puts it on the card, an auto-squash run whose merge
is confirmed by a GET on the pull request rather than by the agent's claim, a run
whose CI fails twice and leaves the pull request open, and the same recording as
that third one with the host answering OPEN — which comes out
merge_unverified. The last two are a pair on purpose: a runner that had stopped
verifying would pass the confirmed one and fail the other.

One environment variable comes with that, and is worth knowing about even though
almost nobody needs it. BITBUCKET_API_BASE_URL points the adapter at an API
root other than api.bitbucket.org — for a host behind a proxy that terminates
Atlassian's API on its own name, and for the conformance suite, which uses it to
reach a fake host on loopback. Unset means Atlassian's own root; a wrong value
fails at the first request, naming the URL it could not reach.

A once_dry run on a project with no assigned work no longer files a bug

once_dry used to have one ending in its prompt: success. So a dry run against
a project whose queue was empty had nothing to answer with — the child fetched
the next piece, was told "No available tasks or stories found", went looking for
one by hand, and then reported status: "error", which is a hard failure. The
runner filed a high-priority piece in its own Errors epic about a project that
was simply idle, and did so again every time that project was dry-run.

An empty queue is now the mode's own answer: the run ends no_more_tasks, the
log says No tasks available, the dry run had nothing to display, and nothing is
filed.

What to do. Nothing on the host — the change is in the prompt the runner
composes, so it takes effect with the new binary. Any pieces already filed under
Runner error: result — <project> for an idle project can be closed.

An API refusal is now named instead of retried

When the coding CLI ends its turn because the API refused it, the runner stops
and says which refusal it was. It no longer reads that as "the child forgot the
result marker", so the three --continue retries thirty seconds apart are gone —
none of them could ever have succeeded — and the ending is no longer filed as
stream_ended.

Four endings, and what each asks of you:

Termination What happened What to do model_unavailable the configured model is not being served — retired, removed, or spelled differently by this provider fix model: in config/mcptask_runner.yml; the message names the model and quotes the provider not_authenticated the CLI has no credentials sign the CLI in as the user the runner runs as on that host usage_limit the account is past a hard cap that clears on a date nothing, until that date — the message carries it verbatim api_error the API refused in words no harness profile recognises read the quoted text; if it is worth naming, add its phrase to the profile's usage_limit / model_unavailable / not_authenticated list

What changes for an operator. usage_limit files no bug piece — it is a
budget, like the day's quota and the usage window before it, and an account with a
monthly ceiling would otherwise file one piece a month. The day still ends,
because nothing can be spent until the cap clears. The other three do file a
piece, and it names the condition and the remedy rather than saying the stream
ended. The dashboard card and the run log's error_message now say the same
thing, where all three of the reports that prompted this carried null.

What does not change. A 429 still gets its eight patient waits and a 529 its
ten — a refusal the profile could name outranks them, an unnamed one does not — a
context overflow is still an overflow, and a turn that failed for a reason a
resume can fix still takes the marker retry it always had.

Recognition is Claude Code's wire shape for now, that being the only one these
conditions have been recorded on. A Codex or OpenCode host hitting the same wall
still spends the marker retries.

The lock guard no longer refuses commands that merely name a test run

check_test_lock — the PreToolUse hook that stops a second suite starting
beside somebody else's — decided what a test run was by looking for a recognised
invocation anywhere in the command text. A space counted as the start of a
command, and every word inside a quoted string has one in front of it, so a
command that only talked about a test run was refused while a suite was
running:

git commit -m 'fix: bin/ci now takes the lock'    BLOCKED, and it runs nothing
grep -q '=== bin/ci exit:' "$LOG"                 BLOCKED, and it is the line inside /ci-wait

The same anchor was wrong in the other direction, which is the half that
mattered: an invocation had to be preceded by whitespace or nothing at all, so
../other-worktree/bin/ci — how a story branch runs the gate — matched nothing
and a genuine second suite was waved straight through.

The command text is now read the way a shell reads it, into words, and only a
word in command position can be a test run: the first word of the line or of
a new command (after &&, ||, ;, |, &, a newline, (, `, $(),
optionally behind a wrapper that is not the command itself (sudo, env,
time, nohup, an option to one of those, a VAR=value assignment) or an
interpreter that introduces one (bash bin/ci, sh -c 'bin/ci', and
python -m pytest, which was not recognised before). Quotes and # comments
are tracked, and a path-bearing invocation counts. The
exemption for the lock's own helpers (test_lock, run_with_log) is decided
the same way, so a commit message naming one of them no longer exempts a real
test run sharing the line.

Blocking is unchanged for everything that actually runs a suite: bin/ci,
bin/ci --fast, cd worktree && bin/ci, and every command the project declares
in .claude/test-commands.json.

What to do. Run mcptask_runner update on each host — the hook is a helper
in ~/.claude/bin, not a file in any checkout, so a project that only pulls the
new binary keeps the old guard until the helpers are reinstalled. Nothing else
changes, and git commit -m with bin/ci in the message works again.

The machine-wide test lock: reinstall the helpers on every host

~/.claude/bin holds the copies of the helper scripts that actually run, so
none of the fixes below reaches a machine until its helpers are reinstalled —
mcptask_runner update on each host. A repository carrying the fix is not a
fixed machine, and the lock is the one piece of tooling where the difference is
invisible: a host still running the old test_lock keeps serialising its suites
exactly as badly as before, and says nothing about it.

Run the update while no suite is in progress on that host. The scripts are
replaced in place, and a run that is mid-flight through the old one keeps the
file it started with.

test_lock records COMMAND_PID and LOGFILE on Linux as well as macOS

set_command_pid and set_logfile edited the lockfile with sed -i and an
empty backup suffix as a separate argument, which is the BSD spelling: GNU sed
reads that empty string as the script, the real script as a filename, exits 4
and writes nothing. Both call sites discard the failure, so on Linux the lock
never gained either field and nothing said so.

Three consequences, all of them silent, all of them gone now: a lock whose
COMMAND_PID stayed empty was reapable while its suite was still running; ci_wait
could not tell a crashed run from a slow one, because the check for that reads
COMMAND_PID; and a LOCKED verdict could not name the other run's log, so
OTHER_LOG was always unknown on that platform.

Both fields are now written by filtering the key out of the file and renaming a
temporary copy over it — no in-place edit, nothing to spell two ways, and a
reader can no longer catch the file half-built. (task
https://mcptask.online/jchsoft/tasks/12402)

Every run through the helpers is ten seconds shorter

run_with_log started its stall watchdog with a command substitution, and a
backgrounded subshell inherits that substitution's stdout — so the line meant to
start a watchdog and move on blocked until the watchdog's next sleep returned.
Measured against a command that exits instantly: the command was gone at 0.03 s
and the Exit code: footer arrived at 10.8 s. Every run paid it, whatever the
command, because the wait, the tee teardown and the footer all queued behind
that line. The same measurement is now about a second.

/ci-wait and /test-wait answer sooner as a result, and so does anything
watching for the footer. ci_wait's 15-second footer grace is unchanged and now
pure headroom rather than a measured need — nothing pays it on a healthy run,
since it returns the moment the footer appears.

Two more corrections in the same file: the stall watchdog sized the log with
stat -f%z alone, which is BSD-only, so on Linux it read every log as zero
bytes and dumped thread traces into a run that was not stalled; and the watchdog
subshell's own stray output now lands in the log instead of on the launcher's
stderr.

run_with_log is also now maintained in this repository rather than vendored
from the retired Ruby gem, which is what made the fix possible here at all.
(tasks https://mcptask.online/jchsoft/tasks/12401,
https://mcptask.online/jchsoft/tasks/12402)

A lock is now held for as long as the run, and released only by that run

This is the behaviour change to read before updating a shared host. The lock did
not protect what it appeared to protect, in three independent ways, and all three
were silent while they happened.

A lock is stale when nothing it stands for is alive, and nothing else makes it
stale.
Age is no longer consulted at all. Before, age >= 900 was tested
first, ahead of any look at the command, so a suite that ran past fifteen
minutes lost its lock while it was working — and a full bin/ci runs to about
that mark. A lock whose command is alive now keeps it at any age; a lock whose
recorded command has died is over immediately.

An empty command pid is no longer a ten-second fuse. acquire used to leave
COMMAND_PID empty and fall through to a pidfile only run_with_log ever
writes, so a lock taken around anything else — a bare bin/ci, an operator
being a good citizen — was reapable from its tenth second, permanently, with no
warning to either side. acquire now records the process that asked for the
lock, and the sentinel's 900 s backs that up, so an unregistered lock lasts as
long as the shell that took it and then as long as the sentinel. The pidfile
branch is gone, and with it a path whose two sides derived the same filename from
different places.

A release now has to prove it is the run that acquired.
release_if_owner compared CALLER and PROJECT — a constant per skill and a
constant per worktree, both reused by a re-run after a rebase — so a run whose
own lock had already been reaped could reach its cleanup and release the NEXT
run's lock, two minutes into a suite that had done nothing wrong. acquire now
issues a token, prints it after ACQUIRED and records it; release_if_owner
<caller> [instance]
takes the token, or the command pid, or the holder pid, and
answers not-owner when the instance is not this lock's. run_with_log and
bin/ci pass theirs. A release with no instance still works, for callers that
predate the argument, and now says in as many words that it matched on the name
alone.

test_lock status says which case a lock is in. HELD_BY=command … (does
not age out), HELD_BY=holder … (no command pid recorded yet), HELD_BY=sentinel
…
with the time left, or STALE=yes. The first line still starts with LOCKED
or FREE, which is what /wait-unlock reads.

The PreToolUse guard, check_test_lock, was rewritten to the same rules in the
same commit. It had its own copy of the old ones, so on both counts above it
would have waved a test command straight into a running suite while test_lock
was still refusing to hand the lock over.

Two side effects worth knowing. run_with_log now records in the log whether the
lock learned its command pid — the line reads Lock: COMMAND_PID set to …, or
Lock: no lockfile — this run is NOT serialised against other suites on this
machine
, which is a legitimate state for a run nobody took a lock for and a
useful thing to be able to check afterwards. And the Windows flake in the
lock-guard tests around the ten-second cliff
(https://mcptask.online/jchsoft/tasks/11897) is gone with the cliff itself.
(task https://mcptask.online/jchsoft/tasks/11852)

Pull requests go through mcptask_runner pr, on whichever host the project is on

The prompts the runner composes used to say gh. Not always out loud, which was
the harder half: the auto-squash spine merged with gh pr merge --squash
--delete-branch
and gated its own success on gh pr view --json state, while
the CREATE PULL REQUEST step named no command at all — and a step that names no
command gets gh anyway, because that is what a model reaches for when nothing
says otherwise. Either way a project on Bitbucket Cloud was being told to run a
binary it has not got.

Every one of those is now mcptask_runner pr create|list|view|merge|checks,
which asks the adapter git_host: resolved for that project (task
https://mcptask.online/jchsoft/tasks/12362 put the adapters in). The command
prints one JSON object on stdout, and the workflow tells the child to read
.pull_request.number out of it rather than from its own recollection. A new
bundled skill, pr, carries the five commands and their output shapes.

The runner reads that JSON too. The pull request a run opened used to reach
the dashboard through an event only Claude Code emits, so a Codex or OpenCode run
opened one and the card said nothing; and its NUMBER reached the runner only when
the child put it in the result marker, which about three quarters of real markers
do not. Both now come out of the command's own output, which every harness hands
back the same way. A result with no pr_number takes the one the run was seen to
open — the child's own answer still wins when it gives one, and nothing is
invented when no pull request was opened at all.

Permissions are now per host. Bash(mcptask_runner pr:*) is in the baseline
every project gets. Bash(gh:*), Bash(gh pr checks:*), Bash(gh pr view:*)
and the three github.com WebFetch domains are installed only into a project
whose git host resolves to github; api.bitbucket.org and bitbucket.org only
into a Bitbucket one. A project whose host cannot be worked out gets the shared
list and nothing else — nothing falls back to another host's approvals. Existing
entries are never removed, so a project that already has Bash(gh:*) keeps it.

The pull-request template default is GitHub's alone.
.github/pull_request_template.md is GitHub's own convention, and telling a
Bitbucket project to follow it meant the agent looked, found nothing, and wrote
whatever description it liked, with nothing reporting that the instruction had
missed. A git_host: github project still gets that path and that line, byte for
byte. Any other project gets whatever it declares under pr_template: path: and,
declaring nothing, gets no template line in the prompt at all.

What to do. Run mcptask_runner update on each host: the pr skill is a new
file under the project's skills directory and the permission split is a merge
into .claude/settings.local.json, neither of which arrives with the binary
alone. Run it while no runner is working that checkout — the new files would
otherwise land inside somebody's task commit. Nothing else is required, and a
project already carrying git_host: needs no config change; one that does not
gets its host read off git remote get-url origin at run time, and
mcptask_runner init writes it down for good.

Still naming gh: the two review executors (review and reviews). They
read review threads and enumerate every open pull request on a project, and
internal/prhost models neither — its six questions are create, list-for-task,
view, by-branch, merge and checks. Rewording them to name no host would leave an
agent told what not to use and not told what to use instead, so they are excluded
by name, with the reason in words, in TestSpineNamesNoGitHost. Extending the
interface to cover them is its own piece.
(task https://mcptask.online/jchsoft/tasks/12364)

/ci-runner reads a green run that prints no summary block

A bin/ci that prints one line per step and no aggregation at all — no CI
SUMMARY
, no TOTAL RESULTS — matched neither shape ci_wait recognised, so
every green run on such a host ended in NO_SUMMARY_RECOGNISED plus twenty
lines of tail. That tail was not empty, which is what made it hard to spot: on
the run that filed this it happened to carry three of the nine step lines and
lost both of the ones a reader triages by, so the orchestrator reported a ragged
excerpt as if it were the run's summary and could give no step count or timings.

ci_wait now recognises a third shape, the per-step lines themselves:

✅ <step name> passed in 4.88s
❌ <step name> failed in 1m31.30s

They are emitted in order, ANSI-clean, in place of the fallback. A red run
carries both spellings — later steps keep running after one fails — and the
failure path already passed them through its tail filter untouched, so the same
shape now reads on both exits. NO_SUMMARY_RECOGNISED stays for a log that
matches none of the three, and now names all three in its message; it also no
longer spends its twenty lines on the log's own ==== rules, which the two
recognised branches had always dropped and it had not.

/ci-runner's skill body documents both shapes side by side instead of the
older one alone.

What to do. The fix is in ci_wait, which is installed onto the host rather
than compiled into the binary, so a new binary alone does not change it: run
mcptask_runner update in the project to refresh ~/.claude/bin/ci_wait and the
/ci-runner skill. (task https://mcptask.online/jchsoft/tasks/12413)


Changelog

Fixes

  • 1b6b96a7af616625a981712fcd6db8537a053382 fix(dry): an empty queue is an answer, not a bug piece (task #12407)
  • 1ac4ad2b9e8cadbecc497c93cf811781e469efd6 fix(eventstream): the throttle's clock is shared with the stream's own goroutines, so a test hands it over under a mutex (task #11688)
  • b463017bb9094c5304ae4315ff932e9a242a9034 fix(executor): the API refusing a run is named, not retried for a marker it will never give (tasks #12405, #12406, #11850)
  • a74dab9409d1133ef38a1a0a4fa2cbb388fd1e58 fix(helpers): a lock is held while its run lives, and given back only by that run (task #11852)
  • 584a71dd709ce416d09035aed18e4a88bf10d492 fix(helpers): the lock records its command pid and its logfile on either userland (task #12402)
  • 3b202a43d607cbc6957ba3faed2d39c1e3d61094 fix(helpers): the stall watchdog no longer holds the footer for a whole interval (task #12401)
  • 50b0446ede6def664d97428f61ce8ac379fa1d66 fix(stress): an orphan may end itself, and the sweep says which ones did (task #12403) ### Other
  • 4867ec85fccea0e4354333da18c647dd2b69ae3a Recognise an API refusal on every dialect, not just Claude Code's
  • 90353f488c2a233cbe78e862f466798b6e278ca3 Skip the 0600 check for the Bitbucket credential file on Windows
  • 0cb0ada0f2c2de1f328769c6bb44210d77f224b3 [#12362] Ask a PRHost, not gh: internal/prhost + mcptask_runner pr
  • e0a591aafc5e606230bddcebdde5d915b68456d3 [#12363] Drive Bitbucket Cloud: a prhost adapter over REST API 2.0
  • 6333beb769954a50eb88a6b3d22bab9c2d2fb6f8 [#12364] Say mcptask_runner pr, not gh: the prompt, a skill, and per-host approvals
  • 0d46b797c5cbf4e178324077b0c412c663eaf486 [#12365] Drive a Bitbucket project in conformance, and say so in the README
  • 1e4f48e9ef1deb729807cf4f4a679691ef1d3aa0 ci_wait: read a step record that comes with no summary block
  • e4418e46f7a757cd60f5264b387b74cecf8cce12 docs(release): v0.3.24 shipped, the next-tag notes start empty
  • 2425e84c9189223aafb21084742468d0f723f5f9 fix(check-test-lock): naming a test run is not running one (task #11674)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.24

Upgrading

Two commands this time: mcptask_runner update --self for the binary, then a bare mcptask_runner update in each host project — it writes the now-required harness: key and rewrites the helper scripts in ~/.claude/bin. Details below.

harness: is now required in config/mcptask_runner.yml

The runner drives a coding CLI named by a profile, and there is no default. A
project whose config does not name one refuses to start:

harness: is not set in config/mcptask_runner.yml, run `mcptask_runner init --cli <name>`; known profiles: claude, codex, opencode

What to do. Run mcptask_runner update on each host. A project this runner
already installed into — one with a .mcp.json and a
.claude/skills/.mcptask_runner_manifest.json — has harness: claude written
for it once, and the update says so in its output. That is the whole migration
and it changes nothing else about how the project runs.

A project update does not recognise (and any new project) needs
mcptask_runner init --cli <name> instead. init without --cli on a project
that has never named one is refused rather than assuming claude: assuming it is
how a host that installed a different CLI gets a Claude command line built for
it, fails at the first launch, and is told about a binary rather than about the
key nobody set.

A host that already sets launcher.command has said which CLI it starts, so
init and update now derive the key from it when exactly one profile claims
that name — printed, written into the file, never inferred at run time — and a
run whose launcher.command contradicts its harness: key is refused before it
starts, while a command whose program cannot be read (a shell, a wrapper script)
is logged and allowed; point launcher.command at a mock and you write
launcher.cli_check: false beside it, and a CLI that really runs under another
name goes in the profile's binary.aliases.

Where the harness shows up: the runner card on mcptask.online carries a harness
label beside the model (an additive snapshot member the server accepts since
2026-09-08; older runners keep working without it); mcptask_runner version and
the scheduled job's per-run log open with a Harness: <name> -> <binary> line
after the version line; a runner bug piece names the harness and attaches the
profile's own config files instead of the Claude Code settings only.

init --cli also adds config/mcptask_runner.yml to .gitignore if it is not
ignored already. The file is per-machine — harness, schedule times, model pins —
and two developers sharing one committed copy overwrite each other's answers.

Codex CLI support

mcptask_runner init --cli codex sets a project up for the Codex CLI instead of
Claude Code. The work loop, the watchdog, the dashboard card and the mcptask.online
protocol are unchanged; what differs is the command line, the words the prompt
uses for tools and skills, and what init writes:

  • .codex/config.toml in the project, with the mcptask.online MCP server marked required = true. The entry names the environment variable holding the bearer token (bearer_token_env_var) and never the token itself, so the file is safe to commit.
  • [projects."<path>"] trust_level = "trusted" in ~/.codex/config.toml (or under $CODEX_HOME), because Codex reads a project-level config only for a project it trusts.
  • Eight skills into .agents/skills, invoked as $test-runner, $ci-runner and so on. Helper scripts stay in ~/.claude/bin, shared with any claude project on the same machine.
  • AGENTS.md, pointing at the project's existing CLAUDE.md rather than copying it.

Both config writes append and never rewrite: an entry that already says what the
runner would have written is left alone, and one that says something else is
reported as a conflict and left as it is.

One thing to do by hand: codex login. The runner never logs a CLI in. It
checks, before a run starts, that this one is logged in already — codex login
status
has to exit 0 — and refuses with the probe's own output when it does not.
An unauthenticated headless Codex produces a failed turn with nothing done, for
as long as the loop runs.

Read the sandbox note before scheduling an unattended Codex run. The
implementing executors launch codex exec with
--dangerously-bypass-approvals-and-sandbox: full access, no approval prompts,
which is the parity with Claude Code's bypassPermissions this runner has always
used. Dry runs, triage and review get --sandbox read-only, plus one -c
override that approves the mcptask-online server's MCP tools and nothing else.
That second token is not optional: codex exec is non-interactive, which means
its approval policy is never, and never REFUSES an MCP tool call rather than
waving it through — without the override, every read-only-class run is cut off
from mcptask.online and cannot fetch the task it was started for. A write from a
read-only class is still refused by the sandbox. The checkout the runner is
pointed at is the only boundary there is.

How far Codex support has been proven. The conformance suite runs on
--dialect codex as its own CI job (twelve of fourteen scenarios; 06 and 07
describe a state Codex cannot reach and say so in the file), and the chaos stress
day runs on it too. Both drive a mock CLI.

Since 2026-09-09 the real codex binary has been driven as well, on codex-cli
0.153.4: a dry run, and a task worked end to end — branch, edit, go test,
bin/ci, pull request — with five progress logs left on the piece along the way.
Two faults that only a real CLI could show came out of that and are fixed here.
The approval-policy one is described above. The other is that continuations
were refused outright
: codex exec resume is a sub-subcommand with a smaller
option set than codex exec, taking neither -C nor --sandbox, so every retry
and every marker retry died with exit 2 and an argument-parser usage message. A
Codex host running an earlier build has never had a working retry. The flags now
precede resume, which is the order codex exec [OPTIONS] <COMMAND> documents.

Still unexercised: the rate-limit and context-overflow literals are source-derived
and no real run has reached either.

Claude Code hosts are unaffected by any of this beyond the harness: key.

Codex asks for its plan tool by name. codex-cli 0.153.4 does not offer
update_plan unless it is enabled, so every Codex command line now carries
-c tools.update_plan.enabled=true. Without it the child could log its progress
but never post a plan, and the runner card showed a run with no steps for its
whole duration. The prompt now also describes the tool the CLI actually has:
pending, in_progress and completed, exactly one step in progress at a
time, the whole list on every call.

Codex runs the test helpers directly, instead of a skill it cannot invoke.
A Codex run used to reach "UNIT TESTS", read $test-runner — codex exec does
load .agents/skills without being asked, so the skill really was there — and
stop. That skill is a state machine which drives itself through a sub-skill tool
Codex has not got, and its own rules forbid running the test command directly
because that would skip the machine-wide test lock, so the child was left with no
legal move. It refused, correctly, and every task_manual run on this harness
died at the same step. The prompt now names the lock-taking helper scripts the
installer already puts in ~/.claude/bin: test_start or ci_start to take the
lock and launch the run detached, then ci_wait polled for the exit code. Same
lock, one wrapper layer down — nothing about how runs are serialised on a shared
machine changes. Claude Code hosts keep invoking /test-runner and /ci-runner
and are unaffected.

A detached run now starts in a session of its own, and a fast one is no longer
called dead.
Two more defects, found by the first Codex run that actually
reached the helpers, both in the helpers themselves and both fixed for every
harness. test_start and ci_start used to detach with nohup and disown,
which survives the launching shell exiting but not its process group being
signalled — and a Codex child runs each command in a group of its own and
kills that group when the command returns, so the run died within seconds,
leaving no Exit code: footer and the lock held by a dead pid. Both scripts now
start the run in a new session: setsid where there is one, otherwise Ruby,
which this helper family already requires, and if neither is present they refuse
and say so rather than launching something that will be reaped. Separately,
ci_wait allowed two seconds between the command's process disappearing and the
footer arriving, where the real gap is about ten — the footer is written after
run_with_log has torn down its stall watchdog — so a sub-second command was
reported FAILED_EXTERNAL on a suite that then exited 0. That window is now 15
seconds, polled twice a second, and such a run is reported by its actual exit
code. The Codex vocabulary also spells the invocation as
test_start 600 cmd arg arg rather than test_start 600 <command>, and
test_start now refuses a whole command quoted into one argument by name
instead of guessing where to split it.

What to do. The same two-step upgrade the next section describes: these are
helper scripts in ~/.claude/bin rather than code in the binary, so
mcptask_runner update --self and then a bare mcptask_runner update in any one
project on the host.

bin/ci takes the test lock itself — reinstall the helpers

A bin/ci started by hand, or by a CLI child that found the script on disk,
used to run beside whoever was mid-suite; only the /ci-runner and
/test-runner skills took the machine-wide lock. The script now takes it
itself, and refuses out loud (exit 2, naming the holder) when somebody else has
it. The skills keep working through a marker: ci_start and test_start take
the lock first, as they always did, and export MCPTASK_TEST_LOCK_HELD so the
script knows the lock is its own.

What to do. The helpers live in ~/.claude/bin, not in the binary, so this
version needs the two-step upgrade: mcptask_runner update --self for the
binary, then a bare mcptask_runner update in any one project on the host,
which rewrites the helper scripts from the new binary. Until that second step
runs, a suite started through /ci-runner finds the lock held by the old
ci_start — which does not export the marker — and refuses itself as if
another agent held it. A hand-run bin/ci is fine either way.

Changelog

Features

  • c4eccf4f65a2cadec7e33fcff2896a27b63dec35 feat(bugreport): the bug piece names the harness and attaches the profile's config files (task #12300)
  • 9362d47e1beb2748c46af30e3c87d3f236d6dc86 feat(capture): capture mode records the child's raw stream for building new dialect fixtures (task #12287)
  • 037ee1d0bed0f907879ce6a9919586abeffaf28f feat(config): the version banner and the launcher log name the harness and its resolved binary (task #12301)
  • 48084db83a10dd77dfd9704861ab26d6f020940e feat(conformance): scenarios in normalized events, mock-cli --dialect with a renderer, chaos on the renderer (task #12286)
  • e804acbfb869f8c34e52a670166d04e7d640171a feat(conformance): the suite, mock-cli, chaos and the stress day run on the codex dialect (task #12291)
  • d23be25ec8018da875a9387ac27970667c2047d0 feat(harness): bundled codex.yml profile; command.order for subcommand CLIs; codex argv goldens (task #12288)
  • 044eea0e434dbc711436eeae289860aea6fb3e73 feat(harness): bundled opencode.yml profile; positional prompt with no token; readonly enforced via config (task #12295)
  • 81ad327cc9c069ee70d34aeb8b1f09f0e229147a feat(harness): codex failure literals from codex-rs 16ff14c, each cited by file and line (task #12293, part 1)
  • 20464a1d8ba009e00e04c6bc06b418bcd10d5197 feat(harness): codex prompt vocabulary: $test-runner/$ci-runner, MCP-tool piece fetch, update_plan block (task #12290)
  • 7fab6b61730ab89a47c14b363d3c7fba067df432 feat(harness): codex_jsonl decoder and renderer; resume by thread id from the captured session (task #12289)
  • ab5e95d01d01983505ecbcd435a7c22808483b89 feat(harness): every progress anchor in the spine spells the log tool through vocabulary.log_tool (task #12370)
  • 991f448a2e463139e69704d7e6d502a83c617c8d feat(harness): launcher.command is cross-checked against harness: at run time and derives the key at install time, never inferred at run time (task #12357)
  • 59bc636ee6aae849f258084538e03ffae4463b50 feat(harness): log lines name the harness display name; the engine delegates binary lookup to the profile (task #12285)
  • 7fcce96d85d09f100f68e1e1cc0b88af47b3047b feat(harness): opencode failure names from the OpenCode source, model tiers from the real CLI, real-run findings (task #12298)
  • 14a937f96d8b3b8dc1b9ad891ed0dd4ad4d98fdc feat(harness): opencode_jsonl decoder and renderer; result from marker and exit, never from step_finish (task #12296)
  • 438efa65eddc437eb5861090cbea78fc7a2689c8 feat(harness): profile schema, loader and the harness: config key with no default (task #12280)
  • 77962344c5bce58bdb178ae91d2958f45ae23a90 feat(harness): the claude_stream_json decoder emits normalized events; every consumer reads events and ToolKind (task #12282)
  • 8c883865a82d75827ac5ce69f437f10acf464419 feat(harness): the current-user fetch comes from vocabulary.current_user, not a literal resource line in the spine (task #12331)
  • 6af102c1f138dc30fc1cb5e44d01ace65c47262a feat(harness): the engine builds the child command from the claude profile (task #12281)
  • 6a989e2dd9db0c6b02d4808d53a76f8e50d56c5f feat(harness): the prompt spine reads its harness fragments from the profile vocabulary (task #12283)
  • 9c94208a607e941e7db261149cda8d11e2f7a19b feat(install): codex install section: .codex/config.toml MCP server, project trust, .agents/skills, AGENTS.md pointer, login preflight (task #12292)
  • 3182be07fffc1e83a7ddd102d4ab2900c7f3bd1b feat(install): opencode install section: opencode.json remote MCP writer, readonly overlay via OPENCODE_CONFIG_CONTENT, real vocabulary (task #12297)
  • 29452358747c03f1348feb98265499694706ce1a feat(install): the helper home is the profile's install.helpers_dir, rendered into skills and helpers at install time (task #12303)
  • 81ec85ca672e6e679e57e7b906ae27d7d30ddc9c feat(install): the installer reads its layout from the harness profile; init --cli, loud update migration, binary preflight (task #12284)
  • 35bb8f0110288dd68d07eee79c8bd36509086cb1 feat(snapshot): the additive harness member on the wire, in the run log and in the handoff note (task #12299) ### Fixes
  • 7d4bffd66f8eb1fe6b5d1efdaa9737e3a126dba7 fix(banner): the Models line names the profile whose ids it prints (task #12381)
  • 55df871aae8eff9c3bd5cdfbdfef45420125fd09 fix(ci): bin/ci takes the machine-wide test lock itself when not started through ci_start, and refuses when it is held (task #12369)
  • 298dc2022dcebf501c96f5a9d1b2fe003820eed1 fix(codex): ask for the plan tool, and describe the one this CLI has (task #12383)
  • b0f4dfa8113e57847d2f58d6fddccd4f0107a75c fix(codex): real runs on the ChatGPT plan — read-only class keeps its MCP server, continuations parse, refusals read from the wire (task #12293)
  • 7e42be2c3800bd4216fe7ac33d64d8912f9943b9 fix(codex): the prompt drives the lock's helper scripts, not a skill this CLI cannot invoke (task #12395)
  • 26a8adfc848505219155fcb529598dcdd78756dd fix(eventstream): a failed dial no longer prints the token it dialled with (task #12379)
  • 1b3a269360341a021ef2b557c95865e54bc96ce5 fix(executor): a run that answered with the marker is labelled result on every dialect, not stream_ended (task #12368)
  • 5815b2545a0afa62b046aef520afb23e56a48f56 fix(executor): the stderr label names the CLI that wrote the line (task #12380)
  • fc4178d720ed8b9483b063995268457491235523 fix(harness): codex exec takes --dangerously-bypass-approvals-and-sandbox or --sandbox read-only, never -a (task #12332)
  • f39750da54afbfb2935f773947d08a69c61a6780 fix(harness): opencode progress anchors name the callable log tool on every line (task #12367)
  • cb5a352cfb9f3ec3096e1baef06e1a90435b4071 fix(harness): the opencode readonly class keeps the shell and refuses writes by pattern (task #12353)
  • 77767eef92910fa8981498d2c5f13151ff21c496 fix(helpers): a detached run gets a session of its own, and a fast one is no longer called dead (tasks #12399, #12397)
  • 7c1d5abae8159d505e245705432c849dd68160ce fix(init): a scripted install is asked of the OS, not of the file mode (task #12382)
  • 5f4360b7d9c46d069093365af33c7be1895c558d fix(opencode): the plan comment stops claiming update_plan has no in_progress (task #12391)
  • 068f188f5dc5ff2adc3fa7bf0e5645dd9fd7390e fix(smoke): the banner checks ask a directory with no config (task #12384)
  • 381260a399407a1a859cf691eed5c8e0181d723b fix(tests): the init refusal test expects the host path, not the git spelling (story #12276) ### Other
  • d34e6c905f24482cf5c5f36c0e5184947023b4cc docs(ci): the codex conformance job comment counts its exclusions correctly (task #12294)
  • c3d811c3b2b7fe14b7c44f33dc9fa2aa5bcfbe45 docs(codex): the plan comment stops asking for a correction opencode.yml already carries (task #12391)
  • b679dd3947a85012cbc7064ab38528f6f62d9a74 docs(harness): README harness section, Codex host specifics, requirements, npm and CLI texts, release notes draft (task #12294)
  • c0583d6f795df1c4f39bc43b8815d490d86aa5ce docs(harness): the opencode skills comment reflects the self-locking local gate from task #12369 (story #12278)
  • 39ae784fb4cc661975f43cebceee59036c3f3c06 docs(release): the plan-tool flag, and the two-step upgrade the self-locking gate needs (story #12277)
  • 8ce7005b0b01d1df78928c5750d8cacf4d8c2e77 docs(release): where the harness shows up: card, version banner, launcher log, bug piece (story #12279)
  • 6dfba41f87808b2d5ec372b487b28cb74739676f merge: story #12276 — harness profile layer: Claude Code behind a data profile, no behaviour change
  • 0aa6545ba0faca27e80c14ccde4042d37f3b50fd merge: story #12277 follow-ups — token-free dial errors, per-harness stderr label, profile-true banner, scripted init, codex plan tool, smoke in a clean dir (tasks #12379-#12384)
  • b058a7d06faf042aaaed16c8d1d2d13ff52ef396 merge: story #12277 round 3 — plan-state comments agree, first real codex plan captured, release notes for the plan flag and the helper upgrade (tasks #12391, #12392)
  • bf250676dd324a0c6698be4bc00f10f0138a413b merge: story #12277 round 4 — codex drives the lock helpers, helpers survive their launcher, ci_wait waits for the footer (tasks #12395, #12399, #12397)
  • d5d8a86187669a6b0a2ab505b15be2173eae1c0b merge: story #12277 — Codex CLI harness (profile, decoder, vocabulary, install, conformance, docs; real-CLI runs pending login)
  • e2ee3141f3f18e6e9400388f27df42a934ee1c80 merge: story #12278 — OpenCode harness (profile, decoder, install, readonly by pattern, real runs verified; self-locking local gate)
  • 59e7ebd7a4066966e9425816201ff6add0225dae merge: story #12279 — harness visibility (snapshot member, bug piece, version banner, helpers_dir, launcher.command cross-check, log_tool vocabulary)
  • 650a002d65c1d9d3dbcc9fa223ef8ffd2bf01ad4 merge: task #12293 part 2 — codex driven for real (read-only MCP approval, resume argv, wire-shaped refusals)
  • 333b34da96f1ef8662e13bae51db95b91cfc0ced test(codex): the first real plan of a runner run, from the capture (task #12392)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.23

Upgrading

The usual one command replaces the binary:

mcptask_runner update --self

The bundled skills-and-helpers pack did not change in this release, so a
follow-up bare update is not required. Two behaviour changes worth knowing
before an unattended host picks this up:

  • --ignore-quota now means ignore, everywhere. The daily mode's day-level scheduler no longer queries mcptask.online about the budget when the flag is set, so a REST outage cannot park an unattended daily host until the next business day. The end-of-workday clock still applies — it is a time rule, not a quota rule (#11891).
  • The work window keeps its minutes. init --until 18:30 stores the new additive work_window.end_of_workday_minute key and the loop stops picking up new work at 18:30, not 18:00. Configs without the key mean :00, so existing hosts change nothing; a value that cannot be stored is refused loudly, never rounded (#11892, #11896).

And one to expect on your next bare update: locally modified helper
scripts
now get the same five-outcome manifest protection as skills — a
hand-edited helper reports conflict-skipped instead of being silently
overwritten, and FORCE=1 leaves a .bak beside anything it replaces
(#11894). Also in this release: the streamable-HTTP .mcp.json entry for
fresh installs (#11873, #11875), the token reaching interactive terminals
(#11874), and the ruby conformance baseline moving to nightly-only CI
(#11893).

Changelog

Features

  • 335d101837b7738278e105c3c46d93ca41802bcc feat(config): the work window keeps its minutes (task #11896)
  • 937d21249d4b1f6af9f16998672ddd731275d19a feat(update): carry a legacy SSE mcptask-online entry forward (task #11875) ### Fixes
  • e3ac5976b98e2c189534e3174803fea81935fc88 fix(install): --until refuses minutes it cannot store (task #11892)
  • 525daab5394e4d244bb5f52434a9195a24eec7fa fix(install): a fresh .mcp.json gets the streamable-HTTP entry (task #11873)
  • a8277924b6ea1dcdb274d9ec4cce211443650067 fix(install): helpers get the same five-outcome protection as skills (task #11894)
  • 22525d047a05382f6842127ab1d5953b6f99a958 fix(install): the token reaches a terminal, not only the scheduled job (task #11874)
  • 232399064826689e4dded85488bb5eed1b90ef80 fix(runner): --ignore-quota covers the daily scheduler too (task #11891)
  • 4c870dcc01c07fdf47f3eea5445e3ca817c434d7 fix(tests): the lock-guard test resets its clock before every subtest (task #11897)
  • 9e7abcc1af46b18293a54eb9fb5d2ebf3cbc2ce5 fix(tests): the shell-file table joins its paths instead of spelling them (task #11880) ### Other
  • aadfcf8d051948d4d9db6a04af39a1e4bcd96cb4 ci: the ruby baseline runs nightly only, and the docs stop contradicting it (task #11893)
  • be020fa81612407e1cfddbdfdb03054d7e1b94f0 docs(comments): the outage window is named by its constants, and childEnv stops teaching the pre-#11670 model story (task #11895)
  • 495223cb63bf2f8e1bee3ddbd3079163a51cc6bd merge: runner-audit fixes — tasks #11891-#11896

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.22

Upgrading takes two commands this time, not one

mcptask_runner update --self     # replaces the binary
mcptask_runner update            # in each host project

The first is the usual one and covers every loop change below. It is not
enough for the skills and helpers: those are a data pack the binary carries, and
only the second command writes them into a project.

Skipping it leaves the old /test-runner and /ci-runner in place — and on a
project that is not Rails those name commands it cannot run, which is the very
thing this version fixes. A host that has locally edited a skill keeps its edit;
FORCE=1 takes the new one and leaves a .bak beside it.

Worth knowing why this is easy to get wrong: the data pack lives inside the
binary, so an update run through the old binary reinstalls the old skills.
The order above is the whole trick.

One runner, every framework

The bundled CI and test skills stop naming a framework. /test-runner and
/ci-runner used to spell out bin/rails test, test:system and RuboCop, and
check_test_lock recognised a running suite by six regexes, five of which were
Ruby. On a Django, Node or Swift host the commands did not run, and the guard
against a second concurrent suite waved every real test command through — blind
on exactly the projects a universal runner exists for.

Both now read what the project declares in .claude/test-commands.json, which
mcptask_runner init writes and update backfills. With no such file the
toolchain is detected from go.mod, Gemfile, package.json, manage.py,
pyproject.toml or Package.swift — and when nothing is recognised the runner
says so rather than inventing a command that will fail later and elsewhere.
(#11853)

Two more places the prompt spoke Rails to everyone. The auto-squash workflow
told every project to run bin/rails assets:precompile RAILS_ENV=test between
its unit and system suites, and the manual screenshot step named a Rails class.
Neither is measured by what it breaks — the first costs a turn or invites a
plausible substitute, the second is silently ignored on a project that has no
such class — but both contradict the claim the product makes.
(#11858,
#11864)

A test now guards it: every executor prompt, both configs, is checked against
two lists — commands only one ecosystem can run, and identifiers that exist in
only one.

The working day

A failed task waits an hour, not the rest of the day. A task that fails is
set aside and the runner carries on with the next one.
(#11846)

The dashboard can take a piece off the skip list. A piece set aside earlier
can be released from the card, and the runner picks it up again without a
restart. (#11847)

The dashboard connection

A closed socket is described by whoever actually saw it close. Two paths
reported a dropped connection — the read loop, which knows the close code the
server sent, and a failed write, which knows only that the socket was gone.
Whichever finished first wrote the notice, so a deliberate close with a code
could be reported as a bare transport error.

The reader's account now wins. That is what makes "the server closed it, code
1001" worth trusting when you are deciding whether the server needs looking at,
and it is the difference between a server letting a connection go and a
connection dying underneath both ends.
(#11860)

Under the hood

The loop's test logger is safe to write from goroutines — a data race the race
detector caught on CI while forty local runs stayed green
(#11859) — and a stress day that
overruns its wall clock no longer leaves a child writing into a temporary
directory that has already been removed
(#11865). Neither changes runner
behaviour.


Verified before tagging: the full local bin/ci green on all ten steps with no
skips, including the stress day and both conformance suites, and GitHub CI green
across three operating systems and the Go floor from go.mod.

v0.3.21

The one that was costing whole days

A deploy of mcptask.online now costs one attempt instead of the working day (task #11832).

While the site is being deployed its MCP server is unreachable for minutes and then comes back on its own. The runner spent its entire recovery budget inside about forty seconds — a child dies in ten, five seconds between restarts, two restarts — so all three attempts lost the same race before the server could possibly be back, and the run ended with error. That status is what ends the day: one deploy could stop a day that had already finished eighteen tasks.

Restarting faster cannot fix a wait. The two situations that reach that path are now told apart where they are detected, rather than by reading the message afterwards:

  • A deferred tool that never loaded is the child's own miss. It keeps its two fast restarts and still ends in an error, because it is not going to fix itself.
  • An outage gets six restarts at twelve times the wait — about six minutes, which spans a deploy — and when that runs out it ends the way a stall does: a verdict the loop carries on from, the task left in_progress for the next triage, and no bug piece, because a deploy is not a defect anybody can fix. The wait keeps the heartbeat going, so a card watched through a deploy says the runner is waiting rather than falling silent.

This was the largest cluster in the Errors epic — five of twelve real failures, across three machines and two host projects.

The one you will see

Today's skip list is visible (task #11835). A task the runner has set aside — out of scope, failed, already taken by somebody else, a merge it cannot finish — is skipped by every later triage that day, and until now the only trace was one log line at the moment it was decided. It now shows on the startup banner, in every wait, and on the dashboard card. A quiet runner can be read rather than guessed at.

Also

main is unprotected on purpose, and the two habits that stand in for a branch rule are now written down (task #11831); a new push into a pull request stops the CI run it supersedes (task #11826).

Upgrading

Run mcptask_runner update --self on each host, or move a project's gem pin to 0.3.21 and run rake mcptask_runner:update. The binary is per machine and the newest one any project offers is the one that stays, so one project is enough.

Until a host actually runs this binary, a deploy of mcptask.online goes on costing it the day.

v0.3.20

Upgrading: run mcptask_runner update --self on each host, and do it for
this one rather than at leisure. The fix below can only reach a runner that is
running the new binary, and until it does, a task moved between two agents keeps
being worked by both. A daily process takes the new copy at its night swap — but
only once that copy is on disk, so the update still has to happen. Nothing to
migrate; no configuration changes.

What changed

A piece assigned to somebody else while this runner is working it now stops
the child.
The dashboard socket has always carried both halves of an
assignment — "given to you" cut an idle wait short, "given to somebody else"
reached nothing — so a task moved between two agents was picked up by its new
owner while its old one worked on to the end. Two agents on one piece, the same
branch name pushed twice, two sets of efforts logged. The runner now abandons
the piece, says so in the log, and goes back to the queue for one that is still
its own. It is not a reason to stop the day, and the piece is deliberately not
set aside: the queue has already stopped offering it, and it stays workable if
it is handed back. Neither quota ending loses its priority — a run killed by a
crossed budget and a reassignment in the same breath still stops the day, since
that is the claim about the whole day and this is a claim about one piece.

The run log is coloured as it is read, and the file itself stays plain. The
two lines that matter in a wall of uniform text — the ones Warn and Error
already mark with a glyph — are painted red and yellow, a child's stderr is
yellow, and the [Component] tags, cost lines and phase separators are dimmed
so the shape of a run is visible in a scroll. Only what the runner itself marked
gets coloured; guessing a severity from the wording would be wrong on the day it
mattered. Nothing is written to the file, which is what bug reports attach and
what people grep. runner-log --paint is the same filter over stdin, so
tail -f any.log | runner-log --paint works for a log read some other way.

The stress harness gained a day that proves the first of these end to end: two
otherwise identical days differ only in whether the broadcast names the piece in
flight, and they come out opposite — one abandons, the other leaves the child
alone. That second half is the one worth having, because killing healthy work
over a stranger's assignment would be worse than the bug this fixes.


Changelog

Features

  • 1fb910ebee499b6f89efc1935806a84ad4c7ce14 feat(stress): a day proves the runner lets go of a piece taken from it, and keeps working (task #11790) ### Fixes
  • 82c2ddddd93abee15fe8c3ed560f436c110128a3 fix(eventstream): a piece given to somebody else stops the runner that was working it (task #11788) ### Other
  • e3de44be2da659b116dcb8d824769cb78211a13a feat(runner-log): the run log is coloured as it is read, and the file stays plain (task #11786)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.19

The one you will notice

Every step that does work says so on the piece before it moves on (task #11784). A run on 2026-08-25 opened its pull request having called LogWorkProgressTool zero times — while its plan moved on the dashboard card the whole time. The card said "working", the piece said nothing, and the hours landed nowhere.

Task #11616 had already established the shape of the fix: the cadence has to be written into a numbered step, because a block sitting past the last step never gets read. It put three anchors on the auto-squash spine — and left the seven steps between CREATE BRANCH and the CI gate saying nothing about logging, while the three manual workflows got no anchor at all.

Now every substantive step carries its own LOG PROGRESS NOW line, and each one names the TaskUpdate that closes the same step's todo item:

4. IMPLEMENT TASK (incremental commits, clear messages)
   - LOG PROGRESS NOW — TaskUpdate this step's item "completed", then
     LogWorkProgressTool at ~45%: implementation written, tests not run yet

That pairing is the whole trick. The child was calling TaskUpdate reliably all along — that is why the card moved — so hanging the log on the call it already makes is what ties the effort trail to the plan. Auto-squash now walks 20 / 45 / 60 / 70 / 75 / 80 and reaches 100% only after gh pr view says MERGED; the manual workflows top out at 90%, because the PR is left for a human; a review run logs at 25 / 70 / 100.

It binds the model rather than forcing it — the runner posts no effort of its own, and has no way to speak to a child mid-run. What changed is that skipping the log is now skipping a step.

The one that was quietly stopping the 08:00 job

The child searches the project, not the whole disk (task #11782). A child that cannot place a file falls back to find / or find ~ — two days of one project's logs hold five such calls, one of them find /Users/<you> -maxdepth 3.

On macOS that walk enters ~/Documents and ~/Desktop, which are gated behind TCC, so the OS raises a permission dialog. The dialog names mcptask_runner — macOS asks on behalf of the responsible process, the launchd-spawned root of the tree, not the find underneath it — which is why the binary gets blamed for folders it never opens. And the job fires at 08:02 with nobody at the keyboard: the dialog blocks the call until it times out, and then denies it.

The prompt now bounds the search to the project directory, and — the part that makes a prohibition stick — names where to look instead: bundle show <gem> / gem which <file> for Ruby, go env GOMODCACHE for Go, then search under the path it prints.

Upgrading

Move your project's gem pin to 0.3.19 and run rake mcptask_runner:update. The binary is per machine and the newest one any project offers is the one that stays, so one project is enough to update the host.

Both changes are prompt text sent to the child, so they take effect on the next run the updated binary starts. Nothing to migrate, nothing to re-install beyond the binary.

v0.3.18

The one you will notice

The install and update output is coloured (task #11779). The update summary was a wall of grey in which the one row that needed attention read exactly like the thirty that did not. Now: green for up-to-date and added, yellow for updated and force-updated, red for conflict-skipped, dimmed section tags, warnings in red on stderr.

It is decided per stream and only when the stream is really a terminal — a pipe, a redirect and a launchd log all get exactly the bytes they got before. MCPTASK_RUNNER_COLOR (1/true/always/force, or 0/false/never) overrides, and NO_COLOR is honoured. Your run logs stay greppable.

The one that was quietly costing you processes

A child's process group is remembered, so what the child left behind is still reachable (task #11773).

A process group outlives its leader: while any member is alive the group id stays allocated and a signal to -group still reaches them. But the kernel can no longer name that group once the leader has been reaped — and killTree looked it up at kill time, got nothing, and fell through to signalling a bare pid that no longer resolved.

So anything a child started and then exited on went on running. Every claude session that launches a dev server, backgrounds a command, or leaves a watcher behind and then finishes its turn is exactly that shape. Nothing raises, so nothing ever reported it — the processes just accumulate on an unattended machine.

This had shipped in every version of the runner there has ever been. It is fixed by reading the group when the child starts, while it is certainly alive, instead of asking about it when it is already gone.

How it was found, which is the part worth repeating

Not by the test suite. bin/ci runs the chaos stress harness against one fixed seed, and the invariant that catches this — the orphan sweep — has existed since the harness was built. They never met, because the fixed seed draws a different misbehaviour at the moment the signal lands.

It took eight random seeds. Two failed on the current code, and a third failed all the way back at v0.3.16. Scheduling that sweep is now task #11778.

Also in here

The stress harness grew to cover the dashboard-card behaviour that landed after it was built (story #11765, seven tasks) — none of which ships in this binary, but it is what turned up both of the above.

Upgrading

Move your project's gem pin to 0.3.18 and run rake mcptask_runner:update. The binary is per machine and the newest one any project offers is the one that stays, so one project is enough to update the host.

v0.3.17

What this is

One story: the chaos stress harness grown to cover the behaviours that landed after it was built (story #11765, seven tasks).

For an operator, nothing changes. No new behaviour, no changed wording, no new configuration. The entire shipping delta against v0.3.16 is 29 lines, and the only one that touches a code path is a refactor that leaves production behaviour identical.

The one real fix

Stream.SetOnAssignment (new), wired unconditionally by the work loop.

The assignment wake-up added in v0.3.16 was wired inside the branch that builds a stream when the caller supplied none. Production always takes that branch, so the feature worked — but every caller that supplies its own Stream got one with no way to fill the wake channel, which meant no test could ever see it fire. The harness proved it: 299 assignments delivered to the runner's own user, not one wait cut short. The wiring now happens on whatever stream the loop ends up with.

If you are on v0.3.16 this changes nothing you can observe. It is here so the next change to that path cannot break it silently.

Everything else is the harness

internal/stress and internal/chaos do not ship in this binary. What grew there:

  • the stub's dashboard socket pushes as well as records, and can end a connection either with a close frame or with none at all
  • a stop names which of the three ended the day, and a reconnect names who closed the socket
  • the card's message must stand, a clean run must say how it went, and a retry must say what it is retrying for
  • an assignment wakes the short waits and must not touch the overnight sleep or the daily restart's anchor
  • a forked task keeps the name that buys it the watchdog's generous ceiling, while a tool that never closes is still killed
  • a plan reaches the card, and every child is proved to have been launched with the tools that make one

The suite now costs about four minutes rather than ninety seconds; bin/ci says so.

Upgrading

Nothing to do. The gem ships at the same version as always — move a project's pin to 0.3.17 and run rake mcptask_runner:update when convenient, or wait for the next release that actually changes something.

v0.3.16

Upgrading

Run mcptask_runner update --self on each host. Everything here is binary
behaviour — no new skills or helpers ship with this version — so plain
mcptask_runner update has nothing to do, and neither does bundle update:
the binary the 08:00 job runs lives in ~/.mcptask/bin and only --self
replaces it.

The wrapper gem mcptask-rails-runner is published at the same version, on all
six platforms.

What this release is about

The runner card on the dashboard is filled in rather than sparse. It names
the binary drawing it, says how a finished run went, counts out the retries the
CLI is riding out, links a pull request the run has just opened, and says what
an MCP tool call is acting on instead of showing a bare tool name. A task forked
inside a subagent reaches the card too, and the notification that ends it is no
longer thrown away. A card whose stream goes quiet keeps saying the last thing
it knew instead of emptying out, and the TODO list fills in on Opus and Sonnet
as well.

An idle wait ends the moment mcptask.online says this runner has been given a
piece
, rather than sitting out the rest of the poll interval.

A stop on the daily quota says which quota, and how many hours it means —
worked_today=8.2h of per_day=8h. Those hours are the user's day on
mcptask.online, not the run's: another session spends the same budget, so a
runner that idled all morning could end its day on hours it never worked, and
the old one-line message gave a reader no way to tell that from a bug. The same
line now separates a spent budget from an endpoint that never answered, and a
stop caused by a failed task no longer claims to be a quota stop.


Changelog

Features

  • e921a308b8fbcf2a25170d045efa21a28661bc59 feat(executor): a finished run tells the card how it went (task #11758)
  • ee0af55aaf5b9c2413245d5f9e5e13d8f53b6869 feat(executor): the card counts out the retries the CLI is riding out (task #11754)
  • 7ab052bb17845a0b24b2f353a6b96f19119f72f4 feat(executor): the pull request a run just opened shows up on its card (task #11753)
  • cfa61d63b0f4d26311cbf1253a5ad4d5bb4dc8e1 feat(runner): an idle wait ends when mcptask.online says this runner was given a piece (task #11741)
  • 55872bda6abec508541898b59a2ccafbb6f152b9 feat(snapshot): the runner card names the binary that is drawing it (task #11747) ### Fixes
  • 82ebb34127201c463e683996d27d755e999b331f fix(eventstream): the reconnect notice names what closed the socket
  • 94c3df906f741e11924f1f9a81238964d1d763e9 fix(executor): a task forked inside a subagent reaches the card, and the notification that ends it is not thrown away (task #11751)
  • 29978c739ec623b14188fc6708e5447ba3120881 fix(executor): an MCP tool call says what it is acting on, instead of arriving as a bare name (task #11752)
  • 1f8c1cf4cd755524d7ff565cda919344dd37bc5b fix(executor): the dashboard TODO card fills in on Opus and Sonnet too (task #11755)
  • cae366f659a7abd6287c0337c1aa387f566a8762 fix(runner): a daily-quota stop names the quota and the hours behind it (#5)
  • 87cbd6820e787192ab590e5d04f44648dafbca42 fix(snapshot): the card keeps saying the last thing it knew (task #11757) ### Other
  • 3a73198713dd540b1fefaf46dec7ec1f3c9b7eb8 docs(ci): a piped bin/ci reports tail's exit code, not the suite's
  • a2555366c751739e2513be0580c793e79c930d30 docs(snapshot): the two session views, and why neither is a superset (task #11759)
  • 3a1056cee599be8b20ba258926e78a8814618f03 test(executor): the branches the two card stories left uncovered (tasks #11750, #11756)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.15

What this release is for

Two things an unattended host feels directly, and one the dashboard does.

A 429 rate limit is recognised. The CLI retries a spent usage window ten
times internally and then stops mid-answer with no result marker. The runner had
no branch for that, so it called the ending a missing marker, spent its three
thirty-second retries against a window that had not reset, failed the task, filed
a bug piece about the marker — and, because the quota Decider stops the loop on
any error, ended the day's remaining work. The window that prompted this came
back about a minute after the last retry. A 429 now has its own flag, its own
budget of eight, and a back-off measured in minutes (5 → 10 → 20 → 30, capped),
classified ahead of the missing-marker path because a 429 always also looks like
one. A spent usage window is no longer filed as a bug: nobody can fix one.

A daily runner adopts a new binary by itself. Everything that notices a new
version — the bundle adopt, the update hint — runs once, at startup, and a
daily process never ends. It kept whichever binary it started with, however
many releases went by, on exactly the host nobody is watching. On its way into
the overnight sleep it now compares its own executable with what is installed and,
if they differ, retires the card, releases the instance lock and replaces itself
— at the configured end of the workday, or at midnight on a host with no window.

A failed task now shows as failed on the card. A run is marked finished the
moment its stream ends, before the ending is classified, and finished → error
was not a legal transition — so a failed task filed a bug piece while the
dashboard went from finished straight back to triage. Anyone watching saw a
healthy runner through every failure of the day.

Upgrading

Nothing to do beyond the usual install. One exception, and it is the reason to
bother: a daily host will not pick this up by itself. The self-restart
above only exists from this version on, so the process now running has to be
stopped once by hand — after that it keeps itself current.

launchctl bootout gui/$(id -u)/online.mcptask.runner-<slug>
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/online.mcptask.runner-<slug>.plist

today hosts need nothing: the next morning's job starts the current binary.
Check which one a machine is running with ~/.mcptask/bin/mcptask_runner version.

Changelog

Features

  • be9e6ef541de78e0cb923fbc32f4c4de13791f37 feat(chaos): a CLI stand-in that starts well, decays through five phases, and records what it did to the runner (task #11715)
  • 9b8b4cc4a50ceca2cf32614e40e554b8a7f11383 feat(chaos): a stub mcptask.online that freezes, lies, drops the dashboard socket, and keeps an honest ledger (task #11716)
  • 184cc296585c53afc787ed3bdb89769a1843ca5f feat(executor): a 429 rate limit is waited out on its own budget instead of misread as a missing marker (task #11731)
  • c5e50fa889774937a76d2fe8a70a2f9253cf977a feat(runner): a daily runner replaces itself with the installed binary at the day's end (task #11729)
  • dcef1a9ef3d1421d5221b85111a76fc927b9af60 feat(stress): a real signal to a real binary, and everything it has to take with it (task #11719)
  • e0eb4f012390d344327db1f970d12b0722212ea4 feat(stress): an in-process driver that compresses a workday into ninety seconds and states how the loop may end (task #11717)
  • e38e3638755ce63536aa8bdcbd3fc5c00a2e57be feat(stress): one fixed-seed day in bin/ci, bin/stress for the rest, and the harness written down (task #11720)
  • 36a22107e88c62e2bf5f3613025a18c08a374550 feat(stress): the card has to keep telling the truth while everything burns (task #11718) ### Fixes
  • f799bfe05a5d6490c2cdf691fc58ff499044df93 fix(eventstream): a first subscription confirmation has nothing to resend, so a fresh session stops duplicating its opening frame (task #11725)
  • ea6fc1839a472f7bbaeb4b7e3507d92c9652215a fix(executor): a marker whose JSON does not parse is not the child's answer — keep looking, and take the retry (task #11724)
  • 74dc7ad64df2f477e6d8306632d25327bbf98add fix(executor): a server that has not finished connecting is recognised as the MCP route failing, not left to become an empty bug report (task #11712)
  • 693d609f24dff2143e40924c879ed26907bf2e11 fix(executor): the card says what it is retrying for, because the runner already knew (task #11732)
  • c8a70925de6f362b0fb156740529ba395cb5faf9 fix(loop): a day that ended on the clock sleeps until the next one instead of spinning until midnight (task #11726)
  • c9644cf59a609d39643aac69134fb839b88a874a fix(runner): a failed task shows as failed on the card, and the check that says so cannot be fooled (task #11734)
  • 8678677b6c093f7e312956b35d762983bb8c95d9 fix(snapshot): the dashboard TODO card follows TaskCreate/TaskUpdate, the plan tools that replaced TodoWrite (task #11710) ### Other
  • 2658215af771f52833587916fc08b23defc13792 docs(config): every wording of work_window says which mode exits and which sleeps (task #11728)
  • bfc76427946075e8fffac21a29db48c7fe01dbd1 merge main into story-11714 before the story merge

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.14

Upgrading

Run mcptask_runner update --self on each host. This release changes loop
behaviour inside the binary, and nothing replaces the binary for you: the copy
the scheduled job runs lives in ~/.mcptask/bin, and neither bundle update
nor a gem bump touches it. mcptask_runner version says what a host actually
has.

No new skills or helpers ship here, so plain mcptask_runner update has
nothing to do in this one.

What changed

An empty queue is asked about every half hour, not every five minutes.
The REST next-task check added in v0.3.13 made an idle round cheap, and cheap
rounds ran twelve times an hour: two GETs each, and two rows in the dashboard's
activity feed each, all saying nothing had happened.

waiting_strategy.short_wait_minutes still governs the first rounds — it is
the right answer in the minutes just after a task finishes, when work is
expected any moment. What is new is that a run of empty rounds counts as
evidence: once those short waits have added up to one long_wait_minutes, the
long wait takes over and keeps it until there is work again. Nothing waits
longer than long_wait_minutes, the ceiling daily mode always had, and the
switch is announced in the log in both directions.

On the installer's 5 and 30 that is 21 idle rounds across an eight-hour
afternoon instead of 96. A configuration whose long wait is no longer than its
short one has nothing to back off to, and is left alone.

Two smaller log fixes in the same path. The quota check made after an empty
round used to end the day on a bare break with no line at any level; it now
says why the runner stopped. And daily mode no longer claims "will wait 1 hour
before retry" while waiting long_wait_minutes.

Changelog

Features

  • adad985d17350cb40bd716dbe08d02fd92b13816 feat(loop): a queue that has been empty for half an hour is asked about every half hour, not every five minutes (task #11699) ### Fixes
  • a65a5fda79fdaf748efb8a6204ef7393f553fb21 fix(loop): the backoff line names its waits the way the waiting strategy does (task #11699)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.13

Upgrading: mcptask_runner update --self (or bundle update mcptask-rails-runner to 0.3.13) upgrades the binary; then run mcptask_runner update in each project — this release ships a new helper (runner-log) and rewritten config sections, and nothing installs those for you. The scheduled job runs the per-machine binary in ~/.mcptask/bin, so check mcptask_runner version there afterwards.

What changes in behaviour:

  • Before paying for a triage session the loop asks GET /api/:account/pieces/next; an empty queue answers no_more_tasks with no model involved. Until mcptask.online ships that endpoint (task #11691) it answers 404 and the log shows next-task REST check failed (HTTP 404 …); falling through to triage every cycle — expected, triage behaves exactly as before.
  • The configured end of workday is the one hard stop, checked after a task finishes. There is no default any more: a config whose work_window.end_of_workday_hour is unset or null has no end of day (the 18:00 default is gone; the quota and the day gate still apply), and a running task is never interrupted by it.
  • The skip list is persisted, so a restart does not re-pick what today already declined; failure, task_already_started, merge_failed, merge_unverified and out_of_scope set a task aside for the day instead of stopping the run.
  • respect_working_hours is read from the user profile once at startup — a change to the checkbox is heard at the next start.

Changelog

Features

  • a1798c3cb6bc2f707d31e01c4d5d5e5f4d89aef0 feat(bugreport): a guard of the runner's own that did not hold files a piece instead of a log line (task #11684)
  • a14790c3995e036d6368c7b5742768246f617fcc feat(install): ship runner-log, the helper an operator types to tail the current run
  • 3fb7637cc491069964281b2f729ecf3c2f949491 feat(install): the install ends with a summary of every config section the runner reads, and the sections it writes carry a comment (task #11690)
  • 8f78435eb604330ea5267a1cd1085c521da4ff08 feat(loop): ask the REST next-task endpoint before triage — an empty queue is no_more_tasks without a model (task #11692) ### Fixes
  • dde40a14bf9a2fefe495970f3accf73a5e5dc9ba fix(loop): a task declined as out of scope goes on today's skip list, and the loop never stops over it
  • dd3a529977087feb39b27a0626c0e6dc63c35596 fix(loop): failure, task_already_started, merge_failed and merge_unverified set the task aside for the day instead of stopping (task #11680)
  • f727a15afec1b66ff369fc56daaa397f763d9040 fix(loop): the configured end of workday is the one hard stop, checked after a task finishes — and there is no default any more (task #11689)
  • 0b18385f30d5a1f429a68be46714ce8763350b6d fix(loop): the skip list outlives the process, so a restart does not re-pick what today already declined (task #11683)
  • 32c00200b5849084d91bb1e0beaa2249fe4293bd fix(quota): read respect_working_hours from the user profile once at startup, not from time_status (task #11686)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.12

This release is mostly about a guard that had stopped guarding, and a setting that could not be right for everyone reading it.

The lock that let tests through. check_test_lock refuses a direct bin/ci while another agent's suite is running. It looked for pidfiles in a flat directory nothing has written in a long time, so every lock older than ten seconds read as stale and the command went through — the guard worked for the first ten seconds of a lock's life and never again. The path is now rebuilt from the lock's own project. Two smaller failures in the same lock went with it: the directory that serialises the moment of acquiring could not go stale, so an acquire killed midway wedged every later run behind a lock reported as unknown; and a lockfile missing its PID= line killed acquire outright before the staleness check could reap it, leaving nothing but a blank error to act on. (#11669, #11671, #11672)

Bundled skills no longer declare a model:, and the install-time rewrite that resolved it is gone.

A model name written into .claude/skills/ can only be right for one backend, and that directory is read by two: the runner's child process and whoever opens the same project in a session of their own. Whichever way the name was written, it was wrong for one of them — and wrong silently, because the CLI returns its model error as the skill's own result and the agent above reads that as the answer. Each fork now runs on its session's own model. context: fork is untouched, and that is what actually keeps a fork's output out of the parent's context. (#11670)

Upgrading

Run mcptask_runner update in each project after upgrading the binary. Nothing does it for you — upgrading the binary does not refresh the skills and helpers already on disk, so without this step every fix above stays installed nowhere. Helpers are replaced in ~/.claude/bin; skills are refreshed per project. Do it in every project you use, not just one. (Making the upgrade reconcile them by itself is tracked as #11673.)

No --force is needed. Hosts whose skills carry a rewritten model name are replaced normally by mcptask_runner update, because the installer recorded the manifest hash from the same content it wrote — so the rewritten form is the baseline, and the updater reads those files as untouched. Verified against three hosts carrying rewritten copies (gemma4:31b-cloud, minimax-m3:cloud and haiku): all eight bundled skills hash equal to their manifest entry on each. A skill somebody edited by hand is still reported as a conflict and skipped, exactly as before.

One known rough edge, not fixed here: now that the lock guard blocks correctly, it also blocks commands that merely mention bin/ci — grepping a CI log, for instance — because the pattern matches the substring rather than the command position. If a refusal looks wrong, check ~/.claude/bin/test_lock status to see whether a run is genuinely in progress. Tracked as #11674.

Changelog

Fixes

  • cccbcab0c3ec7d60f581c594405a91121f9945b6 fix(shutdown): a SIGTERM is not a failed attempt, so it announces no retry (task #11655)
  • 481f3a986071f48906736ef6dc791eb0d4a8329f fix(skills): a pinned model can only be right for one of the two backends reading the same directory (task #11670) ### Other
  • f6d5a9815968c7fa3f2c7443e150928740e698ed fix(ci-lock): the acquiring guard could not go stale, and a nameless lock killed acquire outright (tasks #11671, #11672)
  • bb0b389ebab93a7153d7f439e3d8600367a4a13a fix(ci-lock): the guard against a direct test run looked for pidfiles nobody writes (task #11669)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.11

Changelog

Fixes

  • 230496500ec4a402e8d4bdd4bcff1918147f74df fix(quota): a role that ignores working hours is not stopped by its hour goal (task #11643)
  • 886a3355e74423c1d01e1a367a3d39a613a4731c fix(triage): a disabled fetch tool is a detour, not the end of the day (task #11642)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.10

Changelog

Fixes

  • 5fef8b8808e0f121d83d093370ddc01964e07cb1 fix(helpers): resolve Ruby the same way on rbenv as on rvm (task #11641)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.9

Changelog

Fixes

  • bc156c306eb0ce4f115e07d004fdcd6802851175 fix(ci): bump pinned golangci-lint to v2.13.0 for Go 1.27 support
  • 642002843a54df7469862fd332bae3adb55e1c74 fix(install): print which account a scheduled job will authenticate as (task #11494)
  • b782fbecb65c01aa7072cf72b78b5ea5db07b958 fix(launcher): one log file per run instead of one that grows forever (task #11617)
  • a40f46624f890ca770b34e0181d5d204d35ab7ce fix(loop): a declined task skips, it does not end the day (task #11622)
  • e38b8d6025f8662f11a9a6c2cdb8b69652dfda15 fix(triage): discard a pick that belongs to a different project (task #11619) ### Other
  • 55f52d52fb2f96a49b26aa4c7c2d1361246eabb9 docs(skills): follow the MCP parameter renames ahead of the server change (task #11609)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.8

Changelog

Features

  • bb18d74c555e82f11ac1324d9b3b92ac9b3e4d22 feat(config): make end-of-workday hour configurable at install time (task #11495) ### Fixes
  • 7264ca3901696ef4c861fada80b586d35793394a fix(prompt): anchor the progress-logging milestones inside the auto-squash steps (task #11616)

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

v0.3.7

Changelog

Features

  • b66ced087bcd9d6afae57c80e57cf5ea1ce4283a feat(bugreport): remove the manual bug-report CLI command (task #11481) ### Fixes
  • 66bab4931cd4555bab6cfc0655994d9ae61bd275 fix(bugreport): retarget automatic error reports to this project's own Epic (task #11482) ### Other
  • 0ae13df7ff594e7e7db44aa841892ed4a0ee6df4 chore(claude): each story gets its own worktree to protect against concurrent session conflicts
  • c1af57c30c09d8a8f16ecd8618aaf109d0443325 chore(claude): ignore /.claude/settings.local.json repo-wide
  • 30d47d8edc50821df2fb312e2f9eeccfff6b5075 chore(claude): install discover/mcptask-read/mcptask-write/memory-search skills

curl -fsSL https://github.com/jchsoft/mcptask-releases/releases/latest/download/install.sh | sh

Or brew install jchsoft/tap/mcptask_runner, scoop install mcptask_runner,
or npx @mcptask/cli. Verify any manual download against checksums.txt.

list_altReference CLI

Pět veřejných subcommandů, dva persistent flagy, tři exit kódy

Runner je jediná binárka s jedním cobra rootem: run, init, update, version, pr. Šestý subcommand internal je schovaný před --help, protože ho volají vygenerované launcher skripty, ne operátor: internal launcher-log-name je to, čím si windowsový launcher pojmenuje vlastní log. Skrytý neznamená dočasný – launcher na něm za běhu stojí. Dva persistent flagy --verbose a --ignore-quota se čtou z odpovídající proměnné prostředí bez ohledu na velikost písmen (verbose=true, ignore_quota=true). Flagy jednotlivých subcommandů jsou vypsané u každého příkazu níže. Exit kód je 0 při úspěchu, 3 u neimplementovaného režimu, 1 u čehokoliv jiného a 128+N při signálu. Trojka je navržená pojistka, ne konec, na který dnes dosáhnete: každý režim je implementovaný, takže ErrModeNotImplemented nikdo nevrátí.

menu_bookVšechny příkazy, přepínače a návratové kódy – referenční přehledexpand_more

Subcommandy

run — pracovní smyčkaod v0.1.0

Spustí pracovní smyčku v jednom z patnácti režimů: jeden úkol, vše, co je přidělené na dnešek, celou frontu, podúkoly jedné Story, konkrétní úkol, čekající code review, nebo bezobslužnou smyčku, která běží každý pracovní den. Režimy auto-squash dovedou úkol až ke sloučenému pull requestu; manuální režimy nechají pull request otevřený, aby ho prošel člověk.

  • --story-idod v0.1.0

    Story, jejíž podúkoly režimy story_manual a story_auto_squash zpracují.

  • --task-idod v0.1.0

    Konkrétní úkol, na kterém režimy task_manual a task_auto_squash pracují.

init — nastavení projektuod v0.1.0

Připraví projekt pro runner jediným příkazem: zapíše, který kódovací nástroj projekt pohání a na kterém git hostingu žijí jeho pull requesty, nainstaluje přibalené skilly a pomocné skripty, zapíše konfiguraci MCP a oprávnění, která nástroj potřebuje, uloží přístupový token a vygeneruje naplánovanou úlohu na pracovní dny, kterou si zapnete sami.

  • --cliod v0.3.24

    Určí kódovací nástroj, který projekt pohání — claude, codex nebo opencode. Napoprvé je povinný, výchozí hodnota neexistuje a další spuštění odpověď převezmou.

  • --git-hostod v0.3.25

    Určí git hosting — github, bitbucket nebo gitlab — tam, kde to vzdálený repozitář origin neprozradí, třeba za SSH aliasem, na zrcadle nebo na vlastní instanci GitLabu.

  • --epic-idod v0.1.0

    Epic na mcptask.online, do kterého runner zakládá chybové úkoly o sobě samém; 0 je nechá v kořeni projektu. Odpoví na otázku bez dotazu.

  • --epic-nameod v0.1.0

    Zobrazovaný název uložený vedle --epic-id.

  • --modeod v0.1.0

    Režim pracovní smyčky, ve kterém poběží naplánovaná úloha, například today_auto_squash, nebo none, pokud žádnou naplánovanou úlohu nechcete.

  • --atod v0.3.3

    Čas, kdy se naplánovaná úloha spustí, ve tvaru HH:MM; bez něj v 08:00. Hodnotu, která není denním časem, odmítne, nikdy ji nezaokrouhlí.

  • --untilod v0.3.8

    Čas, po kterém runner nezačne žádný další úkol, ve tvaru HH:MM. Rozpracovaný úkol ještě dokončí. Bez něj pracovní den nekončí a runner zastaví jen denní kvóta.

  • --scheduleod v0.1.0

    Znovu vygeneruje naplánovanou úlohu a na nic jiného nesáhne.

  • --forceod v0.1.0

    Přepíše skilly, pomocné skripty, sekce konfigurace i naplánovanou úlohu, které už existují. Zapíná ho i FORCE=1.

  • --helper-bin-dirod v0.1.0

    Kam se nainstalují pomocné skripty pro CI a testy a co budou skilly volat. Výchozí je adresář, který určuje profil kódovacího nástroje.

  • --home-dirod v0.1.0

    Umístí adresář s tokeny, startovací soubor shellu a naplánovanou úlohu do jiného domovského adresáře, než je ten aktuálního uživatele.

update — obnova skillů nebo samotného runneruod v0.1.0

Srovná nainstalovaný projekt s verzí runneru, na které teď běží. Každý přibalený skill a pomocný skript porovná s tím, co bylo nainstalováno, a ohlásí, zda ho přidal, zda je aktuální, zda ho aktualizoval, zda ho přeskočil, protože jste ho upravili, nebo zda ho na vaši žádost přepsal a zálohu nechal vedle. S --self místo toho vymění samotný binární soubor runneru za poslední vydání.

  • --selfod v0.2.0

    Nahradí binární soubor runneru posledním vydáním ověřeným proti jeho kontrolním součtům; nepovedené nebo nesouhlasící stažení nechá nainstalovaný soubor beze změny.

  • --checkod v0.2.0

    Ohlásí, zda existuje novější vydání runneru, a nic nemění.

  • --forceod v0.1.0

    Přepíše skilly a pomocné skripty, které jste lokálně upravili, a vedle každého nechá kopii .bak. Zapíná ho i FORCE=1.

  • --helper-bin-dirod v0.1.0

    Kde se obnoví pomocné skripty pro CI a testy a co budou skilly volat. Výchozí je adresář, který určuje profil kódovacího nástroje.

  • --home-dirod v0.1.0

    Přesměruje vše, co update zapisuje mimo projekt, do jiného domovského adresáře.

version — co běží a s jakým nastavenímod v0.1.0

Vypíše verzi runneru spolu s konfigurací, kterou pro tento projekt vyhodnotil: kódovací nástroj, modely pro jednotlivé úrovně, spouštěcí příkaz, git hosting a další nastavení, u každého i to, odkud se jeho hodnota vzala.

pr — pull requesty na kterémkoli podporovaném git hostinguod v0.3.25

Zakládá, vypisuje, čte, slučuje a kontroluje pull requesty na tom git hostingu, kde projekt je — GitHub, Bitbucket Cloud nebo GitLab — a na každé volání vrátí jeden JSON objekt. Přibalené skilly i zadání runneru používají jen tento příkaz, takže stejné pokyny fungují na každém hostingu.

  • pr createod v0.3.25

    Otevře pull request pro aktuální větev — na GitLabu merge request — a použije šablonu pull requestu, pokud ji projekt má.

    • --titleod v0.3.25

      Název pull requestu. Povinný.

    • --bodyod v0.3.25

      Popis pull requestu zadaný přímo v příkazu.

    • --body-fileod v0.3.25

      Načte popis pull requestu ze souboru.

    • --baseod v0.3.25

      Větev, do které se má slučovat; bez něj výchozí větev hostingu.

    • --headod v0.3.25

      Větev, na které je práce; bez něj aktuální větev.

    • --draftod v0.3.25

      Otevře pull request jako koncept.

  • pr listod v0.3.25

    Najde pull requesty: ty, které zmiňují úkol z mcptask.online, ten, který vznikl z dané větve, nebo všechny otevřené pull requesty repozitáře.

    • --taskod v0.3.25

      Vyhledá pull requesty, které zmiňují úkol z mcptask.online. Výsledky jsou kandidáti — hledání jednoho čísla najde i delší čísla, která ho obsahují.

    • --branchod v0.3.25

      Najde pull request založený z této větve, ať je v jakémkoli stavu.

    • --stateod v0.3.25

      Zúží hledání přes --task na otevřené, sloučené, zamítnuté nebo všechny pull requesty.

    • --openod v0.3.26

      Vypíše všechny otevřené pull requesty repozitáře — celou frontu na review.

  • pr viewod v0.3.25

    Načte jeden pull request — jeho stav, větve a odkaz — jako JSON.

  • pr reviewsod v0.3.26

    Načte v jednom seznamu vše, co na pull requestu nechali recenzenti: nejdřív verdikty, potom komentáře, ty k řádkům kódu i se souborem a řádkem, a to včetně botů, kteří jsou tak označení.

  • pr mergeod v0.3.25

    Sloučí pull request a znovu ho načte, takže výstup říká, co se opravdu stalo, ne co se požadovalo. Uvnitř běhu v režimu auto-squash sloučení odmítne, dokud CI na hostingu neřeklo ano — nebo dokud hosting žádné kontroly nespouští a projekt má vlastní lokální bránu.

    • --squashod v0.3.25

      Spojí větev do jediného commitu. Ve výchozím stavu zapnuto.

    • --delete-branchod v0.3.25

      Po sloučení smaže zdrojovou větev. Ve výchozím stavu zapnuto.

  • pr checksod v0.3.25

    Ohlásí verdikt CI na git hostingu pro daný pull request: úspěch, selhání, stále běží, nebo žádný — a žádný se nikdy nepočítá jako úspěch, protože repozitář, ve kterém CI vůbec neběželo, nic nesplnil.

  • --git-hostod v0.3.25

    Pro jeden příkaz pr přepíše git hosting bez ohledu na to, co říká konfigurační soubor nebo vzdálený repozitář origin.

Globální flagy

--verboseod v0.1.0

Vypisuje každý řádek, který kódovací nástroj streamuje, nejen shrnutí. Zapíná ho i proměnná prostředí verbose=true.

--ignore-quotaod v0.1.0

Vypne kontroly denní kvóty: runner se mcptask.online neptá na zbývající rozpočet a kvůli němu neodmítne žádný úkol — za útratu pak odpovídá ten, kdo runner spustil. Zapíná ho i ignore_quota=true.

Konfigurační klíče

harnessod v0.3.24

Který kódovací nástroj projekt pohání, podle názvu profilu: claude, codex, opencode nebo váš vlastní profil. Jediný povinný klíč — výchozí hodnota neexistuje a projekt, který žádný neuvede, runner odmítne, místo aby něco předpokládal. Zapisuje ho init --cli.

git_hostod v0.3.25

Na kterém git hostingu žijí pull requesty projektu: github, bitbucket nebo gitlab. Nepovinný — bez něj ho runner vyčte ze vzdáleného repozitáře origin, a ten, který nepozná, odmítne jmenovitě, místo aby hádal.

models.geniusod v0.1.0

Model za nejsilnější úrovní, používaný na náročné programování a na úkol, který je potřeba dotáhnout poté, co předchozí pokus nestačil. Bez něj se použije model, který určuje profil kódovacího nástroje.

models.smartod v0.1.0

Model za prostřední úrovní, používaný na třídění úkolů a na review. Bez něj se použije model, který určuje profil kódovacího nástroje.

models.primitiveod v0.1.0

Model za nejrychlejší úrovní, používaný na jednoduchou práci, která hlavně čte. Bez něj se použije model, který určuje profil kódovacího nástroje. Na backendu, který není vlastním backendem nástroje, nastavte buď všechny tři úrovně, nebo žádnou.

launcher.commandod v0.1.0

Příkaz, který spustí kódovací nástroj místo jeho samotného binárního souboru — třeba přes ollama launch nebo obalený do caffeinate, aby Mac uprostřed úkolu neusnul. Runner ho kontroluje proti klíči harness, takže nemůže potichu spustit jiný nástroj, než jaký projekt uvedl.

launcher.flagsod v0.1.0

Přepíše jednotlivé přepínače příkazové řádky, které profil předává kódovacímu nástroji; přepínač nastavený na null se vynechá úplně.

launcher.cli_checkod v0.3.24

Ve výchozím stavu zapnuto. Na false ho nastavte jen tehdy, když launcher.command ukazuje na testovací náhražku místo skutečného kódovacího nástroje; vypne kontrolu, že příkaz spouští nástroj, který projekt uvedl.

waiting_strategy.short_wait_minutesod v0.1.0

Jak dlouho today_auto_squash čeká, než se znovu zeptá, když není připravený žádný úkol. Výchozí je 30 minut.

waiting_strategy.long_wait_minutesod v0.1.0

Jak dlouho režim daily čeká, než se znovu zeptá, když není připravený žádný úkol. Výchozí je 60 minut.

skip_list.revive_after_minutesod v0.3.22

Jak dlouho zůstane úkol, který smyčka musela odložit — protože selhal, nebo ho už začal někdo jiný — mimo třídění, než ho smí znovu vybrat. Výchozí je 60 minut; 0 ho vyřadí do konce dne.

skip_list.max_revivalsod v0.3.22

Kolikrát za den se odložený úkol smí vrátit k dalšímu pokusu. Výchozí jsou 2; 0 znamená nikdy.

bug_destination.epic_relative_idod v0.1.0

Epic na mcptask.online, do kterého přicházejí chybové úkoly, jež runner zakládá o vlastních selháních. Bez něj skončí v kořeni projektu.

bug_destination.epic_nameod v0.1.0

Popisek tohoto Epicu, zobrazený jen ve shrnutí při startu.

pr_template.pathod v0.1.0

Šablona pull requestu, kterou má agent dodržet, zadaná relativně ke kořeni projektu. Projekty na GitHubu a GitLabu dostanou bez ní obvyklou cestu svého hostingu; na Bitbucketu je tento klíč jedinou možností, jak šablonu určit.

work_window.end_of_workday_hourod v0.3.8

Hodina, po které režimy today, today_auto_squash a daily nezačnou žádný další úkol; rozpracovaný úkol se ještě dokončí. Bez ní pracovní den nekončí. Zapisuje ji init --until.

work_window.end_of_workday_minuteod v0.3.23

Minuta v rámci té hodiny, aby den mohl skončit třeba v 18:30. Nepovinná, bez ní platí 0; nečitelná hodnota zruší celé okno a runner to při startu oznámí.

story_branches.enabledod v0.3.35

Větve Story na tomto hostiteli: úkoly jedné Story se slučují do společné větve Story a poslední úkol Story přenese celou Story do hlavní větve jediným sloučením. Zapnuto, dokud není nastaveno false, a i host s false se drží větve Story, která už existuje; hodnotu, která není true ani false, runner odmítne, místo aby hádal.

debug.captureod v0.3.24

Zaznamená surový výstup kódovacího nástroje do adresáře s logy projektu, s vymazanými tajnými údaji, kvůli diagnostice běhu. Na průběhu běhu nic nemění a proměnná prostředí ho přebije oběma směry.

Kódovací CLI

Claude Codeod v0.1.0

Kódovací nástroj od Anthropicu. Dostane kompletní sadu skillů včetně tří, které vyhledávání předají samostatnému dílčímu sezení, takže jeho surový výstup nikdy nezaplní hlavní kontext, a přes launcher.command a nastavené modely může běžet proti libovolnému backendu kompatibilnímu s Anthropic API.

Codex CLIod v0.3.24

Kódovací nástroj od OpenAI. Runner před startem ověří, že je přihlášený, a pokud není, běh odmítne — místo aby utratil denní kvótu za nástroj, který by selhal hned při prvním volání.

OpenCodeod v0.3.24

Open-source kódovací nástroj, který bere modely od mnoha poskytovatelů. Nemá přepínač pro režim jen pro čtení, a tak mu runner dodá mapu oprávnění, která stejnou záruku zajistí jinou cestou.

Git hosty

GitHubod v0.3.25

Ovládaný přes nástroj gh a jeho vlastní přihlášení. Šablonu pull requestu čte z místa, kde ji GitHub očekává. Instance GitHub Enterprise je jiný hosting a runner ji jmenovitě odmítne, místo aby ji považoval za github.com.

Bitbucket Cloudod v0.3.25

Ovládaný přes REST API Bitbucket Cloudu, buď přístupovým tokenem pro stroj, nebo e-mailem a API tokenem pro člověka — nikdy obojím a nikdy heslem aplikace. init uloží přihlašovací údaj do souboru, který přečtete jen vy, a nikdy ho nevypíše.

GitLabod v0.3.27

Ovládaný přes nástroj glab, který si drží vlastní přihlášení, a to na gitlab.com i na vlastních instancích; projekt na vlastní instanci uvede gitlab v konfiguraci, protože z jeho adresy to poznat nejde. Runner tam zakládá merge requesty.

Exit kódy

0 — běh proběhlod v0.1.0

Příkaz doběhl do konce. U run to znamená, že smyčka dokončila svou práci, ne že se povedl každý úkol — neúspěšný úkol se ohlásí na mcptask.online a v logu a není důvodem, aby selhal celý proces.

1 — příkaz nemohl udělat svou práciod v0.1.0

Chybělo nebo bylo odmítnuto něco, co příkaz potřebuje — neuvedený kódovací nástroj, chybějící token, nepřihlášený nástroj, neznámý režim, nepovedené stažení — a chybová hláška na stderr to jmenuje. Runner raději nezačne, než aby hodinu nic nedělal a skončil úspěchem.

3 — režim není implementovánod v0.1.0

Vyhrazeno pro režim pracovní smyčky, který binární soubor zná, ale neumí provést. Dnes jsou implementované všechny režimy, takže ho aktuální runner nevrací; naplánovaná úloha ho přesto může brát jako „aktualizujte runner“.

130 — zastaveno klávesami Ctrl-Cod v0.3.0

Runner byl přerušen (SIGINT). Zastaví kódovací nástroj, který spustil, označí sezení na nástěnce jako ukončené a skončí obvyklým kódem shellu 128 plus číslo signálu. Úkol zůstane rozpracovaný a příští běh na něj naváže.

143 — zastaveno systémemod v0.3.0

Runner dostal pokyn k ukončení (SIGTERM), typicky od plánovače nebo při vypínání systému. Zastaví se čistě stejně jako po Ctrl-C, takže zastavený běh se nikdy nezapíše jako úspěšný.

monetization_onBrány kvóty

Jedno číslo, tři místa, kde se ověřuje

Runner nikdy nevěří vlastnímu odhadu. Každé rozhodnutí o kvótě čte živé číslo z mcptask.online přes REST a vynucuje ho před úkolem, mezi úkoly i během úkolu.

sourceOdkud se číslo bere

Denní rozpočet se načítá živě z mcptask.online přes REST (GET /api/{account}/users/current/time_status). Runner nikdy neodhaduje rozpočet přes agenta a nikdy nečte cache – účet je jediný zdroj pravdy.

gpp_goodTři brány, v tomto pořadí

KdyCo se kontroluje
PŘED BĚHEM (po fázi triage)Znovu se dotáže živého REST endpointu po triage. Samotná triage trvá minuty; čerstvý dotaz zachytí cokoliv, co uživatel utratil, zatímco runner úkol třídil.
MEZI ÚKOLY (decider)Zastaví smyčku při zabití kvótou uprostřed úkolu, při vyčerpaném denním rozpočtu nebo když úkol skončí statusem error. Selhaný úkol, úkol mimo rozsah, už rozpracovaný úkol nebo nepovedený merge jde místo toho na seznam přeskočených a smyčka pokračuje dalším úkolem.
BĚHEM ÚKOLU (každých 360 s)Přeptá se každý DefaultQuotaPollInterval. Při překročení zabije potomka a ukončí smyčku s quota_exceeded_mid_task – bez retry. Sám se vrátí jen režim daily, další den a znovu od triáže. Kill streak (DefaultQuotaFailureKillStreak = 3) absorbuje asi 18minutový výpadek RESTu, než to runner vzdá.

tuneZáměrné rozdělení na fail-closed a fail-open

Na dvě otázky se schválně neodpovídá stejně. „Byla kvóta překročena?“ selhává do CLOSED, „Můžu dnes pracovat?“ selhává do OPEN. Jedna chyba zablokuje běh, druhá ho nechá začít.

blockFail closed

Když se ptáme „už jsme dnes utratili rozpočet?“ a REST neumí odpovědět, odpověď je ANO – rozpočet se bere jako vyčerpaný a nová práce se nespustí. Lepší je přeskočit úkol než přečerpat.

check_circleFail open

Když se ptáme „zbývá mi dnes rozpočet?“ a REST neumí odpovědět, odpověď je ANO. Lepší je spustit běh než promarnit pracovní den kvůli přechodnému výpadku.

reportVýpadek vs. vyčerpaný rozpočet

Výpadek RESTu, který zabil zdravou práci, je jiný případ než vyčerpaný denní rozpočet. Runner je záměrně rozlišuje (ErrQuotaPollOutage); vyčerpaný rozpočet není ničí chyba, kdežto výpadek RESTu je 18minutový incident, který si zaslouží bug úkol.

ScénářCo runner udělá
Denní rozpočet vyčerpánDecider mezi úkoly zastaví smyčku. Žádný bug.
Výpadek RESTu (pod kill streakem)Zkusí to znovu v dalším intervalu. Smyčka pokračuje.
Výpadek RESTu (nad kill streakem, ~18 min)Ukončí smyčku fail-closed se statusem error a vlastní terminací quota_poll_outage – ne s quota_exceeded_mid_task. Ta terminace není na seznamu měkkých ukončení, takže bug úkol se založí vždy, aby se ten 18minutový incident sledoval.

--ignore-quota přeskočí všechny tři brány. Runner se pak nikdy neptá na stav odpracovaného času a nikdy neodmítne úkol – jediné volání RESTu, které zbude, je kontrola při startu, kterou zjistí uživatele (/users/current); operátor přebírá odpovědnost za útratu.

memoryPřetečení kontextu

Co přežije, když kontext přeteče

SESSION se nedá obnovit. PRÁCE ano – protože práce je větev a commity na disku. Dva nezávislé mechanismy nesou zbytek dál a třetí verdikt říká harnessu, aby přestal zakládat šum.

menu_bookJak runner přežije přetečení kontextu – technické detailyexpand_more
check_circle

Přežije

Co runner předá dalšímu pokusu

  • account_tree

    Git větev a její commity

    Samotná práce je větev s commity na disku. I když se ztratí každý bajt kontextu session, diff jde pořád projít, PR pořád otevřít a změny pořád obnovit.

  • history

    Poslední 3 akce (restart uvnitř procesu)

    Když je rozpočet čerstvý a proces se restartuje na místě, preamble restartu nese poslední tři akce (RecentActionsCap = 3) a naměřené poznatky o ceně kontextu, takže nový pokus navazuje na rozjetou práci, ne naslepo.

  • rule

    Pravidla ContextBudget

    Každý restart zdědí sdílená pravidla handoff.ContextBudget() – co se počítá do limitu, co ne, a jak runner rozhodne, že je bezpečné spustit další pokus.

block

Ztraceno

Co žádný mechanismus nezachrání

  • chat_bubble_outline

    Celý kontext konverzace

    --continue by znovu načetl stejný přerostlý kontext, takže SESSION je neobnovitelná. V preamble restartu jsou jen poslední tři akce; vše předtím je pryč.

layersDva nezávislé mechanismy

MechanismusRozsahCo se nese dál
In-process fresh restartUvnitř téhož procesu runneru (retry.go, handleContextOverflow)Rozpočet = 1 restart na proces (maxOverflowRestarts = 1, záměrně – po přezkoumání ponecháno na 1). Nový pokus čte poslední 3 akce, poznatky o ceně a pravidla rozpočtu.
Cross-process TaskHandoffinternal/handoff/ – jedna poznámka na úkol (engine.go, recordTaskHandoff)Zapíše log/handoffs/task_<id>.json v OBOU větvích – v restartu i v terminální (terminální konec tohoto procesu, ne úkolu). Další NOVÝ proces ji čte jako preamble promptu pro první pokus – nikdy při --continue retry. overflow_count se sčítá napříč procesy runneru. Poznámky se zahodí, jakmile běh skončí jinak, a mažou se po 30 dnech.

gavelTřetí verdikt: overflow_pr_open

Když je rozpočet restartu vyčerpán A existuje otevřené PR k úkolu, status je overflow_pr_open, NIKOLIV error. PR je výstup; otevřené PR se zelenými checky znamená, že práce prošla. Zakládat tady error stojí celou denní smyčku kvůli úkolu, který už je hotový. Skutečný úkol takhle přišel o hotový commit, o otevřené PR i o zbytek dne – přesně kvůli tomuhle špatnému zařazení, které mu navíc automaticky založilo bug úkol, který si nezasloužil.

overflow_pr_open (ne error)error (co to bývalo)
rozpočet restartu vyčerpán + PR otevřené → overflow_pr_openrozpočet restartu vyčerpán + bez PR → error
memoryPro vývojáře

Proč to není kouzlo

Souběžně běží čtyři věci: potomek ve vlastní process group, čtečka streamu, watchdog a event stream. Stavový automat pod nimi má deset pojmenovaných stavů a explicitní allow-list – zamítnutý přechod je varování, nikdy důvod k zastavení.

menu_bookCo běží souběžně – skupina procesů, hlídač a stavový automatexpand_more
lan

Potomek ve vlastní process group

Runner spouští coding CLI v nové process group (syscall.SysProcAttr{Setpgid: true}, internal/executor/process_unix.go, configureProcessGroup), takže SIGTERM a SIGKILL z runneru zasáhnou celý strom a nikdy neprobublají k rodiči. Čekání na dokončení čteček je omezené na 30 s (engine.go, stderrJoinTimeout + defaultStdoutJoinTimeout), aby zaseknutý potomek nemohl zablokovat vypnutí.

stream

Čtečka streamu

Čte výstup CLI (stream-json) řádek po řádku, jak vzniká, v tom dialektu, který deklaruje profil harnessu. Parsuje eventy, hledá TASKRUNNER_RESULT completion contract a krmí StallDetector. Nečeká na konec procesu – reaguje průběžně na každý řádek, který přistane.

timer

Watchdog

Nezávislé hlídací vlákno s 30sekundovým heartbeatem. Zabití po 20 minutách nečinnosti, měkké varování na zamrzlý proces po 3 minutách, stropy pro zaseknuté nástroje (quick / long), absolutní pojistku a živý REST dotaz na kvótu každých 6 minut. Jediný externí dozorce nad procesem potomka.

wifi

Event stream

Perzistentní WebSocket na mcptask.online přes ActionCable (RunnerSessionChannel). Snapshoty odcházejí nejvýš jednou za ~0,5 s; změnu stavu FSM protlačí okamžitě. Při výpadku následuje asynchronní reconnect nejvýš jednou za 30 s a odkladové okno 0,5 s na finální snímek „closed“, aby živá karta nezamrzla na starém stavu. Proměnná MCPTASK_RUNNER_DISABLE není jen vypínač streamu – je to globální kill switch veškerého provozu na mcptask.online: stream, poll kvót i hlášení chyb (internal/mcptask/endpoint.go, DisableEnv).

hubStavový automat pod tím

schema verze 3

Deset pojmenovaných stavů. Každý přechod je na allow-listu; cokoliv jiného se zapíše jako warning a práce běží dál. Frozen, pending, stalled a closed jsou vždy povolené cíle přechodu z jakéhokoliv stavu; vynucuje je watchdog (frozen, pending), detektor zaseknutí (stalled) a uzavření session ve smyčce (closed).

Stavy

startingtriageprocessingwaitingfinishedstalledfrozenpendingerrorclosed

Allow-list (ukázka)

Pár přechodů, které smyčka používá každou minutu – a vynucené přechody, které odpalují watchdogy samy od sebe.

zdokýmpoznámka
startingtriagesmyčkaPo spawnu, než se vybere první úkol
triageprocessingsmyčkaÚkol zvolen, kontext předán CLI
processingwaitingsmyčkaBackoff mezi pokusy o tentýž úkol
processingstalleddetektor zaseknutíOpakované chyby nástroje Edit, opakované stejné chyby Bashe nebo tentýž podpis nástroje v klouzavém okně
libovolnýfrozenwatchdogNečinnost delší než FROZEN_WARN_THRESHOLD, žádné aktivní nástroje
libovolnýpendingwatchdogJeden nástroj přetáhl svůj varovný strop
libovolnýclosedsmyčkaend_session – finální snímek
visibilityPozorovatelnost

Co je vidět za běhu

Tři nezávislé plochy – run log na disku, provozní log procesu a živá karta přes ActionCable – plus explicitní opt-out proměnné. Ze své podstaty fungují bez záruky doručení: samotné pozorování – snapshoty a logy – běh nikdy nepřeruší.

menu_bookCo runner zapisuje do logů a co ukazuje živá karta – technické detailyexpand_more
data_object

Run log – jeden JSON soubor na spuštění procesu potomka

Nejužitečnější plocha ve chvíli, kdy otázka zní „proč už runner 90 minut visí“. Soubor se otevře IHNED při spawnu, takže i okamžitý pád zanechá něco čitelného.

folderKde
log/runs/run_*.json (jeden soubor na spuštění procesu potomka)
  • lock_open

    Otevřen

    Při spawnu – dřív, než se přečte první řádek stream-json

  • favorite

    Heartbeat

    Obnovuje stream_events, inactive_s a stream_quiet_s při každém ticku, takže zaseknutý běh zanechá svůj rozpracovaný stav na disku

  • task_alt

    Finalizace

    Orazítkuje, proč pokus skončil

Změní „proč runner visí už 90 minut“ z grepu nad 268 000 řádky stream logu na přečtení jednoho souboru.

description

Provozní log procesu – širší než run log

Není to lidsky formátovaná kopie JSON run logu. Je to chronologický textový log celého procesu: každý podsystém do něj píše svá volání Debug/Info/Warn/Error, takže obsahuje i to, co se v žádném run logu neobjeví – a naopak nenese pole run logu jako session_id nebo stream_events.

folderKde
log/mcptask_runner_YYYYMMDD_HHMMSS.log
short_textTvar řádku
[timestamp] SEVERITY - message (internal/observe/logger.go)
wifi

Živá karta – jeden ActionCable WebSocket

Jeden perzistentní WebSocket nese celé snapshoty, nikdy jednotlivé události. Ztracený frame stojí čerstvost, ne správnost, a reconnect nepotřebuje replay.

KanálPayloadOmezeníTimeouty
RunnerSessionChannel (internal/eventstream/eventstream.go, channelIdentifier)Pouze celé snapshoty – žádný replay událostísnapshot nejvýš jednou za 500 ms, reconnect nejvýš jednou za 30 s, 500 ms odklad finálního snímku10 s subscribe + dial

Ukončená karta: Zůstane 60 s po konci smyčky (snapshotCloseTTL)

Bez tokenu: Runner spuštěný bez tokenu to řekne jednou, nahlas, při startu – nikdy tiše nepropásne první událost

shield

Hranice – a opt-outy

Každá z těchto ploch je ze své podstaty bez záruky doručení. Odchozí pozorování běh nikdy nepřeruší; nil *Log je fungující no-op. Dvě proměnné prostředí ukládání čistě vypnou.

Run logTask handoff
MCPTASK_RUN_LOG=0 – Vypne JSON run logMCPTASK_TASK_HANDOFF=0 – Vypne poznámku task-handoff mezi procesy

Jediná výjimka – přeřazení: Kanál RunnerSessionChannel je obousměrný: kromě odchozích snapshotů přijímá i řídicí události piece_assigned a set_aside_cleared (internal/eventstream/eventstream.go). Přeřazení úkolu je přes abandonReassigned (internal/runner/loop.go) zapojené do AbandonReassigned v enginu – zabije běžícího potomka a jeho smrt překlasifikuje na přeřazení. Je to jediná, záměrná řídicí role tohoto kanálu.

Co to NEZNAMENÁ: Žádná z pozorovacích ploch – run log, provozní log, odchozí snapshoty – nemůže běh zastavit, restartovat ani klasifikovat. Chybějící záruka doručení je přesná hranice premisy „žádné fallbacky“: je to jediné místo, kde runner ustoupí – a nahlas to říká.

Flotila runnerů: Všechny instance runneru v účtu jsou navíc na stránce Flotila runnerů (uživatelské menu): jeden řádek na stroj a projekt s CLI, modelem, stavem, posledním nahlášeným úkolem, dnešními hodinami proti dennímu limitu a verzí runneru, živě aktualizováno. Manažeři účtu a vlastníci firmy vidí všechny runnery, členové projektu runnery svých projektů.

bug_reportHlášení bugů

Když nastane tvrdý pád, bug úkol se založí sám

Žádný manuální příkaz spouštět nemusíte. Když nastane tvrdý pád, runner za něj automaticky založí bug úkol. Skutečný pád se stane úkolem, který někdo může zvednout – ne tichou mezerou, kterou nikdo nehledá.

menu_bookKdy runner sám založí chybový úkol, co k němu přiloží a jak se vyhne duplicitámexpand_more
rule

Kdy se bug úkol založí

Za tvrdý pád se počítá jen status error, status crash nebo status anomaly. To jsou tři hodnoty v hardStatuses (internal/bugreport/runner_error.go). Každý jiný výsledek – success, stalled_for_genius, urgent_bug_pending nebo graceful quota bail – znamená, že runner pracuje, jak má, a nezakládá nic.

check_circle

Založí se

status error, status crash, status anomaly

Co je anomaly

vlastní pojistka runneru, která v běhu neplatila, a běh přesto pokračoval – nic nespadlo, denní práce jela dál a zvenčí není co vidět. Error se ohlásí tím, že ukončí běh, a crash tím, že zabije proces; anomaly se neohlásí nikomu, a právě proto dostane bug úkol.

horizontal_rule

NEzakládá se

success, stalled_for_genius, urgent_bug_pending, graceful quota bail – to jsou verdikty „pracuje, jak má“

priority_high

Měkká výjimka

zabití kvótou uprostřed úkolu, vyčerpané rate-limit okno, vyčerpaný limit používání CLI i odebrání úkolu během běhu přicházejí se statusem error a žádný bug úkol nezaloží (softTerminations = {"quota", "rate_limited", "usage_limit", "reassigned"}, internal/bugreport/runner_error.go)

attach_file

Co bug úkol nese s sebou

K založenému bug úkolu se přiloží tři artefakty (internal/bugreport/runner_error.go, attachArtifacts), aby ten, kdo ho zvedne, mohl číst pád bez nového spuštění.

description

Run log selhaného pokusu

JSON, se kterým run log na disku skončil

terminal

Konec stream logu

nejnovější stream log, oříznutý na 512 KB (internal/bugreport/runner_error.go, streamTailBytes)

visibility_off

Redigované configy

podle harnessu – .mcp.json, .claude/settings.json a .claude/settings.local.json na Claude Code, .codex/config.toml na Codex CLI, opencode.json na OpenCode – s vyříznutými tokeny (mcptask/piece.go, ConfigFiles + Redact)

fingerprint

Jedna událost je jeden bug úkol

Identické pády dostanou otisk, projdou omezením četnosti a nahlásí se jen jednou. Bez toho by těsná smyčka založila stejný bug 200× dřív, než si toho někdo všimne.

FingerprintNormalizaceOkno omezeníRazítko se nárokuje před založením
sha256 nad termination + normalizovaná zpráva + task id + projekt, oříznutý na 16 hex znaků (internal/bugreport/runner_error.go, fingerprint)odstraňuje cesty, hex řetězce a desetinná čísla – co vypadá jako stejný pád jen s jiným časovým razítkem, se počítá jako stejný pádšest hodin (internal/bugreport/runner_error.go, throttleWindow + claim) – druhý stejný fingerprint v okně se zahodírazítko s fingerprintem se nárokuje PŘED založením úkolu a UVOLNÍ se, pokud se úkol nikdy nezaloží. Pád, který nejpravděpodobněji rozbije CreatePiece, je přesně ten, který by retry jinak tiše potlačil
warning

Paniky – vyhozené znovu, se stack trace

Panic je tvrdý pád se stopou pro forenzní analýzu. Bug úkol dostane trace; smyčka si panic ponechá (loop.go, reportPanic). Jinak by zotavení z paniky trace spolklo a ten, kdo úkol zvedne, by o něj přišel.

shield

Fail-safe všude

Každá reportovací cesta polyká vlastní chyby, včetně vlastní paniky (internal/bugreport/runner_error.go, MaybeReport defer/recover). Chyba reportéra nikdy nesmí ukončit běh. Celá premisa stojí na tom, že nejhorší reportér je ten, který stojí nejméně – spolknutý řádek v logu, ne promeškaný pád.

visibility_off

Slepé místo, otevřeně řečeno

Reportér se autentizuje stejným tokenem jako vše ostatní, takže NEMŮŽE založit bug, který říká, že token chybí. Chybějící nebo nenastavený token reportéra tiše přeskočí – OutcomeSkippedDisabled nezapíše žádný chybový řádek, takže to nese jen vlastní exit kód a log runneru. Token, který server odmítne, je naopak hlasitá cesta: OutcomeRefused zaloguje chybu, která token pojmenuje (internal/bugreport/runner_error.go). Alternativou bylo mezeru skrýt; raději ji pojmenováváme.

swap_horizPro vývojáře

Model-agnostický přes konfiguraci

Tři úrovně si namapujete na libovolného poskytovatele i na libovolné ze tří coding CLI. Claude Code nasměrujete na jakýkoli endpoint kompatibilní s Anthropic API (doloženým příkladem je lokální Ollama) změnou launcheru a modelové sekce. OpenCode bere ID ve tvaru provider/model od libovolného poskytovatele, Codex CLI má vlastní výchozí modely a `ollama launch <cli>` funguje pro všechna tři.

menu_bookUkázky konfigurace modelů v YAML – technické detailyexpand_more

Sekce models: je volitelná. Bez ní každý harness používá své vlastní obecné aliasy, které si dané CLI přeloží za běhu: opus / sonnet / haiku na Claude Code, gpt-6-astra / gpt-5.6-terra / gpt-5.6-luna na Codex CLI a ID ve tvaru provider/model na OpenCode. Konkrétní ID připínejte jen tehdy, když chcete deterministické retry, nebo když runner provozujete přes jiný backend než Anthropic – a jakmile to uděláte, nastavte všechny tři úrovně genius / smart / primitive. Runner neúplnou sekci models: neodmítne, jenže práce, kterou CLI dělá samo na pozadí, pak spadne s chybou „model may not exist“.

terminalVýchozí nastavení podle harnessu
# models: je volitelná. Bez ní každý harness použije obecné aliasy,
# které si jeho vlastní CLI přeloží za běhu.

# Claude Code
models:
  genius:    opus
  smart:     sonnet
  primitive: haiku

# Codex CLI
models:
  genius:    gpt-6-astra
  smart:     gpt-5.6-terra
  primitive: gpt-5.6-luna

# OpenCode – každé ID je provider/model
models:
  genius:    <provider>/<nejsilnější>
  smart:     <provider>/<střední>
  primitive: <provider>/<nejrychlejší>
terminalPříklad: jiný backend než Anthropic
models:
  genius:    minimax-m3:cloud
  smart:     kimi-k2.7-code:cloud
  primitive: deepseek-v4-flash:cloud

launcher:
  # Lokální Ollama
  command: [env, "ANTHROPIC_BASE_URL=http://localhost:11434", "ANTHROPIC_AUTH_TOKEN=ollama", "claude"]

ID níže jsou jeden příklad, ne doporučení – stárnou nejrychleji ze všeho na téhle stránce. Důležitý je tvar.

Flagy v launcher.command (launcher.go, Flag* constants): null flag se z argv VYNECHÁ, value flag bez tokenu se předá POZIČNĚ.

auto_fix_highLoader nikdy neselže

Chybějící, prázdný nebo poškozený config/mcptask_runner.yml se bere jako „žádná konfigurace“ (config.go, LoadFile) a použijí se vestavěné defaulty – samotný loader nikdy neselže. Předstartovní kontrola run je ale samostatná a start odmítne: když není nastavený harness:, když chybí token nebo když projekt nemá CLAUDE.md.

check_circleŽádný pin

Žádný dodávaný skill nedeklaruje `model:` – ForkModelEnv proto dosáhne na každého forkovaného subagenta, včetně těch, kteří běží na jiném launcheru než Anthropic.

info

Konfigurace žije v config/mcptask_runner.yml a hledá se relativně k pracovnímu adresáři launcheru (config.go, FileName; README.md) – je to stejný soubor, jaký čte Ruby gem, se stejnou sémantikou. Můžete tam nastavit i další volby, například cílový Epic pro automaticky nalezené bugy nebo prodlevy mezi běhy.

GEM – přesná cesta k souboru, přesná sémantika

system_update_altSamoaktualizace

update --self a záměrně chybějící automatika

Runner je statická binárka instalovaná stažením, ne gem, který se posune při každém bundle. Binárkou hýbou dva příkazy: update --self stáhne novější vydání a vymění za něj tu stávající, update --check oznámí, jestli nějaký existuje, a nic víc neudělá. Třetí rozhodnutí – jestli se runner vůbec smí povýšit sám před naplánovaným během – je to, které záměrně neautomatizujeme.

menu_bookJak runner vymění sám sebe za novou verzi – technické detailyexpand_more
swap_horiz

Jak update --self vymění binárku

Nová binárka přistane ve stejném adresáři jako ta stávající a přejmenování ji posune na místo (selfupdate.go). Ten postup je důležitější, než vypadá: přejmenování je atomické jen v rámci JEDNOHO souborového systému a /tmp jím obvykle není, takže se nový soubor připraví vedle, nepřepisuje se z jiného místa. Když stažení selže nebo nesedí kontrolní součet, k výměně vůbec nedojde – binárka na disku zůstane nedotčená a příkaz skončí nenulově.

file_download

1. Stažení z veřejného zrcadlového repozitáře s vydáními

binárka pro dané GOOS/GOARCH se stáhne z jchsoft/mcptask-releases (selfupdate.go, DefaultRepo). Zdrojový repozitář je soukromý, takže veřejný zrcadlový repozitář je jediné místo, odkud může stahování přijít bez tokenu

verified

2. Ověření proti checksums.txt vydání

sha256 assetu se porovná s hodnotou, kterou zveřejnilo vydání (selfupdate.go, verify). Neshoda zruší běh dřív, než se do instalačního adresáře vůbec něco zapíše

file_copy

3. Připravit vedle, běžící binárku přejmenovat stranou

nová binárka se zapíše vedle běžící, pak se běžící přejmenuje na sousední název. Tenhle krok je zároveň to, díky čemu celé funguje na Windows – přepsat soubor, který je namapovaný pro spouštění, na Windows nejde, a přejmenování stranou to obchází

published_with_changes

4. Připravenou binárku atomicky přejmenovat na místo

závěrečné os.Rename na stejném souborovém systému přesune připravenou binárku do instalační cesty. Atomické v rámci toho souborového systému. Příkaz skončí nulou až po výměně

visibility

update --check – jen dotaz na tag, žádné stažení

update --check zjistí tag nejnovějšího vydání, porovná ho s běžící verzí, oznámí výsledek a skončí. Nestahuje žádný asset a neověřuje žádný kontrolní součet – ani jeden ze čtyř kroků výše se neprovede. Binárka na disku zůstane nedotčená. Je to správný příkaz, pokud si chcete do nočního jobu přidat upozornění bez instalace.

shield

Záměrně chybějící funkce

Runner se před naplánovaným během NEaktualizuje sám. Pokud existuje novější vydání, vypíše jednořádkové upozornění a nic jiného neudělá, protože jedno špatné vydání nesmí v 08:05 zasáhnout každého hostitele bez obsluhy. Upgrade zůstává rozhodnutím operátora, i když necháváte binárku běžet bez dozoru.

notifications

Co upozornění dělá

selfupdate.Hint vypíše jeden řádek, když existuje novější vydání (hint.go, hintTimeout + hintInterval). Limit 2 s, cache 24 h, ignoruje každou chybu – výpadek sítě ani rate-limit se k běhu nikdy nedostanou

block

Jak upozornění vypnout

nastavte MCPTASK_NO_UPDATE_CHECK=1 v prostředí hostitele. Hodí se to hostiteli v izolované síti nebo s měřeným připojením; kontrola se umlčí dřív, než se vůbec sáhne na cache

inventory_2

Jedna ze tří cest, kdy se runner znovu spustí – bundle adoption

BUNDLE ADOPTION: u run, pokud Gemfile.lock projektu sám jmenuje novější verzi wrapper-gemu, než je běžící binárka – pak se run na té binárce znovu spustí. NIC NESTAHUJE – operátor to rozhodl už tím, že ten lockfile zmergoval. Hostitel v 08:05 čte lock, který si stáhl z mainu, a běží na tom, co lock říká. Třetí cesta leží mezi úkoly: iterující režim před dalším úkolem převezme nově nainstalovanou binárku – včetně té, kterou update --self uložil do ~/.mcptask/bin – a na téže hranici znovu zkontroluje lockfile kvůli adopci.

repeat

MCPTASK_ADOPTED

nastavený v prostředí znovu spuštěného potomka zastaví adoption, aby se z ní nestala nekonečná smyčka execů (adopt.go, EnvMarker). Binárka, která hlásí starší verzi, než říká lock, by jinak adoptovala, spustila se execem, znovu se shledala zastaralou a zacyklila se

push_pin

MCPTASK_NO_ADOPT

nastavte na 1, ať zůstane nainstalovaná binárka připnutá bez ohledu na to, co nesou projekty. Hostitel, který chce jednu verzi runneru napříč všemi projekty, to nastaví jednou

bedtime

Další cesta – daily se před noční pauzou předá novější binárce

Cestou do noční pauzy daily běh zkontroluje, jestli je binárka na disku pořád ta, se kterou vyběhl. Pokud ne, uvolní instance lock a execem spustí tu nainstalovanou – restartOnCurrentBinary (restart.go). Ani tahle výměna NIC NESTAHUJE – jen přebírá to, co na disk už dřív dal update --self nebo operátor. Na pořadí kroků záleží: denní práce je hotová, takže neběží žádný potomek a žádný pracovní strom není rozepsaný, a proces, který se zítra probudí, je ten, co dnes večer vyběhl.

help_outline

Proč to existuje

všechno, co novou verzi zaznamená, se odehraje jen jednou, při startu. Daily proces se spustí jednou a nikdy neskončí, takže by hostitel bez obsluhy jinak držel binárku, se kterou náhodou vyběhl, ať vyjde jakkoli mnoho vydání – přesně ten hostitel, na kterého se nikdo nedívá. Zvažovalo se místo toho daily ukončovat, a bylo to zamítnuto: runner, který zmizí, se přes chybějící nebo nenačtený naplánovaný job už nikdy nevrátí

lock_open

ErrLockLost – třetí konec

restart.go deklaruje ErrLockLost – výměnu, která se vzdala instance locku a pak si ho nedokázala vzít zpět. Tenhle konec smyčku SKUTEČNĚ ukončí, záměrně a nahlas – pokračovat bez locku by druhému naplánovanému startu dovolilo pustit do stejného checkoutu druhý runner

autorenewÚdržba

Aktualizace skillů bez ztráty vašich úprav

Nainstalovaný projekt se rozchází s runnerem, jak runner přibírá skilly. Jeden příkaz ho sesynchronizuje: mcptask_runner update. Porovná každý skill a helper proti instalačnímu manifestu, u každého oznámí, co s ním udělal, a odmítne přepsat cokoli, co jste si upravili sami. Je to taky způsob, jak nainstalovaný projekt přestane volat `gh`: přibalené skilly přešly na `mcptask_runner pr` a právě `update` tuhle migraci donese do projektu nainstalovaného předtím.

Příkaz

mcptask_runner update
menu_bookCo příkaz update udělá s vaším projektem – pět výsledků a --forceexpand_more
fact_check

Pět výsledků, jeden na soubor

Updater si drží obsahový hash každého souboru, který nainstaloval. Každý skill a helper proti tomuto manifestu zařadí a výsledek vypíše – pět výsledků, ne tiché přepsání.

add_circle_outlineadded

Soubor na disku není. Updater ho zapíše. Takhle přistávají nové schopnosti runneru, aniž byste si o ně museli říct jménem.

check_circle_outlineup-to-date

Soubor na disku má stejný hash jako záznam v manifestu. Nic se nezapisuje. Hostitel nainstalovaný přes gem se této binárce jeví jako aktuální – ověřeno přímo proti souboru, ne odvozeno z instalačního kanálu.

syncupdated

Soubor odpovídá staršímu záznamu v manifestu, takže je to dodávaná verze a vy jste se jí nikdy nedotkli. Je bezpečné ho nahradit – a nahradí se.

blockconflict-skipped

Hash souboru neodpovídá ani aktuální, ani žádné známé dodávané verzi – upravili jste ho. Updater ho nechá přesně tak, jak je, a řekne to. Tohle je výchozí chování a právě proto je bezpečné příkaz spustit na projektu, který jste si přizpůsobili.

restore_pageforce-updated

Tentýž konflikt, vyřešený opačně, protože jste o to požádali. Dosažitelné jen s --force nebo FORCE=1.

restore_page

Přepsání konfliktu s cestou zpět

mcptask_runner update --force

--force (nebo FORCE=1) lokálně upravené soubory místo přeskočení přepíše – a ke každému přepsanému nechá .bak. Vaše úpravy se nemažou; odsunou se stranou, kde si je můžete porovnat.

shield

Čeho se prosté update záměrně nedotkne

update sahá jen na skilly a helpery – a na jednu zastaralou položku v .mcp.json, viz níže. Všechno ostatní, co instalátor kdysi zapsal, zůstává přesně tam, kde je. Ta zdrženlivost je přednost, ne opomenutí: sesynchronizování skillů nesmí potichu znovu aktivovat plánovanou úlohu ani přepsat konfigurační soubor, který jste si ručně doladili.

schedule

Plánovaná úloha

LaunchAgent, uživatelský systemd timer ani úloha v Task Scheduleru se negenerují znovu a znovu se nezapínají. Když ji chcete opravdu vygenerovat znovu, použijte init --schedule.

data_object

.mcp.json

Vaše konfigurace MCP klienta se znovu negeneruje. Název tokenové proměnné, který jste deklarovali, aktualizaci přežije nedotčený; jediná změna, kterou update udělá, je převod staré položky mcptask-online z transportu SSE na HTTP, kterým server mluví dnes.

tune

Konfigurační sekce

config/mcptask_runner.yml si ponechá hodnoty, které jste nastavili – pracovní okno, příkaz launcheru, hodinu konce pracovního dne. update na ně nemá názor.

system_update_alt

Aktualizace binárky je samostatné rozhodnutí

update sesynchronizuje projekt. Nehýbe binárkou, která tu synchronizaci provedla. To dělají dva jiné příkazy a ani jeden z nich se nespustí sám: mcptask_runner update --self vymění binárku za novější vydání a mcptask_runner update --check jen oznámí, jestli nějaký existuje, a nic nezmění.

person

Runner se před během nikdy neaktualizuje sám

Na začátku běhu runner může vypsat jeden řádek o tom, že existuje novější vydání. To je celá ta funkce – strop 2 s, cache na 24 h, každá chyba ignorována a úplně se umlčí přes MCPTASK_NO_UPDATE_CHECK=1. Rozhodnutí o povýšení zůstává na operátorovi, protože runner, který by se před plánovaným během posunul sám, by bez zeptání měnil to, co tu práci dělá. To je podstatné, pokud necháváte binárku běžet bez dozoru.

verified

Vadné stažení se k výměně nikdy nedostane

update --self stahuje z veřejného zrcadlového repozitáře jchsoft/mcptask-releases a ověřuje soubor proti checksums.txt daného vydání. Neúspěšné nebo nesouhlasící stažení skončí dřív, než jakýkoli zápis sáhne do instalačního adresáře: binárka na disku zůstává nedotčená a příkaz skončí nenulovým návratovým kódem.

verifiedDůkaz

Konformanční scénáře a matice CI

Důkaz za každým tvrzením o spolehlivosti na této stránce: obnova, watchdog, detekce zaseknutí i kvóta mají každá vlastní zaznamenaný scénář.

playlist_play

Čtrnáct konformančních scénářů

Každý scénář je zaznamenaný rozhovor s claude CLI plus chování, které musí runner při přehrávání předvést. Scénáře jsou implementačně neutrální: čisté YAML plus kontrakt v conformance/README.md, ani řádek Go nebo Ruby. Spouští se přes `go run ./cmd/conformance {list,run --impl go}`.

Každý režim pádu, o kterém runner tvrdí, že ho přežije, má vlastní zaznamenaný scénář v conformance/scenarios/. Obě implementace spouští stejné soubory; Ruby gem je vysloužilá, zmrazená referenční implementace, proti které se dá suite porovnat, a bránou je zelená Go suite.

menu_bookVšech 14 scénářů shody, kterými runner musí projítexpand_more
Scénáře 
01 – successšťastná cesta: čistý běh skončí zeleným verdiktem, bez retry, a větev je připravená k odeslání
02 – missing_marker_retryrun log nikdy nezafixuje marker; runner to zkouší znovu, dokud mu nedojde rozpočet, a pak skončí verdiktem místo zaseknutí
03 – context_overflow_fresh_restartpřetečení kontextu, když engine ještě běží – práci nese dál in-process restart a běh zůstane naživu
04 – context_overflow_terminalpřetečení kontextu poté, co engine skončil – carry-forward mechanismus předá soubor novému procesu a jedině tak se běh dostane ven
05 – api_overload_529upstream vrátí 529 – runner couvne, běh zůstane na správném úkolu a marker na přechodném stavu nepostoupí
06 – tool_not_enabled_fresh_restartnástroj, který běh potřebuje, chybí na prvním pokusu; `--continue` ho spustí jako fresh restart a běh se zotaví
07 – tool_not_enabled_terminaltentýž nástroj chybí i po fresh restartu – běh skončí verdiktem, ne zaseknutím
08 – stall_edit_failuresengine edituje soubor, který harness sleduje; edity se rozcházejí a watchdog to chytí dřív, než marker postoupí na základě lži
09 – inactivity_kill_then_recoverproces potomka přestane komunikovat – watchdog ho zabije, debug dump zůstane na disku a další pokus zvedne tentýž úkol
10 – recoverable_retries_exhaustedvyčerpají se všechny zotavitelné pokusy a běh pořád není hotový – to je verdikt, a verdikt je jediný bezpečný výsledek
11 – quota_mid_taskrozpočet se vyčerpá uprostřed editu – runner zabije potomka, pokus skončí jako quota_exceeded_mid_task bez retry a smyčka končí; jen režim daily začne další den znovu od triáže
12 – stream_closedupstream zavře stream v půlce tahu – runner zavření detekuje, běh vyhodnotí jako neúplný a další pokus odvodí stav z disku, ne ze streamu, který zmizel
13 – hung_tool_killnástroj, který engine zavolal, přestal odpovídat – watchdog zabije potomka signálem kill-on-hang a běh jde spustit znovu
14 – triage_unverified_picknext-task picker vrátí úkol, který nebyl ověřený – runner ho odmítne nastartovat, nezaloží nic a počká na ověřeného kandidáta, místo aby slepě něco vybral
science

Binárka nenese žádný vlastní test hook

Runner je nasměrovaný na mock CLI přes běžný override `launcher.command` (README.md) – je to stejná konfigurace na hostiteli, jakou vývojář použije pro řízení skutečného CLI. Suity běží ve všech třech dialektech harnessu (Claude Code, Codex CLI, OpenCode) a konformanční projekt na Bitbucket Cloudu a jeden na gitlab.com pokrývají zbylé dva hostitele PR, takže zelený výsledek platí pro všechny tři dialekty i všechny tři hostitele. V binárce není žádná test-only code path, o kterou by se konformanční suite opírala, takže zelený běh tady je stejný kód, jaký si uživatel nainstaluje.

grid_view

Matice CI

Každý commit projde maticí tří OS × dvou verzí Go (ubuntu / macos / windows × 1.25.x / stable), dále go vet, race-detector testy na Unixu + stable, bin/smoke na SESTAVENÉ binárce, gofmt, golangci-lint, kontrolou goreleaser configu plus cross-compilem na šest cílů a celou konformanční suitou (.github/workflows/ci.yml). Matice existuje, protože runner se instaluje na vývojářské stroje – každá desktopová platforma musí projít buildem i testy.

fact_check

Co lokální běh nepokryje

Zelený LOKÁLNÍ běh vypíše, co NEpokryl – „Tenhle stroj jen: <goos>/<goarch>“ – protože jeden OS, jedna architektura a jedna verze Go není to, co běží v CI (bin/ci). Zelená lokálně je nutná, ne dostačující. Porovnání proti Ruby je referenční kontrola, ne brána: když vysloužilý gem není checkoutnutý, přeskočí se, neselže – stroj, který se k referenci nedostane, by měl umět zkontrolovat aspoň svou vlastní práci.

visibility_off

Hranice, pojmenovaná

Konformanční harness spouští jeden proces runneru na scénář, takže nedokáže pokrýt tu půlku přežití přetečení, která jde napříč procesy; to pokrývají testy, které pustí dva enginy za sebou a sdílejí jen soubor (README.md). Porovnání proti Ruby se přeskočí, neselže, když vysloužilý gem není checkoutnutý, a merge tak jako tak nedrží.

rocket_launch

Ráno zapněte stroj

Jeden statický binární soubor, jeden instalační řádek, a runner začne brát úkoly z vaší fronty.

verified_user30 dní zdarma. Bez kreditní karty. Kdykoli zrušíte.