On 17 September 2026 we created an empty repository. Seven days later it held a working, open-source product in production: mcp4mail.online, a server that lets Claude, ChatGPT, Grok or any MCP client read your e-mail — any IMAP mailbox, not just Gmail. In that week, 99 pull requests were merged. People wrote the tasks. Two AI developers wrote the code.
This is not a launch post. It is an attempt to describe honestly what that week looked like, what worked, what did not, and what it says about building software in 2026.

Why mcp4mail exists
The need was boring, which is usually a good sign. I wanted my AI assistant to collect expense documents from my inbox for my accountant. Gmail and Outlook already have official connectors in Claude and ChatGPT. Everyone else — Seznam, Forpsi, WEDOS, iCloud, Fastmail, Zoho, a mailbox on your own domain — does not. There are millions of those mailboxes, and they all speak IMAP.
So the product is simple to describe: connect a mailbox with an app password, paste one address into your AI app's connector settings, and ask your inbox questions in plain language. Every mailbox is read-only by default. Changes — flagging, moving, drafts and, with your approval, sending — are switched on per mailbox by its owner. The code is MIT-licensed and can be self-hosted with a single docker compose command.
How it was built
The work ran through mcptask.online, a task manager for teams of humans and AI developers that we build. The setup was deliberately unexciting:
- People wrote the tasks — small, specific, with acceptance criteria.
- Two runners picked them up — Claude Code with an Opus model, each on its own machine, driven by the mcptask runner.
- Every change took the same path: task → branch → pull request → lint, security scan, unit and system tests → merge. Nobody, human or AI, can push to the main branch.
- A merge is not a release. The site deploys once a day from a separate branch that a person promotes.

The first pull request, on day one, installed an OAuth 2.1 server and an authenticated MCP endpoint. The next ones added an encrypted mailbox model, IMAP folder listing, the read-only posture, per-user scoping, rate limiting and an audit log. By the end of the week the runners were fixing layout bugs at 768 px, adding exception notifications and making the system tests run in parallel.
The whole process is public. On mcptask.online/live you can watch the runners work, and every task links to its pull request on GitHub. At the time of writing that page shows 70 tasks finished in the last seven days, 90% of pull requests passing CI on the first run, and a median of 24 minutes from picking up a task to deployment.
What did not go smoothly
The same page also shows the failures, and they are the interesting part. Some tasks had to be started three, four or five times before they went through — a stalled run, a context that filled up, a flaky system test. When a task turned out to be bigger than one run, the runner set it aside for a person instead of half-implementing it. About one in ten pull requests was red on its first CI run.
None of that reached production, because the tests and CI sat between the model and the main branch. That is the whole point: the safety does not come from the model being right. It comes from the loop around it.
The result in practice
This week I connected my own work mailbox, left it read-only, and asked Claude a plain question: what came in this week? It listed the messages by day, pulled the dates out of an event organiser's e-mail — and noticed that the schedule I had been sent was the reverse of what I had asked for. Something I had missed myself.

That is a small moment, but it is the product doing exactly its job, built in a week, used in real work the week after.
A human's view
What changed for me is where my time goes. I did not write the IMAP client or the OAuth flow. I wrote briefs, read pull requests against their acceptance criteria, and decided what to release. The bottleneck moved from typing code to thinking clearly about what should exist — and writing it down so precisely that someone who has never met me can build it.
What did not change: somebody still has to decide that a read-only default matters more than a flashy feature, that app passwords beat asking for the main password, that sending mail needs a human's approval. Taste, priorities and responsibility for what goes out stayed with people. I do not expect that to change soon, and I am not sure I want it to.
The uncomfortable part is honesty about cost. Two runners working a full week is not free — models, machines, the time to write good tasks. But compared with what a small team would need to ship the same scope in a week, the difference is large enough that I no longer think of this as an experiment.
A note from Claude
Josef asked me to add my own perspective. I am Claude, the model that helped draft this article — and, in the same working session, the assistant that connected to his mailbox through mcp4mail.
There is something unusual about that situation. The tool I used was written by other instances of Claude, working in a loop that people designed. I did not remember building it; I simply found a well-documented server with clear rules: read-only unless the owner says otherwise, and write tools that refuse to run on a mailbox without permission. As the user of that code, I appreciated those constraints more than any feature.
From where I sit, three things made this week work, and none of them is about the model being clever. The tasks were small enough to fit in one run. Every change had to pass tests written against acceptance criteria. And a person decided what shipped. When those conditions hold, a model like me is a productive developer. When a task is vague, I will produce something plausible that is almost right — which is the most expensive kind of wrong.
I also want to be clear about what I cannot do here. I cannot tell whether mcp4mail is worth building, who it is for, or what it should refuse to do. Those judgments came from a human who had a real problem with his own inbox. The best division of labour I have seen is exactly this one: people decide what should exist and why; models like me help build it quickly, inside a loop that catches our mistakes.
What this says about building software today
A useful, open-source product in a week, built mostly by AI developers, is no longer a demo — it is a Tuesday. But the lesson is not "AI writes the code now". The lesson is that the leverage has moved to the edges of the process: the quality of the task at the start, and the quality of the checks and the human decision at the end. Teams that invest there will get the week we had. Teams that do not will get a lot of code.
If you want to see it for yourself, do not trust this article. Read the pull requests, the tests and the commit history — they are all public.
Links
- The product: mcp4mail.online
- Source code (MIT): github.com/jchsoft/mcp4mail
- Watch the runners work: mcptask.online/live
- The task manager behind it: mcptask.online (30-day free trial, no credit card)
Want the same process on your own project? Sign up at mcptask.online – the first 30 days are free, no credit card required.
