Last Wednesday I was breaking down a cross-file refactor. The task split into five parts, three terminals open at once: Claude writing the shared types, Codex changing the call sites, omp running the dependency scan.
Forty minutes later I was tracking in a notes app who had finished. Claude on the left was stuck at a permission prompt, Codex in the middle was scrolling too fast to read, omp on the right had finished ages ago and nobody told me. Token spend got reconciled after the fact. By the time I wrapped up that night, two of the three sessions were dead, context gone.
That wasn’t “the model isn’t strong enough”. That was nobody doing the scheduling.
The agent-mcp README is clear: the point isn’t “open more agents” — it’s turning any CLI into a dispatchable, monitorable, resumable, killable work pool.
The main agent does exactly two things: split and merge.
Dispatch, waiting, steering, timeouts, queueing, resume, and downgrade all go to the control plane. Model inference still happens in each CLI’s native runtime; agent-mcp doesn’t rewrite the agent loop or lock you to a single model.
The other principle: match the model to the task.
Read-heavy exploration goes to fast models (omp / pi / grok); deep-reasoning planning goes to strong ones (claude). Cost and quality get matched on the spot, not “one model for everything”.
Problem
Opening more terminals is easy. The hard part:
- What if a subtask runs wild?
- How do you queue when all slots are full?
- How do you steer when you realize the direction is wrong mid-way?
- The main agent can’t idle while waiting for results — but dumb polling would blow up the context.
- Who’s running, where, and how much is burning — all tracked by a human.
And the CLIs don’t speak to each other: some emit JSONL, some only plain text, usage field names differ completely, and resume support is hit or miss. Ignore those differences and multi-agent just amplifies the chaos.
Approach
From the outside, agent-mcp is a set of MCP tools. Any MCP-capable host can plug them in. The ones I use most:
estimate_complexity: grades S/M/L locally, zero tokens, no spawn — do it directly by default, split only when neededspawn_agent: dispatches a sub-agent, takes CLI / model / timeout, hands you anagent_idimmediatelywait_agent: blocks briefly until termination, returns a summarysteer_agent: steers mid-flight — kills the current run, continues from the same nodefollowup_task: merges pending messages, triggers the next turnorchestrate_task: task graphs with dependencies — parallel where independent, ordered where not
Underneath is a daemon control plane. A Run is the only unit of execution. Full slots mean automatic queueing. A timed-out task kills the whole process tree. Blow the token budget and it reruns at a lower tier. session_id is the ownership boundary.
The adapter layer normalizes event streams, usage, and sessions for 11 built-in CLIs (claude / grok / opencode / omp / codex / kimi / copilot, and more). A CLI not on the list just needs a JSON config — no code changes.
DeepSeek Harness gets native integration too: one insert patch registers all 34 tools under mcp__agentmcp__*; the daemon auto-starts if down and reconnects with exponential backoff.
The constraints are practical. Don’t split everything by default. Early versions spawned a pile of sub-agents for tiny things, and coordination cost more than the task itself. That’s why the complexity gate exists.
The adapter layer is messier than you’d think. Headless modes, event streams, resume, permission flags — every vendor has its own. Normalization is the real engineering.
A thick control plane makes debugging heavy. When something breaks you have to look at the main agent, daemon logs, and subprocesses at once. That’s heavy for a personal project; still trimming.
Status
After building it, my workflow changed: from “I watch everything myself” to “split, dispatch, wait in a loop, merge”.
A cross-file refactor: first estimate_complexity grades it L, then orchestrate_task declares dependencies — shared types first, call sites next, review last. Independent scans run in parallel. Each subtask can name a different CLI.
Scope crept mid-way; steer_agent narrowed it. Something hung; the timeout caught it, and the session could still resume.
Currently v4.0.0a1: 34 MCP tools, 589 tests passing. It didn’t make me “faster, cheaper, better” — but it did free me from staring at terminals.
Limits remain. Still alpha. Adapter test coverage is uneven; the sandbox is mostly “one policy translated into each CLI’s own flags” — not real isolation; cross-vendor review in complex merges needs more mileage.
Next up: fill in adapter test coverage, enforce policy at the actual execution point, then simplify onboarding.
curl -fsSL https://raw.githubusercontent.com/37chengshan/agent-mcp/main/install.sh | bash
Main repo: 37chengshan/agent-mcp. Issues welcome — complaints too.
FAQ
How is this different from just opening more terminals?
Terminals have no scheduler: runaway subtasks, slot queueing, mid-course correction, token reconciliation — all on a human. In agent-mcp, a Run is the only unit of execution: full slots queue automatically, timed-out tasks kill the whole process tree, blown token budgets rerun at a lower tier, and session_id draws the ownership boundary.
Does it lock me into one model?
No. Inference still happens in each CLI’s native runtime; agent-mcp doesn’t rewrite the agent loop. The principle is match-the-model-to-the-task: read-heavy exploration to fast models (omp / pi / grok), deep-reasoning planning to strong ones (claude), cost and quality matched on the spot.
Do small tasks need splitting into sub-agents too?
No. Early versions spawned a pile of sub-agents for tiny things, and coordination cost more than the task. Now there’s estimate_complexity: local S/M/L grading, zero token cost, do it directly by default, split only when needed.
My CLI isn’t one of the 11 — what then?
One JSON config plugs it in, no code changes. The adapter layer normalizes each vendor’s event streams, usage fields, and sessions; the differences get ironed out there.
Is it stable now?
v4.0.0a1, still alpha. Adapter test coverage is uneven; the sandbox is currently just “one policy translated into each CLI’s own flags” — not real isolation; cross-vendor review in complex merges needs more mileage. The author says adapter coverage comes first.
上周三,我在拆一个跨文件重构。任务拆成五份,三个终端同时开:Claude 写公共类型,Codex 改调用方,omp 跑依赖扫描。
四十分钟后,我在备忘录里记谁跑完了。左边 Claude 卡在权限确认,中间 Codex 输出刷得看不清,右边 omp 早就跑完,但没人告诉我。token 花了多少,事后对账。那天晚上收工时,三个会话里有两个已经废了,上下文全丢。
这不是「模型不够强」。这是没有人做调度。
agent-mcp 的 README 写得很清楚——核心不是「多开几个 Agent」,而是把任意 CLI 收成一个可派发、可监控、可续接、可终止的工作池。
主 Agent 只做两件事:拆解和汇合。
派发、等待、插话、超时、排队、续接、降档,全部交给控制面。模型推理仍发生在各家 CLI 原生 runtime 里;agent-mcp 不重写 agent loop,也不锁死单一模型。
另一个原则是:按任务匹配底座。
读密集探索丢给快底座(omp / pi / grok),深推理规划丢给强底座(claude)。成本与质量现场匹配,而不是「全家都用同一个模型」。
问题
多开几个终端很简单。难的是:
- 子任务跑飞了怎么办?
- 槽位满了怎么排队?
- 中途发现方向错了怎么插话?
- 主 Agent 等结果时不能空转,也不能傻轮询把上下文撑爆。
- 谁在跑、跑哪了、烧了多少,全靠人盯。
各家 CLI 还互不相通:有的吐 JSONL,有的只吐纯文本,usage 字段名完全不同,resume 有的支持有的不支持。如果不管这些差异,多 Agent 只会把混乱放大。
做法
对外 agent-mcp 是一组 MCP 工具。任何支持 MCP 的宿主都能挂上。常用的几件:
estimate_complexity:本地判 S/M/L,零 token,不 spawn——默认直接做,按需才拆spawn_agent:派发子 Agent,指定 CLI / 模型 / 超时,马上拿回agent_idwait_agent:短阻塞等终止态,返回摘要steer_agent:中途插话,终止当前 run,在同一节点接着跑followup_task:合并挂起消息,触发下一 turnorchestrate_task:有依赖的任务图,无依赖并行,有依赖按序
底层是 daemon 控制面。Run 是唯一执行单位。槽位满了自动排队。任务超时会终止整棵进程树。token 预算超了可以降档重跑。session_id 是所有权边界。
适配器层内置 11 款 CLI(claude / grok / opencode / omp / codex / kimi / copilot 等)的事件流、usage、session 归一化。不在列表里的 CLI,写一份 JSON 配置也能接,不用改代码。
DeepSeek Harness 也做了原生接入:一行 insert patch,34 个工具以 mcp__agentmcp__* 全量注册;daemon 未起自动拉起,断线指数退避重连。
约束很实际。不要默认什么都拆。 我早期版本遇到小事也 spawn 一堆子 Agent,协调开销比任务本身还大。后来才有复杂度分级门。
适配层比想象脏。 headless 模式、事件流、resume、权限参数,每家一套。归一化才是真正的工程量。
控制面做厚了,调试会变重。 出问题时要同时看主 Agent、daemon 日志、子进程。这对个人项目偏重,还在砍。
现状
写完之后,我的工作流变了:以前是「我自己盯」,现在是「拆完、派出去、循环 wait、汇合」。
一次跨文件重构:先 estimate_complexity 判成 L,再 orchestrate_task 声明依赖——先改公共类型,再改调用方,最后 review。无依赖的扫描并行跑。每个子任务可以指定不同 CLI。
中途范围写大了,steer_agent 收窄。跑挂了,超时兜住,session 还能 resume。
当前 v4.0.0a1:34 个 MCP 工具,589 个测试通过。它没有让我「多快好省」,但确实把我从盯终端里解放出来了。
局限仍在。 项目仍是 alpha。适配器实测率不均匀;沙箱目前主要是「统一策略翻译到各 CLI 自己的参数」,不是真隔离;跨厂商审查在复杂合并场景还要再踩。
下一步优先补适配器实测、把策略真正拦在执行点上,再简化上手路径。
curl -fsSL https://raw.githubusercontent.com/37chengshan/agent-mcp/main/install.sh | bash
主仓库:37chengshan/agent-mcp。欢迎提 issue,也欢迎直接骂。
常见问题
跟直接多开几个终端有什么区别?
多开终端没人调度:子任务跑飞、槽位排队、中途改方向、token 对账全靠人盯。agent-mcp 里 Run 是唯一执行单位,槽位满了自动排队,任务超时会终止整棵进程树,token 预算超了可以降档重跑,session_id 划清所有权边界。
它会锁死我用某个模型吗?
不会。模型推理仍发生在各家 CLI 原生 runtime 里,agent-mcp 不重写 agent loop。原则是按任务匹配底座:读密集探索丢给快底座(omp / pi / grok),深推理规划丢给强底座(claude),成本与质量现场匹配。
小任务也要拆成子 Agent 吗?
不要。作者早期版本小事也 spawn 一堆子 Agent,结果协调开销比任务本身还大。现在有 estimate_complexity:本地判断 S/M/L,零 token 消耗,默认直接做,按需才拆。
我用的 CLI 不在那 11 款里怎么办?
写一份 JSON 配置就能接,不用改代码。适配器层负责把各家的事件流、usage 字段、session 归一化,差异由它抹平。
现在稳定吗?
v4.0.0a1,还是 alpha。适配器实测率不均匀;沙箱目前只是”统一策略翻译到各 CLI 自己的参数”,不是真隔离;复杂合并场景的跨厂商审查还要再踩。作者说下一步先补适配器实测。