I opened the bill. The number was climbing, but I couldn’t say what it was spent on.
The same files I’d fed the model last round got sent again this round, byte for byte. Tool logs, search results, duplicated README snippets… context leaks like a bucket with a hole. And whether any “optimization” actually worked — there was no receipt to check against.
d-token answers three questions:
- Where did these tokens actually go?
- Which provider really handled the request?
- Did compression actually happen — or does it just feel cheaper?
The README is blunt about what it is: a local context control plane for AI coding agents.
It sits between “the agent you’re using” and “the model provider you configured”. It compresses redundancy before context leaves your machine, keeps routing explicit, and makes config changes recoverable — not a pile of scripts and guesswork.
AI coding agent
↓
d-token connect · optimize · route · observe
↓
your configured model provider
Three hard rules:
- Local-first: secrets never end up in plain logs; the control plane runs on your machine.
- Explicit routing: every request takes a stated path, never silently switches providers.
- One receipt per request: source agent, route, transforms, Local / Provider Tokens, recovery status.
Problem
AI coding agents keep resending the same files, logs, and tool outputs. The token bill quietly grows, but:
- Can’t tell which session burns the most
- Can’t tell whether an “optimization” actually worked
- Break a config, and there’s no way back
No prompt-engineering trick fixes this. What’s needed is infrastructure that is visible and roll-back-able.
Local is the premise, not an option. Context carries source paths, secret boundaries, internal company structure. Put compression and routing in a third-party black box and you hand over visibility too.
d-token keeps settings, diagnostics, request metadata, and rollback history on your machine. Plain logs never store full prompts, full responses, source text, or API keys. Your chosen provider still receives the requests you send — but now there’s a door you control in between.
Approach
The desktop app is Tauri 2 + Rust core. The Rust side handles connection state, routing decisions, the compression pipeline, and receipts; the frontend only does display and the safe-config flow.
Config changes aren’t “save and it applies”. They go: preview → redacted diff → backup → verify → roll back. When something breaks, you diagnose, fix, recover — the README calls it “the config doctor”.
Status semantics are deliberately spelled apart:
detected≠configured≠connected≠verified- passing tests ≠ “passing real compression”
The deeper lesson: when you build middleware in the AI toolchain, honesty beats a feature list.
Results
The README records one number: 33,705 fewer tokens in a single real routed request (measured, 2026-08-04).
Not “save 30% on average” marketing — a reproducible measurement. The receipt shows Local Tokens vs. Provider Tokens side by side. How much you saved is on paper, not a feeling.
Currently at 0.2.0-beta.1, a public unsigned pre-release (macOS ARM64 / Windows x64). Unsigned builds may trigger OS security warnings — verify the repo and hashes; don’t bypass system protections.
On the roadmap:
- v0.3: cross-agent Context Relay, Session Cache, cc-switch follow / lock / observe
- v0.4: multi-agent adapters, low-risk auto-fix, stable signing
Deep compression, Headroom-style integrations, and full signed distribution are still on the way.
git clone https://github.com/37chengshan/d-token.git
cd d-token/code/apps/desktop
npm ci
npm run tauri:dev
Main repo: 37chengshan/d-token.
FAQ
What exactly is d-token for?
It’s a local context control plane for AI coding agents, sitting between “the agent you use” and “the provider you configured”: it compresses redundancy before context leaves your machine, keeps routing explicit, and makes config changes recoverable.
Does it really save tokens? How many?
The README records one measurement: a single real routed request on 2026-08-04 sent 33,705 fewer tokens. Note this is a single data point, not an average; the receipt shows Local Tokens vs. Provider Tokens, so the savings are on paper, not a feeling.
Will it upload my code or secrets?
No. Local-first is a hard rule: settings, diagnostics, request metadata, and rollback records stay on your machine, and plain logs never store full prompts, full responses, source text, or API keys. Your chosen provider still receives the requests you send — but now there’s a door you control in between.
What if I break a config?
Changes aren’t “save and it applies” — they go through five steps: preview → redacted diff → backup → verify → roll back. When something breaks, you diagnose, fix, and recover; the author calls it “the config doctor”.
Can I use it now?
It’s at 0.2.0-beta.1, a public unsigned pre-release (macOS ARM64 / Windows x64). Unsigned builds may trigger OS security warnings — after downloading, verify the repo and hashes; don’t bypass system protections.
打开账单,数字在涨,但说不清花在哪。
同一批文件,上一轮刚喂过模型,这一轮又原样重发。工具日志、搜索结果、重复的 README 片段……上下文像漏水的桶。所谓「优化」有没有用,没有回执可以对账。
d-token 要回答三个问题:
- 这些 token 花在了哪里?
- 哪个 Provider 真的处理了请求?
- 压缩是否真的发生了,而不是「感觉更省了」?
README 的定位很干脆:面向 AI 编码 Agent 的本地上下文控制面。
它运行在「你正在用的 Agent」和「你配置的模型服务商」之间。在上下文离开本机之前压缩冗余,保持路由显式,并让配置变更可恢复——而不是一堆脚本和猜测。
AI 编程 Agent
↓
d-token 连接 · 优化 · 路由 · 观察
↓
你配置的模型服务商
三条硬原则:
- 本地优先:秘密不进普通日志;控制面跑在本机。
- 显式路由:每个请求走明确路径,绝不静默切换服务商。
- 每请求一张回执:来源 Agent、路由、变换、Local / Provider Token、恢复状态。
问题
AI 编码 Agent 会一遍又一遍重发相同的文件、日志和工具输出。token 账单悄悄增长,但:
- 说不清哪个会话最烧钱
- 说不清「优化」是否生效
- 配置改坏了,回不去
这不是再写一个 prompt 工程技巧能解决的。需要一层看得见、可回退的基础设施。
本地是前提,不是可选项。 上下文里有源码路径、密钥边界、公司内部结构。压缩和路由如果放在第三方黑盒里,你就把「可见性」也交出去了。
d-token 把设置、诊断、请求元数据、回退记录都留在本机。普通日志不会主动保存完整 prompt、完整 response、源代码正文、API key。你选的服务商仍会收到你主动发出的请求——但中间多了一层你自己掌控的门。
做法
桌面端是 Tauri 2 + Rust core。Rust 侧负责连接状态、路由决策、压缩管线、回执;前端只做展示与安全配置流程。
配置变更不是「保存即生效」,而是:预览 → 脱敏 Diff → 备份 → 验证 → 回退。出问题可以诊断、修复、恢复——README 管这叫「配置的医生」。
状态语义也刻意分开写:
detected≠configured≠connected≠verified- 测试通过 ≠ 「真实压缩通过」
更底层的教训是:在 AI 工具链里做中间层,诚实比功能表重要。
结果
README 记录的数字是:单次真实路由请求物理减少 33,705 token(实测,2026-08-04)。
这不是「平均省 30%」的营销话术,而是一次可复现的实测点。回执里能看到 Local Token 与 Provider Token 的对比——省了多少,写在纸上,不靠感觉。
当前 0.2.0-beta.1,公开未签名预发布(macOS ARM64 / Windows x64)。未签名构建可能触发系统安全警告——请核对仓库与哈希,不要绕过系统保护。
路线图上还有:
- v0.3:跨 Agent Context Relay、Session Cache、cc-switch 跟随 / 锁定 / 观察
- v0.4:多 Agent 适配、低风险自动修复、稳定签名
深度压缩、Headroom 类集成、完整签名分发,都还在路上。
git clone https://github.com/37chengshan/d-token.git
cd d-token/code/apps/desktop
npm ci
npm run tauri:dev
主仓库:37chengshan/d-token。
常见问题
d-token 到底是干嘛的?
它是面向 AI 编码 Agent 的本地上下文控制面,跑在”你正在用的 Agent”和”你配置的模型服务商”之间:在上下文离开本机之前压缩冗余,保持路由显式,并让配置变更可恢复。
它真能省 token 吗?省多少?
README 记录了一次实测:2026-08-04 单次真实路由请求物理减少 33,705 token。注意这是单次实测点,不是平均值;回执里能看到 Local Token 与 Provider Token 的对比,省了多少写在纸上,不靠感觉。
我的代码和密钥会被它上传吗?
不会。本地优先是硬原则:设置、诊断、请求元数据、回退记录都留在本机,普通日志不会主动保存完整 prompt、完整 response、源代码正文、API key。你选的服务商仍会收到你主动发出的请求,但中间多了一层你自己掌控的门。
配置改坏了怎么办?
配置变更不是”保存即生效”,而是走五步:预览 → 脱敏 Diff → 备份 → 验证 → 回退。出问题可以诊断、修复、恢复,作者管这叫”配置的医生”。
现在能用吗?
当前是 0.2.0-beta.1,公开未签名预发布(macOS ARM64 / Windows x64)。未签名构建可能触发系统安全警告,下载后请核对仓库与哈希,不要绕过系统保护。