Someone asked: should we give students an AI coding assistant?

Ask a model directly and it’ll give you something that sounds smooth: efficiency gains, personalization, the future is here. The literature tells a different story — CHI 2023 has controlled studies at 1.15× / 1.8×, but also “no significant difference after a week”; PNAS 2025 measured −17% in unsupervised exams.

AI doesn’t lack search. It lacks evidence discipline.

TL;DR EduEvidence is an AI Agent Skill built for one problem: when AI advises on education decisions, it talks smoothly and cites nothing. It locks in three rules: retrieved snippets are only pointers, not evidence; no study gets designed without an evidence-grounded knowledge gap; the real fact layer is a versioned, immutable Evidence Graph — reports are just projections. At the end you don't get a yes/no, you get four outcomes — ADOPT, PILOT, REJECT, INSUFFICIENT EVIDENCE — plus how to pilot and how to evaluate.

What EduEvidence does is turn a “decision question” into an answer that is evidence-backed, pilotable, and evaluable — not a longer summary.

The README’s product line is: From Research Questions to Evidence-Based Decisions.

It ships as an AI Agent Skill; inside the Skill runs the EduEvidence Research Engine — persistent, auditable, compressing decision questions into evidence-grounded answers.

Three frozen principles:

  1. Retrieved snippets are locators, not evidence.
    OpenAlex / Semantic Scholar / Sciverse give you leads; to become evidence, something must enter a traceable citation chain.

  2. Without an evidence-grounded Knowledge Gap, no new study gets designed.
    A frozen scientific rule: study design must explicitly reference a grounded Gap ID.

  3. result.json / HTML / Markdown are projections, not the fact store.
    The real fact layer is versioned, immutable revisions of the Evidence Graph.

The output isn’t a binary “allow / ban” either. It’s four states:

ADOPT / PILOT / REJECT / INSUFFICIENT EVIDENCE, plus an executable intervention and evaluation plan.

The Problem

Education decisions have three traps:

九阶段协议
Research Core 六阶段 + Decision Extension 三阶段。
  • Citations that look real: the year, journal, and conclusion all check out, but the details don’t survive scrutiny
  • Correlation passed off as causation: “scores went up after adoption” is usually just selection bias
  • Skipped applicability boundaries: who it works for, under what conditions it fails — often erased

Real studies like CHI 2023 and PNAS 2025 remind us: AI-assisted coding may speed up short tasks, but it can also drag in transfer and unsupervised settings. “Faster across the board” with no boundary conditions is unusable as decision input.

The Method

The engine splits the full research cycle open:

  • Research Core: frame the question → retrieve → grade evidence → Gap → synthesize
  • Decision Extension: option → pilot design → evaluate → update the decision

Three public workflows:

WorkflowAnswers
Evidence ReviewWhat the existing evidence supports — and what it can’t support
Decision & PilotHow to pilot, how to verify
Evaluate & UpdateHow new data rewrites the decision

Multi-domain support works through contracts, not engine copies: education and policy each declare their frame schema, outcome taxonomy, and method list; unknown tokens fail closed.

四态决策
ADOPT / PILOT / REJECT / 证据不足。

How It Works

After install, hosts like Claude Code / Cursor / Codex / OMP can auto-load the Skill on “teaching decision” type questions.

npm install -g eduevidence
eduevidence skill --host claude

Native Core only depends on the Python standard library — no Agent MCP, no daemon. That’s deliberate: a research engine shouldn’t be welded to one orchestration framework.

Retrieval channels come in zero-config (OpenAlex / Semantic Scholar / CrossRef…) and keyed (Sciverse etc.). Sciverse is for citation-level locating; then things go into the Graph for structuring.

Results

  • Benchmark simulation ≠ evidence. benchmarks/results/ is just a simulation harness; the first real empirical round is B2 vs B3 (10 questions × 3 repeats) — it can’t be written up as “dominates across the board.”
  • Four states are harder to sell than two, but more useful. “Insufficient evidence” is a legitimate conclusion, not a failure.
  • Projections can change; the fact layer can’t. Reports can be reordered; once a Graph revision lands, it shouldn’t be quietly rewritten.

Currently at 6.2.0. The Landing / Research Studio / Deep Research comparison pages all open directly. Sample report:

open examples/ai-coding-assistant-evidence/EduEvidence_Report.html

It won’t make education decisions for you. It only guarantees: before you say “go” or “no go” — where the evidence comes from, what the Gap is, how the pilot gets verified — all written down.

Main repo: 37chengshan/eduevidence.

FAQ

How is EduEvidence different from just asking an AI?

Ask an AI directly and you get something that sounds smooth. EduEvidence compresses a decision question into an answer that is evidence-backed, pilotable, and evaluable — ending in ADOPT, PILOT, REJECT, or INSUFFICIENT EVIDENCE, plus an executable intervention and evaluation plan.

Does a retrieved paper abstract count as evidence?

No. Retrieved snippets are only locators — leads. To become evidence, something must enter a traceable citation chain. That’s the first frozen principle.

Will it just tell me whether students should get an AI coding assistant?

It won’t make the decision for you. It only guarantees: before you say “go” or “no go” — where the evidence comes from, what the knowledge gap is, how the pilot gets verified — all written down.

Do I have to install Agent MCP to use it?

No. Native Core only depends on the Python standard library and isn’t welded to any single orchestration framework; hosts like Claude Code, Cursor, Codex, and OMP can all auto-load the Skill.

Is “insufficient evidence” a failure?

No. It’s a legitimate conclusion, not a failure. Four states are harder to sell than two, but more useful.

有人问:要不要给学生上 AI 编程助手?

让模型直接答,它会给你一段听起来很顺的话:提效、个性化、未来已来。翻文献又是另一回事——CHI 2023 有 1.15× / 1.8× 的对照,也有「一周后无显著差距」;PNAS 2025 在无监督考试里测到 −17%。

AI 不缺搜索,缺的是证据纪律。

一句话总结 EduEvidence 是一个 AI Agent Skill,专门治"AI 给教育决策出主意时满口顺话、没有证据"的毛病。它定死三条规矩:检索片段只是线索、不算证据;没有证据接地的知识缺口,就不许设计新研究;真正的事实层是版本化、不可变的 Evidence Graph,报告只是投影。最后不给你"行/不行"二选一,而是 ADOPT、试点、驳回、证据不足四种结论,外加怎么试点、怎么验证。

EduEvidence 要做的,是把「决策问题」变成「有证据支撑、可试点、可评估」的答案,而不是一段更长的摘要。

README 的产品句是:From Research Questions to Evidence-Based Decisions.

交付形态是 AI Agent Skill;Skill 里面跑的是 EduEvidence Research Engine——持久、可审计,把决策问题压成证据接地的回答。

三条冻结原则:

  1. 检索片段只是定位器,不是证据。
    OpenAlex / Semantic Scholar / Sciverse 给你的是线索;要成为证据,必须进入可追溯的引用链。

  2. 没有证据接地的 Knowledge Gap,就不允许设计新研究。
    冻结科学规则:study design 必须显式引用已接地的 Gap ID。

  3. result.json / HTML / Markdown 是投影,不是事实库。
    真正的事实层是 Evidence Graph 的版本化、不可变 revision。

输出也不是「允许 / 禁止」二值,而是四态:

ADOPT / PILOT / REJECT / INSUFFICIENT EVIDENCE,外加可执行的干预与评估计划。

问题

教育决策有三个坑:

九阶段协议
Research Core 六阶段 + Decision Extension 三阶段。
  • 引用像真的:年份、期刊、结论都对得上,细节却经不起核对
  • 把相关当因果:「用了分数提高」往往只是选择偏差
  • 跳过适用边界:对谁有效、在什么条件下失效,常常被抹掉

CHI 2023 和 PNAS 2025 这类真实研究提醒我们:AI 辅助编码在短任务上可能提效,在迁移与无监督场景也可能拖后腿。没有边界条件的「全面提效」是不可用的决策输入。

方法

引擎把完整研究周期拆开:

  • Research Core:框定问题 → 检索 → 证据分级 → Gap → 综合
  • Decision Extension:方案 → 试点设计 → 评估 → 更新决策

三条公开工作流:

工作流回答什么
Evidence Review现有证据支持什么、不能支持什么
Decision & Pilot怎么试点、怎么验证
Evaluate & Update新数据如何改写决策

多域靠契约而不是复制引擎:education 与 policy 各自声明 frame schema、outcome taxonomy、方法清单;未知 token fail closed。

四态决策
ADOPT / PILOT / REJECT / 证据不足。

做法

安装后,Claude Code / Cursor / Codex / OMP 等宿主可以在「教学决策」类问题上自动加载 Skill。

npm install -g eduevidence
eduevidence skill --host claude

Native Core 只依赖 Python 标准库,不需要 Agent MCP,也不需要 daemon。这是刻意的:研究引擎不该绑死在某一个编排框架上。

检索通道分零配置(OpenAlex / Semantic Scholar / CrossRef…)与密钥通道(Sciverse 等)。Sciverse 用来拿引用级定位;再进 Graph 做结构化。

结果

  • Benchmark 仿真 ≠ 实证。 benchmarks/results/ 只是模拟 harness;第一轮实证是 B2 vs B3(10 题 × 3 重复),不能写成「全面碾压」。
  • 四态比二值难卖,但更有用。 「证据不足」是合法结论,不是失败。
  • 投影可变,事实层不可变。 报告可以重排;Graph revision 一旦落地就不该被悄悄改写。

当前 6.2.0。Landing / Research Studio / Deep Research 对比页都可以直接打开。示例报告:

open examples/ai-coding-assistant-evidence/EduEvidence_Report.html

它不会替你做教育决策。它只保证:在你说「上」或「不上」之前,证据从哪来、Gap 是什么、试点怎么验——写清楚。

主仓库:37chengshan/eduevidence。

常见问题

EduEvidence 和直接问 AI 有什么区别?

直接问 AI,你得到的是一段听起来很顺的话;EduEvidence 把决策问题压成有证据支撑、可试点、可评估的回答,最后给出 ADOPT、试点、驳回、证据不足四种结论,外加可执行的干预与评估计划。

检索到的论文摘要算证据吗?

不算。检索片段只是定位器,是线索;要成为证据,必须进入可追溯的引用链。这是第一条冻结原则。

它会直接告诉我该不该给学生上 AI 编程助手吗?

不会替你做决策。它只保证:在你说「上」或「不上」之前,证据从哪来、知识缺口是什么、试点怎么验——都写清楚。

一定要装 Agent MCP 才能用吗?

不用。Native Core 只依赖 Python 标准库,不绑死在某一个编排框架上;Claude Code、Cursor、Codex、OMP 等宿主都可以自动加载这个 Skill。

「证据不足」算失败吗?

不算。它是合法结论,不是失败。四态比二值难卖,但更有用。