EduEvidenceDataLab [Light]
语言
EduEvidenceDataLab [Light]数据来源:真实文献 · 人工精编

我们准备在大学一年级 C 语言课程中允许学生使用生成式 AI 编程助手。它到底会不会提高学习效果?应该怎样引入?

模式:平台原生 · 生成时间:2026-08-24 · 证据 12 条 · 来源 8 个

先看结论

该不该做、置信度多高、最关键的证据边界在哪。

建议决策 试点验证 置信度 · 中

任务表现的正面证据 + 有据可查的无护栏风险 + 混合的质量/可用性信号 + 大学层面学习证据缺失 → 有界、护栏化、带评估的试点,而非全面采用。

最强支持结论

AI 编程助手在训练期稳定提升练习效率:69 名新手的随机对照中完成率 1.15 倍、用时 0.57 倍。

关键不确定性 / 反例

缺少大学层面的直接学习证据;唯一大规模试验显示,无护栏使用 GPT-4 的学生独立考试成绩下降 17%。

主要风险

无护栏使用会抬高练习表现却压低独立考试表现,而学习者往往意识不到这一落差。

下一步

开展分阶段 CS1 试点:给提示而非答案、每周实验课使用、并以无 AI 迁移考试作为可叫停的验收条件。

证据 / 来源 · 12 / 8

任务表现 ≠ 学习效果

只展示真正有解释力的结果分离;正向、负向与零效应按 effect_direction 编码。

结果分离 · 任务表现 ≠ 学习效果

将不同结果类型分开裁决,避免把训练时更快、更高分直接等同于真正学会。

效应方向来自 evidence.effect_direction;“支持某个主张”不等于“结果是正向”。

任务 / 近端表现

  • 完成时间正向效应 2
  • 代码质量零效应 1
  • 作业成绩正向效应 2

学习 / 保持 / 迁移

  • 知识获得正向效应 1
  • 记忆保持零效应 1
  • 独立问题解决负向效应 1零效应 2

风险 / 依赖

  • 过度依赖负向效应 1

其他结果

  • 元认知正向效应 1
各结果类型证据效应分布各结果类型的正向 / 负向 / 零效应证据条数(基于 effect_direction,不等同于 Claim 是否被支持)。知识获得1记忆保持独立问题解决-1完成时间2代码质量作业成绩2元认知1过度依赖-1正向效应负向效应零效应

这意味着什么:各 Outcome 的正向 / 负向 / 零效应证据条数对比。这里展示的是 effect_direction,不是“这条证据是否支持某个主张”。

证据裁决

支持、不确定、被反驳与缺失证据分开放置,不把长段落平铺在同一层。

决策:试点验证 · 置信度:中

可以主张

7
  • AI 编程助手在训练期间可靠地提升任务表现(完成速度、正确性)—— E-001、E-006。

    E-001E-006
  • 无护栏的生成式 AI 访问在移除工具后可能损害独立问题解决能力 —— E-004。

    E-004
  • 护栏设计(给提示而非给答案)能大幅缓解负面学习效应 —— E-005。

    E-005
查看其余 4 条
  • 任务表现提升并不自动等于学习提升 —— E-004 与 E-006 的研究内对照。

    E-004E-006
  • 工具能力可观:Codex 能解出约半数至四分之三的 CS1 考试风格题目 —— E-010。

    E-010
  • 职业开发者 RCT 显示 Copilot 带来约 55% 任务提速;但职业人群限制直接性 —— E-008。

    E-008
  • LLM 代码讲解的质量评级与学生自撰讲解相当,可作支架材料 —— E-011。

    E-011

尚不能主张

4
  • AI 编程助手能否真正改善或保持大学新手的编程学习——本证据集中没有大学层面的直接 RCT [无直接证据]

  • Kazemitabaar 2023 的一周中性保持性能否延伸到一个学期 —— E-003。

    E-003
  • 基准质量结论(E-009)与讲解质量评级(E-011)能否转化为课堂学习收益。

    E-009E-011
查看其余 1 条
  • 可用性研究所记录的理解/所有权困难(E-012)在整学期护栏条件下会如何演变。

    E-012

被反驳的主张

2
  • 'AI 工具总能提高学习'被 E-004 反驳(无护栏访问,独立考试 −17%)。

    E-004
  • '速度收益等于学习收益'被 E-001/E-006/E-008 与 E-004 之间的任务-学习分离所反驳。

    E-001E-006E-008E-004

缺失证据

4
  • 在大学编程课程中带保持与无 AI 迁移测试的 RCT。

  • 同一课程内变化 AI 使用政策的研究。

  • 跨越一门课的 AI 依赖纵向数据。

查看其余 1 条
  • 职业提速 RCT 的同行评审重复(Peng 等仍为预印本)。

从证据到行动

适用性、护栏、停止条件与评价连成一条可执行路径。

证据

AI 编程助手在训练期间可靠地提升任务表现(完成速度、正确性)—— E-001、E-006。

适用性

在大一 C 课程以护栏化使用政策开展试点

决策

试点验证

护栏

AI 使用分三档明确分级(解释 / 协作 / 无 AI 迁移)。照抄未审视的 AI 输出属学术诚信违规,并通过推理痕迹要求核查。

停止条件

迁移测验成绩显著低于基线同届预期; 推理痕迹中出现普遍诚信违规; 风险指标中 AI 依赖信号超阈值; 助教/教师工作量不可持续

评价

实验班独立问题解决非劣(差异在 5% 以内)且保持相当或更优、AI 依赖指数低于阈值;若独立问题解决下滑超过 10%,无论任务收益如何,试点均判为失败。

Lieflat 实证手作画廊

AI 按数据形状从 Lieflat 目录选型编排;每张图的数字都可溯源到 result.json。

关键来源

摘要页只列最关键的来源;完整溯源在完整报告中展开。

完整报告中还有 4 个来源可展开追溯。

EduEvidenceDataLab [Light]Data: real studies · manually curated

我们准备在大学一年级 C 语言课程中允许学生使用生成式 AI 编程助手。它到底会不会提高学习效果?应该怎样引入?

Mode: Platform native · Generated: 2026-08-24 · Evidence: 12 · Sources: 8

Decision first

What to do, how confident we are, and the most important evidence boundary.

Recommended decision Pilot Confidence · Moderate

Positive task-performance evidence plus documented unguarded-access risk, mixed quality/usability signals, and missing university-level learning evidence → bounded, guardrailed pilot with evaluation, not full adoption.

Strongest supported conclusion

AI coding assistants reliably speed up practice work: completion rate 1.15x and time 0.57x in a randomised trial of 69 novices.

Key uncertainty / contradiction

No university-level RCT measures learning directly, and the one large trial that did - unguarded GPT-4 - saw independent exam scores fall 17%.

Main risk

Unguarded access can raise practice performance while lowering independent exam performance, and learners may not notice the gap.

Next action

Run a phased CS1 pilot with hints-not-answers guardrails, weekly lab use, and a no-AI transfer exam that can stop the pilot.

Evidence / sources · 12 / 8

Task performance ≠ learning

Only informative outcome separation; positive, negative and null effects use effect_direction.

Outcome Separation · Task performance ≠ learning

Outcomes are adjudicated separately so faster training performance is not silently treated as evidence of learning.

Effect direction comes from evidence.effect_direction; supporting a claim does not imply a positive outcome.

Task / proximal performance

  • Completion timePositive effect 2
  • Code qualityNull effect 1
  • Assignment scorePositive effect 2

Learning / retention / transfer

  • Knowledge gainPositive effect 1
  • RetentionNull effect 1
  • Independent problem solvingNegative effect 1Null effect 2

Risk / dependency

  • Over-relianceNegative effect 1

Other outcomes

  • MetacognitionPositive effect 1
Outcome evidence effect balancePositive / negative / null effect counts per outcome (based on effect_direction, not claim support).Knowledge gain1RetentionIndependent problem solving-1Completion time2Code qualityAssignment score2Metacognition1Over-reliance-1Positive effectNegative effectNull effect

What this means: Positive / negative / null effect-direction evidence counts per outcome. This visual encodes effect_direction, not whether evidence supports a claim.

Evidence tribunal

Supported, uncertain, contradicted and missing evidence stay separated instead of flattened into long prose.

Decision: Pilot · Confidence: Moderate

Can claim

7
  • AI coding assistants reliably increase task performance during training (completion speed, correctness) — E-001, E-006.

    E-001E-006
  • Unguarded generative AI access can harm independent problem solving when access is removed — E-004.

    E-004
  • Guardrail design (hints instead of answers) substantially mitigates the negative learning effect — E-005.

    E-005
View 4 more
  • Task performance gains do not automatically imply learning gains — E-004 vs E-006 (within-study contrast).

    E-004E-006
  • Tool capability is substantial: Codex solves roughly half to three-quarters of CS1 exam-style questions — E-010.

    E-010
  • Professional-developer RCT shows ~55% faster task completion with Copilot; directness limited by professional population — E-008.

    E-008
  • LLM code explanations rate comparable to student-authored explanations, viable as scaffold material — E-011.

    E-011

Cannot yet claim

4
  • Whether AI coding assistants improve or preserve actual programming learning in university novices — no direct university-level RCT in reviewed set [无直接证据]

  • Whether one-week neutral retention (Kazemitabaar 2023) extends to a semester — E-003.

    E-003
  • Whether benchmark quality findings (E-009) and explanation-quality ratings (E-011) translate into classroom learning gains.

    E-009E-011
View 1 more
  • How comprehension/ownership difficulties documented in usability studies (E-012) behave over a full semester with guardrails.

    E-012

Contradicted claims

2
  • The claim 'AI tools always improve learning' is contradicted by E-004 (unguarded access, -17% independent exam).

    E-004
  • The claim 'speed gains equal learning gains' is contradicted by the task-vs-learning separation across E-001/E-006/E-008 vs E-004.

    E-001E-006E-008E-004

Missing evidence

4
  • RCT of AI coding assistants in university programming courses with retention and no-AI transfer tests.

  • Studies varying AI usage policy within the same course.

  • Longitudinal data on AI dependency beyond one course.

View 1 more
  • Peer-reviewed replication of the professional speed RCT (Peng et al. remains a preprint).

Evidence to action

Applicability, guardrails, stop conditions and evaluation form one executable path.

Evidence

AI coding assistants reliably increase task performance during training (completion speed, correctness) — E-001, E-006.

Applicability

pilot in first-year C course with guardrailed usage policy

Decision

Pilot

Guardrails

AI usage is allowed in three explicitly graded modes (explain / collaborate / no-AI-transfer).…

Expand full explanation

AI usage is allowed in three explicitly graded modes (explain / collaborate / no-AI-transfer). Copying unexamined AI output is an academic integrity violation and is assessed via the reasoning-trace requirement.

Stop conditions

transfer-test scores drop significantly below baseline cohort expectations; widespread integrity violations in reasoning traces;…

Expand full explanation

transfer-test scores drop significantly below baseline cohort expectations; widespread integrity violations in reasoning traces; AI dependency signals exceed threshold in risk metrics; TA/teacher workload becomes unsustainable

Evaluation

treatment group shows non-inferior independent problem solving (delta within 5%) AND superior or equal retention AND AI dependency index below thresho…

Expand full explanation

treatment group shows non-inferior independent problem solving (delta within 5%) AND superior or equal retention AND AI dependency index below threshold; if independent problem solving declines >10%, the pilot is judged unsuccessful regardless of task-performance gains.

Lieflat Editorial Gallery

Charts selected and composed by AI from the Lieflat catalog; every number traces back to result.json.

Key sources

Only the key sources in the brief; full traceability expands in the full report.

4 more sources are traceable in the full report.