/tdd 技能
红-绿-重构循环的规则。
安装此技能
npx skills@latest add mattpocock/skills --skill=tdd然后输入 /tdd 来调用它。
本页内容
它的作用
tdd builds a feature or fixes a bug test-first: one failing test, then just enough code to pass it, then the next behaviour. It carries the standards that make that loop produce tests worth keeping: what a good test is, where tests go, what mocks are for, and the three anti-patterns that make a suite worthless.
It writes no test at a seam you have not agreed to first. Before any test exists, it names the public boundaries it intends to test at and stops for your confirmation. Testing effort is finite, and this step spends it on the critical paths instead of on every edge case. tdd is also a reference, not a driver. It contains the rules of the loop, and something else (you, or implement)运行 会话 应用它们。
何时使用
输入 /tdd,或这个 agent reaches for it automatically when a task fits: building a feature or fixing a bug test-first, or when you say "red-green-refactor".
当有具体行为要构建、有输入和可观察的输出、并且你想要经得起重构的测试时使用它。
| 你的情境 | 去哪里 |
|---|---|
| A behaviour with defined inputs and outputs (business logic, a request/response contract, a transformation, validation) | tdd |
| 行为还没被钉死 | to-spec,它还会在任何代码写出之前就约定测试接缝 |
| 问题其实是接口的形状,而不是测试 | codebase-design |
| 你有一个 spec or tickets 并想让整个构建替你完成 | implement,它驱动 tdd 每个任务 |
| 配置、接线、胶水、类型注解、直白 CRUD 委托 | Nothing here fits well; see the open gap below |
That last row is a real gap. The skill decides where the seams go, but nothing in it decides whether a change is worth the loop at all. If you run it on a change with no independent source of truth to assert against, you get a test that restates the implementation. That is the tautological anti-pattern the skill warns about, reached from the other direction. It is issue #746, and it is open. Until it closes, you make that call yourself, or write the rule into your CLAUDE.md.
前置条件
codebase-design 需要被安装。 tdd used to carry its own deep-module and interface-design notes; v1.0 deleted them in favour of the shared skill, and tdd now uses its interface-design vocabulary. Nothing else; the skill is 无状态 并且不写自己的文件。
循环,以及它运行的接缝
The skill rests on three terms.
红-绿。 Write the failing test, then only enough code to pass it. Do not write code for the test after next. There is no refactor phase. The skill dropped it in June 2026 because agents almost never performed it, and because review and implementation work better as separate sessions. Refactoring belongs to code-review.
垂直切片。 Write one test at one seam, then the minimal implementation, then repeat. The first cycle is a 曳光弹 that proves a single path end to end. The opposite is horizontal slicing: all the tests first, then all the code. Tests written in bulk verify imagined behaviour. They check the shape of things rather than what a user does, and they commit you to a test structure before you understand the implementation.
预先约定的接缝。 A seam is the public boundary you observe behaviour at without reaching inside. The rule has no exceptions. No test goes at an unconfirmed seam. In the full chain the seams are agreed earlier, during to-spec: "/tdd 被告知只在预先约定的测试接缝处工作, /code-review 检查只使用了约定的测试接缝。」单独调用时, tdd 直接问你。
它写来预防的三个反模式:
| 反模式 | 征兆 |
|---|---|
| 实现耦合 | 重命名内部函数时测试破裂,尽管行为没有改变。被 mock 的内部协作者、被断言的调用次数、用来验证的数据库查询——而不是接口。 |
| 同义反复 | The expected value is computed the way the code computes it, so the test passes by construction. Expected values have to come from somewhere else: a known-good literal, a worked example, the spec. |
| 横向切片 | 一批测试在实现之前就落地了。 |
Mocks are for system boundaries only: external APIs, time, randomness, sometimes the filesystem or the database. Never mock your own modules.
常见问题
它为什么不重构?描述说「红-绿-重构」。
Because the refactor step was removed and the description was not. The removal was deliberate. Agents almost never did the step, and keeping implementation and review in separate sessions works better. Whether the result still counts as TDD by the book matters less than whether the loop produces better code. The mismatch between the trigger phrase and the body is filed as issue #589 and is still open, so "red-green-refactor" continues to work as a phrase that fires the skill. What you get is red-green, with refactoring in code-review.
它让我选测试接缝,我完全不知道该选哪个。
这是该技能被报告最多的摩擦(issue #607). The prompt lists candidate seams by name only, with nothing about what each one catches or misses, so you are choosing between labels. There is no fix shipped yet. The practical workaround is to ask the agent for the trade-offs before answering: what does the component-level seam miss that the integration seam catches, and how much slower is it. It is also why the chain agrees seams up front in to-spec,在那里你能看到整个功能,而不是一个提示词。
它先写了实现,尽管技能说先红。
确实会发生。一位用户按下了 model on it and got an unusually honest answer: "I knew the skill said 'one test at a time, watch it fail for the right reason'. I read it. I just defaulted to my normal habit." The skill accepts this. No instruction makes an agent comply 100% of the time, and stricter wording restricts the agent's creativity for little gain. The loop is worth running even when the agent does not follow it strictly, because the results are still better overall. If strict adherence matters for a particular slice, watch the run rather than trusting the skill to enforce it.
它应该先写浏览器测试还是端到端测试?
Usually not, and the skill will not stop it. A user reported the agent writing a Playwright test first, then spending a long loop re-running it and concluding the test was broken for a feature that did not exist yet. Browser tests are slow enough that the red-green feedback loop stops paying for itself. State in your repo's CLAUDE.md that browser tests come after the behaviour works.
是否 /tdd replace /implement,或课程的 /do-work?
不。 /tdd 记录方法论; /implement is a simple work→feedback→commit loop and is the direct stand-in for /do-work。课程的唯一 /do-work 步骤现在被拆在 /implement, /tdd and /code-review。如果你在问对某个任务该运行哪一个,答案几乎总是 /implement.
深模块和接口设计指南去哪了?
进入 codebase-design 在 v1.0 中,经过泛化让多个技能共享一套词汇。 refactoring.md 同时离开;重构现在 code-review的职责,而那个技能带有 Fowler 坏味道基线。
它知道我其他的任务吗?
No. Run against one ticket, it can propose work that belongs to a sibling ticket, because it has no view of the rest of the issue graph (issue #129). This is not tdd的职责。把规格说明与任务一起传递会有帮助;从一开始就把任务切成合适大小更有帮助。
做到以下就算成功
- 在任何测试文件存在之前,它停下来点名打算测试的接缝,然后等待。
- One test appears, goes red, gets just enough code to pass, and only then does the next test appear, not a batch of tests followed by a batch of code.
- 测试名读起来是能力(「用户可以用有效购物车结账」),而不是内部实现(「checkout 调用 paymentService.process」)。
- 断言中的期望值是你能追溯到规格说明的字面量,而不是按代码的计算方式重算出来的值。
- 重命名内部函数不会破坏测试套件中的任何东西。
- Mocks appear only at external boundaries (the payment API, the clock) and never around your own modules.
在流程中的位置
tdd runs inside the build step of the main chain; it is not a step of its own:
grill-with-docs → to-spec → to-tickets → implement → code-review → retro
to-spec 预先约定测试接缝, implement drives tdd 每个任务,而且 code-review checks afterwards that only the agreed seams were used, and owns the refactoring tdd 不再如此了。它的另一个邻居是 codebase-design, the shared source of the seam and deep-module vocabulary that tdd uses. You can also reach for it on its own, whenever there is a concrete behaviour to build and no full spec in play. When you are unsure which skill fits your situation, ask-matt 为你指路。
技能操作
npx skills@latest add mattpocock/skills安装整套技能,然后在智能体中输入 /tdd 来调用它。