AIHero
    13 / 27AI Skills for Real Engineers · 9 min read · Updated Aug 24, 2026

    /diagnosing-bugs 技能

    从一个可复现的失败场景出发,诊断疑难 bug。

    Matt Pocock
    Matt Pocock
    下一页

    安装此技能

    npx skills@latest add mattpocock/skills --skill=diagnosing-bugs

    然后输入 /diagnosing-bugs 来调用它。

    本页内容

    它的作用

    diagnosing-bugs 对疑难 bug 或性能回归运行六阶段诊断:构建复现、最小化、假设排序、插桩、用回归测试修复、清理。

    在 tight feedback loop exists: one named command, already run once, that goes red on this bug and green when it is fixed. Given a bug report, a coding agent by default reads code and guesses. This skill blocks that. If no red-capable command exists, there is no Phase 2. That single gate is what the skill is for. Once the loop exists, everything after it (bisection, hypothesis-testing, instrumentation) is mechanical.

    何时使用

    输入 /diagnosing-bugs, or the agent reaches for it on its own when a task fits. It is model-invoked, and fires on "diagnose" or "debug this", or on a report that something is broken, throwing, failing, or slow.

    Reach for it on the hard ones: a bug you can't solve at first look, an intermittent flake, a regression introduced between two known-good states. It is slow and thorough on purpose, so it is the wrong tool for a question you want answered in one message.

    你的情境去哪里
    一个你能用症状描述的具体缺陷这个技能
    一个变慢的端点或带有已知前后对比的时序回归This skill. It has a performance branch (measure a baseline, then bisect)
    "Where are the bottlenecks in this codebase?", no specific symptom不是这个技能。它诊断一个已知故障,不审计
    来自他人的原始 bug 报告,尚未确认或整理triage first
    回答设计问题的一次性代码,不是追查缺陷prototype
    以测试先行构建计划好的行为tdd
    Asking what would have prevented the bug, once it is fixedretro, run in the same session
    没有好的接缝来锁定 bugimprove-codebase-architecture, which you start yourself

    紧密循环就是技能

    Phase 1 gets the most effort because it is the only hard phase. The skill lists ways to build the loop, roughly in order of preference:

    1. 在能触达 bug 的接缝处有一个失败的测试。
    2. 针对运行中开发服务器的 curl 或 HTTP 脚本。
    3. 带 fixture 输入的 CLI 调用,与已知正确的快照做 diff。
    4. 一个对 DOM、控制台或网络进行断言的 headless 浏览器脚本。
    5. A replayed capture: a saved request, payload, or event log, run through the code path in isolation.
    6. 一个一次性的驱动装置:系统的最小子集,一次函数调用。
    7. 针对「时好时坏的输出」的属性测试或模糊测试循环。
    8. 一个可以交给 git bisect run.
    9. A differential loop: same input, old version against new.
    10. A human-in-the-loop bash 脚本,最后的手段。该技能自带 scripts/hitl-loop.template.sh 为此:智能体运行脚本,你在终端里跟着提示走,你的回答以可解析的输出返回。

    A 循环不是目标。 紧密 is: fast (seconds), deterministic (same verdict every run), sharp (asserts your exact symptom, not "didn't crash"), and runnable by the agent without you. A 30-second flaky loop is barely better than none. For a bug that only shows up sometimes, the target is not a clean repro but a 更高的复现率: loop the trigger, parallelise, add stress, inject sleeps, until the flake rate is high enough to debug against.

    When the agent cannot build one, the skill tells it to stop and say so, list what it tried, and ask you for 环境 access, a captured artifact, or permission to add temporary instrumentation. It must not go on to form hypotheses anyway.

    阶段之间的门控

    The phases are gates, not a checklist. The agent cannot enter a phase until a specific condition is true.

    门控什么必须成立
    进入第 2 阶段一条已运行并连同输出一起粘贴的具名命令,它能在这个 bug 上变红
    进入第 3 阶段复现被复现 and minimised: every remaining element is needed to reproduce the bug
    进入第 4 阶段有 3–5 个排序后的可证伪假设,每个都陈述其预测,并在测试任何假设前展示给你
    进入第 5 阶段Probes map to a specific prediction, one variable at a time, every debug log has a tag like [DEBUG-a4f2], so one grep finds them all for cleanup
    完成The original repro no longer reproduces, the instrumentation is gone, and the commit message names the hypothesis that turned out correct

    Phase 5 has one exception. The agent writes the regression test before the fix, but only if a 正确的接缝 exists for it: one where the test exercises the real bug pattern as it occurs at the call site. Where the only available seam is too shallow, the skill tells the agent to say so instead of writing a test that gives false confidence. The missing seam is itself a finding, and the agent records it instead of hiding it.

    常见问题

    它在我只想要直接答案的快速问题上触发了。 This is the most-reported problem with the skill. On GPT-5.6-Sol especially, users report it triggering on a plain description of a problem: "the model triggers the rather formal diagnosing-bugs skill instead. It then goes on to construct a reproduction scenario (often building a mock scenario with limited value) before giving me a response or suggestion. This results in considerable reply delays." Four separate people reported the same problem on issue #578. The accepted fix is to start with a lighter approach and move to the full diagnosis only when the problem needs it, but that change has not shipped. The skill is tuned for Claude Code's invocation behaviour, and a model with a lower activation threshold fires it too often. Until the fix ships, the workaround is to say what you want ("just answer this, don't diagnose") or to disable model invocation for it in your 运行框架.

    我能让它指向代码库并问性能问题在哪里吗? No. It diagnoses one failure you can already name. Its performance branch is for a regression with a symptom (establish a baseline measurement, then bisect, measure first and fix second), not for a proactive sweep. A skill for the proactive version was 被提议并关闭; no skill does it now.

    它在写下修复之前会停下来问我吗? No. Only Phase 3 has a human checkpoint: the ranked hypothesis list is shown to you before any is tested, and the agent continues with its own ranking if you are away. There is no gate between instrumentation and the fix, so the agent can start writing code before you have agreed with its root cause. Issue #124 要求那个门控,且仍然开放。如果你想要它,调用技能时就说出来。

    我已经运行过 /triage 在这个 bug 报告上。这是同一个工作再次出现吗? Partly, and neither skill says so. As one reader put it: "Triage's step 3 is essentially a shallow, bounded instance of diagnosing-bugs Phase 1–2, but neither file mentions the other." Triage does a bounded "is this actually a bug, and what is the surface" pass; this skill does the thorough version. Running triage first is not wasted, because its verification often gives you most of Phase 1's raw material. But expect this skill to redo that work in full, and expect no cross-reference to tell you so.

    它粘贴的复现输出会泄露密钥吗? It might. The skill asks the agent to paste the invocation and its output, and to request artifacts like HAR files, log dumps, and core dumps. No instruction tells the agent to sanitise them. Issue #674 raises exactly this (credentials, tokens, cookies, and personal data copied into a chat, an issue, or a PR) and proposes a redaction guardrail. It is open and unimplemented. Treat redaction as your job for now, particularly before the output goes anywhere public.

    我的安全扫描器把这个技能标记为高风险。 Snyk 标记了它,而标记是误报。它是套装里唯一发布可执行 shell 脚本的技能(hitl-loop.template.sh) alongside instructions to run it and to curl a dev server. A shipped .sh file, instructions to run it, and outbound HTTP together are enough to trigger a static scanner. The script itself is about 30 lines of read -r -p prompts that pause for human input. The scanner rates what the skill could do, not a proven exploit.

    发生了什么 /diagnose? v1.0.0 renamed it to /diagnosing-bugs. The old name no longer exists. Anything of yours that chains /diagnose (a wrapper skill, a saved prompt) needs updating.

    做到以下就算成功

    • 它先给你看一条命令和它的红色输出,然后才提出任何理论。如果理论先到,说明技能没有在运行。
    • 它复现的失败正是你报告的那个,不是它路上发现的近似物。
    • It shrinks the repro before it starts guessing, and can tell you why it needs each remaining piece.
    • It shows you a ranked list of 3–5 hypotheses, each with a prediction you could falsify, before it tests any of them.
    • 它添加的每条调试日志都带一个像 [DEBUG-a4f2],而当它宣布完成时,对那个标签的 grep 返回空。
    • 提交或 PR 消息点名哪个假设是对的。
    • When it cannot lock the bug down with a test, it says so instead of writing a shallow one.

    在流程中的位置

    diagnosing-bugs is a reach-for-it-anytime standalone. You start it when something is broken, and it ends when the fix and its regression test are in. It keeps no state and needs no prior setup. ask-matt 把「有东西坏了」路由到这里。

    两个邻居很重要。 retro comes after it: once the fix is in, run it in the same session to ask what would have prevented the bug, while the session has more information than it had at the start. diagnosing-bugs never invokes retro itself, because retro is user-invoked. triage comes before it for bugs that arrive as raw reports from other people, and does a shallower version of the same first two phases.

    技能操作

    安装技能

    Live Skills.sh install count
    npx skills@latest add mattpocock/skills

    安装整套技能,然后在智能体中输入 /diagnosing-bugs 来调用它。

    用以下命令更新: npx skills updateSkills.sh