AIHero
    12 / 27AI Skills for Real Engineers · 9 min read · Updated Aug 24, 2026

    /improve-codebase-architecture 技能

    以可视化报告的形式,找出值得重构的模块。

    Matt Pocock
    Matt Pocock
    源代码下一页

    安装此技能

    npx skills@latest add mattpocock/skills --skill=improve-codebase-architecture

    然后输入 /improve-codebase-architecture 来调用它。

    本页内容

    它的作用

    improve-codebase-architecture 调查代码库寻找 深化机会. These are places where a shallow module (an interface nearly as complex as the thing it hides) could become a deep one. It writes them up as a self-contained HTML report, and then grills you through the one you pick.

    It never changes the code. The whole run produces a conversation and one HTML file in your OS temp directory. You do the refactor later, in a separate 会话, through the normal build flow. This makes it a survey, not a refactoring tool, so you can run it on a codebase you are not ready to touch yet.

    Two filters stop the report from becoming generic cleanup advice. First, every candidate must pass the 删除测试: if you removed this module, would its complexity move behind a smaller interface, or spread across the callers? Only the first case gets a card. Second, unless you point it at a specific area, it reads recent commit history first and focuses the scan on paths that change often. A deepening in code that nobody touches is a refactor that never pays back.

    何时使用

    你通过输入 /improve-codebase-architecture; the agent 不会自动调用它。

    It is not a step in the main build loop. You run it periodically to queue up more work that improves the codebase. People use it in four situations:

    情境如何使用
    日常维护Run it every few days, or when you have a spare moment, so the structure does not decay between features.
    大型构建之前把它指向 spec and ask "how can we make this change easy?" This is the most effective prompt for it.
    存量代码审计在大型、无结构或 vibe-coded repo to find out what shape it is in.
    遗留测试工作Use it to find the missing seams before you write tests against untestable code.

    它与同类容易混淆之处:

    • To design one module you have already chosen, use codebase-design. This skill finds the module to work on, and codebase-design is where you design it.
    • 对于大到一次会话装不下的整个任务,用 wayfinder.
    • 「这个具体的东西坏了」,用 diagnosing-bugs. It sends you back here when the real finding is that there is no good seam to lock the bug down.

    前置条件

    None to run it. If GLOSSARY.md or ADRs in docs/adr/ exist, it reads them and uses your domain's own nouns. A candidate then reads as "deepen the Order intake module," not "refactor the FooBarHandler."

    它写在两个地方。报告去往 <tmpdir>/architecture-review-<timestamp>.html,在仓库之外。在 追问审视 loop it adds or sharpens terms in GLOSSARY.md, and creates that file if it does not exist. It also offers to record a rejected candidate as an ADR, so a future run does not suggest it again.

    深度,以及追猎它的报告

    The skill rests on one idea: depth. A deep module puts a lot of behaviour behind a small, stable interface. A shallow module exposes its implementation through an interface nearly as wide as the code beneath it. The report looks for three forms of shallowness:

    • Pure functions extracted only for testability, while the real bugs are in how callers use them (no locality).
    • Modules that leak across their seams.
    • A concept you cannot understand without opening five files.

    For each one, it proposes the deepening that fixes it.

    Each candidate is a card with the files involved, the friction, a plain-English solution, the benefit in terms of locality and leverage,一张前后对比图,和一个强度徽章。

    徽章它对你意味着什么
    Strong删除测试明显通过,摩擦真实存在。认真对待它们。
    Worth exploring看似合理的深化,但回报取决于代码接下来走向哪里。
    SpeculativeIncluded for completeness. You can ignore most of these.

    报告以 首选推荐, the candidate it would do first. Then the skill stops and asks which candidate you want to explore. At that point you have decided nothing, and no code has changed.

    选定一个之后会发生什么

    When you pick a candidate, a 追问审视 session starts on it. It covers the constraints, what goes behind the seam, which tests survive, and what the deepened interface should look like. The output of that session is a decision, not a diff. From there the normal flow applies: take the decision into to-spec,然后 to-tickets,然后 implement.

    常见问题

    它为了一个想法审问了我一小时,而不是给我看选项。能关掉吗?

    Yes. Say so when you invoke it ("don't grill me, just show the report"). This is the most common complaint about the skill. One user liked it as "a convenient way to get a thorough analysis of improvements," but after the grilling loop was added, found it "borderline unusable." In their sessions it proposed a single solution and then asked "10's or 100's of questions." The design intent is that the report comes first, and the grilling starts only on a candidate you chose. But weaker models skip straight to an interview about the first idea they had. Results in that thread vary a lot by model. It is an open issue, and the skill does not yet have a documented no-grill mode.

    报告以无样式原始 HTML 打开,没有图表。发生了什么?

    The report loads Tailwind and Mermaid from CDNs, so it needs network access when you open it. If something blocks those scripts, the page breaks with no error. In the reported case, a security hook required SRI hashes. The agent added them, but the CDN served different bytes to the browser than to the curl that computed the hash, so the browser blocked the script. Offline and locked-down environments have the same problem. The agent cannot see it, because it never renders the page. As a workaround, ask for inline CSS and hand-built SVG diagrams instead of the CDN setup. This is an open issue.

    它给了我十二个候选。我在同一个会话里处理它们,还是开一个新的?

    Use one candidate per session. If you work through several in one conversation, the 上下文窗口 fills with the report, the grilling, the domain-model edits and the code changes all at once. The report is only a temp file, so carry the candidate forward, not the file. Pick one, grill it, and take the decision into /to-spec. Turn the rest into tickets that you can pick up separately later. Put the chosen improvement into a spec instead of going straight to implementation. People ask this often, and the skill itself documents no workflow for it.

    我该怎么向它提问?

    Prompt it with the next thing you are building. If a big build is coming up, point it at the spec and ask "how can we make this change easy?" A run with no prompt looks for hot spots on its own. That is fine for routine upkeep, but a direction makes the report actionable.

    它能在大型遗留代码库上工作吗?

    Partly. It works well on big existing codebases that lack consistent structure, and it is the recommended upkeep tool after any one-time structural setup. But users with out-of-control projects report it "helped a little but still doesn't seem to cut it." One developer with an eight-year-old legacy codebase reported that the model went in circles, though the same skill produces a clean graph on a tidy repo. There is no dedicated /refactor skill for that case yet. If the codebase has no shared vocabulary, run grill-with-docs first to create one. That usually makes this skill's output much better.

    这与 /codebase-design?

    /codebase-design is a reference, not a session driver. It supplies the vocabulary (module, interface, depth, seam, adapter, leverage, locality), and this skill uses it. If you give a fresh agent /codebase-design as the task, it fails in a known way. It has no process of its own to follow, so the agent invents one, explores the code again, and runs for a very long time before it asks you anything. Run this skill, and let it use that one.

    它会告诉我代码库没问题吗?

    Rarely, so know that before you start. The skill exists to output findings, so it tends to produce candidates instead of concluding that nothing is wrong. Use the strength badges to correct for this. If every candidate in a report is Speculative, the skill found nothing.

    它能在 Codex 或其他运行框架中工作吗?

    Partly. The exploration step names Claude Code's Agent 工具搭配 subagent_type=Explore directly. A 运行框架 without that tool may skip the parallel exploration instead of using its own equivalent. The skill still runs, but the scan is less thorough. Someone has proposed a harness-neutral rewrite, but it is not merged.

    在 TypeScript 中到底怎么实现深模块?

    The skill does not ship a good answer. People often ask for a TYPESCRIPT.md with concrete file and module layouts for the principles, and it does not exist. The skill tells you where a deepening belongs and what goes behind the seam. You must turn that into a package or directory structure yourself.

    做到以下就算成功

    • The candidates name your domain's concepts, not invented class names: "the Order intake module," not "the FooBarHandler."
    • The candidates are in files you edited recently, not in parts of the repo nobody touches.
    • 运行期间没有任何代码改变。唯一的新文件是临时目录里的 HTML 报告。
    • It stops after the report and asks which candidate you want. It does not continue on its own.
    • Each card explains the payoff as locality or leverage, and says which tests get simpler, not just "this is cleaner."
    • When you reject a candidate for a lasting reason, it offers to record an ADR, so the next run does not suggest it again.

    在流程中的位置

    improve-codebase-architecture is 定期维护. You run it every few days, outside any chain, to queue up work, not to do it. Its neighbours:

    • codebase-design owns the depth-and-seam vocabulary that every candidate uses.
    • 追问审视 walks the decision tree after you choose a candidate.
    • domain-modeling keeps GLOSSARY.md and the ADRs current as you make the decision.

    Its output is an idea, which goes back into the main build flow at grill-with-docs or to-spec. Its counterpart at the end of the main flow is retro. This skill improves the code the agent works in, and retro improves the environment around it (checks, standards, steering files) after a build. For which skill fits a situation, ask-matt 是整套技能的路由器。

    技能操作

    安装技能

    Live Skills.sh install count
    npx skills@latest add mattpocock/skills

    安装整套技能,然后在智能体中输入 /improve-codebase-architecture 来调用它。

    用以下命令更新: npx skills updateSkills.sh