Claude Code Superpowers: Building a PSI Monitor the Careful Way

By Hafizh Yuwan Fauzan · 2026-10-08

The first article in a series on the Claude skills I use every day.

In August, my monthly stability check said a characteristic had shifted. The flag was right that something had changed. It was wrong about what.

Utilisation's characteristic stability index (CSI) jumped from 0.03 in July to 0.30 in August, well past the 0.25 line for a "significant shift". More than a third of that number came from one place: 2.1% of August's applicants had no utilisation value at all, against none in the development sample. A data-quality change was being read as a change in the population.

What I like about how that got caught is that nobody patched it. The tool was built with a Claude Code plugin called Superpowers, and its debugging skill stopped at "this is a design decision, not a code bug" and handed the decision back to me. Like my Claude in Excel and Excel-to-PowerPoint write-ups, this article walks through one whole session, from the first question Claude asked to the review that found two more problems I hadn't thought of. All data is synthetic.

What Superpowers is

Superpowers is a free plugin for Claude Code, made by Jesse Vincent and Prime Radiant and released under the MIT licence. Its own description is "a complete software development methodology for your coding agents": a set of skills that Claude checks before it starts work, so that it brainstorms before building, plans before coding, writes tests first and debugs from evidence. The README puts it bluntly: "Mandatory workflows, not suggestions."

If "skill" is new to you: in Claude Code, a skill is a file of written instructions that Claude loads only when a task calls for it, and a plugin bundles several skills so you install them in one step (Anthropic's Claude Code documentation on skills, as of October 2026). Superpowers' skills appear as superpowers:brainstorming, superpowers:writing-plans and so on.

The facts below come from the Superpowers repository on GitHub, checked on 8 October 2026.

As of October 2026
Install in Claude Code /plugin install superpowers@claude-plugins-official
Licence MIT
Main skills brainstorming, writing-plans, executing-plans, subagent-driven-development, test-driven-development, systematic-debugging, verification-before-completion, requesting-code-review
Also runs in Codex, Cursor, Gemini CLI, GitHub Copilot CLI and several other coding tools (each installed separately)

The workflow is simple to describe. Brainstorming refines the idea through questions and presents a design. A plan breaks the work into small tasks with exact files and tests. The tasks are built test-first. A review runs at the end. Each stage (design, spec, plan, execution) waits for your approval before the next.

I use it for almost anything bigger than a quick fix, including this site: the seven-part credit risk scorecard series started as a Superpowers brainstorm, with a written design I approved before any article was drafted.

The task: a monthly PSI and CSI check

The population stability index (PSI) compares how this month's applicants are spread across score bands with how the development sample was spread. The characteristic stability index (CSI) does the same for each input, such as age or utilisation. Both use the same formula:

PSI (or CSI) = Σ over bins (Actual% − Expected%) × ln(Actual% / Expected%)

where Expected% is a bin's share in the development sample and Actual% its share this month. I covered what the numbers mean in Part 7 of the scorecard series; here the job is building the check itself.

My opening prompt was as short as most of mine are:

i want to make small tool to check PSI and CSI monthly for my scorecard, use superpowers

Step 1: Brainstorming, one question at a time

The brainstorming skill started by classifying the job. A new tool with no existing code to change counts as "architectural", which means the full path: questions, approaches, a design, a written spec, then a plan. It said so out loud, so I could have overruled it.

Then came the questions, one per message:

It offered three approaches (a small Python package with a command line, a notebook, an Excel template), recommended the package, and I agreed.

Claude Code brainstorming: the tool classified as architectural, three questions with the answers chosen, and the approach selected

None of these questions is hard. That's the point. They are the questions I would otherwise answer halfway through the code, after the wrong default was already baked in.

The design came back in four short parts, and after I approved it, Superpowers wrote it to a spec file and asked me to review that file before going further. This is the part I value most. I'm not reviewing code at this stage; I'm reviewing a page of decisions, and correcting a decision is cheap.

Step 2: A plan that found a bug before any code existed

The writing-plans skill turned the spec into five tasks and 26 steps. Each step is small: write this failing test, run it and expect this error, write this code, run the tests again, commit. The plan has to name every file and include the actual code, with no "add error handling here" placeholders.

Its most useful section is "Review Focus": the inputs the spec never mentioned that are most likely to hurt someone using the tool. It listed five, and each became a test:

Then the plan reviewed itself, and caught a real bug. The first draft counted a bin as "empty" when it was empty in either sample, including bins empty in both. An unused Missing bin would have triggered a "bins were floored" warning on every single run, which is the kind of noise you learn to ignore and then miss the one time it matters. The fix was one character, | to ^ (count bins empty in exactly one sample), and an assertion pinning it.

The plan's Review Focus section listing five edge cases, and the self-review fix changing the empty-bin count from bins empty in either sample to bins empty in exactly one

This surprised me most in the whole session. The bug was found by reading a plan, before a line of the tool existed.

Step 3: Tests first, every time

Superpowers then asked how to run the plan: "subagent-driven" (one fresh agent per task, with a reviewer after each) or "native" (everything in this one session, with a single fresh review at the end). My answer:

proceed, native

The test-driven-development skill governs every task. The test is written first and run first, and it has to fail for the right reason: "module doesn't exist yet", not a typo. Only then comes the smallest code that makes it pass. Across the four code tasks the suite grew from 5 tests to 16, with a commit per task; the fifth task generated the data.

Test-driven development in the terminal: the test fails because the module does not exist, then passes after the minimal code; the suite grows from 5 to 16 tests across four tasks

Step 4: The real run, and a number that looked wrong

The last task generated synthetic data: a development sample of 20,000 applicants and three months of 3,000 each, with planned drift. Utilisation creeps upward month by month, some utilisation values go missing from August, and in September the lowest score band is empty.

The tool's output for July, August and September: utilisation CSI rising from 0.0299 to 0.2969 to 0.6778, and the score PSI reaching 0.2866 in September

The September result matched an independent hand calculation exactly (0.6778, recomputed with a different method). But August's jump from 0.03 to 0.30 looked too big for the drift I had planned.

The systematic-debugging skill has one rule above the others: no fix without a root cause. So instead of adjusting anything, it broke August's number down bin by bin.

Systematic debugging: August's utilisation CSI broken down by bin, with the Missing bin contributing 0.1118 of the 0.2969 total

The five utilisation bands contribute 0.185 of it: real drift, and on its own an "investigate". The Missing bin contributes the other 0.112. In August 2.1% of applicants had no utilisation, against 0% in development, and because an empty bin is floored at 0.0001, that one bin's contribution depends heavily on an arbitrary constant.

The code was doing exactly what the spec said. The spec just hadn't anticipated the effect. So the skill stopped and asked me, rather than quietly changing the formula to make the number look reasonable.

Step 5: A fresh reviewer finds what I didn't

Before calling the work finished, Superpowers sends the whole branch to a fresh reviewer that has not seen the conversation. Its verdict was "ready with fixes", and two of its findings were things I would not have tested:

The final review's verdict and findings: an empty month file reporting a shift on every variable, a missing column giving a bare KeyError, and the Missing-bin effect, followed by the fixes passing 19 of 19 tests

Both were fixed the same way as everything else: a failing test first, then the fix. The reviewer also agreed with the decision not to change the Missing-bin behaviour without me, and listed five minor issues, which were noted rather than fixed.

Step 6: My decision, then the tool says what it means

I chose to report both. The main index stays as it is, because a block of new missing values is a real change I want flagged. But each line now also shows the index with Missing set aside and the missing rate in each sample, and the empty-bin warning names the bin:

The tool's output after the change: August utilisation shows 0.2969 shift, 0.1867 excluding Missing, and missing values rising from 0.0% to 2.1%

(The 0.1867 is a little above the 0.185 from the debugging table because setting Missing aside re-spreads the shares over the five bands, so each band's share is compared as a share of non-missing applicants.)

Now August reads as what it is: a genuine drift worth investigating, plus a data-quality change to raise with whoever owns the extract. Two different conversations, two different owners. The suite ends at 21 tests.

Where Superpowers gets in the way

It is too heavy for small tasks. For a one-line fix, three questions, a spec and a plan are pure overhead, and I skip it. Everything else I say about it comes with that caveat.

Why I use it anyway: the plans. A spec and a plan are documents I can read in five minutes and correct before anything is built. Reviewing a plan is far cheaper than reviewing code, and in this session the plan stage alone caught a bug that would have made the tool's warnings meaningless.

Two more honest notes. The approvals are real gates, so a session takes longer than letting Claude write code straight away; I trade speed for fewer surprises. And the final review is only as good as what it is shown. It caught the empty-file case because it was asked to think about the person using the tool, not just the code.

A checklist to try it this week

  1. Install it: /plugin install superpowers@claude-plugins-official.
  2. Pick a task with a real decision in it, not a one-liner.
  3. Answer the brainstorming questions honestly, including "I don't know yet". Those answers are the spec.
  4. Read the spec file and the plan's Review Focus before saying "proceed". That is where your domain knowledge earns its keep.
  5. When a number looks wrong, ask for the evidence before the fix.
  6. Treat a "ready with fixes" review as a gift.

Next in this series: consulting-pptx-skill, which turns a slide rulebook, two checkers and a fresh-eye reviewer into a board deck. Which skill would you want to see after that? Tell me.

Hafizh Yuwan Fauzan (Hafizh Fauzan) is a credit risk data scientist in Jakarta, Indonesia, building scorecards and machine learning models on national-scale credit data.