⚡ Powered by Finn · Day 117 of 365
117

The Skill That Went Red While I Slept

The verdict was red before I opened the laptop this morning. A skill I installed had read the overnight write log on a client's books, found three entries by an author it did not recognise, and refused to sign off. Nothing was changed. It held everything and left me five questions for the bookkeeper. That is what I want from Claude Code skills now. Not that they write code faster, but that they catch the AI in its own mistakes before those mistakes reach a client.

Claude Code runs skills, small installable instruction sets that change how the agent behaves. Most people reach for the ones that go faster: scaffold this, ship that. The essential ones do the opposite. They slow the agent down at the exact moment it is about to lie to you with a green light. I lifted most of mine from an open toolkit called gstack and pointed them at the bank reconciliation I run across several legal entities for a client, which is about the least forgiving place to let an AI be confidently wrong.

The first is a guard that audits the books while I sleep. It reads the overnight write log and flags any entry it cannot attribute to itself or a named human. This morning it went red on exactly that: three writes by an author it did not recognise, and, to its credit, one of its own writes missing from the machine's audit log. It changed nothing. Each red line became a question for a human, not a fix the AI talked itself into. I wrote up the wider version of this in eighteen findings from deploying AI agents safely.

The second is a review skill, and it earns its keep by reading the whole record, not the last sample. I had a rule for how to treat one recurring line on the books, mined from a single month where it looked obviously right. Run against years of ledger history instead, it did not hold: across the months it contradicted itself, and two other rules built the same quick way fell over with it. One month is a guess. This skill will not promote a rule until the same treatment proves out across at least three separate months, with the months named as its evidence.

The third is an investigate skill with one rule it will not bend: no fix without a root cause. When my reconciliation once ran clean and still skipped nearly half the statement, the tempting move was to widen the matcher until the exceptions went quiet. Investigate refuses that. It made me find why the lines dropped, which was a supplier list from the month before, before it would touch a thing.

The fourth is the one I trust most, a completeness check, and it comes straight out of that same clean-but-wrong run. Green used to mean the script finished, not that the books were right. Now every line has to be posted or sitting on a named held list with a reason, and posted plus held has to equal the statement to the cent or the run stops and names what it cannot account for. It is the same instinct behind testing AI-generated code: never let "it ran" stand in for "it is correct."

An essential skill is not the one that makes Claude quicker. It is the one that stops Claude handing you a result that looks clean and is not. The fast skills are nice. These are the ones that let me sleep while an AI touches a client's books.

I do this kind of thing for clients most weeks, building the AI and then building the things that catch the AI. New here? The start of all this explains what I am building in public, one day at a time, and the fund it exists to pay for explains why.

Monthly Revenues $11,000 | Clients 2 | Prospects 1 outbound live, Meta and WhatsApp still down

Day 117 of 365.

Get the next build-in-public post by email

One short dispatch most days — the AI ops builds, what's working, what isn't. No spam.

← Day 116 All posts

Follow the BIP

AI Deployment as a Service. One workflow at a time.

Book 15 minutes. We see if your workflow is one we'd build.

Schedule a call