Book of Errant Pages

AI has a pickaxe polishing problem (P3)

Les Wilhelm

A miner holds a gleaming golden pickaxe aloft while a crowd of astonished miners looks on and the actual digging goes on behind them

The Pickaxe Polishing Problem (aka P3)

Imagine your task is to mine for gold. You grab a pickaxe and start swinging. Soon you notice a fellow miner who, instead of mining, is polishing their pickaxe. You ask if the polish makes a difference. They explain it helps them mine faster. The polished pickaxe certainly looks more impressive but you ask how much faster it mines. The miner replies they do not have any numbers, but it feels faster. In fact, others who have polished their pickaxes say the same thing. And with so many people coming to the same conclusion there must be merit to the idea.

Imagine your task is to create software. You setup your tool stack and AI harnesses and start working on a ticket. You notice a fellow developer who, instead of working a ticket, is developing an AI SDLC framework. You ask if the framework makes a difference. They explain it decreases token cost and prevents the AI from writing bad code or tests (Gloaguen et al., 2026). The scale of their framework with all of its skills, hooks, and other harness modifications certainly looks impressive but you ask if there are metrics describing its benefits. The developer replies they do not have any numbers but it feels better. In fact, other team members who have used the framework say the same thing. And with so many people coming to the same conclusion there must be merit to the framework.

Tool work has value, but only if the result is a tool which provides more value. A distinction between a sharper pickaxe and a polished pickaxe must be made.

The Spread

I have seen numerous presentations from devs who are convinced they have corrected all of the issues of AI development with their frameworks. Their slide decks contain incredible estimates of the framework's value, but notably there are no reproducible claims, for example a documented A/B test showing decreased token consumption. In each presentation I ask if they have reproducible metrics to back up their claims. They all say they do, but the numbers are not part of the presentation. This is a red flag. I ask for the numbers to be provided later. They say they will, but then they forget. I follow up and they say they are busy. If I am lucky I get a tiny table which is a rearrangement of their non-testable claims.

At one point in the past devs displayed their AI skill by consuming massive numbers of tokens. Once that fell out of fashion another method to impress their bosses was needed. AI harness changes, AI SDLCs, skill repos, and other such activities became avenues of easy recognition. With a technology changing so rapidly even if their claims were investigated they could claim changes to the underlying models altered the performance of their solutions.

I do not believe most devs do this knowingly. AI is legendary for convincing people they have created something worthwhile (Sharma et al., 2024). Suppose you are in this position. You truly believe you have created something of value for your organization (Norton, Mochon & Ariely, 2012). If your organization does not require formal testing then why should you expend the effort? If the tests show genuine benefits then they have only confirmed what you know. If they show no benefit then all of your labors are laid waste (Arkes & Blumer, 1985). Between positive recognition by your org and realizing your work was pointless what would a rational actor choose?

The Correction

Your organization should already have a set of standards for the adoption of tools. If you do not then you should correct that first. Then, create a clear set of standards for AI tooling (Anthropic, 2026). This will be more complex than normal since AI tools are non-deterministic, but it is possible (Miller, 2024). Tooling should state the models it is validated against. If new models are released the tool should be revalidated against the new model.

Leadership needs to provide clear accountability. Showering devs with praise, prizes, and promotions for unverified work will establish an incentive which will quickly lead to your company being blinded by a horde of polished pickaxes (Kerr, 1975). If you want to succeed then go and dig!

Additional Citations