Skip to content
Back to the shelf

Seven expensive problems.

Front matter · The problem

Seven expensive problems.

The expensive part is never the work. It is the delay, the rework, and the decision taken without the number.

The problem

Every problem on the shelf above has the same shape. Something that should take an afternoon takes a week; something that should be checkable is taken on trust; and a decision that will cost real money gets made before the figure that should inform it arrives. None of that shows up as a line item, which is exactly why it persists.

Skilled time spent assembling, not deciding

A finance lead spending the first week of every month reconciling four systems is the most expensive data pipeline in the company, and the least maintained one.

Decisions made before the number arrives

When the pack lands after the meeting, the decision was made on instinct and the reporting became a record of what already happened rather than an input to what happens next.

Numbers nobody will stand behind

A figure two departments compute differently is worse than no figure: it gets argued about instead of acted on, and the argument recurs every month.

The clause found after it renewed

Documents read by one person, under time pressure, produce misses that surface as a cost months later — with no record of why that document was treated the way it was.

How an engagement runs

Four stages, each ending in something you can hold rather than a status update.

Process · 01

We look at the actual decision and the actual data, and agree what a good outcome would be before anything is built. Often this is where a project gets smaller.

Discovery
A written problem statement, and an honest view of whether it is worth doing.

Process · 02

One workflow, one document type, one forecast. Small enough that finding out it does not work is a cheap answer rather than a sunk cost.

Bounded pilot
A working prototype and an evaluation on your own data.

Process · 03

The thing that survived the pilot gets built where the work actually happens, with the review step in place from the first day rather than added later.

Delivery
The deployed workflow, its documentation, and the checks that show when it drifts.

Process · 04

Your team runs it without me. That includes knowing how to re-evaluate it when the inputs change, because they will.

Handover
Documentation, the evaluation harness, and a walkthrough with whoever owns it next.

What the work looks like when you check it

Three things I can show you rather than assert. Each links to the artefact and the numbers behind it, including where they fall short.

Evaluation set29 claims

Method
Environmental claims extracted from corporate communications and rated for greenwashing risk by human reviewers, forming the reference set the detector is scored against. Human expert ratings collected via a structured questionnaire, then used as ground truth.
As of
2026-01-29
Limitations
Twenty-nine items is a small evaluation set. Every figure derived from it carries wide uncertainty and none of it should be read as a population estimate.

Evidence

Krippendorff's α0.69

Method
Agreement between the human raters who produced the reference ratings. Krippendorff's alpha across raters; pairwise agreement ≈ 0.852.
As of
2026-01-29
Limitations
α = 0.69 is substantial but not strong agreement: the humans themselves disagree on roughly a third of the signal, which caps how well any model can be expected to match them.

Evidence

Pearson r (v1)0.906

Method
Correlation between the v1 detector score and the human reference rating, over the 29-claim set. Pearson correlation on continuous scores. Spearman 0.811, Kendall 0.627, MAE 0.206, RMSE 0.237.
Compared with
v3 rule-based detector, r = 0.793; v2 semantic detector, r = 0.557.
As of
2026-01-29
Limitations
Correlation on 29 items. It says the score moves with reviewer judgement, not that the score is calibrated as a probability of greenwashing.

Evidence

Recall / precision (v3)0.40 recall at 1.00 precision

Method
Binary flag / no-flag performance of the v3 rule-based detector against the human reference labels. Thresholded score compared with binarised human labels on the same 29-claim set.
Compared with
v1 baseline: recall 0.15 at precision 1.00 (F1 0.261). v3 reaches F1 0.571.
As of
2026-01-29
Limitations
Recall of 0.40 means roughly three in five flaggable claims are missed. Precision of 1.00 on a set this small is a handful of correct positives, not a guarantee. The tool is a reviewer aid; it cannot be a filter that runs unattended.

What decision or workflow are you trying to improve?

Tell me the problem in a few sentences. I reply personally, usually within a couple of working days, and I will say plainly if it is not something I should take on.

Discuss your project
100%

Drag the page to turn · Drag the glass across it

Nothing on this site is accounting, audit, tax, legal or investment advice, and no figure here is a promise of a result. Where a number appears, its source and its limits appear with it.