ATROPOS Evidence-based AI cost reduction, by Ozymind

Measure the thread. Then cut.

About 95% of enterprise AI pilots never show up in the P&L, and rarely because of the model. ATROPOS closes the gap: five gates, each with a test that can fail, effect sizes from peer-reviewed research rather than vendor decks, and kill criteria agreed before anything is built.

Book a Phase-0 diagnostic Try the arithmetic ↓
01

The problem

The savings are real. They are just not everywhere. Enterprise studies and macro estimates point the same way: value concentrates in specific, findable tasks, and finding it is engineering work, not a licensing decision.

~95% of enterprise GenAI pilots show no measurable P&L impact MIT Project NANDA, The GenAI Divide, 2025
−19 pp correctness on a task chosen beyond the AI frontier Dell'Acqua et al., Organization Science, 2026
+34% throughput for novice agents where AI lands well Brynjolfsson, Li & Raymond, QJE, 2025

Knowing which task you are standing on is the whole game. You can test the arithmetic below.

02

The equation

Five factors, multiplied task by task, minus what the system costs to run. Drag the sliders and watch what a headline claim becomes.

Ozymind's Razor

A savings claim that does not decompose is a slogan. A saving that changes no budget line is a story.

Two blades, applied five times: that is the whole framework. After William of Ockham, who taught the oldest rule of cutting: do not multiply claims beyond what you can measure.

One task. Adjust the factors.

Salaries, contractors, licences, plus the cost of errors and rework
The programmes that fail assume 100%
Taken from an instrumented pilot, never from self-report
The share of eligible work that actually goes through the new path
Effects fade. We re-measure at month twelve.
Inference, engineering, review, governance, refresh. In our experience, underestimated by a factor of two to four.

What survives measurement

€59,213 Modelled net annual saving on this task 8.7% of the task's annual cost

A deck promising “40%” would have put €272,000 on the slide. 78% of that would not have survived measurement.

The headline claim (40%)
€272,000
Automatable and measured
€187,000
After adoption and retention
€119,213
Bankable, net of TCO
€59,213

The defaults show a realistic accounts-payable task. Modelled, not banked: Gate 5 is what turns a number like this into a budget line. Everything runs in this page; nothing you enter is sent anywhere.

03

The method

The Moirai of Greek myth worked in strict order: Clotho spins, Lachesis measures, Atropos cuts. So does the framework. Work passes one gate at a time, and every gate ends with a test that can fail.

Act I · Clotho spins Act II · Lachesis measures Act III · Atropos cuts

Gate 1. Decompose the cost base into tasks, not headcount

Take the function's fully loaded cost: salaries, contractors, licences, external spend, the cost of errors and rework. Break it into 20 to 60 discrete tasks, each with a measurable unit of production such as tickets, invoices, reports or dossiers. Give every task a volume, a unit time, a unit cost, an error rate and a cycle time.

Exit criterion

At least 85% of the function's cost is allocated to named tasks with a unit of production.

Why this gate exists

Every credible effect size in the literature is task-level, not role-level. Brynjolfsson, Li and Raymond measured issues resolved per hour, which is a task unit. Bottom-up beats top-down FTE targets.

Gate 2. Map every task against the jagged frontier

AI capability is jagged. Inside its frontier, consultants in the BCG and Harvard field experiment completed 12.2% more tasks, 25.1% faster, at around 40% higher rated quality. On a task chosen to sit beyond the frontier, the same people were 19 percentage points less likely to get the right answer. Every task therefore gets a zone label: inside (automate and measure), edge (human in the loop), or outside (do not automate, do not even pilot). Regulatory class (EU AI Act, GDPR Article 22, DPIA) is settled here too, because it changes the verification layer, and the verification layer changes the cost.

Exit criterion

Every task carries a zone label, a verification method, and a named cost for a silent error.

Why this gate exists

The negative result is the expensive one. The wrong deployment does not produce zero savings. It produces negative savings plus a quality problem.

Gate 3. Baseline from systems of record, never from self-report

In METR's randomized trial, experienced developers expected to be 24% faster with AI and reported afterwards that they had been 20% faster. Measured, they were 19% slower: a 39-point gap between perception and reality. A 2026 follow-up cohort measured roughly minus 4%, which METR itself calls weak evidence. Either way, self-report cannot be the baseline.

Exit criterion

A pre-intervention baseline exists for each task, from systems of record, with a staggered rollout designed in before anything is deployed.

Why this gate exists

A staggered rollout gives you a comparison group, and a comparison group is what makes a savings claim stand up in front of an auditor.

Gate 4. Rebuild the workflow, do not bolt a tool on

This is where the 95% failure rate gets decided. Four layers have to exist before a saving is real. A data layer: the inputs the task needs, retrievable, current and permissioned. A task layer: deterministic logic stays in code, and routing, scheduling and allocation stay optimization problems, not LLM prompts. A verification layer: every output gets a check sized to the cost of a silent error. And a process layer: handoffs, approvals and SLAs redesigned, because if the process still needs the same three approvals, the saving never reaches the P&L.

Exit criterion

The workflow runs end to end in production, on real volume, for one full cycle, with verification active and an exception path defined.

Why this gate exists

In our experience, stalled pilots stall at the data layer and vanished savings vanish at the process layer. The tool was never the constraint.

Gate 5. Bank it: turn time saved into money, explicitly

Time saved is not cost saved. The conversion has to be named in advance, and there are exactly four options: absorb growth with the same team, reduce external spend on agencies, contractors and licences, reduce the cost of errors and rework, or reduce headcount. The last one is the slowest, the most regulated and the riskiest. The first three are what the payroll data actually supports.

Exit criterion

Each validated saving maps to a named budget line, a named owner and a date. Finance countersigns.

Why this gate exists

Time freed with no budget line changed is the most common way a technical success becomes a financial nothing. A cut made here does not creep back, because quarterly re-measurement is part of the deal. That is why the framework carries this gate's name.

04

The evidence

The strongest available effect sizes, each labelled for what it is: peer-reviewed, working paper, or preliminary. We use them as starting priors, never as promises. Two patterns run through all of it: the gains concentrate at the bottom of the skill curve, and they can turn negative at the top. Pointing AI at your senior specialists to save money is the most reliable way to destroy value that we know of.

Customer support +15%

Issues resolved per hour, across 5,172 agents. Novices gained 34%. The most experienced gained almost nothing, and lost a little quality.

Peer-reviewed · Brynjolfsson, Li & Raymond · Quarterly Journal of Economics 140(2), 2025
Knowledge work +12.2%

More tasks done, 25.1% faster, at roughly 40% higher rated quality inside the frontier. On one task chosen beyond it, 19 points less likely to be correct. 758 BCG consultants.

Peer-reviewed · Dell'Acqua et al. · Organization Science 37(2), 2026
Software engineering −19%

Experienced developers were slower with early-2025 tools, while believing they were faster. A 2026 follow-up cohort measured about minus 4%, which METR itself calls weak evidence with wide intervals. We quote both.

Preprint RCTs · METR · 2025 & 2026
Macro ceiling ≤0.66%

Economy-wide total factor productivity gain over ten years under task-based accounting. An order of magnitude below vendor projections.

Peer-reviewed · Acemoglu · Economic Policy 40(121), 2025
Enterprise programmes ~95%

of GenAI pilots showed no measurable P&L impact. The failures were integration and measurement, not model quality. A directional figure from a small sample, and we treat it that way.

Preliminary industry report · MIT Project NANDA · The GenAI Divide, 2025
Employment effects −19%

Employment of workers aged 22 to 25 in the most exposed occupations sits 19% below the path of less-exposed peers. The authors call it early, descriptive evidence; experienced staff show no gap, and Danish data shows aggregate effects near zero. The realistic lever is reduced backfill, not cuts.

Working papers · Brynjolfsson, Chandar & Chen · Stanford Digital Economy Lab, 2025, rev. 2026 · Humlum & Vestergaard, 2025
05

The leak test

Six statements. Tick what is true for you today; each maps to the gate that closes it.

Reading

No leaks ticked. Either impressive, or unmeasured.

Both are possible. A Phase-0 diagnostic settles which one you are, in three to four weeks.

Get the diagnostic
06

Headcount

The evidence supports reduced backfill, reduced external spend and absorbed growth, not structural reduction. We do not sell headcount targets.

Payroll data shows narrow displacement, concentrated in early-career hiring; aggregate effects sit near zero. In Belgium, structural reduction can trigger works-council information and consultation, and at scale the collective-dismissal procedure, with lead times measured in quarters depending on size and thresholds. And entry-level roles are both where AI gains are largest and the pipeline that produces your future seniors: cutting them books a five-year cost against a one-year saving. When a client asks for headcount scenarios, we model them, costs in the open.

07

The engagement

  1. Phase 0. Diagnostic (3 to 4 weeks, fixed price)

    Cost decomposition, task inventory, frontier classification, data-readiness check, and a plan for baseline instrumentation. You get a ranked opportunity register with a go or no-go per task. You can stop here and keep a usable asset.

  2. Phase 1. Instrumented pilot (6 to 10 weeks)

    Two or three of the highest-ranked tasks, built properly: data, task, verification, process. Rolled out staggered, against a comparison group. You get a production workflow, a measured effect size, a validated TCO and a scaling decision. Measured, not asserted.

  3. Phase 2. Scale and bank (ongoing)

    Extension to adjacent tasks, monitoring in place, savings mapped to budget lines and countersigned by Finance. We re-measure quarterly, because effects fade and long tails are where cost sneaks back in.

Kill criteria, agreed upfront: if the measured effect does not clear the pre-registered threshold net of TCO, the task is retired. A programme with no kill criteria is a programme with no measurement.

Six things we will not do

  • No savings claim without the five-factor decomposition behind it.
  • No deployment on outside-frontier tasks, no matter how good the demo looked.
  • No baseline from surveys. Systems of record, or we do not start.
  • No pilot without a pre-registered threshold and kill criteria signed before the build.
  • No “time saved” booked as money until Finance countersigns a budget line.
  • No headcount targets sold as deliverables.
08

The name

Hesiod names three Fates who handle every mortal thread. Clotho spins it, Lachesis measures it, and Atropos, whose name means “she who cannot be turned”, cuts it. The detail we like: the shears never move before the measuring rod has passed. Cost work deserves the same order of operations, and the same finality. A saving cut this way, banked, countersigned and re-measured, does not creep back.

The Greeks also wrote down the first autonomous machines in literature: Hephaestus' self-moving tripods in the Iliad, wheeled devices Homer calls automatoi. People have been thinking about automation and its governance for twenty-eight centuries. We try to keep both halves of that conversation.

Hesiod, Theogony 904–906 · Homer, Iliad XVIII 373–379 · Adrienne Mayor, Gods and Robots, Princeton University Press, 2018

Antique tailor's shears and a thread spool on pinstripe suiting
09

Questions we get asked

What is Ozymind's Razor?

The two-sentence rule above the simulator. The first blade cuts inflated promises before the build (Gates 1 to 3); the second cuts unfinished victories after it (Gates 4 and 5). Ockham's parsimony, pointed at business cases.

Can you guarantee a percentage saving?

No. Any percentage promised before Gate 1 fails the Razor. What we do guarantee is the discipline: pre-registered thresholds, staggered rollouts, and kill criteria.

Will this eliminate jobs?

The levers the evidence supports are reduced backfill, reduced external spend and absorbed growth. We model headcount scenarios on request, costs in the open, but we do not sell targets.

How long before savings reach the P&L?

Savings count as banked when a named budget line changes and Finance countersigns, typically after the first full production cycle. Expect a dip before the gain: process redesign is an investment that comes before the measured effect.

What happens if the pilot shows no effect?

The task is retired, as agreed before the build. That is the framework working, not failing. You keep the opportunity register, the instrumentation and the measured result.