A new problem

Your AI coding bill is a black box.

Teams now spend more on AI coding than on the machines it runs on. Almost nobody can say what the spend actually bought. This page explains the problem, the process behind it, and the way out.

The problem

Token counts are the new lines of code

For decades the industry knew that counting lines of code rewarded the wrong thing. Then AI coding arrived, and overnight the same mistake came back wearing a new unit: the token.

Token dashboards went up everywhere. Leaderboards followed. And volume became the score: more tokens looked like more productivity, so more tokens is what people produced. Sessions restarted cold, context regenerated endlessly, code generated blind and thrown away. The bill grew either way.

The number on the invoice cannot tell the difference between a developer who spent a dollar understanding a system before changing it, and one who spent the same dollar regenerating the same function eleven times. One of those dollars was excellent engineering. The other was waste. The invoice reports them identically.

What you know today

$48,913

tokens: 1,204,558,102

One number. No story. Pay it again next month.

What was underneath it

understanddesignbuildtestreview

58% blind generation, 6% design. Same bill, very different meaning.

The industry noticed

This conversation is already happening

Industry research

The engineering research community's position is now explicit: token counts are the new lines of code. Cost per outcome, rework rates and team-level measures are what matter.

Public backlash

Internal token leaderboards at major tech companies leaked, drew public criticism, and were dismantled. Volume-based scoring of developers did not survive contact with daylight.

Engineering media

Engineering leaders report throwing out raw token tracking entirely, and say the tooling to connect AI spend to shipped outcomes still barely exists. That missing tooling is exactly the gap phase attribution fills.

Finance teams

AI coding spend has become a standing line item in engineering budgets, growing monthly, with less explanation attached than any comparable cost center.

We keep this list current as the conversation evolves.

The process

Software has a shape

Building software was never one activity. Every meaningful change moves through phases: you understand the system, you design the change, you build it, you test it, you review it. That is as true with AI as it was without it.

Which means AI spend has a shape too. Every token spent belongs to one of those phases. Seen that way, the bill stops being a number and becomes a distribution, and distributions can be read.

Understand

Reading and exploring before changing. Cheap insurance against expensive mistakes.

Design

Deciding the approach before generating. The phase most often skipped, and most missed.

Build

Producing the change. Healthy when it follows thought, wasteful when it replaces it.

Test

Verifying behavior. Its share of spend predicts how often work comes back.

Review

Judging finished work before it ships. The difference between velocity and churn.

Reading the shape

What a healthy shape looks like

There is no single correct distribution. A prototyping week should not look like a hardening week. But unhealthy shapes are recognizable at a glance: build towering over everything, design barely visible, review missing entirely. That is blind generation, and it is where AI budgets go to die.

A team whose spend shows visible understanding, deliberate design and real verification is using AI the way strong engineers use any tool: thought first, generation second, judgment last. The distribution does not judge people. It starts the right conversation.

Team A

a healthy shape
unddesbldtstrev

Visible thought before generation, real verification after it.

Team B

needs a conversation
unddesbldtstrev

Generation towering over everything, design and review starved.

The way out

What to do about it

1

Measure by phase, not by volume

Attribute spend to the work it did. The unit of insight is the distribution, not the total.

2

Pair spend with outcomes

A distribution next to delivery and rework tells you why a number is high and what to change.

3

Never rank individuals

The moment a measure becomes a leaderboard, it gets gamed and stops measuring. Read teams, coach people.

At scale

And it works at the scale you actually run

This is not a single-developer dashboard problem. The organizations that feel it most run offices on multiple continents, remote engineers, Windows fleets next to MacBooks, and a different mix of coding tools per team.

Phase attribution only becomes a management instrument when all of it lands in one governed picture: every device measuring locally, only metrics traveling, one server your organization owns.

metrics only · never codegoverned readsAmericasoffice · SFOpaymentsplatformwinmacagents: 46Europeoffice · BERmobiledatawinmacagents: 46Asiaoffice · KHIinframlwinmacagents: 46RemoteanywhereBill-y serveryour infrastructure · one orgLeaders' consolephases · spend · forecastsevery device classifies locally · agents dial out only · one listener to secure

This is what Bill-y does

Phase attribution for AI coding spend, across every major coding tool, on Windows and macOS, entirely on your own infrastructure.

Get started

Get Bill-y

Bill-y is available now for teams coding with AI at scale, on Windows and macOS, with every major coding tool. Tell us a little about your team and we'll get you set up.