Research
We are writing the standard for AI coding spend.
Phase attribution is a new discipline, so someone has to define how it is measured honestly. This is where we propose the metrics, the standards and the guardrails, and where we share what we learn as teams put them to work.
Metrics and standards
What we propose, and hold ourselves to
Each of these is a position we have taken and shipped. They are meant to be cited, argued with, and adopted.
Cost per phase
Attribute every unit of AI spend to a phase of software work: understand, design, build, test, review. A bill stops being one number and becomes a distribution, and a distribution can be read, questioned and acted on. Every standard below builds on this one.
Read more→State the basis, always
A phase share is meaningless without its basis. By tokens shows where the money went; by queries shows what people asked for most. The two disagree in useful ways. We hold that any phase figure must carry its basis as part of the figure.
Read more→Abstains are surfaced, never redistributed
When confidence is low, a turn is labeled an abstain and shown as one, not quietly folded into the nearest phase. A hidden abstain rate is contamination dressed as signal. Measurement integrity means the uncertainty stays visible.
Read more→Fresh share as a waste signal
Split paid input into fresh versus cache. A consistently high fresh share usually means long sessions restarted cold instead of continued: the same context regenerated and paid for again. We propose fresh share as a leading indicator of avoidable spend.
Read more→There is no single correct shape
A prototyping week should not look like a hardening week. We reject target distributions and phase quotas: the moment a share becomes a goal, it stops measuring anything. The value is in reading the shape and starting the right conversation, not hitting a number.
Read more→Measure teams, never rank individuals
A system that can rank individuals will be gamed and feared. We hold that AI-usage analytics must be team-level by construction, with individual data gated, small groups suppressed, and no leaderboard anywhere. Discipline that helps the people doing the work is the only kind that lasts.
Read more→Forecasts must disclose their method
A projected number without its method is a guess in a suit. We hold that every AI-spend forecast should state how it was made, carry a confidence band, and decline to invent a number when history is too thin. Budgets read the band, not a false point.
Read more→In the field
What teams see when they measure
Piloted on their own team, and changed real decisions within the week.
“I run engineering at one company and the business at another, so I put Bill-y in front of my own team. The patterns were the surprise. As the lead engineer, I had the most balanced distribution across understand, design and build, and seeing that laid out changed how I read the whole team. When I asked Bill-y a plain-language question, the answer was grounded in my own numbers, and I made real calls off it that week.”
CTO at one company, CEO at another
Have a metric to propose, or a result to share?
We are building this discipline in the open. If your team has found something worth measuring, we want to hear it.
Get started
Get Bill-y
Bill-y is available now for teams coding with AI at scale, on Windows and macOS, with every major coding tool. Tell us a little about your team and we'll get you set up.