Skip to content
By use case · Workflow impact measurement

Measure what AI changed in the work

Compare similar work with and without AI, keeping the category, time window, quality measures, and cost definitions consistent.

The records behind the decision

See what moved, and what is still open.

A useful comparison follows the work beyond the AI-generated draft.

What changed in maintenance fixes, and what has not closed?

412 changes, one category
OximyOximy compares the same kind of finished work with AI and without it, in one window.
Lead time to mergeWithout AI, 198 changes0 hWith AI, 214 changes0 h9 h sooner
Review rounds per changeWithout AI, 198 changes0.0With AI, 214 changes0.00.3 fewer
Change failure rateWithout AI, 198 changes0.0%With AI, 214 changes0.0%0.5 points lower
Cost per merged changeWithout AI, 198 changes$0With AI, 214 changes$0$8 lower
Finding

Maintenance fixes merge 9 hours sooner at $8 less per change, and the change failure rate did not rise.

Proposed actionExpand with evidence

The comparison supports taking maintenance fixes to the next eligible team with the same definitions attached. The ninety-day window on customer bug reports is still open, so the customer result stays under measurement.

Owner
VP Engineering
Next step
Name the next eligible workflow

Attribution supports an inspectable comparison; it does not by itself prove that AI caused the change.

How to read the evidence

Keep the rules beside the result.

Define the comparison before interpreting the difference.

One maintenance fixCompletion record
Draft
Review
Merged
Follow-up
The event the business already acceptsMerged change

What counts as done?

Use the event the business already accepts as completion. A draft, a suggestion or an agent run is an input to the workflow unless it is itself the agreed finished unit.

Comparison rules412 changes
Work categoryMaintenance fixes onlyHeld
Time windowCompletion and follow-upHeld
QualityRollbacks and reopensHeld
CostShared chargesRecorded apart

Is the comparison like for like?

Keep category and scope consistent: maintenance changes are not compared with feature work because both reached the same repository. Complexity, open windows and shared costs are recorded beside the result rather than folded into it.

What else movedSame window
AI useStaffingComplexity
The same 412 changesOne comparisonAttribution, not causal proof

Can we say AI caused it?

Only when the method supports it. A before-and-after comparison can also reflect changes in staffing, complexity or process, so those explanations stay visible beside the finding.

Next step