Skip to content
By team · Engineering

Measure AI in engineering

Attribute coding-tool use to reviewed changes, then assess delivery, correction and cost in comparable work.

GitHub CopilotCursorCursorClaude CodeClaude CodeWindsurfWindsurfGeminiGeminiGitHubGitHubJiraJiraLinearLinear
From activity to production

The change is the thread through the work.

Start with a defined work category and the records available for its review.

AI coding tools
GitHub CopilotSeat and activity
CursorCursorAgent usage
Claude CodeClaude CodeSession and cost
Oximy
Engineering systems
GitHubGitHubPull request and review
JiraJiraWork items
ChecksFailed checks and reruns
Completed workMerged change
Beyond one workflow

Answer the engineering questions behind an AI rollout.

Extend the work that earns it, improve the process that needs it.

Repeat workflow coverage
WorkflowAI useMaintenance fixesRepeat activityRegression testsConsistent activityDependency upgradesOccasional activityFeature changesSparse activity

Where is use repeatable?

Maintenance fixes and regression tests carry AI week after week; feature changes barely do. Separate durable maintenance, testing, upgrade and feature workflows before extending the rollout.

Review burden, like-for-like completed changes
Comparable workAI-assisted
Review comments58%72%
Requested updates48%52%
Failed checks43%64%
Reverts34%36%

Where does review work rise?

Comments and requested updates move a little on like-for-like changes; failed checks move a lot. Keep comments, correction, checks, reverts and incidents visible instead of folding them into one quality score.

Cost per merged change, comparable work
CursorCursor
Claude CodeClaude Code
GitHub Copilot
OpenAI CodexOpenAI Codex

Which tool lowers unit cost?

Each tool's charge lands on the changes it actually helped merge, so four coding tools rank against comparable work. Compare tool cost against similar merged work, not against seats or tokens.

Decision docket, current review
WorkflowWhat the record showsNextMaintenance fixesRepeat use, review steadyScaleDependency upgradesFailed checks +21 pointsImproveTool cohortsFour tools, same workConsolidateFeature deliverySparse use so farKeep measuring

What should change next?

Every workflow leaves the review with one action and the record it rests on. Scale, improve, consolidate or open another evidence window: here, improve the CI handoff before dependency upgrades expand.

Next step