A

adoption

A framework from span research

Five dimensions for measuring how efficiently and effectively your team leverages agentic development tools

Beyond the Model: What Distinguishes Effective Agentic Development

Beyond the Model: What Distinguishes

 Effective Agentic Development

A shared, observable way to see how well your organization converts AI usage into shipped, working code, grounded in what we can read directly from agent traces today.

WHY NOW

WHY NOW

DORA measures your delivery pipeline.
It was never built to measure your toolchain.

The measurement gap

For a decade, DORA gave leaders a research-backed way to assess their delivery system, and it remains the right framework for exactly that.

But DORA was built for a human-driven era. When a growing share of merged code is written by agents, the unit of performance is no longer just the team and its pipeline. It also includes the agentic toolchain: the model, the harness around it, and the codebase it operates in. Questions appear that DORA was never designed to answer:

  • How much of what ships is AI-produced?

  • How much human steering did it require?

  • Does the quality hold up after merge?

And a newer pressure: cost

Token spend is rising fast. For many organizations, AI usage is becoming a meaningful line item with almost no visibility behind it.

Total spend on its own tells you almost nothing. It can't distinguish a team getting more done from a team burning tokens on sessions that go nowhere.

Understanding how effectively your organization converts AI usage into shipped, working code is becoming as important as understanding cloud spend was a decade ago.

THE FRAMEWORK

THE FRAMEWORK

The five dimensions, and what each one measures.

A — Adoption

How much of what we ship does AI write?

Licenses bought and tools installed measure intent. This measures what actually reached merge. A wide spread across teams or repos usually points at something structural.

AI code rate

↑ higher is better

↑ higher is better

C — Capacity

Where is our AI capacity being invested?

The other dimensions tell you whether AI is producing code efficiently. This one tells you what that newly cheap capacity is being spent on: building new things, and the things the business said matter most.

Turnover rate

↑ higher is better

↑ higher is better

AI-weighted project rate

↑ higher is better

↑ higher is better

C — Cost efficiency

What does each unit of shipped code cost?

Not what you spent, but what you spent per unit of code that actually shipped. Only merged work reaches the denominator, so it separates a team getting more done from one burning tokens on sessions that go nowhere.

AI merge cost

↓ lower is better

↓ lower is better

E — Excellence

Does AI-assisted code hold up after merge?

The guardrail. Every other dimension can be improved by pushing more code through faster; this is the one that tells you whether that code held up, or whether the gains are being borrowed against future rework. Read it as a rate, not a count.

AI defect rate

↓ lower is better

↓ lower is better

L — Leverage

How much human attention does shipped code consume?

Where most of the human time in AI-assisted work actually goes: the author steering the agent, and the reviewers downstream. It's invisible in PR-level metrics, and effort that disappears from one side often reappears on the other, so both are measured.

Turns to merge

↓ lower is better

↓ lower is better

AI review load

↓ lower is better

↓ lower is better

About the Framework

About the Framework

Adoption, capacity, cost, excellence, leverage: the questions every engineering leader will need answered.

The example metrics are our current best instruments, grounded in what we can observe directly from agent traces today. We expect them to get sharper, and we'll iterate in the open as the industry's understanding matures.