New research: Leading indicators of AI coding agent effectiveness
New research: Leading indicators of AI coding agent effectiveness
Frameworks
5 Dimensions for Productive Engineering
5 Dimensions for Productive Engineering
Stephen Poletto
•
Engineering leaders owe their teams and their businesses a clear-eyed view of how work flows: Where is friction accumulating? Are new hires ramping successfully? Are we trading too much quality for speed? Is a team's trajectory changing, and why? These are questions that deserve data-supported answers, not because data replaces judgment, but because data prompts the right conversations and grounds them in shared reality rather than anecdote. However, the answers are scattered across the tools developers live in, and in raw form it lacks the depth and context to support decision-making.
This is why Span exists. Span integrates with GitHub/GitLab, Jira/Linear, and the other systems where development work actually happens, then layers analytics and AI on top to turn scattered activity data into insight - in aggregate, by team, or down to the individual.
The purpose is not to score people. It is to demystify the flow of engineering work, identify trends and opportunities for continuous improvement, and surface anomalies early - the same way engineers instrument their production systems to catch regressions before customers feel them.
The Goal is Customer Value
Before going any further, it's worth stating the thing that every healthy metrics program keeps at its center, and a belief we hold strongly at Span: the goal is customer value. Everything else, such as PR counts, cycle time, deployment frequency, AI adoption rates are intermediary measurements of the machinery that produces that value.
This distinction is easy to state and easy to lose. When an industry leader announces that "90% of our PRs are committed with AI," they are describing how well they've adopted a trend, not whether they've delivered anything customers care about. Customers don't care how many tokens you spent, how many PRs you merged, or how sophisticated your tooling is. They care whether their problems got solved and how quickly.
A useful mental model is the distinction between outcomes and activities. If your goal is to run a 5k under a target time, that finish time is the outcome. It’s a lagging indicator you only learn at the end. Running an hour a day, fueling properly, and sleeping well are activities - leading indicators that predict whether you'll hit the goal. You'd be foolish to ignore them, and equally foolish to confuse "I ran an hour today" with "I achieved my goal."

Software works the same way. The outcome is shipping something that delights customers, visible in signals like NPS, adoption, and retention. The activities are producing PRs, reviewing them quickly, and deploying them without incident. These activity metrics are leading indicators of the outcome: they tell you whether the machinery is healthy long before the lagging outcome metrics can. That's precisely what makes them valuable, and precisely why they must never be mistaken for the goal itself. Tracking both - outcomes to steer by, activities to act on - is what produces healthy results.
Productivity is Multi-dimensional: Span’s 5 Dimensions of Productive Engineering
Because no single number can describe a healthy engineering organization, Span organizes measurement around five dimensions of productive engineering.
Dimension | Core question | Key metrics | What they tell you |
|---|---|---|---|
Velocity | How fast are we moving? | PR cycle time (first commit → merge); time to merge (PR open → merge); deployment frequency; PR volume per dev per week | End-to-end delivery speed; review-process efficiency and code-review bottlenecks; CI/CD maturity and ability to deliver continuously; throughput and ability to break work into manageable chunks |
Quality & Stability | Are we building well and staying reliable? | Incident frequency; change failure rate (incidents per deployment); mean time to resolution; defect rate, code rework rate | How well teams maintain service continuity and prevent defects, and how well they respond when production issues occur. The counterweight to velocity — speed that requires rework or causes negative customer issues is not speed at all |
Engagement | Are people happy and supported? | Focus time (heads-down hours per person per week); developer sentiment via benchmarked, research-backed surveys | Whether systems and meeting culture support deep work, and the qualitative side activity data can't capture: does the team feel effective, supported, and set up to succeed? |
Capacity | Do we understand where our effort goes, and can we commit it with confidence? | Inferred investment mix (maintenance vs. platform improvement vs. new feature development); innovation ratio; workstream alignment; planning accuracy (planned work vs. actual delivery) | Whether execution is predictable and the organization genuinely understands its own capacity — what makes roadmaps trustworthy |
AI Effectiveness | Are we leveraging modern tools effectively? | AI code ratio; turns to merge; AI review load; context and prompt effectiveness; spend efficiency | Who is leveraging modern tools, how effective the harness is, and whether adoption is improving delivery or merely generating activity. |
Different Metrics for Different Altitudes
A second principle keeps metrics programs sane: metrics don't work uniformly across different levels of an organization.
Picture a pyramid. At the top, a CXO is accountable for a handful of financial and business outcomes: revenue, retention, cost. These are highly aggregated, lagging, and nearly impossible to attribute to any single cause. At the bottom, teams track hundreds of granular delivery and service metrics (e.g. cycle time per PR, uptime, latency) which are leading, specific, and much easier to attribute. Each layer's metrics are leading indicators for the layer above and lagging indicators relative to the layer below. A CEO gets nothing from staring at token spend, and an individual engineer can't move annual revenue; each level needs metrics broad enough to matter and narrow enough to act on.

Span is built for this reality, serving each altitude differently.
Executives use Span to stay informed through active work summaries, track org-wide trends across the five dimensions, and plan resourcing allocation via AI investment mix.
Managers use Span to identify and understand the barriers affecting individual and team performance — excessive meetings, lengthy PR cycles that produce stale code and context switching, inefficient workflows, technical debt, resource constraints, communication gaps.
Individual engineers use Span to reflect on the quality of their own work, discover opportunities for professional growth, and speed up self-reviews with brag sheets that combat recency bias in performance cycles. Same underlying data, sliced to the level where each person can actually act.
How Should Metrics Be Used?
As proxies, in context, at the right altitude - engineering metrics do real work. Three uses stand out, and Span is designed around all three.
Unblocking teams
The most common and least controversial use: identifying the friction that keeps good engineers from doing good work. Rexview bottlenecks visible in time-to-merge, focus time eroded by meeting load, PR cycles long enough to produce stale code and context switching. Managers use the data to find and remove these impediments; engineers use it to reflect on their own workflow. This is metrics in service of the people doing the work, not the other way around.
Establishing working norms
Metrics give teams a shared, concrete vocabulary for how they want to operate. Span's team targets turn a vague aspiration ("let's review each other's code promptly") into a working agreement — a PR review SLA the team sets for itself and can see itself keeping — while benchmarking contextualizes where a team stands and where focus would pay off. Crucially, these are team norms about process, not individual quotas about output. That distinction is the entire difference between a healthy program and a Goodhart trap.
Prompting the hard conversations
Trends and outliers — a sudden change in someone's activity pattern, a new hire whose time-to-first-PR suggests a slow ramp, a team whose quality signals are drifting — are signals that something deserves attention. The metric is never the conclusion; it's the instigation of inquiry. It's leadership's job to notice early and approach with curiosity: ask what's going on, listen, and don't take the question personally, whichever end of the conversation you're on. Catching a struggling new hire in week four instead of month six is a kindness, not an indictment.
Goodhart's Law and the Gaming Trap
Goodhart's law states “when a measure becomes a target, it ceases to be a good measure”, and it definitely applies to engineering metrics.
Every activity metric can be gamed. PR counts can be inflated by slicing work into meaningless fragments. Cycle time can be juiced by rubber-stamping reviews. Deployment frequency can be pumped by shipping trivial changes. The moment individuals believe a single number determines their standing, they will optimize the number, and the number will stop telling you anything true. This can have negative cultural consequences too: measurement gets experienced as surveillance, data gets weaponized, and a fear-driven dynamic replaces the curiosity that made the data useful in the first place.
This is why Span takes an explicit position on what the platform is not for. We do not believe any single metric can properly capture the impact an engineer delivers, and we advise against any kind of leaderboard dynamic or individual goal-setting on a single metric.
Span provides data-driven context, one signal of many that organizations use to achieve their people, process, and technology outcomes. Context matters enormously with nuanced metrics like PR volume or cycle time: a low number might reflect a hard problem, a mentoring-heavy week, or a genuine issue, and only conversation can distinguish them. Reading into a single metric without shared context isn't just unhelpful; it's counterproductive.
We also recommend information transparency between managers and ICs. Everyone is on the same page about what is measured, why, and how it will be used. Transparency is what turns "surveillance" into "self-improvement."
(One exception worth noting: cycle time is unusually resistant to harmful gaming. The main way to "game" it is to slice work into smaller batches, which genuinely improves flow. A metric whose exploit is a best practice is a good metric.)
Conclusion
The difference between a metrics program that succeeds comes down to a few commitments.
Measure activities, but aim at outcomes — activity metrics are leading indicators of customer value, never substitutes for it.
Watch all five dimensions, because productivity is a system, not a speedometer.
Match metrics to their altitude, and don't ask a number to do a job it can't do at that level.
Never rank individuals or set goals on a single metric.
Be transparent with managers and ICs alike about what is measured, why, and how it will be used.
Treat every surprising number as the beginning of a conversation, not the end of one.
Everything you need to unlock engineering excellence
Everything you need to unlock engineering excellence