New research: Leading indicators of AI coding agent effectiveness

New research: Leading indicators of AI coding agent effectiveness

The ledger for your AI software factory

Engineering intelligence for turning effort and tokens into better software outcomes.

Engineering intelligence for turning effort and tokens into better software outcomes.

Don't spin your wheels with AI. Get prompt-to-prod visibility
to raise effectiveness and drive real progress.

Don't spin your wheels with AI. Get prompt-to-prod visibility
to raise effectiveness and drive real progress.

  • PR #4871 · PAYMENTS · MERGED$31.60 · 6 TURNS
  • MOBILE · RELEASE 2.23$84.20 · 12 TURNS
  • PR #5012 · PAYMENTS · COMPLETED$16.40 · 3 TURNS
  • ENG-2170 · CLOSED$52.10 · 9 TURNS
  • PR #5103 · BILLING · IN REVIEW$27.90 · 5 TURNS
  • INC · NONE @7D · DAGSTER$5.17 · 5 TURNS
  • WEB · RELEASE 2.26$12.00 · 17 TURNS
  • PR #4958 · INFRA · REVIEW 1 ROUND$2.39 · 7 TURNS
  • PR #5001 · API-GATEWAY · REVIEW 1 ROUND$8.38 · 10 TURNS

trusted by high growth startups and the fortune 500 alike

trusted by high growth startups and the fortune 500 alike

Search
Organization
Playbooks
Catalog
Src
AI effectiveness
Overview
Beta
AI insights
Traces
Beta
Env readiness
Beta
AI spend
AI Investment
Insights
Allocation
Productivity
Sentiment
Dev FinOps
Jared Erondu
jared@span.app
Traces
Updated 30m ago
Ask Src
Q2 2026
Apr 1 - Jun 30
418 sessions
Apr 1Apr 16May 1May 16May 31Jun 15Jun 30
Trace
Cost
PRs
Duration
Task type
Agent / model
Eval
Started
Update dev config for webhook delivery worker
$5.26
#670
2.1h
+2
GPT-5.6 Sol
1m ago
Optimize notification delivery pipeline
$1.74
#1356
+1
0.6h
Sonnet 5
Analyzing...
2h ago
Review workspace form component changes
$6.96
#1482
+1
3.1h
Code review
GPT-6 Astra
3h ago
Review CSV export changes
$1.82
#1640
+1
1.3h
Code review
GPT-6 Astra
3h ago
Review session timeout middleware update
$1.22
#369
2.7h
GPT-5.6 Sol
5h ago
Draft approach for role-based workspace permissions
$1.11
#170
0.2h
Fable 5.1
1d ago
Tune billing API deployment health checks
$4.04
#716
0.2h
Claude Opus 5
1d ago
Prototype product usage dashboard filters
$2.77
#803
0.7h
Claude Opus 5
1d ago
Add tests for subscription proration edge cases
$3.14
#509
1h
Testing
Sonnet 5
1d ago
Fix local billing API bootstrap script
$2.00
#862
0.2h
GPT-5.6 Terra
1d ago
Document authentication policy exceptions
$2.60
#1926
+1
3.6h
Documentation
Claude Opus 5
1d ago
Document subscription proration edge cases
$3.05
#807
0.4h
Documentation
GPT-5.6 Luna
1d ago
Agent spend analysis
What type of work are we spending the most on?
Traces › Trace sessions
Ask a follow-up question
Auto
0
Search
Organization
Playbooks
Catalog
Src
AI effectiveness
Overview
Beta
AI insights
Traces
Beta
Env readiness
Beta
AI spend
AI Investment
Insights
Allocation
Productivity
Sentiment
Dev FinOps
Jared Erondu
jared@span.app
Traces
Updated 30m ago
Ask Src
Q2 2026
Apr 1 - Jun 30
418 sessions
Apr 1Apr 16May 1May 16May 31Jun 15Jun 30
Trace
Cost
PRs
Duration
Task type
Agent / model
Eval
Started
Update dev config for webhook delivery worker
$5.26
#670
2.1h
+2
GPT-5.6 Sol
1m ago
Optimize notification delivery pipeline
$1.74
#1356
+1
0.6h
Sonnet 5
Analyzing...
2h ago
Review workspace form component changes
$6.96
#1482
+1
3.1h
Code review
GPT-6 Astra
3h ago
Review CSV export changes
$1.82
#1640
+1
1.3h
Code review
GPT-6 Astra
3h ago
Review session timeout middleware update
$1.22
#369
2.7h
GPT-5.6 Sol
5h ago
Draft approach for role-based workspace permissions
$1.11
#170
0.2h
Fable 5.1
1d ago
Tune billing API deployment health checks
$4.04
#716
0.2h
Claude Opus 5
1d ago
Prototype product usage dashboard filters
$2.77
#803
0.7h
Claude Opus 5
1d ago
Add tests for subscription proration edge cases
$3.14
#509
1h
Testing
Sonnet 5
1d ago
Fix local billing API bootstrap script
$2.00
#862
0.2h
GPT-5.6 Terra
1d ago
Document authentication policy exceptions
$2.60
#1926
+1
3.6h
Documentation
Claude Opus 5
1d ago
Document subscription proration edge cases
$3.05
#807
0.4h
Documentation
GPT-5.6 Luna
1d ago
Agent spend analysis
What type of work are we spending the most on?
Traces › Trace sessions
Ask a follow-up question
Auto
0

Track AI code from prompt to prod. 

Connect coding agent traces to tickets, PRs, and production outcomes, showing what the work cost and where to improve.

Traces capture the complete picture

Traces are the new unit of work. Accurately capture prompts, tool calls, model choices, file edits, and token spend across your full agentic stack.

Weave related traces together

Link related sessions that planned, investigated, implemented, tested, reviewed, and debugged the work, including dead ends and efforts that never became a PR.

Layer in engineering context

Build on a mature context graph to connect traces to the metrics, teams, tickets, PRs, projects, and outcomes around it.

platform overview

One ledger maps tokens and effort to outcomes.

Report on ROI and progress with confidence.

Engineering leaders

Compare your AI program against industry benchmarks and show stakeholders how it improves each quarter. Map every number to the evidence that produced it.

"Span gives us a clear view of our engineering time and budget. We now have real data showing how our teams and processes are performing, and we use those insights to ship better products."

Scott Woody

CEO & Founder

metronome

Stop guessing what to improve.

AI platform leaders

Rank improvements to your agent harness and development workflows, then measure which ones increase leverage.

"Span gives us the clarity and alignment we need to operate at clock speed. It’s become a core part of how we run and grow our engineering team."

Waseem Alshikh

Co-Founder & CTO

writer

Point the factory at what matters.

Business alignment

Map wages and token spend to the project and work type they served. Get intelligence so you know what to scale up or down. Set the priorities, then monitor budgets vs. actuals to improve alignment.

"Span is vital to making our teams better, stronger, and faster. And the team’s speed inspires us, too."

Dan Gill

Chief Product Officer

carvana

Claim R&D credits on agent spend.

Finance

Accurately attribute wages and AI spend for R&D capitalization and tax credits. Generate simple, auditable evidence to get more from DevFinOps with less manual work.

“Span gives us a level of visibility we’ve never had before. It shows exactly what our teams are working on and how that effort ladders up to our biggest priorities, so we can focus talent where it matters most.”

Oran O'Dowd

VP of Engineering

fin

Engineering leaders

Report on ROI and progress with confidence.

Compare your AI program against industry benchmarks and show stakeholders how it improves each quarter. Map every number to the sessions that produced it.

"Span gives us a clear view of our engineering time and budget. We now have real data showing how our teams and processes are performing, and we use those insights to ship better products."

Scott Woody

CEO & Founder

metronome

AI platform leaders

Stop guessing what to improve.

Rank improvements to your agent harness and development workflows, then measure which ones increase leverage.

"Span gives us the clarity and alignment we need to operate at clock speed. It’s become a core part of how we run and grow our engineering team."

Waseem Alshikh

Co-Founder & CTO

writer

Business alignment

Point the factory at what matters.

Each session lands on the project and work type it serves. Set the priorities, then see what each initiative consumed in human effort and agent capacity, and what it shipped.

"Span is vital to making our teams better, stronger, and faster. And the team’s speed inspires us, too."

Dan Gill

Chief Product Officer

carvana

Finance

Claim R&D credits on agent spend.

Attribute metered agent spend and engineering hours to the projects they supported, each with sessions and a confidence tier. One record serves ROI, capitalization, and Form 6765 Section G.

“Span gives us a level of visibility we’ve never had before. It shows exactly what our teams are working on and how that effort ladders up to our biggest priorities, so we can focus talent where it matters most.”

Oran O'Dowd

VP of Engineering

fin

More use cases

Not just software. The whole package.

88%
turn yield per point
of harness engineering
Know what good looks like

Benchmarks from 103 teams, so you know where good is and which practices lower cost and rework.

Guidance and FDE

Our engineers build the deliverable with you: the board report, the harness fixes, the capitalization schedule.

Slack
mcp
skills
api
cli
Meets you where you work

Ask Src in Slack, get a Monday brief from Playbooks, pull the numbers into Claude Code through MCP.

security

A secure and private system of record.

Enterprise-grade security

Audited controls across infrastructure, access, and the models Span runs.

SOC 2 Type II

Regular Service Organization Controls audits.

Audit logs

Full audit trails, including every read of trace content.

GDPR compliant

Clear data retention and deletion policies.

Zero-retention LLMs

Models that analyze traces keep nothing.

Identity and access

Least privilege, provisioned from your identity provider.

SSO

Sign in through your identity provider.

Audit logs

You set who can read whose sessions, per team and per repo.

SCIM

Provision and deprovision users automatically.

Admin controls

Restrict system-level access to named administrators.

Trace privacy

Private first. Reports are about repos and teams, never people.

Private by default

An engineer's sessions are theirs. Span reports 

in aggregate at the repo and team level.

Secret and PII redaction

Span strips credentials and personal data before storing a trace.

Not a proxy

The recorder sits on the developer machine, rolled out through MDM. It never touches the request path.

Sensitive-topic screening

Span flags sensitive-topic traces and limits them to approved reviewers.