AI Is Writing More Code. Now Engineering Leaders Need to Rethink Everything Around It.

At a recent Enrich event, Stephen Poletto – Enrich Member and Field CTO of Span – shared new research on how engineering teams are adapting to agentic software development.

Drawing on a survey of approximately 100 engineering leaders and an analysis of real-world AI coding activity across more than 100 engineering teams, Stephen explored what's working, where teams are struggling, and which practices are emerging among organizations further along in their AI transformation.

Read on for the biggest takeaways from Stephen’s session, and ​download the full report from Span for a deeper dive into the findings.

1. AI adoption is no longer the goal. Effectiveness is.

AI coding tools have moved well beyond experimentation.

According to Span's Q3 2026 survey:

  • 70% was the median reported share of merged code that was AI-assisted or AI-generated.

  • 38% of respondents reported at least some agent-generated pull requests merging without human review.

  • Nearly 80% reported agents picking up work directly from tools such as tickets or pull request events.

The shift is significant. Engineers aren't simply using AI to autocomplete code. They're increasingly delegating entire tasks to agents, sometimes with minimal intervention.

But many organizations are still measuring success through adoption, usage frequency, and token consumption.

Those metrics tell leaders how much AI is being used, not whether it's delivering value.

The lesson: Move beyond measuring adoption. Start asking what your AI investment actually produces.

Track outcomes such as cost per completed task, engineering time saved, rework, and the quality of delivered software.

2. Code is getting easier to produce. Reviewing it is becoming the bottleneck.

AI agents can generate code faster than ever, but that doesn't mean engineering teams can ship it at the same pace.

In Span's survey, review capacity emerged as the most pressing concern.

Among teams running two or more agents in parallel, nearly 70% cited review volume as their biggest issue.

The problem is straightforward: when agents produce more pull requests, someone still needs to verify that the changes are correct, secure, and ready for production.

And that responsibility often falls on senior engineers.

Stephen described an emerging shift from traditional code review toward verification engineering—designing systems that can demonstrate whether an agent's work meets its intended requirements.

One example comes from Fin, formerly Intercom. According to the case study Stephen shared, approximately 20% of its pull requests are automatically approved without human review.

The company uses risk scoring and automated checks to determine which changes can be safely approved and which require human attention.

The lesson: Don't just optimize how quickly your team generates code. Invest in how effectively it verifies that code.

Start with lower-risk changes where automated testing and approval may be appropriate. Reserve human review for work that genuinely requires deeper judgment.

3. Better instructions can make AI dramatically more efficient.

One of the most actionable findings from Span's research was surprisingly simple:

The clearer the instructions, the less expensive the work.

Span analyzed AI coding sessions across more than 100 engineering teams, examining what contributed to lower costs, greater autonomy, and less rework.

One consistent relationship was between prompt clarity and the cost of producing a pull request.

When engineers clearly defined their task, agents were able to complete work more efficiently.

Stephen recommended approaching agent instructions much like a well-written engineering ticket.

A strong prompt should establish:

  • Context: What problem are we solving, and why?

  • Goal: What outcome should the agent produce?

  • Definition of done: How will we know the work is complete?

  • Constraints: What requirements must be respected?

  • Out of scope: What should the agent avoid changing?

This reduces ambiguity before the agent starts working, rather than requiring repeated clarification and correction.

The lesson: Treat prompt writing as a skill your team can develop.

Create reusable templates, share examples of effective instructions, and refine them based on actual results.

4. Your development environment matters as much as your AI model.

Even a highly capable agent will struggle if it doesn't have access to the tools and information needed to complete its work.

Span's research found that better-prepared development environments were associated with greater agent autonomy.

In practical terms, agents performed more work between human interventions when they could independently access documentation, run tests, build software, and troubleshoot problems.

Stephen offered a useful analogy: think about how you onboard a new engineer.

You want that person to quickly understand the codebase, configure their environment, access the right tools, and run tests without constantly asking for help.

AI agents need many of those same capabilities.

That means investing in reproducible development environments, accessible documentation, testing infrastructure, and tools agents can use independently.

The lesson: Before upgrading to another model, look at the environment your agents are working in.

Identify where they repeatedly get stuck. Give them the tools and context to resolve those issues independently, then measure whether their autonomy improves.

5. Quality needs to be built into the process, not checked at the end.

When AI produces code, it's tempting to treat the generated pull request as finished work and pass it directly to a reviewer.

But Span's research suggests a better approach.

Engineering teams that addressed quality considerations during development experienced fewer downstream review cycles.

That means thinking about questions traditionally raised during code review before the agent completes its task.

For example:

  • What happens if this change fails?

  • How will we monitor its behavior?

  • What logging is required?

  • How can we roll it back safely?

  • What tests would demonstrate that it works?

Engineers can incorporate these requirements into the initial instructions or ask agents to evaluate their own work before opening a pull request.

The lesson: Make verification part of the development process.

An agent shouldn't simply deliver code. It should also help demonstrate that the code works as intended.

6. Your best AI workflows shouldn't stay with your best engineers.

One of the more interesting findings was how unevenly AI productivity gains are distributed.

Nearly 70% of surveyed engineering leaders reported that the productivity gap between their strongest and weakest engineers had widened since adopting AI tools.

This was a self-reported assessment, rather than an independently measured productivity gap, but it points to an important leadership challenge.

Some engineers are developing sophisticated agent workflows and seeing substantial productivity improvements. Others are still using AI primarily for autocomplete or simple questions.

And even when engineers discover highly effective practices, those discoveries often remain isolated.

Only about one-third of surveyed organizations reported maintaining reusable artifacts or investing in centralized sharing of effective workflows.

Stephen noted that some companies are beginning to establish dedicated AI developer experience or AI infrastructure teams to identify, test, and distribute these practices.

The lesson: Turn individual AI expertise into organizational capability.

When an engineer discovers a workflow that consistently improves results, capture it, validate it, and make it available to the broader team.

The opportunity isn't just to make your strongest engineers faster. It's to help everyone benefit from what those engineers have learned.

7. AI spending needs the same discipline as any other engineering investment.

AI tooling costs are rising, but many organizations haven't developed effective ways to manage them.

Span's survey found that typical spending varied considerably, with 17% of respondents reporting costs above $1,000 per developer per month.

Many leaders also reported approving spending increases to keep teams productive, without firm limits or clear measures of return.

Stephen highlighted Uber as an example of a company taking a more systematic approach.

After experiencing significant AI tooling costs, Uber examined what was happening inside agent sessions and introduced several changes:

  • Selecting less expensive models for well-defined tasks.

  • Managing model defaults rather than automatically adopting the newest versions.

  • Improving the tools and context available to agents.

  • Giving engineers greater visibility into their own AI consumption.

According to the case study Stephen shared, these changes helped Uber expand agentic workloads while keeping spending flat.

The lesson: Cost optimization doesn't have to mean restricting experimentation.

Start by understanding where money is going, which workflows produce value, and where better instructions, tooling, or model selection could improve efficiency.

8. The future of engineering is moving toward judgment, not just code production.

The session ended with a thought-provoking question from the group:

If every engineering organization has access to increasingly capable AI coding agents, what will differentiate great engineering teams?

Stephen's perspective was that engineering has never been solely about writing code.

Great engineers identify the right problems, make thoughtful architectural decisions, evaluate tradeoffs, and design solutions customers actually want.

Two engineers using the same AI coding tools can still produce very different outcomes.

The differentiator is their ability to determine what should be built, why it matters, and whether the resulting solution is any good.

Stephen also shared his prediction that engineering organizations will increasingly shift human attention toward specifications, design decisions, and higher-level verification rather than reviewing every line of generated code.

That future isn't guaranteed, but the direction creates an important opportunity for leaders today.

The lesson: As AI takes on more implementation work, invest in the human capabilities that shape the quality of the outcome.

Problem-solving, technical judgment, product thinking, and the ability to define success may become even more valuable.

The through line: Better AI outcomes require better engineering systems.

The biggest takeaway from Stephen's research is that engineering teams won't unlock AI's full potential simply by adopting more powerful models.

The real opportunity lies in improving everything around them.

Clearer instructions help agents work more efficiently. Better development environments allow them to operate more independently. Stronger verification reduces rework. Shared workflows spread productivity gains across the organization.

And better measurement helps leaders understand whether those improvements are actually working.

The next phase of AI transformation isn't just about making agents more capable. It's about building engineering organizations that know how to use those capabilities effectively.

For engineering leaders, that's where the opportunity lies: not simply producing more code, but delivering better software with greater speed, confidence, and efficiency.

—-

Want to explore the ideas shaping technology leadership alongside the people putting them into practice? Join Enrich for interactive discussions with industry leaders and experienced peers—where you can bring your questions and help shape the conversation.

Apply to join Enrich.

Next
Next

From a Spreadsheet to FanDuel: Lessons on Building, Pivoting, and Scaling as a Founder