
What is the real bottleneck?
Writing code is the smallest part of the cycle. Clay's sharpest challenges are verification of changes and review bottlenecks, not generation, so the company stood up a dedicated team to tackle them.
Observability for agents is the piece most companies skip. Clay treats it as the foundational layer: visibility into how teams use agents and whether the output holds up, which Mark notes other companies aren't prioritizing the same way.
How Clay approaches internal AI adoption
Adoption beats efficiency, for now. Any tool an engineer wants to try is supported (most land on Claude Code and Cursor), with no leaderboards or usage targets, because the goal right now is building intuition rather than optimizing spend.
Closing the loop is already running in production, carefully. A Devin-based "fuzzer" builds a workflow against an MCP server, runs tests, finds issues, fixes them, and opens a PR; it runs once a day and turns up something every day.
Senior engineers' roles are changing
Engineering becomes metaprogramming. You stop building the system directly and start building the system that builds it, setting the guardrails and the container the agent works inside.
Specification is the new core skill. The leverage moves to being very precise about exactly how you want something built, then letting the agent figure out how to get there.
What Mark is watching
The junior-to-senior pipeline may be breaking. The on-the-job pattern recognition seniors lean on when directing agents is exactly what juniors no longer get to build, which raises the question of where capable engineers come from next.
Build-vs-buy shifted, but domain depth is still a moat. Prototypes are cheap now, yet knowing what to actually build (scalability, observability, the non-obvious edge cases) still protects specialized products, at least until agents close that gap.
timestamps
0:00 Intro
0:39 Mark's role and what Clay does
1:44 How "go-to-market engineering" got its name
3:46 The trash-can detection agent
5:18 Clay's hands-off approach to AI tooling
6:47 The real bottleneck: verification and review
7:29 Standing up a team for "agentic DevEx"
9:36 Standardize, or let a thousand flowers bloom?
11:44 Why observability has to come first
13:22 Closing the loop and the "software factory"
14:52 Clay's closed-loop agent: a Devin fuzzer
16:23 The future of the senior IC engineer
17:14 Metaprogramming for software engineering
20:46 Adoption over efficiency, and the human side
22:48 When roles blur
25:34 The junior-to-senior pipeline (and apprenticeship)
27:53 Build vs buy and the "SaaSpocalypse"
31:13 A tip: don't take "I can't" from an agent
32:52 Wrap-up
Transcript
Stephen Poletto (0:00)
Mark, how are you doing on this rainy New York afternoon?
Mark Hahnenberg (0:04)
Great, doing great.
Stephen Poletto (0:06)
Thanks for hanging out with us. Just so our audience gets a sense of who you are and what Clay is all about. Tell us about the role you're playing and what Clay does.
Mark Hahnenberg (0:17)
I've been at Clay for about four years now, from back when we were in a little apartment in Williamsburg. I've watched how things developed internally on our engineering team, and I'm currently the Head of Technology at Clay. That means I'm a high-level IC. Sometimes I'll step into a management role to drive a project forward. Basically, I look for the highest-impact thing facing Clay the company or Clay the product, target it, and make sure we drive some kind of solution. Right now, one of the biggest things everybody's thinking about is AI software development, agentic software development lifecycle stuff. So that's what I'm looking into these days.
Stephen Poletto (1:10)
We definitely want to dig into that, because it's very relevant to our audience. Clay has pioneered this whole new job title, this whole new field of go-to-market engineering. You've been with the company since the early days, since the apartment in Williamsburg. Tell us about what go-to-market engineering is and how you pioneered it. It's been incredible to watch the growth you've had over the last couple of years.
Mark Hahnenberg (1:45)
It's an evolution of a number of patterns that had been happening in sales orgs throughout the tech industry and wider. It's maybe an evolution of the RevOps role, if people are familiar with that, you have people behind the scenes doing a lot of orchestration of the data that powers their sales teams. Clay has become an integral tool in connecting a CRM to various data providers and piping data so it's in the right place, clean, and fresh. If you squint, that starts to look like an engineering problem. Maybe you're not writing code per se, but you're designing what looks like a data pipeline, storing things in various places, routing data, stitching a lot of different tools together. That's essentially engineering. So that's where the term came from: you're doing engineering for go-to-market, you're a go-to-market engineer.
Stephen Poletto (2:47)
Incredible. And the leverage of LLMs has been a big part of Clay's growth too, right? You're able to do a lot of automation and cool AI wizardry to help streamline that job function. Any cool examples recently, things you've shipped where LLMs are having a big impact on the product?
Mark Hahnenberg (3:13)
We have a web research agent that can take screenshots of web pages. For one of our customers, they wanted to see, for local businesses, what waste management vendor they were using, and also whether their trash cans were overflowing.
Stephen Poletto (3:39)
Send boots on the ground, send people in, the old way.
Mark Hahnenberg (3:43)
Right. You'd have to drive around with a list of businesses and go check out their trash cans. Or maybe you wouldn't do it at all, because it's just not scalable to action on. But we designed a workflow in collaboration with them: take Google Street View photos, have the web research agent navigate to the business's address, take a screenshot, and figure out the color of the trash cans. You can guess the vendor based on the color. You can also ask the agent whether the trash can looks like it's overflowing: is there trash next to it, is the lid open? That was a complete new unlock for them, a workflow they could add to their go-to-market strategy. It's a particularly fun one. And it comes down to the creativity of the go-to-market engineers actually building these workflows to come up with that idea in the first place.
Stephen Poletto (4:45)
Very cool. It sounds like right now you're working on the internal enablement of how your engineering team uses AI to ship products faster and better. Let's talk about some of the challenges you're facing. What are some of the cool problems you're working on?
Mark Hahnenberg (5:02)
Everybody's still trying to figure this out. Some companies have been very vocal about their successes and the things they've tried. We've taken a very laissez-faire approach to AI adoption internally for engineering. Basically any tool anybody has wanted to try, we've supported. We're trying to see patterns. Most people have landed on a couple of tools (primarily Claude Code and Cursor seem to be what our engineering team gravitates toward), but people are open to experimenting with whatever they want. Some challenges we're running into are around verification of changes and review bottlenecks. Writing code isn't the only part of the software development lifecycle. In fact, it's probably not even the majority of it. Those are the things we've been more challenged with, so we've created a new dedicated team to take on some of these challenges and get things working really well, so people can just keep shipping.
Stephen Poletto (6:14)
Very cool, and something we're seeing across the industry. Code has sped up, but the specification of what to build, the process of aligning on direction with EPD, and then the verification that what you've built is good (high quality, passes all the functional tests, satisfies the requirements). That's where the new bottlenecks are emerging. Folks are starting to think about how they AI-fy those parts of the process too, to gain efficiency and speed up the overall cycle. This new team you're investing in. Tell me about it. What's the hypothesis behind building a team to focus on this?
Mark Hahnenberg (6:56)
There's traditional DevEx, where you're trying to speed up humans, and then there's what we're calling agentic DevEx (a provisional term for the team). It's really about building the foundations on which your entire agentic engineering lifecycle can happen. Agents are a little different from humans when they build software. They don't always tell you what they struggled with or what could be better. Humans are very vocal: "oh man, this thing is slowing me down." They do a bit of self-evaluation. Not to say agents can't; you could ask them to. So we want to build the system that makes sure we're using them effectively, that our engineers are effective in their usage, and that lets us see what's working and what's not. A big part of that is an observability layer, so we can see how much teams are using various agents, how successful they are at shipping code that doesn't need to be rewritten immediately, and where they're struggling with parts of the codebase.
Our codebase exists from a pre-agent era (at least non-trivial chunks of it do), and agents can struggle with those legacy parts. We want to see where they're struggling and identify those parts for refactoring or additional context, via an AGENTS.md or something like that. There's also an evangelization component: we want knowledge sharing between the humans about how best to do this. So we're establishing an AI guild internally, with each team having a member in the guild, so we can disseminate best practices.
Stephen Poletto (9:03)
Very cool. I'm curious, and you might not know the answer yet, as you establish a central team to think about agent experience and optimize those flows (checking when they call the same tool multiple times with erroneous arguments, documentation, code review feedback that could be incorporated into docs, linters, tests), there are all these tools you might invest in to improve the agent experience. Do you think you'll need to do more standardization (these are the golden paths, the golden roads), or do you think you'll be able to continue that philosophy of open experimentation, bring your own tool, let a thousand flowers bloom?
Mark Hahnenberg (9:50)
I think the experimentation will probably continue for now. What we're waiting for is natural fault lines (people gravitating toward the best thing) versus standardizing too early and then people fighting the tooling, really wishing they could use this other thing. For example, with cloud development environments, we want to enable engineers to use them. Right now we're primarily confined to our local laptops for development. We even had a guy get a second laptop so he could do more agentic stuff, just to unblock him. We see cloud development environments as the future, but we're not sure which one is best right now. They're all kind of similar, but there are trade-offs, and there's no clear winner yet, so we're letting people continue to experiment there, just like they would with their models, until there's more clarity.
Stephen Poletto (10:52)
Makes sense. Given that the scope and mission of the team is fairly broad, and it's in this whole new arena of agent experience. How are you thinking about where to focus first? Is there an acute need the team is tackling first and foremost?
Mark Hahnenberg (11:11)
We have some ideas around what's actually taking people a long time in their day-to-day. Things like having to babysit PRs to make sure they address all the bot feedback, the linters, the type checkers. You don't want a human sitting there checking whether the lint passed. That's a pretty robotic thing, perfect for an agent to fix. Those are some of the things we're targeting. We're also trying to close the loop, a term I think OpenAI published in a blog post, where an agent can build, test, post a PR, and maybe even ship and deploy automatically, especially for things deemed not super risky. That's the full-blown version, plus the augment-the-human version. We're targeting both. But fundamentally, the observability layer is the foundation. We've talked to a few other companies, and I feel like they're not focusing on that as much, which is interesting to me. I do think you need to establish visibility to see whether what you're doing is effective. So that's where we're targeting that foundational layer, in addition to all the AGENTS.md context-hierarchy stuff. We're doing all the things, I guess.
Stephen Poletto (12:48)
Totally. If you think about traditional developer platforms and developer experience, you'd survey and ask the engineers, or they'd speak up and tell you what's broken. It's really smart to apply that same philosophy to your agent workflows. Assess where in the traces they're getting stuck, where things can be optimized. So I'm totally on board with you there. You've probably seen some of the blog posts: Stripe talked about Minions, Ramp has their background agent system, WorkOS just put one out a couple of weeks back. It does seem like that's the direction a lot of companies are investing in: cloud environments and the ability to offload or delegate wholesale tasks, maybe even eventually wholesale features, to these agents so they can close the loop: build, self-verify, open a PR, maybe have additional agents auto-reviewing those PRs, and then the human is really just doing high-level user acceptance testing.
This idea of a software factory, agents doing all these different parts of the process. How do you see that? Where are you on that journey? Do you have any background agents you're developing? You mentioned cloud environments as the future. How's the team thinking about that?
Mark Hahnenberg (14:19)
We're experimenting with a few different things. For example, the project I'm currently spinning off of is a new kind of workflow builder feature inside Clay, and it has an MCP server in front of it, so you can actually build workflows using Claude Code. That was a big unlock for us. We set up a Devin workflow that's almost like a fuzzer: it builds a random workflow using the MCP server we built, then runs tests with that same MCP server. Everything runs locally in the cloud, and it can look at the logs, see what happened, identify things that might need fixes, make the fix, test whether it worked, and post a PR as a result. That's our first foray into a full closed-loop development cycle. We want to expand it to other teams and other parts of the product. That's basically what we're going to start working on next. We only run it once a day right now, but I feel like we could turn up the knob and see how much it can find. It's finding things every day. It's new software, so it's good we're ironing out the bugs.
Stephen Poletto (15:50)
Very cool, so it's almost like an automated QA agent that goes and tests, finds errors, and then tries to fix them on its own. As you look forward to a world with more background agents, more autonomous systems, closing the loop. How are you thinking about your role as a senior IC engineer? You've been building software for a long time; before Clay you were working on critical infrastructure systems at Meta and Apple. What do you think senior IC engineering looks like? I know you don't have a crystal ball, but what does it look like six months, a year from now?
Mark Hahnenberg (16:40)
My background is in programming languages. That's what I studied, type systems and so on, in college. At Apple I was working on their JavaScript VM inside WebKit. So my brain naturally tends toward thinking about things from those perspectives. To me, it feels like we're entering a higher level of abstraction, where you're still operating within the bounds of the system but with less detail about the low-level things going on. And there's another version of abstraction here, which is metaprogramming (Lisp macros, C++ template metaprogramming, a Ruby DSL), where you can treat the components of the system you're running as first-class objects and directly manipulate them. I think what engineering ends up looking like in the future is metaprogramming for software engineering.
You're not building the system directly anymore; you're building the system that builds the system. It's a meta version of software engineering. You're building the guardrails, building the container in which the agent can build the thing you wanted to build, and you're offering guidance. It figures out how to do it to your specification. So it's about getting really good at being very specific about how exactly you want something built. I think that's what it's going to look like in the future.
Stephen Poletto (18:25)
There's an analogy I heard maybe six months ago that stuck with me, and I think it's analogous to what you're saying. If you look at the automobile manufacturing industry and the assembly lines, there historically were a lot of manual-labor tasks in the assembly process. Then robotics and automation came into the plants, and now many of those tasks are done by robotics and machines. But it's not like the factory floor is empty. You have people tuning the robots and doing quality assurance over the work being done. The jobs have changed dramatically; it's this meta overseeing and tuning of the overall process. I see that as people talk about harness engineering and AI-native developer platforms, thinking about how you scaffold your CI/CD, your verification processes, your tests, how you specify and how you validate. Those become the new art forms of bringing software to life. How do you feel about all this? Clay is a well-respected, innovative engineering culture, you've got a brand people love. How are folks internally navigating this sea change at the more human level? How is the team stomaching it? What's going well, and what are some of the challenges as people think about these new job requirements and shifts to their roles?
Mark Hahnenberg (20:13)
In general, people seem pretty excited. We've had a lot of organic adoption, so people aren't kicking in their heels. Because we haven't had anything more structured or formal, they've been able to explore. It's important: we're really optimizing for adoption right now, more than efficiency.
Stephen Poletto (20:41)
Your token-max era.
Mark Hahnenberg (20:42)
Kind of, a little bit. And we don't gamify it. Finance would have words for me if I said...
Stephen Poletto (20:51)
You don't have a leaderboard incentivizing it, but you're also not constraining and optimizing.
Mark Hahnenberg (20:57)
Exactly. We're in this period where you need to build intuition about how to use the agents effectively, so we just want people to get in there and start doing it. People are generally responding pretty well. I heard an analogy the other day. It's kind of like Civilization. If you've played, you always want to do one more turn. It's like that working with an agent: "oh, we could actually do a little bit more here," and you always want to pull another turn. People are getting their dopamine hit out of it, so they haven't lost that aspect of creating new stuff. On the GTM side of the company, people are very jazzed, not about the engineering team, but about their own ability to build software for themselves. Somebody demoed something they'd built the other week, a Chrome extension for something they wanted to do, and I was like, this is pretty impressive, and they're not an engineer, not on the engineering team. So people across Clay are excited about the possibilities.
Stephen Poletto (22:15)
The democratization of bringing an idea to life is pretty cool. Are you seeing that? There's this image I saw in a blog post a couple of weeks ago, the Spider-Man emoji of PMs and designers and engineers all pointing at each other, because they can all step into each other's roles: engineers doing crude designs, PMs building prototypes as the new version of their spec. That's awesome because it allows more free-form expression. What are you seeing with the muddying of roles, PMs shipping code, or using prototypes as spec? How is that evolving?
Mark Hahnenberg (22:56)
It hasn't gotten super muddy, but there's definitely a lot more overlap now. For example, on the project I'm working on right now: when I first started, I built out a prototype by myself. I could do all the design and everything myself, didn't need to involve anyone else, which was great for moving quickly. Over time, as it formed into a cohesive concept, we started involving more people. Now I have a designer on the team, and to demonstrate some of her design ideas, she just uses Claude Code to build a front-end version of what she wants it to look like, much higher fidelity than what we can generate in Figma. It's interactive too, which is great. You can see how it's supposed to look without her getting into the weeds of actually shipping the code. She's not necessarily shipping code, but she's building stuff to get her ideas across, which is pretty cool. The PMs generally aren't shipping any code, at least at Clay as far as I'm aware. I haven't seen a lot of demo-building from them, but I may not have enough insight there.
Stephen Poletto (24:13)
It's very cool to think about. A picture's worth a thousand words. If we can reduce the amount of PRD writing and static design mocking, and increase free-form expression of ideas that more fluidly capture what you're trying to express, then use that as the input into a requirements process that turns into a spec, that an agent can break down into a tech design and then into tickets, and now we've got our software factory flowing. It's a pretty fun world to be able to express ideas more cleanly. We've talked about a lot of the positives. Any worries on the horizon? Any anxieties about all this AI in the SDLC?
Mark Hahnenberg (25:01)
In the back of my mind there's this voice. It took me a good while on the job, even after I graduated from college, to learn a lot of things and build my pattern recognition for how to build good systems. I can use that very effectively with an agent as my copilot. But I wonder: where is that training happening now? This is the junior-to-senior pipeline. Have we cut that significantly, where junior engineers aren't actually learning the things they need to learn to be effective with agents? Maybe agents get good enough to cover that gap. But that's something I worry about a little: where are the engineers going to come from who know how to effectively use agents?
Stephen Poletto (26:03)
I wonder if we'll see more of an apprenticeship mindset in our industry. If you look at how plumbers and electricians work, there's an apprentice, a junior working with a master plumber, a senior who's doing a job but teaching them, delegating little bits of it, pair programming in some sense. That's a very different model from how we've thought about growing juniors for the last decade, where there would be very modular, discrete, well-specified tasks they could execute autonomously on, then get review and feedback and learn the scope. But now it's almost like the agents are doing that (taking the well-defined tasks), and we're growing the scope of what the agents can do. So I think it's ripe for experimentation, for cultures to try new styles. Maybe we'll see some resurgence of pair programming and apprenticeship style. But that's a whole different model that companies need to be willing to invest in, because unlike the well-defined tasks, which are immediately accretive, the apprenticeship model isn't immediately accretive, it's more of a long game. So, zooming out, any other stuff about AI?
Mark Hahnenberg (27:20)
It looks pretty gloomy for general-purpose SaaS businesses, I'd say. The build option has gotten a lot cheaper, a lot easier, a lot more accessible, so that alters people's choice in the build-versus-buy question. I don't think it's all gloom, though. Especially from our perspective, there are a lot of things you don't know to do when you're building a prototype if you're not immersed in that problem space, including things around scalability, observability, and so on. Maybe that's a temporary moat, where agents can eventually figure out how to generate all that for you, but I don't think we're quite there yet.
Stephen Poletto (28:19)
I always think about the mind-share tax. When you build something yourself, even if you can prototype quickly, and you can prototype and demo stuff really fast right now...
Mark Hahnenberg (28:31)
Right, yep.
Stephen Poletto (28:32)
To actually have it satisfy all the edge cases and the nuanced, detailed requirements you didn't think of in the first go, it's always-on. A little bit of your mind share is always going toward that thing, and mind share is often the limiting factor in companies. The more you can direct it toward solving problems for your customers, the better off you are. But the build-buy equation is definitely changing. We see it in the SaaS public markets. The SaaSpocalypse, it's a real thing.
Mark Hahnenberg (29:04)
To your point, there's a kind of fractal amount of detail, or care, that you have to put in, so you naturally slow down as the thing becomes more real. The prototype is very quick, you can fire that off. But if you're truly vibe coding, you tend to accrete what's essentially agentic tech debt, and eventually you get to a point where the agent can't figure out how to fix something, and you don't know anything about how you built the thing. Agents are getting better, so maybe that point gets high enough that you could build a pretty scalable product before you hit it. But I'm not as doom-and-gloomy as some people are. More broadly, things tend to go in hype cycles, and I feel like we're deep in the middle of one. Maybe the hype is justified, we're definitely investing a lot in this right now. In general, I'm fairly optimistic. People will become more efficient, and they'll build more and cooler stuff. That's great.
Stephen Poletto (30:17)
That's a very positive note to end on. This has been a really fun conversation. One closing question: any cool techniques, hacks, or workflow setups you've gotten a lot of leverage out of in your own usage of AI: a tip or trick for one of our listeners?
Mark Hahnenberg (30:39)
Good question. There was a funny thing the other day. With this Devin setup we're using to close the loop, you can chat with the agent after it tries to do whatever task you've asked. I asked it to build a workflow and test it locally, and it ran into a bunch of problems. I Socratic-method-ed it to fix them: I didn't actually change anything about the configuration of the environment, I just asked it questions about what it needed. It suddenly started realizing it actually could do the thing: it could do a workaround for something, or it did have the key it thought it didn't have. I literally changed nothing about the configuration, and it worked. So sometimes agents just take a little convincing that they do know how to do something. They're a little eager to give up on tasks. So maybe that's the tip: don't assume that what the agent is telling you it can't do, it actually can't do. It might actually be able to.
Stephen Poletto (31:48)
That's wild. Sometimes there are these moments in human-agent interactions where it really feels like you're talking to a human, because that's a very human thing too, right? If you have a junior engineer and they get stuck, you Socratic-method them through it and they get through it. It's kind of wild that's the experience you can have with one of these.
Mark Hahnenberg (32:10)
And then we just save the playbook and say, all right, in the future just do this. Learn from your past mistakes.
Stephen Poletto (32:19)
Very cool. Well, Mark, thanks so much for spending time with us and sharing some of the cool projects you're working on. Really fun conversation.
Mark Hahnenberg (32:27)
Thank you so much.
Let's talk about where your AI program is.
The ledger for your AI software factory
platform
© 2026 Attuned Inc.
New research: Leading indicators of AI coding agent effectiveness



