
The harness is the new craft
When agents write the code, the engineering job moves up a layer. The work becomes building and maintaining the harness, the deterministic hooks, schemas, and judges that keep a probabilistic model on the rails and escalate when it drifts, because a model convinced of something will otherwise just keep running with it.
Old-school fundamentals came back, not away. Upfront specs, test-driven design, and human-in-the-loop checkpoints turned out to matter more once an AI was doing the typing, not less.
Code review is the bottleneck now
Laccetti's team fights it with more agents, not fewer. Reviews run in a fresh session, never the agent that wrote the code, with separate security, adversarial, and database-change passes, since different models catch different things, though the approach is verbose, token-heavy, and admittedly not the final answer.
Unit tests stopped proving the thing works. Real bugs only surfaced when they spun up the service and exercised it end to end, so the move is to write the Playwright test that actually clicks the button.
Cost is about to get real
Usage-based pricing is coming, and no single budget fits. A well-funded startup can absorb the overrun and a public company can't, so the plan is an LLM gateway that sends cheap work to cheap models and saves the expensive one for problems that need deep thinking.
ROI is the question, not raw spend. The team runs hot by design, ingesting hundreds of millions of events through ML pipelines, so the point was never whether it's expensive, it's whether it pays off.
The role, and the people in it, are being redrawn
Directing code replaced writing it, and not everyone signed up for that. "I didn't come here to write documentation" is a real reaction Laccetti hears, and his answer is ordinary management: match people to the work that still plays to their strengths.
Curiosity beats tool fluency. Specific tools shift too fast to hire on, so what carries is a team that wants to experiment, iterate, and say plainly what isn't working.
timestamps
0:00 Intro
4:32 What still holds, and what's new, in engineering
7:33 What Gorgias is building
9:58 Unit tests and why "skills" aren't enough
12:14 Building the harness: hooks, judges, and keeping AI on the rails
15:54 Meta-engineering and shared learning
18:45 The token cost question
21:18 Going hands-on: demos built on real data
24:05 The costs of agentic development
30:57 Rethinking code review
35:15 What's next: gateways, validation, and memory
Transcript
Stephen Poletto (0:00)
Hey, Michael, nice to see you. How are you doing today?
Michael Laccetti (0:05)
Good to see you, Stephen. It's been a hot minute since we had dinner here in Toronto. Was that five years ago now?
Stephen Poletto (0:11)
That's what it feels like in the AI era. Well, you're leading engineering teams over at Gorgias, and you've been in engineering leadership for quite the career. You've worked at a mix of more established companies and a lot of startups along the way. How did you end up at Gorgias? What brought you into the role you're working on now?
Michael Laccetti (0:34)
That's an interesting one, and very much an outcome of the times. The engineering career progression for the first 20, 25 years of my career was pretty linear. You're an IC, you get very good at your craft, and at some point you try to become that T-shaped engineer, where you're broad and deep and you know stuff. At some point, probably in 2014, 2015, I realized the problems I was trying to solve were not just technical problems, they were also people problems. There's a system that everything operates within. That was the transition to more of a leadership role. I went from principal engineer to director, to VP, and ran an entire engineering org. One realization I had, maybe five years ago, was that I kept inheriting other people's engineering orgs and trying to change their cultures, and changing a culture is kind of like turning a tanker: it takes you years. I wanted to do this from the ground up. So I joined a much smaller startup, went through that journey, and found that the sweet spot for me is an engineering team of 100 to 150 people. That's where my skills really landed and helped scale things up.
Then I had the opportunity to join Gorgias, which was the startup within the scale-up. There's an established organization here. We'd been building a help desk, which became more of a conversational commerce platform, and there was an opportunity to build a marketing platform, which was something I'd always wanted to do. I was given the opportunity to focus on building a product from zero, building the team, building the culture, but doing it with the support of an organization instead of frantically trying to find PMF before the money runs out. And doing all that while we're also in the midst of this transition to AI, not just using AI but also building products that are AI. So it all came together quite nicely.
Stephen Poletto (2:59)
Great, I love that. Having the resources of a broader parent company but the autonomy to execute quickly on a new product sounds like the best of both worlds. Given that you've seen this longitudinal study of engineering leadership, the people problems, the organizational problems, the technical problems, right now it feels like a really dynamic time to be leading a team and leading software engineers, because there are so many questions about the shape of the role and how it's evolving, and what it means to be an engineer in 2026 and 2027, and what kind of skills matter. As you've taken on this role and guided the team through those questions, what are some of the things that feel tried and true, good old-fashioned engineering leadership from the past decades, and what feels new and novel that you're figuring out with your team?
Michael Laccetti (3:51)
Every day is a new day is definitely the mental model here. As we've gone through the vibe coding extravaganza, where it was just like, let the LLMs do stuff for you, trust it, it turns out prompts matter. Garbage in, garbage out. And then harness engineering, all that sort of stuff. As we go through that evolution, we realize that a lot of the software engineering fundamentals still apply. If we go back a long time ago, there was the waterfall process, where you tried to define upfront what a thing was supposed to do, and you had to be very prescriptive. You'd have these 250-page docs with very granular detail. We sort of ended up going back to that, in essence, because for an AI to do things to the level of quality we want, we're doing that. But at the same time, there's a completely new area of experimentation and exploration, where as engineers we also need to be very close to our customers. Gorgias has always embodied that. One of our things is that every employee is also a CSM. You have to manage a relationship with a merchant and stay close to them.
Stephen Poletto (5:13)
Now your engineers, too, are kind of CSMs. That's cool.
Michael Laccetti (5:17)
But now we also get to be a PM and a CSM, where we can have an idea, mock something up, and bring it straight to a customer and say, hey, we're thinking about this, does this resonate? What do you think? So there's still a lot of deep architecture that's critically important to get the right quality out and to have something scale. You still have to think deeply about that. You worry less about the code itself, and I'm sure there will be challenges in the coming months to years, as some of those quality decisions we've pretended won't be there pop up and we have to deal with them. But on the experimentation side, it's so much fun.
Stephen Poletto (6:09)
You said before we hopped on this call that you're using a lot of AI to build an AI product. Before we get into how you're using AI, maybe tell us a little about what your team is building. I'm sure that's moving quickly too, with the model capabilities evolving and the market around AI-driven customer conversation moving quickly. I'm curious about some of the innovation you're doing for your customers.
Michael Laccetti (6:40)
Sure, I appreciate the chance to plug Gorgias a little here. As I mentioned, we started as that help desk, where when a support agent has to connect to a human, we built that platform over time. It evolved into the little chat bubble that pops up on a website, where you're like, I don't want to email somebody, I want support now. And then, in the last three years, we asked, why wait for a human? That doesn't scale linearly. You have to find the people, they have to be online. AI can solve a solid 80 to 90% of this, and obviously we want to get that number higher. So we started doing that and driving deeply, and we've iterated on our own AI agentic architecture to version three now. That's where Gorgias as a whole has been focused. My team is more focused on the customer engagement side of things.
If you look at a lot of the email service providers or marketing tools in the world, while they're starting to integrate AI, it's more in the vein of, we'll generate the template and the content for you, but it's the same content everybody gets. From our perspective, we know conversational commerce has much better ROI, but we also know that hyper-personalization at the right time drives better ROI. So, given all the context we have about the shopper, the actions, what they're doing, what they like, how do we weave all of that in and create that hyper-personalized messaging? And then, if they have questions, we can keep the conversation going and really drive that engagement forward. Being able to unlock a whole new market has been pretty awesome.
Stephen Poletto (8:31)
Very cool. Let's get into how your team's using AI to build. At the dinner we were both at, you told me about some of the cool things you're doing to reimagine the whole process, going from spec into working software. How did you get there, and what's the current status? What's working, what's not working, as you've pushed the team to reimagine the way of working and bringing software to life?
Michael Laccetti (9:00)
I think we're on the same journey most companies are on, though we might be a little further ahead in some regards. Our team is also a little faster and easier, because we don't have a lot of legacy to bring along. One of the challenges everybody in the industry is acknowledging is that a brownfield project trying to have AI work on it comes with a lot more overhead, because there's context, there's the system architecture, there are all the interactions. We didn't have that, which made it easier for us to get up and going. I think we started the same place everybody does, trying to map out the system. What are the things we do? Where do we spend a lot of time? What feels like not a good use of time? If I could automate something, what would it be? And I think we did what everybody did, which was: I hate writing unit tests, I don't want to deal with that, I want to write the cool stuff, you write the not-cool stuff. And then slowly learning that, wait, if it doesn't understand the code, it can't write the right tests. If it doesn't understand the business problem we're trying to solve, is it writing tests that are valuable at all?
So, coming back to some of those tried-and-true things, we started with test-driven design: build the test, then build the software, use the spec up front to define it. We basically started building the skills that would take an idea and turn it into a product design document, turn that into a Linear project, and start breaking it down, very human-in-the-loop, though. We had to drive all of that at every step. Part of it was that the models weren't quite there yet. Part of it was that trust wasn't there yet. And part of it was that we didn't really know what the right guardrails were to make that work correctly.
Stephen Poletto (10:59)
Makes sense. How has that trust grown over time? Are you finding opportunities to use the latest and greatest models for more of this?
Michael Laccetti (11:10)
Part of it has been around the harness itself. Instead of just building discrete skills, which are just words, and as we all know words can be ignored or misunderstood or misinterpreted, we started building actual plugins that used deterministic hooks, forcing certain schemas, certain approaches, so that when you triggered a specific workflow it was basically kept on the rails, with very specific boundaries of what we were trying to do. It's a very granular thing. A skill does a specific unit of work and moves along. We have judges in there all the way across, and it's made it easier for us to check in periodically to validate the outcome and say, yes, that matches the understanding, and to have it escalate when something goes wrong.
One of the biggest challenges with LLMs is that when they're convinced of something, they'll just keep running with it, and putting those deterministic checks in place keeps it from spinning out of control. So the harness has been a big unlock. The other one, and this was something one of the team members here, Jan, started, is a little agent called Maya. Maya is the thing you use to take a Linear ticket and turn it into a PR, but it intentionally starts on small units of work. As an engineer, I can say, we need to move this five pixels over, or all these little things that are just nits, small cosmetic things. Or a test is flaking. This happens a lot, and nobody wants to go investigate it. By having these small, granular things, we can build trust and validate that it works as expected, that the outcome is there. So for us it's about being engineering-minded: have an idea, have a hypothesis, implement it, validate it quickly. Does it work? Yes. And that's how we've continued building on it.
Stephen Poletto (13:37)
I heard a good definition of harness recently: a set of deterministic software that wraps a probabilistic model. I really like it, because these models, like you said, can start making stuff up or go down a bad path, and if you rerun them with the same input conditions, sometimes you'll get different outcomes. So the harness, the plugins, the tests, the scaffold, constrains it and gives you confidence that where it lands is going to be pretty good. And if it's not good, then hopefully you find a way to improve your harness. I'm curious, because it sounds like you're investing a lot in your harness, what that practice looks like. I've heard people start to call it meta engineering, where you're engineering the factory that then produces the work, building the system that builds the system. It sounds like you're developing shared skills and plugins. What does that collaboration look like in practice with your team, for people to improve the harness and make things better?
Michael Laccetti (14:41)
One of the early learnings we had was that just making a central repository for skills doesn't enable communal learning. And one of the challenges is that it's still very engineering-centric. My partner works in legal tech, and they're deep into this as well. She works on the R&D ops side, and she said, we have Claude, and we have five PMs who've all created variations of a PM skill. How do we collaborate on that, and how do we figure out what to cherry-pick from each and smash together into our canonical PM skill? And it's like, well, have you heard of GitHub? And instantly it's, I can't ask PMs to open PRs. You can ask Claude to do it, but it becomes such a pain. So, baking some of that functionality directly into the tool itself, so people can be abstracted away from the technical stuff, is one huge part of it.
The other, for engineers, because we do tend to live in GitHub, and that's fine, is that by making it a plugin repo we also build skills to reinforce the learnings. As you execute the plugin, things will go wrong. That's fine. How do we extract that and say, what was the root cause of this? Now go update the plugin, update your context and memory, so we don't have this happen again. And now it's distributed and shared, and everybody keeps layering on top. It doesn't just solve your problem, it solves it for everybody.
Stephen Poletto (16:29)
Makes sense. I've sometimes heard people call these golden paths, or canonical paths. You allow some localized differences in workflow and setup because people want autonomy in how they work, but then where are those centralized investments, where some learning or reusable path gets shared, so we all get increased leverage out of the system? So that's cool to hear that folks are actively collaborating and doing that. The other part of the conversation we've been hearing a lot about is cost. Folks are starting to use AI everywhere, and it's definitely making us faster, but token bills are starting to add up. Is that a concern for you? How are you thinking about cost management and efficiency? Is it an active priority, or something on the back burner for now?
Michael Laccetti (17:25)
It is very interesting. As an industry we all know this, and I think we're all on different journeys, with different time horizons. A lot of folks on GitHub got the early surprise about token costs. I just got the email from Atlassian that their AI stuff is moving to usage-based in a week or two. Our Anthropic contract renews, I think, in late July, and moves from our subscription to usage-based. We know it's coming, but we don't know what it looks like yet, so it's hard to start making some of those judgments. I've heard a lot of people say, we budgeted a thousand per person, or this per person. Then you have companies like Ramp saying, why would we budget, it unlocks everybody. So I don't think there's a singular answer. It's contextual to the business and how they're doing. If you're on a huge growth curve, like any startup, you could paper over some of that with fresh funding. If you're a more mature organization, a public company, it gets a little harder.
So it's situational. It is top of mind, and our team is a very strong user of AI, because we're very experimental. We're trying to figure out what resonates with merchants, how we get things in front of them. We have to ingest hundreds of millions of events, feed them through ML pipelines, all that. We're building a lot, and for us it's been a multiplier to get things there faster. So, yes, I know we're expensive. I can fire up some tools to estimate my usage, and I see numbers that aren't cheap, and then I argue for the ROI, which is the biggest question we're trying to answer.
Stephen Poletto (19:36)
Right. If I remember correctly, you're using AI yourself a ton. What are you doing with it, leading the organization?
Michael Laccetti (19:47)
It's interesting, and this goes back to why I joined Gorgias. A lot of it was to be in a position where I could experiment more, be a bit more hands-on, less focused on people and systems and more on the product and the engineering. So I am very hands-on. I'm building the harness with the team specifically so I can keep using it and keep things going. One interesting challenge I've had to address is that a lot of my team is globally distributed, and I don't want to wait for tomorrow. I don't want to wait for reviews. So Claude can reach into the code base, help me understand it, share context, all that, so we can keep the builds going, do a handover, and it just goes on. But the better part for me is when I talk to merchants and I can come to them with a demo. I was talking to one of our merchants today about the autonomous customer engagement platform we're envisioning, and with Claude, being able to reach in through our AI agent running on top of all our data to pull out some context, I could say, this is how we're thinking about your business, these are the numbers we see, and this is how we want to build the product, using real data. That took like zero effort. I didn't have to go bug data engineers and data scientists just to get that together and in front of me. And they said, wow, we've never thought of it that way, but that's where we want to go. Being able to do that so quickly has been a game changer.
Stephen Poletto (21:31)
I see this too with the blending of roles between engineering, product, and design. These tools help everybody feel a little more empowered to get to 80%, maybe not to 100%, because there's a real craft to each of these functions. Building scalable architecture on the engineering side is very different from a prototype. Building a high-fidelity design that considers accessibility and internationalization is a very different skill from slapping a prototype together that you can get feedback on. But it does empower people not to wait for coordination, to take that next step, to build higher conviction, and then put it into something you want to productionize at a high quality bar. So it's cool that it's doing that for you personally, that you're able to generate quick prototypes to get feedback and have conversations about the direction you want to head.
Michael Laccetti (22:28)
It does. The flip side is that it also highlights the systemic issues that predated AI. These are engineering and design challenges that have always been there, where you have fixed capacity and have to decide where to invest the time. We thought code was the most expensive thing. Now AI writes our code, and it's like, okay, but now we spend a lot of time writing specs, and some engineers say, I didn't come here to write documentation. The craft was solving the problem, and it feels like I'm losing the fun part. Others say, I just love solving problems and seeing things land, so whatever, this is awesome. But the bottleneck now moves to code review. And yes, we're automating that too, and for small, low-impact things we'll let it slide, but what's the judge of what's a small, low-impact thing?
I'll give you an example. We intentionally separate our database changes from our code changes, so we can run the migration and then run the code on top of it. To AI, it's a one-file change, it looks small and self-contained, so it's just thumbs up, ship it. But for the rest of us, it's, whoa, wait, I learned that this could lock a table, and if we're ingesting 500 million rows and I just locked a table, somebody's going to be very unhappy. So tuning what's actually small versus what's not is one of the ongoing things. And then I feel pretty bad for PMs, because you used to be able to think deeply about a product, go deep, and spend a lot of time figuring out the surface area and the problem space, and now they just have to be spitting out so many ideas that I don't know if they're going to go as deeply as they used to. It's hard to tell if that truly impacts the roadmap in a positive way, the top line, all of that. So we're still trying to tune the right way to use some of this new stuff.
Stephen Poletto (24:41)
I don't know if you follow the Linear CEO, Karri. He's been talking a lot about how just because you can build it doesn't mean you should, and how craft, taste, judgment, and product cohesion matter more than ever, because you can ship a lot of stuff, but does all of it add up into the whole being greater than the sum of the parts? I think that relates to what you're talking about with the craft of PM, where it's, how do you curate and steward all the different product components into one cohesive vision? It feels like it matters more than ever, but it's also harder than ever, because everything is moving so fast. I'm curious, on the people front, you talked about how some engineers didn't think this was the job they were signing up for. I know some engineers personally who love getting into flow state and crafting code, and that was the joy of the role. It was fun to write code, and now we're saying, well, the code can be written by agents, your job is to design the system, work on the harness, write the specifications. How are you navigating some of those hard people changes, because it's a big revamp of the role definition for a lot of folks?
Michael Laccetti (25:58)
I don't think there's a singular answer to this, and different things work for different people and teams. Part of it is just leadership lessons. People are people. Everybody has a different area of focus. Traditionally, there was some person who, if there was an on-call, would know exactly where to go in the system to fix it. You just knew they could tune things. There are other people who are great at shipping that very first product and iterating rapidly, but they're not good at the middle bit of making it scale. I don't think that's really changed. So tailoring the experience and finding the thing that works for them is still key. The one thing that has changed is that we can't not use AI. It feels like a tool, so why would you say no? Much like we used to have IDEs, and at some point we didn't, and when we started using IDEs we were like, oh my god, what a game changer, I hit play and it just runs. So it's a tool, similar to, I have a hammer, so everything's a nail. Is it always the right tool? We're still figuring that out. But in the general sense of people, it's about trying to understand what brings them joy and tuning things so they have the space to learn and grow.
I feel very lucky that on the team I work with, everybody is always excited, not just to try, but to experiment, to learn, to iterate, and to provide feedback on what doesn't work, so we don't have to keep stumbling blindly through the maze, face first. So I haven't run into a situation where somebody's so reticent they just utterly refuse, because I don't think that would end particularly well.
Stephen Poletto (28:02)
Certainly the cultures that have collective learning and a growth mindset seem to be navigating this a lot better, because every day, like you said, it's a new day, there are new things to learn, we adapt and evolve, which is fun. I wanted to ask you about code review. It sounds like you're doing a lot to streamline, trying to assess the riskiness of a change and fast-path things that are lower risk. I imagine code review is a big bottleneck now, and a lot of your engineers are spending a lot more time reviewing work instead of creating it. Have you found anything that's worked well? What are the experiments you're running that might be interesting to share with our audience?
Michael Laccetti (28:48)
We're deeply invested in this, because it's one of the biggest bottlenecks, and I think it applies to every organization. One thing that's worked well for us right now is using multiple models, because they tend to pick up on different things. We make sure the code review runs in a completely fresh session, ideally not the agent that wrote the code, very similar to how it used to be. We run a different set of reviewers with different perspectives. There's the security reviewer, an adversarial one, one that looks for database changes, all that, and we've tried to codify what each reviewer is looking for. It works very well. It can be very verbose. It's also very token-heavy, because it's many agents. So right now it's okay, but the question is, when we start having to pay, what happens then? So I don't think that's going to be the final answer.
Stephen Poletto (29:58)
Does all that happen before it goes to a human? The reviewers provide the feedback, the authorship agent integrates it, and it's not until that process completes that a human gets their eyes on it. So you're kind of shielding that human attention.
Michael Laccetti (30:17)
I don't know if shielding is the right word, but yes, that's the flow, where the human is the final check at the end of the loop. I'll be honest, because we're on a more experimental side, quality is not my primary focus, so I'm more forgiving. If it breaks for an alpha product, that's fine, we can fix it pretty quickly. So my bar is calibrated a little differently from other parts of the organization. For me I'm okay with that. If I were at a different org or in a different role, would it be the same? Probably not. If I look at the part of our system that ingests all the webhooks from Shopify, that's an enormous event stream. Would I YOLO things to prod with no true human review? Probably not. And the other part that still consistently surprises me, in a wonderful way, is when somebody reviews after the code reviewers, and they ask that one thing: why would you do this? Don't you remember that the system five levels up behaves this way, and this contraindicates everything? And you're like, I didn't have that context, Claude didn't have that context, and trying to map that out is almost impossible. So there's still incredible value in it.
Stephen Poletto (31:50)
And what I hear in your answer is a risk-adjusted profile for what the consequence of defects escaping in a given system is, and how quickly you can remediate those issues, which defines the level of heaviness you need around the human review process, and what you can streamline and what you can't. So, adjusting for risk, which I really like. Michael, this has been a really fun conversation. If you look ahead at the next couple of months, I feel like the world will be totally different by the time we talk next. What are some of the primary challenges on the horizon that you're pushing your team to keep investing in? Developing the harness, automating and streamlining code review. Where do you feel you need to invest to stay ahead of that next chapter?
Michael Laccetti (32:47)
That is the most incredibly insightful and challenging question I've been asked in a long time. Part of it is that we know something is going to happen with token prices, and we know our approach will change. We just don't know how or to what. So it's coaching people to be ready and okay with the fact that whatever we're doing today might not be what we're doing tomorrow, and that's fine, we'll figure it out. We'll put an LLM gateway in place to start routing cheap things to the cheap model and the things that truly need deep thinking to the expensive model. We're going to do what everybody else is doing there, because we have to. The other two areas that are still incredibly important to invest in are platform engineering, the things that make the developer experience better and also ensure the AI can validate things correctly. The biggest lesson over the last year is that unit tests don't truly validate functionality. Making sure you can say, write the Playwright test to actually go to that page and click the button, has uncovered more bugs than anything else. Or, spin up the microservice, hit the endpoint, pass in green data and red data, and make sure it succeeds or fails correctly. That's critically important for true validation success. And the other side is tied to my partner's perspective, that when we think about meta engineering, it's, how do you use AI to improve how you use AI? As models improve and as you learn, how do you encode that learning so it just sticks, so you don't have to remember it, because we drift just as much as AI does, all while trying to manage token costs, because that's going to be the ongoing challenge.
Stephen Poletto (35:01)
It's like a whole new game on the field. How do you distill learning effectively into a memory layer that sits in your harness and governs the way the agent produces code, but do it in a way where you're not bloating context windows, you're managing the memory state effectively, and you've got the right deterministic tests around things, so you can be confident in reducing your human attention on code review? A lot of it is just good old-fashioned engineering. If you were operating an engineering team with thousands of engineers, you'd want a lot of guardrails and safety nets in the process. But it's almost like every company, even if you weren't at scale, now needs to invest in this platform and foundation from day one. So it's a very interesting time.
Michael, this has been really fun. Thank you so much for joining us today, and good luck with all those challenges you're working on.
Michael Laccetti (36:00)
Thank you so much, Stephen. And the next time you're in town, I'd love to catch up and hear what's new.
Let's talk about where your AI program is.
The ledger for your AI software factory
platform
© 2026 Attuned Inc.
New research: Leading indicators of AI coding agent effectiveness



