
Why a newer model isn't always an upgrade
A newer model is not automatically better for your product. When the first reasoning models shipped, they underperformed on some of Vanta's use cases, and the team rolled back to the non-reasoning ones.
Public benchmarks don't settle the call; your own evals do. Vanta scores each model in a large internal table, then rolls promising ones out to a slice of customers and sometimes runs shadow mode before committing.
Why AI adoption is a tax before a multiplier
More tokens don't convert to proportionally more shipping. Like adding headcount, adding AI carries a learning tax first, so output comes back as a fractional multiplier rather than a clean one.
The hump was cleared with enablement, not mandates. An AI Champions program, paired AI wins and AI fails channels, and managers writing a few PRs a month did more than the top-down doc on its own.
What token spend does and doesn't tell you
Spending to look busy rewards the wrong behavior. Some of Vanta's most effective engineers sit nowhere near the top of usage, because they've worked out when one file of context beats twenty tool calls.
Cost conversations are about awareness, not clamping down. Threshold alerts to engineers and their managers open a "what are you doing here?" conversation that sometimes surfaces a great use case and sometimes work that should have stopped days ago.
Where the real leverage showed up
A homegrown tool beat the off-the-shelf ones. An engineer drowning in reviews built his own code review bot, trained on reviews he respected and given a warm personality, and it saw adoption the generic tools never earned.
Internal tooling earns the same rigor as production. Vanta tracks its own AI usage on a Datadog dashboard, down to the thumbs up and thumbs down on review comments.
What good engineers look like now
Vanta hires for three distinct team shapes. Product, platform, and AI teams each call for a different profile, from product-minded generalists to deeply technical architects to zero-to-one experimenters.
Curiosity and growth mindset are the constant across all three. Specific tool fluency shifts too fast to hire on; willingness to learn is what carries across roles.
Iccha Sethi
VP of Engineering
Vanta
Iccha Sethi, Senior Vice President of Engineering at Vanta, on model upgrades, token spend, and how the engineering discipline is changing with AI. The full audio episode is available on Spotify.
timestamps
0:00 Introduction
1:29 Why Vanta built an AI quality framework
5:05 Golden datasets and rolling back reasoning models
8:39 Weekly eval reviews
11:31 The most important metrics to watch
14:14 Getting the org over the AI adoption hump
18:32 Why input multiplier is not output
19:33 Tokenmaxxing and building cost awareness
22:04 The code review bot one engineer built for his team
25:29 Where engineering, product, and design are converging
29:19 What skills matter for engineers now
Transcript
Stephen Poletto 00:00
Good morning, Iccha. How are you?
Iccha Sethi (00:02)
Good morning, Stephen. I'm doing good. Is it super early for you right now?
Stephen Poletto (00:07)
It's relatively early. My brain doesn't totally wake up until the afternoon, so I'm a little groggy, but that's okay. This is the best way to start the day. Just to introduce you to our audience: you lead the engineering team at Vanta, which is a compliance product, and before that you've had a long career working at places like GitHub, and before that at InVision. So you've been leading technical teams for a long time. I've noticed you've been writing a lot recently about AI products and how the team is approaching building AI into what Vanta does, so I'm wondering if we could start there and learn a little bit about some of the functionality you're shipping and how you're thinking about the development process of getting the most out of AI.
Iccha Sethi (00:56)
The reason I've been posting a lot more about how we're building AI products at Vanta, and we actually launched an engineering blog series called the Trustcraft series talking about this too, is there's plenty of content out there on how we're changing engineering with AI, but not enough on how to build AI products well. This is especially important to us as Vanta, because trust and the quality of our products are very important in the compliance and security space, as you can imagine. So we want to be very thoughtful about how we're building AI products. We have an entire AI org dedicated to focusing on this craft.
So a couple of things we're doing that are interesting. One is, if you're really focused on this, you've realized that the cost of maintaining AI features is extremely high. It's because software is never done. New models are always shifting, there's drift in existing models, and your customers are changing too. If you're a growing team, you're growing new types of customers in your customer base, so things are constantly shifting, and software is never done.
We've internally developed this AI eval, or quality, framework. Stage one is: are you actually having traces? Stage two is: do you have golden datasets that you can run evals on? Stage three is: do you actually run evals? Then next is: do you run experiments? And finally: do you take this feedback loop and tie it back into improving your AI product features overall? So we developed this maturity matrix internally, done by one of our AI platform engineering managers, and we see how teams are faring across the board on these things. Maybe the controversial take here is the goal isn't to get to green everywhere, but more to highlight where we want to get to green and where we're okay with the yellow.
So that's one of the things we're doing. There's plenty of other stuff. We've developed internal philosophy and principles on how to build our agents. So a lot of thought is going into how we develop our AI product features.
Stephen Poletto (03:11)
Very cool. I'm reminded of that saying, the map is not the territory. So you might have this framework, you might have evals, they might all get to green, but that doesn't necessarily mean you're done. That's a set of inputs and signals around what can be improved, but it doesn't necessarily mean the AI performance exactly matches what your customers want. There's constant iteration there.
Just in practice, talk me through a little bit what happens when a new model comes out and you're thinking about upgrading your production stack to run it. How does the team think through the implications of that shift, how customers will receive the change, and what the evals are used for in making that decision around the migration? Because I think that's a challenge a lot of folks face as they're now working with this foundation underneath them, and the foundation is shifting, as you said. So it's this: how do we do this change management thoughtfully? Would love to hear a little bit about how you approach, say, a model upgrade as an example.
Iccha Sethi (04:15)
Yeah. So one is, we have offline datasets, which we call our golden datasets, that you can run your evals against. It's like tests, right? We used to call them tests back in normal development. Evals are basically the more non-deterministic version of tests.
So we run offline evals. Anytime a new model comes out, we say: here's what we think, here's the latency, here's the accuracy, the similarity score. We have a bunch of metrics we evaluate against, and we have this huge table internally. We've tested against all the models: Anthropic's models, OpenAI's models, Gemini, so on and so forth, each version. Anytime a new model comes out, we're basically adding a row there and asking: how does it perform across these different use cases?
Something else I'm sure companies are thinking about is that different product features probably need different models. So it's not about using one model across the board. We want to optimize for a set of use cases: this is the specific version and specific model. A newer model is not always better than the older model, which is again counterintuitive to how we've thought about software development, where anytime a new version is out there it's better, and so on. That's not always true for all the models coming out.
So once we've run those benchmark tests and we realize this is better, do we want to move to it or not for these certain use cases? Once that determination is made, then we need to do our online evals: basically roll this out to a percentage of our customers to evaluate how these are performing and gather more signal over there. At times we've also leveraged something called shadow mode, which is you continue to run on the old model in production but also do a comparison against the newer model, and compare which of the two models is performing better in production before making a decision. So these are all different toolkits, and each team gets to pick which of these tools they want to use to make that determination. Then we finally make the switch to the new model if we decide this is the better decision.
A great example is the reasoning models came out, and for certain of our use cases the reasoning models were actually not better. We realized that during our online evals, and then we swapped back to the non-reasoning models.
Stephen Poletto (06:59)
Yeah, interesting. So a new model comes out, you obviously see the public benchmarks, and you effectively have your own internal Vanta benchmark that you've developed with offline datasets. If that looks good, you A/B test to verify in production that it is better, that it represents an improvement, and then you're constantly iterating. So maybe in that example where the reasoning models performed worse in prod, my mental model for the evals is that they're a leading indicator of real-world customer performance, right? You're trying to have that test suite match production workloads as closely as possible. So when they drift and you find these things in prod, how do you think about iterating on the evals and updating them?
Iccha Sethi (07:49)
We have weekly operational reviews on teams. You can call them eval reviews, whatever you want to call them. Back in the non-AI world, you'd do this for on-call handoff, to look at your alerts and so on. It's a similar parallel, except what the team is doing is reviewing your production evals. So what are the traces that are failing our quality thresholds, or the various scores we have out there? They're reviewing them to come up with a hypothesis of why we think these certain use cases aren't performing well in production, and then they take a couple of actions from there.
Update your dataset is one of those. The other action is to develop hypotheses around why this is failing. What should we tweak? Is it in our harness? Is it in our context? Is it something else? And run experiments based on that. Again, there are going to be tons of data points over here, so it's about the most common patterns we see that we want to go respond to. Every week, that's the discipline: it's changing, it's spending time on it.
Something interesting I think Vanta is doing is we're not putting this load only on engineers. Our PMs are very much involved in this too. Our AI team just developed alerts for this, where we're tagging PMs to go look at some of these quality data points, because I think it's very important for PMs to understand how much more the load on engineering is increasing as you invest in more AI product features. The KTLO and the framing around it is changing a lot.
Stephen Poletto (09:39)
Definitely.
Iccha Sethi (09:40)
Yeah.
Stephen Poletto (09:40)
Definitely, yeah. The fact that we have these non-deterministic elements of our production stack now introduces a ton of variance: the harness, the evals, the tests, the response to customer feedback. All these things are in service of constraining that non-determinism to get to the best outcome, but it's a whole bunch of new work, right? And it's cool to hear that you're making it a cross-functional team sport, not just an engineering duty. I'd love to learn a little bit more about these operational reviews, and maybe get a lens into how you think about running teams at scale. You've got a large number of engineers at Vanta, and you want to keep tabs on how the teams are performing. It sounds like these operational reviews are one of the ways you keep close tabs on what's happening. Tell me a little bit about those and other key management interfaces you've designed with the teams to stay hands-on at scale.
Iccha Sethi (10:38)
At the top level, we have our engineering MBR, our monthly business review, where I look at top-level metrics for the engineering org as a whole, and this gets bubbled up all the way to the board. Then there are more detailed metrics I look at internally.
So to talk about the top-level metrics, there are basically three categories. One is how we're doing in terms of delivery, so that includes how many product announcements we had, how many ships, what percentage was delivered on time, those sorts of metrics. The second category is quality, which is how many incidents did we have, did we stay on top of our incident action items, the number of support tickets, and so on. There are metrics that correlate to these in terms of people: how many ramped engineers we have, so we can normalize as we grow as an engineering org. Is our percentage of throughput increasing, and is our quality staying steady state and not dropping? And then finally, what is our lead time to production, so that's around developer experience, kind of proxy metrics around those.
There's a bunch of detailed metrics that I call my input metrics, which lead up to these output metrics and help you diagnose if an output metric isn't performing well. We also cut this across different pillars in my org, so every pillar is going to have a different flavor of these metrics they need to be held accountable for. For example, my platform org is more responsible for lead time, and they may not have as many public product announcements, and that's okay, but they should be having more engineering ships and engineering announcements. Versus a product team, which has a different set of per-ramped-engineer metrics they're held accountable for.
Stephen Poletto (12:30)
Given that you have such a robust metrics framework and operational rigor around these metrics, you seem like the ideal person to talk to about how AI is changing things. So I'm curious what you've seen in those numbers as the team has embraced AI. And maybe, just for a little background context, you could tell us where the team is in its journey with incorporating AI into the development process. Is it well adopted across teams? Is it still pockets of the org? How much is the team bought into AI-centric development? And then I'd love to learn, as you've been reviewing these numbers, what you've seen, what kind of dynamics and second-order effects have been taking place.
Iccha Sethi (13:15)
I feel like pretty much everyone uses AI now in the org. We're well over the hump of "I'm not going to use AI."
Stephen Poletto (13:23)
Was there a hump at earlier points in time?
Iccha Sethi (13:27)
Yeah, there was a hump. Even after the models got good at the end of last year, there was still a perception problem on the models. People were like, "Oh, the models still aren't great." Obviously they're not going to be perfect on a brownfield codebase. There has to be a lot more iteration on a brownfield codebase when it comes to using LLMs, the whole loop to make them better, still having a human in the loop reviewing certain things, especially when you have many patterns in an aged codebase.
Stephen Poletto (14:03)
Totally.
Iccha Sethi (14:04)
So there were points in time where LLMs would lose trust with our engineers, because they'd be like, "Oh my god, it's not using the most perfect pattern, even though I put the instruction in CLAUDE.md." So really, navigating the org to say, "Okay, don't be one-shotting everything. Work through your loops. Can you play around with the instructions? Can you go update the prompts?" And then, going beyond just using it for code, using it for improving our incidents and our quality and other areas too.
So there's a bit of a hump around getting people to just play with the tools and not give up, keep iterating. We launched a bunch of internal programs, like our AI Champions program, where we have reps from each team meet every week, and in cohort they're solving problems related to AI development in our org. So for example, somebody developed a tool to fix flaky tests in our CI pipeline, and we're going to roll this out to everybody. Really pushing for more ground-up sharing of what's working and not working. We had an AI wins channel and an AI fails channel, so we're sharing both ways.
Stephen Poletto (15:12)
Oh, smart. I've heard the AI wins thing a lot, but I haven't heard the AI fails, and that's just as important to learn from.
Iccha Sethi (15:17)
Exactly. And that also helps prevent the perception that we have rose-colored glasses on when it comes to AI, that we're completely up in the clouds, right? It doesn't work for some use cases. I've spoken to Anthropic engineers, and they'll tell you these are the things it's really good for, and here are the things we're working on, but it's not always published publicly.
Stephen Poletto (15:39)
Sure.
Iccha Sethi (15:39)
Right, that's the perception.
Stephen Poletto (15:42)
Left to the enterprises to figure out.
Iccha Sethi (15:44)
Exactly, right. People often talk loudly about the wins, but not enough about the learnings.
So yes, there was that hump. Going back to your original point, we really pushed for people to just play with it, share their learnings, share their fails. I also encouraged all managers to start coding more, at least a couple of PRs every month, so you experience the pain points and the learnings. And I think we're well over the hump. Now it's more about optimizing. I kicked off an AI developer experience team within the org earlier this year too. It's really small, but it's the most seasoned engineers in the org, who've been around the codebase long enough to understand its quirks, to build common tooling within the org. So lots of ground-up enablement and investment helps, versus just top-down "eat it." I still wrote a top-down doc saying this is what we're doing for AI for engineering, but paired with tons of enablement and support and investment is what really helped us get over that.
Stephen Poletto (16:54)
That's very helpful background context, and awesome to hear some of those tactics that worked to get people on board, because I'm sure some folks are still going through that journey themselves. Now that you have pretty universal adoption, I'd love to learn about how it's influenced those operational metrics, and the first-order and second-order effects you're seeing. Is velocity up? Are teams shipping more announcements? Are there any quality implications? Just, how's it going from your perspective, reviewing this data?
Iccha Sethi (17:30)
Our shipping is up, but I wouldn't say it's proportionally up. So for example, X number of more tokens doesn't mean X number of more ships, same as X number of more engineers doesn't mean X number of more ships. It's not multiplicative. That's one thing I think people outside of engineering sometimes have a hard time grokking: input multiplier is not output. There's a tax when you add more people, this coordination tax, this onboarding tax. When you add AI, there's this learning tax. You can't immediately jump to optimizing and have a multiplier. So that's one thing.
Stephen Poletto (18:10)
Yeah.
Iccha Sethi (18:10)
It's up, but it isn't a direct multiplier. It's more like a fractional multiplier. The second thing is, when we look at where our tokens are going, it's very interesting to look at those numbers, because there's the whole token scoreboard hype...
Stephen Poletto (18:28)
Tokenmaxxing.
Iccha Sethi (18:29)
Tokenmaxxing hype phase. I'm glad we didn't buy into that phase, because it would just reward the wrong behaviors, which is "I'm going to spend money and I'm going to look good." We've always been very practical about output being output, right? Ultimately that's what you're held accountable for. If you're using AI for it, great, but don't spend tokens just to show us you're doing the work, because a lot of times, if you look at where a lot of the token spend is going, it's not always where the higher output is happening. People sometimes are figuring out what a good prompt is, or "should I even be leveraging 20 MCP calls, or just uploading a file with some context here?" Some people have really figured out how to do that, and they're not at the top of tokenmaxxing, and that is completely fine. The output and throughput is still there.
On cost, since that comes up a ton right now, the philosophy is more about being aware of your cost. You're going to start getting notifications as you reach different thresholds, and then your managers will get notifications once you reach certain thresholds, so you can have conversations around, "Hey, what are you doing? I'm just curious." Sometimes there are very valid use cases, interesting learnings, and there's a lot of times where it's not worth the ROI. You could have stopped a couple of days ago. You don't need to continue down the line of this investment. So starting to build that awareness is where we're at right now.
Stephen Poletto (20:05)
Great. So just playing back some of what I'm hearing: you index the team on outcome metrics. What are you shipping? What's the quality of what you're shipping? That's what has always mattered.
Iccha Sethi (20:17)
Yes.
Stephen Poletto (20:18)
Token expenditure is a means to an end to get there.
Iccha Sethi (20:22)
Yes.
Stephen Poletto (20:22)
And it sounds like there is some inefficiency in the organization today with how folks are spending, but that's okay from your standpoint, because people are learning. So you're spending that money in the spirit of organizational learning and iteration, but now starting to facilitate conversations around cost consciousness with the team.
Iccha Sethi (20:40)
That is correct, yes. Fable's out, you know. Don't use Fable for everything.
Stephen Poletto (20:47)
I think that's a really good approach for the moment that we're in, so I appreciate you sharing. What is that use case of high token to high value that you alluded to? I'd love to learn.
Iccha Sethi (21:00)
Yeah. So one of our engineers developed a code review slash incident bot for all his teammates. He's actually going to write a blog post on it on our Trustcraft series. What he basically did is he said, "Oh, I'm getting too many code reviews today, and we have this generic code review bot. We've tried Codex and Claude and everything, but that doesn't cut it for me."
So what he ended up doing is he built his own code review bot. He called it Janbot internally, which learns from his past code reviews and the code reviews of engineers he basically looks up to as good code reviewers. Then he put a warm personality to it, to be a more friendly automated code reviewer. He built that bot, and it's seen astronomical adoption, because it turns out giving a personality to a code review bot has made it less mechanical and stoic, and people really are like, "Okay, Janbot, give me a review on my PR." So that was a good spend of his tokens, to go develop this. He's also developed things around alerts, resolving them before people look at them.
Again, there are generic products around this. We use PUP and other tools too. But personalization, this being built around how an engineer thinks within this company or within this org, has seen more success in preventing incidents, or a higher response rate to code review comments versus ignoring them. It's been an interesting use case.
Stephen Poletto (22:53)
Yeah, that's awesome. We were talking at the beginning of this conversation about developing your own benchmarks and evals for prod, and I think a lot of internal AI development platform teams are realizing those same techniques apply for internal development effectiveness.
Iccha Sethi (23:10)
Yes.
Stephen Poletto (23:11)
Because off the shelf will get you only so far. These models keep getting better, and SWE-bench keeps improving, and who knows how far they'll get, but the investment in the harness, the investment in these custom tools, these custom code reviewers, that's where you really get that gain, because it's catered to your company's codebase and your culture. So that's a great example.
Iccha Sethi (23:37)
Yeah, we have a Datadog dashboard, and we track all the same metrics, like you said, for internal usage: what skills are invoked the most, what are the thumbs up and thumbs down on code review comments, evaling for them. So it's very true, internal efforts require the same kind of discipline.
Stephen Poletto (23:53)
Yeah, totally. Especially with the amount of money we're starting to pour into these tokens, it becomes a system you want to optimize and measure, just like your production application. You talked about the team going through the learning curve, and maybe you're spending a little bit inefficiently so people can experiment. Where are you hoping to go if you look three to six months out from now, and what are some of those investments your team is making that you're excited about?
Iccha Sethi (24:20)
I'm going to take a step back and talk about EPD in general. I think more and more mind-meld is happening between EPD, that's engineering, product, and design, in how we approach product development. Our PMs are prototyping within our agent. They're writing prompts, they're figuring out how we want our AI products, or the Vanta agent, to interact with customers. Similarly, I'm thinking about how engineering evolves and comes to the table. We want engineers to be thinking about what the customer needs, how we can get there, what's a long-term investment, a medium-term investment, a short-term investment.
So I'm very much thinking about how team composition needs to evolve over time, so that we're coming to the table as co-business owners or partners. You're always going to need two categories of engineers: those who are very product-minded, and those who are deeply technical. They need to be really partnering together within the EPD triad in bringing this to the table. Also, setting clear expectations that there's always going to be certain long-running efforts, like platform investments and codebase investments, to make it better for both agents and LLM-driven development, as well as for humans. And our designers are trying to make small PRs for code quality and things like that. So I'm thinking about how, as engineering, we facilitate our collaborators and make this more accessible to them, while feeling comfortable ourselves and having the right guardrails in place.
So I don't think specialization is going to go away. Engineering has depth to it. What is a good architecture? What makes certain things good for LLMs? How does it scale with our customers and business needs? That's not going away. There's still going to be that specialization. There's going to be a lot more meeting in the middle. It's already happening today, but continuing to reduce the friction in how that happens. And going into more engineering-specific things, just continuing to invest in our codebase and architecture to really reap the gains of LLMs more. We've been doing that this year. I think there's more to be done over here.
Stephen Poletto (26:45)
Common patterns, common ways things are done, so that the agent repeats those common ways of doing things.
Iccha Sethi (26:51)
Exactly, that's one of them. The second is also our AI DevEx team really investing in our harness, so that we can accelerate all teams versus everyone figuring out the same problem over and over again.
Stephen Poletto (27:08)
And kind of golden paths, golden workflows, common tooling. Yeah, makes sense.
Iccha Sethi (27:13)
And some of this is done, but there's so much more to do as we're learning more.
Stephen Poletto (27:18)
One final question. As you were talking about the evolution of teams and the evolution of roles, a question I'm hearing a lot is, what does this all mean for me as an engineer? What skills should I be honing? What's important now to develop your career? I know you're hiring, you're growing, you're looking for an evolving skill set in how you hire. If maybe a mid-level engineer were to come to you and say, "Hey, what should I be developing in terms of my craft for this moment in time?" how do you think about the foundational skills that are going to matter for the next couple of years?
Iccha Sethi (27:58)
There are three types, or shapes, of teams. There are the platform teams, there are the app or product teams, and there are the AI teams. For my product teams, I'm looking for T-shaped or E-shaped engineers, more so, who are curious about the product and the business and the customers, and still have good engineering practices, bring curiosity to the table, bring a growth mindset to the table. That's one category of engineers.
The second category of engineers is on my platform teams, who are still very much, "I care about the architecture, I care about the technical details," but with the same curiosity and growth mindset. That's going to be common across the board, so that they're thinking ahead about evolving our internals ahead of this massive growth we're continuing to see in what we're building in our product teams.
And my AI teams care about how LLMs work and the details of that. They are very curious, very experimental, very zero-to-one. They're on X all the time, really applying that to what we're building, thinking ahead.
So those are roughly the three categories of teams we're thinking about internally, and hiring different profiles for them. Curiosity and growth mindset are going to be common across all three of them, and then you tend to specialize a little bit more on what gives you joy as an engineer.
Stephen Poletto (29:36)
Great. I love thinking about it in terms of those archetypes, because different product domains and business domains require different skills, different focus. So really helpful. This has been a really fun conversation. Thank you so much for spending time with us.
Iccha Sethi (29:53)
Thanks. Likewise, Stephen. Thanks for all the thoughtful questions. Great way to spend your Monday morning, huh?
Stephen Poletto (29:59)
Definitely. A good kickoff to the week.
Let's talk about where your AI program is.
The ledger for your AI software factory
platform
© 2026 Attuned Inc.
New research: Leading indicators of AI coding agent effectiveness



