
What tokenmaxxing costs you
Running the best model on everything quietly bleeds money. Assembled did exactly this in production until they looked up and realized they were losing real money on a token-by-token basis, and it took four to five months to retrain the team to even ask whether a task needs a frontier model.
Rewarding token spend produces slop, not progress. John's view is that optimizing for tokens burned incentivizes "slop cannons," high volume at low quality, when a long tail of work like summarization runs fine on a far cheaper model.
How non-engineers ship production code
Roles plus a review gate do the heavy lifting. A "builder" (as opposed to "engineer") has to run /review on a coding agent before they can even open a PR, and the system can auto-check whether a change is reasonably scoped or touches a heavily used endpoint, using Assembled's own stats on how often each endpoint is hit.
The real bottleneck is running the code, not writing it. The biggest activation-energy difference is still getting someone to actually run their change and confirm it works, which is why the internal platform (just a login, with Claude Code and Codex in sandboxes) exists to collapse that friction.
Where the internal platform creates leverage
Team-visible automations beat lonely personal agents. A scheduled job in someone's own agent container is invisible to everyone else, so Assembled built a shared loop that mines past sessions, proposes AGENTS.md updates, and opens a PR to the file's owner.
Less tooling with more context works better. They pulled their MCP after it blew up the agent's context and replaced it with a single CLI of scoped sub-commands, built on Go, Next.js, gVisor, and Docker by one or two engineers as a side project they may open-source.
What 'good' looks like now
The quality bar is "would this have shipped before 2025?" Assembled now runs evals on its coding agents (a commit SHA, a prompt, and LLM-judge notes on what a good implementation looks like), wired into CI so that changing an AGENTS.md file triggers them.
Hiring shifts toward product judgment. A new open-ended "build us something" exercise runs before the technical screen to test how candidates reason about quality and trade-offs, and Friday demos reward showing an eval's before-and-after impact over raw activity.

John Wang
Co-founder and CTO

Assembled
John Wang is the Co-Founder and CTO of Assembled, a Series B company ($70M raised) building AI agents for customer support and the orchestration layer that powers them. His work focuses on the intersection of AI and human workflows, determining when an agent acts, when to escalate, and how to execute seamless handoffs.
timestamps
0:00 What is Assembled?
2:07 How customer support changed from pre-LLM to the AI era
6:46 Bringing AI into the engineering build process
9:45 Letting non-engineers ship code: builder vs engineer roles
12:54 Why they pulled out the MCP and gave agents a CLI instead
13:51 The stack behind their internal agent platform
15:54 Team-wide automations and a self-improving AGENTS.md loop
18:10 Why "token maxing is stupid"
22:47 Picking the right model for the job, and the eval trap
24:01 Slop cannons: speed without a quality bar
25:23 Activity vs progress, and hiring for product judgment
28:08 Running evals on your coding agents
30:05 What's next: delightful agent experience
Transcript
Stephen Poletto (0:00)
Good afternoon, John. How are you doing today?
John Wang (0:02)
Good, good. How are you?
Stephen Poletto (0:04)
You survived the rainstorm in New York this weekend?
John Wang (0:07)
Yeah, in SF I don't have to worry about it, but I'm going to be in New York next week.
Stephen Poletto (0:09)
Okay, right on, right on. We'd love to talk a little bit about what you're up to at Assembled. You're one of the co-founders, the co-founder and CTO at Assembled. You worked at Stripe, where I think you met your co-founders and originated the idea. What's Assembled all about?
John Wang (0:32)
So Assembled, we build tools to make customer support teams more effective. It started, as you mentioned, at Stripe, where one of the last things my co-founders worked on was an internal tool at Stripe, a customer support ML system that tried to categorize the incoming issues and give you results on how you should answer them. It actually improved customer satisfaction by 40% and saved a ton of money. The reason we'd built it was because it was so hard to find good customer support tools, so we decided to go do it outside of Stripe. From there it's been building out a set of systems, a set of tools that make your team as effective as possible, given what level of cost you're actually looking for. There's a whole lot that goes into running a good support team, and we help you solve a large chunk of those.
Stephen Poletto (1:44)
Very cool. And you all started the company before Gen AI commercialization, right? Pre-ChatGPT. So I imagine your world has changed dramatically in the past couple of years as LLMs have become part of the customer support scene. How do you view that inflection moment, pre-LLM to post-LLM, and how the company and the world of support is changing post Gen AI?
John Wang (2:16)
It's an exciting time for us. When we started, one of the reasons we started the company is we were talking to a very high-up product manager at Stripe about this support tooling system that had been built at Stripe, and they were like, why are you working on this? I've got this Bitcoin project I want you guys to work on, come do this thing. That's when we decided, oh my god, people don't care about this, there's opportunity here. It's very different now, because now people care about it. The chairman of OpenAI has a startup in the space. Everyone and their mom has a startup in the space, which is exciting for us, because it means there's a lot of innovation happening. It's also meant we've changed a little in the way we think about things. We've always built what's called a workforce management system, which is how do you figure out how many people to hire, where to put them, and how to staff them in terms of lunches and breaks and starting times to match your demand.
That can be very complex when you have 10,000 support agents. It looks a little different in the AI era, because you might have fewer people, and you might care more about how to send volume to AI. So what we've done is, A, we've built out AI agents ourselves, but B, we're adopting how do you build a team that sends the right questions and the right issues to AI, and the right issues to a human. You want to provide a very different level of support for a million-dollar-per-year user versus a free subscription user. And when you exit out of an agent during an escalation, it should escalate to the right person, and you want to make sure that when you escalate the high-value users, you always have that person immediately available, so that experience is really, really smooth. It's almost a joke at this point: if you go to United, actually United can sometimes have good support at this point, but if you go to an airline, if you're 1K...
Stephen Poletto (4:36)
If you spend enough, if you're 1K, you get the right people, you can get very good customer support.
John Wang (4:42)
You don't want to be waiting on hold for 45 minutes. If you go through an AI and the AI can't solve your question, and then you also have to wait 45 minutes, you're going to be really angry. So for us, it's how do we improve that experience? A lot of it comes down to how you build good systems on top of AI that use agents to help you build the right rules, to help you understand how your teams are performing, how to make changes intraday. A big thing we see is that in the old era you'd have a person looking at graphs going, oh my god, there are a lot of tickets coming in right now, I've got to move people from this side of the business to that side, or put up a bunch of overtime, or do all this stuff to move things around, or make my AI agent a little harder to escalate to a human. Now you can do a lot of that in a much more automated way, which is very exciting, because it gives you that next layer of control on how you think about the whole customer support operation.
Stephen Poletto (5:56)
Very cool. So you're doing a mix of observability of the system as well as optimization, helping define the rules, the routing, when to involve a human, when to use agents, and giving analytics and insight into that whole experience from beginning to end.
John Wang (6:16)
Exactly, exactly.
Stephen Poletto (6:18)
Very cool. There are a lot of parallels between what you're doing in customer support and what's happening in engineering, as we're increasingly trying to figure out which tasks can be delegated to agents. We're even seeing work now on routers, trying to figure out which workloads, which tasks can be delegated to agents. Can you use a cheaper model? Do you need a frontier model? So there's a lot of exciting stuff happening in engineering. I'd love to peer inside your engineering team a little bit and talk about how you've been incorporating AI in your build process. Where are you on your journey adopting the tools, AI-ifying your product development loop? Where's the team at today?
John Wang (7:00)
Pretty much every engineer uses a coding agent at this point. We've spent a good amount of time making sure people have access to whatever tools they want. We basically give people a Ramp card and say, try these things out. That has some upsides and some downsides. We've also invested a lot in an internal tool, similar to Stripe and Ramp, for how we make automations happen, how we run our coding agents in a kind of cloud that has access to all of our systems and can read those systems. That's been very exciting. It's still early, but it's unlocked the ability for non-engineers to write a bunch of code too. Previously a non-engineer would have to download Cursor, Claude Code, or Codex and set up their dev environment, and that takes hours, it's a huge activation energy. Putting it in the cloud in an internal system means anyone has access to this, and it's all team-based. So that's been very exciting for us. One of the areas we're also thinking about is how do we make our systems reliable. We have customers like DoorDash, Stripe, Robinhood, where if we go down, their support operations go down too, so it's quite important for us to maintain reliability and not ship a ton of bugs.
We do want to ship really quickly, but we have a somewhat nuanced view: for our core systems, the important thing is high performance, high reliability, and then for our new builds it's go wild within this little sandbox of stuff you're building out new. And there can definitely be different types of people who excel at different parts of those.
Stephen Poletto (9:09)
Very cool. So you've almost defined zones of engagement, different rules depending on the risk to the business and how core and how critical a given subsystem is. That's really cool. When you talk about folks outside of engineering being able to ship code, I'm guessing this looks like authoring a ticket or maybe firing off a workflow from Slack. What kind of parameters do you have around that? Is it kind of a free-for-all, or is there a very specific workflow folks need to follow to spin up and do work in these cloud environments outside of engineering?
John Wang (9:47)
We originally started by writing a little doc that was just, hey, if you're not an engineer by profession, you need to go through these steps. The steps were relatively straightforward: you need to test it yourself in the PR description, you need to post a screenshot or a Loom video of you running it, and you need to get a PM to say, does this make sense? And only after all of those things, get an engineer to review. It worked pretty well. Maybe the thing all of those steps did was make it unlikely for anyone to ship a PR, because now I have to go through these multiple steps of review. So as time went on, we marginally relaxed some of these constraints. A PM or an engineer can now review, instead of it having to be a PM, but we still want you to run the code, and you need to be able to see this change is doing the right thing. That's still the biggest activation energy difference in how our folks actually run code. The reason we went about building this internal system is you just need a login, and the system has Claude Code and Codex running in sandboxes, so you can go and run something.
The nice thing is we can also set up constraints for non-engineers. There's a builder role and an engineer role, and if you're a builder, you need to run review, like any of the coding agents, you need to run slash review on it before you can even create a PR. It also means we can build on top of that, so we can put in something like, it needs to be a reasonable size, and you can run a coding agent to ask, is this a reasonably scoped PR for someone who's not an engineer? You can ask, is this touching a super highly used code path? That's actually also easy to pull out, because we have stats on how often any particular endpoint is hit.
Stephen Poletto (12:08)
So I'm guessing the agent has an MCP they can query, or something that tells them that.
John Wang (12:12)
Yeah. The interesting thing is we set up an MCP, and then we pulled it off, and we just gave it access to a CLI that has a little more context. The MCP basically blew up our context, because we connected it to so many different things. Now we're giving it access to one tool that has a bunch of sub-commands, and that's been working better, because previously you'd have a bajillion tools and we had no control over which tools it had access to. Now we're being a little more prescriptive: okay, for our Mezmo logging system, you can get logs, you can search logs, you can filter, but that's it. Just being slightly more prescriptive on that.
Stephen Poletto (13:03)
Tell me about your stack a little bit. How did you guys set up this remote cloud environment? Is it all built yourself? Are you using any third parties? What does it take to do this if other folks are interested?
John Wang (13:16)
It's relatively straightforward. It's built similar to our stack. It's on Go. Next.js is the front end. The containerization is using gVisor, and really it's just a bunch of Docker containers. So we try to make it super simple. We actually might open-source it in the next few weeks, because the reason we built it is we saw Stripe, we saw Ramp, we saw all these people doing these things, and I was like, I want this.
Stephen Poletto (13:49)
Yeah. WorkOS just did one two weeks back, they blogged about it. I'm talking to a lot of companies right now who are investing in this type of setup.
John Wang (13:55)
They're all doing it, and I find it almost infuriating that no one's open-sourcing this. When I was starting out, I did a bunch of work on Ruby on Rails, and learning from folks like Tenderlove and DHH, core members of the Rails team, they taught me a ton about software engineering. Having this available open source is probably the right thing to do, so we'll be open-sourcing this shortly. If people want to take it, extend it, that would be great. It will make it easier for us to move faster. There's nothing in there that's too novel, it's mostly something that's been built by one or two engineers, so it's not that big of an investment for us. Over the course of a few months, as a side project, we've been able to build this out. So it's quite exciting.
Stephen Poletto (14:57)
Very cool. You talked about the benefits to non-engineers, but are engineers using it as well? Has it become another tool for them in their development, where certain workloads get offloaded to this system?
John Wang (15:10)
Yeah. The big area we're seeing there is the automations. The main area we had problems with automations before was that, let's say you're on Codex or Claude Code and you run a scheduled job, you're the only person who can see that scheduled job, and a PR pops up and you're doing that PR. So the probability that you're going to run an automation to make your team's code better, for example, check to make sure all seven of the engineers on my team have enough tests on their PRs, is relatively low, because you're mostly concerned about your own code. Most of these coding agents were built by engineers thinking about their own workflows, so none of them are really on a team basis. That's where we've seen a lot of uptick on usage of this. Now we can set up something that everyone across the team can see, that everyone has access to and can adjust and make better.
The automation side of things is really helpful. One of the automations we're working on right now that I'm quite excited about is this self-automating loop, where it has access to all of the sessions that have run before. You can take a look and ask, what are all the corrections that have been made? Should we put that into AGENTS.md? Should we make a hook from that? And it creates a PR when it sees that, and then sends it to the keeper of our AGENTS.md. Stuff like that is really easy to build on top of once you have something that just runs for you. But when everything is Codex or Claude Code running in everyone's little containers, it's really hard to have the full picture.
Stephen Poletto (17:03)
Yeah, I've heard from some other teams that there are certain structured workflows you can kick off automations for. So if you have a bug intake process that goes through some level of validation from your support team, and it's well specified, maybe the agent can just auto-draft the PR, the fix can just show up in somebody's inbox for review. There are lots of new workflows that can be ideated like that, running continuously in the background. Another good one I've heard is a refactoring agent that's constantly trying to seek out patterns of tech debt and find ways to clean it up and reduce complexity.
But let's talk a little bit about cost management, because it sounds like you're fairly sophisticated, you've got a lot of these background systems running, your team is very heavily using the tools. There's been this theme of tokenmaxxing lately as a way of encouraging usage and experimentation and adoption. Where are you on the tokenmaxxing journey? How are the costs going? Is cost a concern for you and the team right now?
John Wang (18:20)
My take on tokenmaxxing is, I think it's pretty dumb. I don't know what kind of show this is, but if I were to give my real opinion, I think it's absolutely stupid.
Stephen Poletto (18:33)
Yeah, we want real takes from real operators. We want to combat the headlines, and we want to learn from practitioners and what they're doing in the real world. So you can be candid that it's dumb, and we welcome that on the show.
John Wang (18:46)
So the reason I think it's dumb is not because it's necessarily misguided. Yeah, you want people to use AI, but the way to do that is not via tokenmaxxing. The backstory is, we used to have a tokenmaxxing idea for our production systems. It was, use the absolute best AI model you can, run that for the AI agents we're actually shipping to our end customers, and just do as much as you can. Cost doesn't matter, go ahead and use the best intelligence available. That was great until we looked around and we were like, we are losing money, a lot of money, on a token-by-token basis for everything we do. Also, we have no muscle in any of our bodies that says think about whether you need this model. That was the worst part of it. Now you have to convert an entire team that's used to just throwing the best model at things into a team that thinks about this, and it took us probably four or five months to actually ingrain that into people's working cadences. Switching the models themselves is relatively straightforward, but there's a really long tail of areas where you're going to spend a ton of money for not that much extra gain. If you're doing summarization tasks, for example, you don't need Opus 4.7 or GPT-5.5. Honestly, Llama 7 will probably do a good job for you. It just takes a really long time to get that muscle back, and if you don't have that muscle from the beginning, you're going to look around at your token costs and be like, what the heck is going on. I have a friend at YC who was like, yeah, we spend an ungodly amount of money per person on tokens. I was like, well, man, hopefully you're building a lot with that, but most people are just using 4.7 for their open Claude project.
Stephen Poletto (21:19)
To write this email, please. I need extended thinking for that. I want the best email ever. Having that discretion and that judgment, how has your organization built up a knowledge base of which models to use for which jobs, and how to think about evaluating the effectiveness of a given model for a given task? I'm sure this is an ever-ongoing thing, as you're now being more cost-conscious. How do you think about that skill development and evaluating what tool is the right tool for a given job?
John Wang (21:57)
Probably what everyone else says, which is just evals. For our production AI, we put a bunch of effort into figuring out what we needed to eval and how to eval it well, because the big downside of evals is that if you have the wrong evals, you're more confident in something you shouldn't be confident in. And that's optimizing in...
Stephen Poletto (22:26)
...the wrong direction, because the score. Yeah.
John Wang (22:28)
Yeah, it's the worst place to be. That's an important thing to put into your engineers' minds: what are you actually optimizing for? For our AI agents, it's, is this a good response, is this actually resolving our customer issues? We've been putting effort into evals of our coding agents now, which is, is this providing a high-quality PR? Would this be shippable in 2020, pre-2025? That requires you to think about it a little harder, which is not something a super AI-pilled engineer really wants to do. They want to ship a bunch of stuff. That's also why I think tokenmaxxing is kind of dumb, because you're feeding the tokenmaxxer, hey, it's okay to just ship a bunch of stuff, you're just going to have a ton of slop cannons. If you want slop cannons, that's great, but we don't really need slop cannons. We need people who are shipping high-quality product at a fast pace, and that's actually not what a slop cannon is going to output for you.
Stephen Poletto (23:39)
And you're rewarding that when you look at tokens burned, tokens used.
John Wang (23:43)
Exactly.
Stephen Poletto (23:44)
I've also heard, I don't know if you're experiencing this problem with your team, but because it's now so easy to prototype and build demos, there are a lot of demos, a lot of stuff being built, but actually aligning that activity to the strategic roadmap and pointing the cannons in the right direction is an increasing challenge that I'm hearing. You can have a lot of activity, you can be running really fast, but is it aligned and pointed in a direction that's going to move the business forward for your customers?
John Wang (24:17)
Totally.
Stephen Poletto (24:18)
Is that something you're experiencing, the product specification and deciding what to work on? Is that becoming a new challenge?
John Wang (24:26)
I think there's a lot more emphasis on good product management now, for sure, and it's actually changed the way we think about recruiting. In our interviews, we're hiring much more specifically for product judgment. The other day we launched a new interview that, before you even get to a technical screen, is a quick take-home question, a simple prompt, something relevant to Assembled. Build us something here, very open-ended, allows you to think about anything, and really the question is, did you think well about how to build a high-quality system, and how do you think about trade-offs, which I think is the operative thing these days. The other thing we're trying to do with our team is, we have Friday demos, so every team demos what's going on, but we really encourage people to demo something like an eval. Hey, I've created these evals, here's the before-and-after graph, here's the performance graph, here's the impact.
Stephen Poletto (25:41)
The outcome, instead of just the activity or the demo.
John Wang (25:46)
Exactly, exactly. So that's been going pretty well. Obviously there are still times when it's just, ah, I went off and did this thing, and that can sometimes be really good. But we've basically given people a general sense of, you should spend 80% of your time working on core important features and 20% trying out new things. People will spend that either on fixing tech debt, building out something completely new, adding some polish items, or fixing bugs, things like that.
Stephen Poletto (26:24)
Right on. Very cool. And then you're bringing the evals idea to your coding agents themselves, which is also pretty cool. How are you thinking about your harness and the effectiveness of the harness? I'd love to learn a little bit about the journey you've had there, to give you confidence that the coding agents you're developing are generating good results and good outcomes for the team, hopefully spending not a ton of time just pushing back on slop in code review and things like that.
John Wang (26:57)
Yeah. What we have right now is, you basically have a commit SHA of some point in our code base, and then you have a spreadsheet of, here's a prompt, and here are notes for an LLM judge evaluator of what I expect a good implementation to look like. It's very simple, nothing too crazy, but we wanted to start with one simple thing we can do well. We don't have a lot of evals yet, so I'm curious how other teams are doing this, but at least having a few was better for specifically running and making updates to AGENTS.md. We've hooked it up into our CI, where anytime you change any AGENTS.md file, it'll run these evals, and you actually have to look at it and check. It's pretty simple stuff, but a lot of the simple stuff has bought us a lot on the production AI agents front, so we were like, okay, we might as well put some of the simple stuff into our own internal builds too.
Stephen Poletto (28:18)
And one of the reasons it's cool to learn from your experience is because you have been optimizing real production workloads around agent effectiveness, so the learnings you have from that applied internally to your own dev team makes a ton of sense. Is there anything on the horizon you're investing into internally, with how your team leverages AI in the PDLC, that you're really excited to talk about or share with our audience?
John Wang (28:47)
Probably the most exciting thing for us is just these internal tools we're building. It's making as many automations available as possible. For us, the most exciting thing has been iteratively decreasing the activation energy it takes to kick off a PR from a Linear issue. Doing the simple things to hook up a bunch of our systems, doing the simple things to make it easier so that when you put up a PR, that PR has to be reviewed, it's reasonable, and then you get a human to review it. Back in 2024 you would set up your CI/CD systems, and now it feels like, okay, now you have to set up your AI agent system in the same way.
Stephen Poletto (29:44)
Your local developer environment. I remember when we were going through hypergrowth, both when I was leading engineering at Lattice, and before that at Dropbox, there was really heavy investment in how do we get an engineer to make a commit and deploy to prod on their first day. That meant you had to have a good local environment, good CI/CD, good install scripts to get everything set up. Now it's that same idea of delightful developer experience, we need delightful agent experience, where they have the tools, the friction is really low for them to be able to achieve the tasks we delegate to them. So it's this new game of how to invest in the internal developer platform to make it agent-friendly. You're doing some really cool stuff, John. Appreciate you sharing some stories and some wins with us. Good luck out there with everything you're working on.
John Wang (30:36)
Thanks, you as well.
Let's talk about where your AI program is.
The ledger for your AI software factory
platform
© 2026 Attuned Inc.
New research: Leading indicators of AI coding agent effectiveness

