
Why the groundwork matters more than the agent
Geni wasn't a weekend hack; it was six months of compounding work. Before managed agents even existed, the team had invested in Claude skills, AGENTS.md across roughly 80 codebases, and about 75 skills in an internal marketplace, so the agent plugged into a foundation instead of starting from zero.
The scaffolding is also what keeps spend in check. Usage guardrails are rarely hit, because engineers are already working inside shared skills and repos, boundaries set by engineering rather than finance.
Why toil is the right first target for agents
Point the agent at the reactive work nobody wants. Geni took on the vulnerability backlog and it's all but gone, from criticals down to the lows, which is why the security team is the one celebrating rather than panicking.
Humans still own the review. Geni opens a draft PR, the delegator does the first pass, and a teammate runs the formal review before merge. Changes stay small, no 10,000 line drops with a note that says good luck.
Why velocity beats tokenmaxxing
More spend is not the path to moving faster. Chasing token usage optimizes output instead of outcomes. The goal is velocity, and the efficiency, a two-to-three-nines cache hit rate, followed from good design rather than from burning tokens.
Clean context is the actual mechanism. Sub-agents each own one step and hand off through markdown files written to disk, so no single session compounds into a giant, slow context window.
Why adoption is a culture problem
Enablement beat mandates. A "Power Up" squad of power users dogfooded daily for two months, then rolled the system out team by team like on-call training, and the usage data went vertical.
It still took air cover and a trust battery. Leadership had to absorb short-term delivery tradeoffs, and the headroom came partly from velocity gains the company was already making before AI entered the picture.
Brady Allchin
VP of Engineering
Glossgenius
Brady Allchin has spent nearly 20 years in engineering, through hyper growth at Uber and across startups big and small. He's now VP of Engineering at GlossGenius, where the team has gone all in on agentic development. They built an internal autonomous coding agent called Geni, running on Anthropic's managed agents and wired into Linear: assign it a ticket and it plans, implements, and validates the work. Three weeks into production, it had already all but cleared a vulnerability backlog the team never had time for, and GlossGenius's usage had put it on one of Anthropic's leaderboards.
timestamps
0:00 Intro: 20 years across Uber, startups, and GlossGenius
1:19 Geni, and how GlossGenius landed on Anthropic's leaderboard
3:14 What it took to build: six months of groundwork
5:54 Early results: clearing the toil and the vuln backlog
7:12 Code review when agents are opening PRs
9:06 When PMs and designers ship their own fixes
10:36 What goes to Geni vs. local dev, and the road to bigger tasks
14:48 Sending a PRD, and engineering moving to the edges
17:50 The changing shape of an engineering team
20:08 Why Brady despises tokenmaxxing
23:48 Guardrails set by engineering, not finance
26:37 Getting from zero to one: the Power Up squad
29:55 Air cover and the trust battery
32:06 The personal AI tool Brady built in a weekend
Transcript
Stephen Poletto 0:00
Happy Friday, Brady. How are you doing?
Brady Allchin 0:02
Doing well. How are you?
Stephen Poletto 0:04
Yeah, doing pretty well. I'm also in New York right now. It's funny that we're in separate Zoom portals in the same city, but the weather is beautiful after a rainy weekend, so I really can't complain.
Brady Allchin 0:13
Very nice.
Stephen Poletto 0:14
You're leading engineering over at GlossGenius, but you have a pretty extensive background in engineering leadership. Maybe give our audience a quick intro to who you are.
Brady Allchin 0:24
Thank you. So I'm Brady Allchin, VP of Engineering at GlossGenius. I've been in the industry coming up on 20 years, worked at companies big and small. I got my career start in the Bay Area, and about 10 years ago made it out east to New York. I worked at startups, saw hyper growth at Uber, followed Travis over to CloudKitchens, now known as Atoms, and now I've really found my comfort zone at GlossGenius, a Series C company based here in New York, with technology teams in the US, Canada, and Colombia. We're here to help support small business owners be successful.
Stephen Poletto 1:09
Incredible. This is going to be a fun conversation, because you've had a good longitudinal study of software development practices over the last 20 years. You've seen hyper growth, you've seen small companies, you've seen big companies, and now things are changing so quickly in front of us with AI, so I think you'll have some interesting perspectives on that. But before the call, you were telling me Anthropic reached out to you because you're showing up on one of their leaderboards. What's that about? Tell me more.
Brady Allchin 1:39
Oh, this is one of our proudest moments. So we're very much AI-pilled here at GlossGenius. We've really embraced agentic development, and a few months ago, when we saw Anthropic release this capability called managed agents, we really latched onto it. We created an internal autonomous agent we call Geni, sort of a double entendre for Genius as well as Aladdin, and we built a platform that leverages managed agents. To an end user, to a developer, it lets you open up Linear, assign a ticket to Geni, and Geni will do the ticket. Under the covers, it's spawning up multiple sub-agents, it's creating a plan, it's doing implementation, it's doing validation, and all that orchestration is happening within Anthropic's platform. What we built with Geni is effectively a bridge, a coordinator between Linear, a handful of MCPs and internal systems that we have, and the Anthropic platform. It's been fascinating to see it come to life. We're just wrapping up week three of it in production, and the amount of capability it's brought us is phenomenal.
Stephen Poletto 3:03
Very cool. I work with a lot of companies and startups who are using Claude Code heavily, but in a local session environment. I think they see the blog posts from Stripe and WorkOS and Ramp, who have talked about these background agent systems, but they seem like a lot of work to get off the ground, and I think there's some inhibition about the investment required. I'm curious, for your experience with that, what did it take to get this Geni up and running?
Brady Allchin 3:41
Oh yeah, I pored over those articles about Stripe's Minions, Ramp's Inspect. They were great, but they dedicated entire teams to it, so that fear is not unfounded. I would say it wasn't just a sprint to put it together, it's the work that was compounding over the last six months. We had done quite a bit of investment in Claude skills, making sure that our local developers were successful and had all the tools to help with major refactorings or migrations and to scaffold new endpoints appropriately. We invested in those, and that was a great ground truth to just sort of wrestle agentic development. When managed agents came out, it was this lightbulb moment where we realized we did not have to manage the compute. Now you can manage the compute, they just released a new version, but we didn't have to manage the compute, so we didn't have to deal with our AWS environment or Cloudflare or something like Modal. We could use them out of the box. We'd also recently done a migration off of Jira into Linear, and Linear provides a really great agentic interface out of the box, so they're really emphasizing this workflow where Linear could be that single pane of glass for you to run your PDE organization, and we've embraced that as well. There are these companies out there building the tools and technologies that we've just embraced wholeheartedly, so we're not needing to reinvent what Linear has, we're not needing to reinvent what Anthropic has, we're just marrying the two together and adding a little bit of our own special sauce.
Stephen Poletto 5:41
What kind of results are you seeing? You said it's been three weeks, which probably means it's still early, but what are some of the early signals around the impact it's having on the organization?
Brady Allchin 5:50
I'll open with this: our security team's overjoyed, which is probably a strange thing to say, because usually security teams are freaking out about agentic development and vibe coding and all that. But our security team is overjoyed because we've effectively eliminated our vulnerability backlog. When we released Geni internally, the pitch was that this is going to eliminate toil for engineering. As an engineer who goes on call, you're dealing with reactive work, you're dealing with bugs, you're dealing with a lot of investigations, and that's where Geni shines. I, as an engineer, can use my brain power on complex problem solving, and all this toil I just want to ship elsewhere. All of that reactive work, the "hey, we need to bump this package version, but actually it means we have to upgrade these three other libraries, and then we have to change these function definitions," Geni can just take care of it. So our vuln backlog is all but eliminated, and we're not talking just criticals and highs, we're talking the mediums and lows.
Stephen Poletto 6:56
The stuff that's historically been tricky to justify resourcing for, right? Because the severity is not that high, and you've got a lot of other stuff to do, so it can be difficult to close those quickly. What does the review process look like? Is this creating a challenge downstream on the code review side now that you've got this system cutting all these new PRs?
Brady Allchin 7:20
Yeah, it's a great question. It was a challenge we were facing even before Geni. The workflow we've aligned on is that Geni creates a PR in draft, and it's the responsibility of the ticket author, the Geni delegator, to review and do that first pass on the PR. You might have a back and forth with Geni. You might be like, "Ah, that's actually not what I meant, let's go change this." So you've done a first pass, then it's ready for review, and then you tap a co-worker on the shoulder, another team, to do the formal review, and from there merge it. It did increase our volume, but this is a challenge that we and many other companies are facing. I don't have an answer for you today. We're truly in the midst of figuring out: do we want to lean toward the stacked diff model, which some companies have, and which I've come from as well, or do we want to focus on blocking that next PR until the previous one is fully complete, versus creating a full stack? It's an interesting topical discussion right now. But we're not believers in "here's a 10,000 line change, good luck." We want to keep things small. We still require a human in the loop.
Stephen Poletto 8:46
With Geni online now, are you starting to see folks outside engineering using it? How do you create rules around that? Can PMs and designers cut PRs using the system? What does that look like?
Brady Allchin 9:00
I mean, this is evolving as we speak. We had a tech town hall earlier today, and one of the things shown off was how a PM just stepped in and fixed a problem they wanted to fix themselves. They used both Claude Design and Geni to make that happen. That is the goal: to empower engineering to not be the bottleneck anymore.
Stephen Poletto 9:25
How does engineering feel about it? Is there any concern about inheriting code that they'll be responsible for owning from those cross-functional contributions? Or is it all good?
Brady Allchin 9:39
No, I'd say it's the same. It's accelerating discussions we were already having around harness-driven development, agentic development, where there's a burden, and I think burden is the right word, for the air-quote author of the agentic work to do that review ahead of time and understand what's happening before putting their name on it. This is accelerating it, but it's a challenge that we have today, frankly.
Stephen Poletto 10:11
So if I'm an engineer at GlossGenius today, I have this system, this Geni, that I can use to do background tasks. What's your current mental model around the type of things that should go to Geni versus the type of things I should do in my traditional local development environment with an interactive session?
Brady Allchin 10:32
So right now, as of today, it's toil. We want to eliminate toil, so let's get rid of the bugs, the nitpicks, the things that require heavy investigation. But if we fast forward weeks, months, it's going to change, and the scope and complexity will simply grow. Where I see us trending toward is engineering focusing so heavily on the planning stage, whether that's in a local session or helping craft technical design documents. That's where the time and energy needs to go. The code part is just going to be a smaller and smaller time spent when it comes to the overall SDLC.
Stephen Poletto 11:15
Yeah, so you're thinking in the factory mindset, maybe, where you can eventually take heavier-weight specifications and offload them to an agentic factory with multiple stages, which gives you this fleet of agents to go implement something larger in scope. What do you think the steps will be to get there? I know you're not there, but if you just think about the path forward, what's on the horizon? What are some of the next things you're thinking about tackling to enable that?
Brady Allchin 11:48
A few things. One is just context setting, making sure these harnesses, whether in the cloud or locally, have all the context necessary. Can I pull the right information from our internal wiki? Can I pull the right data from Figma? Can I look at a production data store? So context setting is very critical, and I'm not wanting to just slap a bunch of MCP servers on it and call it a day. Context setting was the number one thing on the horizon right now, making sure Geni has the right information to make the correct decisions. I think that's entirely possible. I think that's the world we're headed toward at a very fast clip, and I think it does change the nature of where engineering fits into this. I still believe engineering is critical at the top of the funnel and the bottom of the funnel. It used to be very critical in the middle of the funnel and less so on the edges. At the top of the funnel you had your marketing team, your product managers, and so on, and at the other end of the funnel you have QA, and engineering sat squarely in the middle. Now I feel like the role of engineering, you're still dealing with validation at the end, but you're also playing the role of a builder. I think about what brought me into engineering in the first place, and it's the love of building. I wanted to build games, I built Legos. I love the idea of building, and being a builder at heart lets me move up the funnel and think about how I can drive outcomes, because at the end of the day, a business is not successful based on how many lines of code you've written, a business is successful based on the outcomes you're driving. So as an engineer, I think it makes our role almost even more important and more powerful. I've spoken about it internally at GlossGenius: everyone hired in engineering was hired to be a problem solver, so let's focus on the problems worth solving. When I think about engineering leadership, what I've preached over the years is that as you automate yourself out of a job, you will have more work to do. There is never a shortage of work to be done, and I believe that wholeheartedly here as well. We're moving toward those edges, and there will be more challenges to solve.
Stephen Poletto 14:32
Yeah, well said. The coding time is compressing, which increases the importance of specking and choosing what to work on, solving the problems well with good designs, and then validating that they work, they land, they make customers happy. Has that reframe, and the expanding role of engineering, changed how you think about hiring, developing your team, who you're promoting, the internal people, the soft bits?
Brady Allchin 15:02
We're right in the middle of it as we speak, thinking about what the shape of engineering of the future is. We used to talk about, at former places I've worked, the different archetypes of an engineer: the entrepreneur, someone who's great at the zero to one, the scalability expert, the investigator. There were these engineering archetypes, and I remember having discussions about promotions or performance reviews, and it was so hard. These were always apples to oranges. Everyone was critical, everyone was needed, but with these archetypes you simply could not compare the person who was great at prototyping but terrible at scale. You needed that person, but you also needed the person who could scale. It was hard to find the unicorn who could do them all. So when I think about the shape of engineering going forward, I almost think of it like different strata in an organization. You're going to have your infrastructure organization that does care about scale and reliability and guardrails coming from security and all of that, that's very important, but at the top level it is the builders. I don't know if it was Meta or Amazon that talked about some potential role changes, some role definition changes, where there might be a title called something like AI product builder, and that might be a title that people have who had a technical background at one point. So I think it's an ongoing conversation, and I think engineers will naturally gravitate toward moving closer to the customer or closer to the platform.
Stephen Poletto 16:47
Very interesting times, as everybody's jobs and their descriptions are evolving very quickly. Shifting gears a little bit: the fact that you showed up on Anthropic's leaderboard is awesome, because it means you're doing innovative stuff. It could also potentially be worrisome from a cost standpoint, in the sense that we're seeing headlines across the industry around token usage and the increasing need to budget for it, and people overrunning their budgets as they develop all these background agentic systems. How are you thinking about the cost dynamic, token maxing, encouraging your team to use Geni, but also knowing that you have some level of predictability in your Anthropic bill? Talk to me a little bit about that, the financial side of all this agentic development work.
Brady Allchin 17:41
I'll open with this: I despise token maxing. It's such a buzzword. I can imagine the discussions happening in boardrooms across the country about how "we need more agentic spend, it's the only way we can move faster." It's something that transparently drives me nuts, because it's focusing on the wrong thing, it's focusing on output and not outcomes. So, tying it to the discussion about managed agents and getting on said leaderboard, it wasn't because of token usage. Our token efficiency is pretty ridiculous, actually. It's two to three nines in terms of our cache hit rate. Having these sub-agents is so critical because they're each doing a piece of the puzzle. It's almost like when you read the Hacker News posts about "hey, I've created a bunch of agents that represent the company, and this is the CEO, and this is the CTO." It's kind of like that's what we were doing here with these sub-agents, and the handoff between each one is a markdown file. It's "hey, I'm the planner, I'm taking in all this context, I'm writing it to disk," and then I hand it off to the next step, and they go pick it up. They kind of have a different context window from there. They're not inheriting, it's not compounding like one giant Claude Code session. That's very important from both a token spend perspective, but also speed. What drives me wasn't thinking about budgets or "we must be efficient on token maxing." I was obsessed, or still am obsessed, with velocity. I want to make sure we're moving as fast as possible, and I've seen sessions get unwieldy and slow and thrashy when you're just siphoning the world. So velocity is what drives me. It's what I instill from a cultural perspective, and being anti-token-max is kind of a positive synergy there, where we can focus on the right thing. We can focus on making sure we have a process that is debuggable. We need to understand what's working and what's not, be able to replay sessions, be able to see where it may have gotten confused. When we do this, when we analyze how Geni is operating, we're able to do that in a fairly efficient manner.
Stephen Poletto 20:19
Very cool. I agree with you wholeheartedly on the token maxing thing. Goodhart's law in practice, where the measure becomes the goal and people start doing things that are misaligned with delivering customer value. So totally with you there. You shared some really good tips and tricks around how to manage context windows and facilitate handoffs and write roll-ups of summary state to disk so that another agent can process it and keep your context window clean. Those are really good tips and tricks. How do you make sure your team's following them? How do you manage the spend within your organization? Are you trying to constrain it at all, or do you let people run pretty free rein? What are the enforcement mechanisms or processes you have in place around token use?
Brady Allchin 21:12
So there's always a guardrail. People don't realize there's a guardrail, but there's always a guardrail. Funny enough, we haven't hit said guardrail. The only time we've hit the guardrails is in the agentic parts of the product itself that we've released to customers, because we just didn't know how things were going to react in the market, so that's where we had to do investigations and right-size accordingly. But when it comes to our development processes, frankly, we haven't hit it. Someone could be thinking, "well, it sounds like you're not using it enough," and that sounds like the obvious response, but I would actually credit the team's internal investment in setting context files. We have AGENTS.md sprinkled all across our 80-some-odd codebases, we have about 75 or so skills in a shared internal marketplace that we've done trainings on to encourage developers to use, and so if you're using these skills, if you're using our repos, you're already within guardrails that aren't set by finance, they're set by engineering. That upfront investment has absolutely paid off from a velocity standpoint and from a FinOps standpoint, I'll say unintentionally so. That investment is something some companies are still not prioritizing, or they're just giving raw access to these amazing tools but then not putting the scaffolding in place in their repos. That's a big mess, in my opinion.
Stephen Poletto 22:59
Yeah, sometimes when I read the headlines, I wonder how that much money can be getting spent and burned, but I think a lot of it is a lack of maturity in the developer infrastructure, in the foundation. For folks who might be going through the adoption journey, they've got Claude Code or Cursor or Codex licenses distributed to their team, folks are experimenting with a lot of individualized workflows, and they want to start investing in a centralized platform foundation for this. How do they get started? What was your journey like to get to that place where you've got this library of skills, this documentation, this harness that's now working so well? What does it look like to get from zero to one on that?
Brady Allchin 23:44
Looking back, it'll seem like such a brilliant set of steps, but at the time we were just figuring it out as we went. Essentially the workflow was this, about a year ago now: "hey, here are all these cool tools, let's issue Ramp cards to folks to go spend money and figure out some of these tools they could play with." A lot of engineers were playing with different tools. We collected a lot of feedback, and we identified the power users, the people who were like "hey, this is really cool," and not just from a toy perspective, but "hey, this is changing the way I work." So my boss and our head of TPM here at GlossGenius got together and formed a squad they called Power Up. It was a virtual team, and it took a bunch of those super users, brought them together, and said "clear your plates, let's figure this out for GlossGenius." So the simple answer is, you've got to create space for this. We created space for it, and that's where they learned by doing. They figured out "hey, this is how we should do a skill, let's go try this out," and they were dogfooding every single day, having stand-ups, testing each other's work. After about two months of that, they turned around and said it's time to bring it to the organization at large. Of course there were people on the outside who were seeing this and experimenting, so the company's still evolving, but then we did a very deliberate team-by-team training to announce "here's how you get set up," the same way a company might do an on-call training. Here's how to use your paging software. We did the same thing here. Here's how you set up the marketplace, here's how you get the skills, here's how you trigger the skills, here's how it impacts the codebase. Now, together, let's go through what it means to remove a stale feature flag, which is a very common sort of toil we may have to do. We walked through team by team doing this, and we came away, stepped back, looked at the data, and said, "wow, we just went vertical as a result of it."
Stephen Poletto 26:04
I think that's the hard part for a lot of organizations. You have to make a short-term trade-off of predictable execution for this long-term benefit, where now suddenly you'll have all these velocity gains and your whole organization will be working more effectively. Did you have to negotiate that with your product team, with your CEO, with executive stakeholders to get the air cover to do this? What did that look like?
Brady Allchin 26:31
Of course. If you end up in a discussion where one person is so critical to the delivery of some project, and "oh, if we don't have this person it delays this project by three weeks, oh my goodness, this is going to be terrible," we absolutely had to go through those negotiations. At some point we just had to put our foot down and say "this is how it's going to be." I think what gave us that headroom was actually things unrelated to AI, which is just us as a company growing, maturing, bettering our processes overall. So removing AI from the equation, our velocity was already improving, and that afforded some of that headroom. But that's a little bit unique to us as a growing startup.
Stephen Poletto 27:22
Yeah. A lot of these things come down to trust battery, right? If the engineering team has a high trust battery with stakeholders, they can say "trust us, we've got to take this risk, we're going to invest in the foundation, it's going to pay off." If you can do things to build up that trust battery to then make that big push, you can create the space to do these transformative investments. So, Brady, this has been a really fun conversation. Thank you so much for sharing some of the learnings and some of the experiments you're running at GlossGenius. Is there any personal AI thing that has recently excited you, maybe with your own workflow, or something really novel and cool you've seen one of your team members do, a new development that's come out in the past couple of weeks that you're excited about? Just curious, what's exciting you today?
Brady Allchin 28:11
This is such a managerial response that I'm going to have, but I was sick last weekend, sitting in bed, and I was like, "sure, I use Claude every day and it helps me get my day done, but there's something more I could do." I spent the weekend hacking around and rethinking that end-to-end process. I'm still kind of testing it out, but I kind of feel like I built my own EA. For me, going from meeting to meeting to meeting, I really need to create these big empty blocks of focus time to just get in deep on a technical challenge, a people challenge, helping move the company forward. I'm really excited that, as someone who context switches maybe too much in the job, I've created something that's going to help create space for me to be heads down and focus again, because I miss it. Yeah, that's what I'm looking for.
Stephen Poletto 29:16
As someone who runs between meetings way too much, that resonates a ton, so I hope it works for you. Thank you again, Brady. This has been a really fun conversation.
Brady Allchin 29:25
Likewise.
Let's talk about where your AI program is.
The ledger for your AI software factory
platform
© 2026 Attuned Inc.
New research: Leading indicators of AI coding agent effectiveness



