The AI wrote a beautiful case for the rewrite. Then we asked it to argue the other side, and it wrote a better one.
- A decision log that got longer before it got shorter
This spring, one of our teams nearly rewrote a login service in a new language. We had the research. We had the design. We had a new repository with a database schema, a data layer and passing integration tests.
Five days after the recommendation, we shelved all of it. The service is still running on the language we almost left, and it's still shipping.
I want to walk through how that happened, because it isn't a story about AI getting it wrong. It's a story about how easy it is to trust a well-written recommendation, and what it took to test one.
You don't need a team or a login service for this to apply. Any time you ask an AI "should I use X or Y?" or "is it time to rewrite this?", you're in the same position we were.
The Case for Leaving
The login service is written in Ruby on Rails. It handles sign-in and issues the tokens every other service on the platform trusts. We needed it to support single sign-on with Microsoft, Google, Apple and the SAML identity providers our customers run.
When we started building that, we hit what looked like three hard limits in the Rails ecosystem:
- The Ruby library for the OpenID Connect server side had a "looking for maintainers" banner on its repository.
- The main Ruby SAML library had shipped six high-severity security fixes in eighteen months, including a full authentication bypass shown at Black Hat.
- One token-exchange standard we already relied on wasn't built into the main Ruby OAuth library. We'd have to write and maintain it ourselves.
AI-assisted research compared the options library by library, with citations, and came back with a clear verdict: port to Python. Python's libraries covered all three gaps.
It was a good document. We started building.
Verdict Two, Then Verdict Three
A few days in, someone asked a reasonable question: why not .NET? Other services on the platform already used it.
The comparison came back leaning .NET. And it surfaced something the Python recommendation had missed entirely: the Python library it was built around had published ten security advisories in eight months. One was critical. Several were authentication bypasses in exactly the code paths we'd depend on.
Later that same day, a follow-up looked at how well each option fit our actual database design, and the call swung back to Python.
That's three recommendations in five days. Each one was well-sourced, clearly argued and stated with confidence. Each one was reversed by a fact the previous one hadn't weighed.
Then the Team Pushed Back
When the team met to review the plan, two objections came up.
First: we had a go-live date seven weeks out. Switching the language of a security-critical service mid-project is the classic rewrite risk. Joel Spolsky's "Things You Should Never Do" is 26 years old and still gets cited for a reason.
Second: AI changes the math. If the Ruby ecosystem has gaps, AI assistance makes it realistic to fill them ourselves in Ruby, in a language the team already knows how to run and debug at 2 AM.
Neither objection had been weighed as heavily as it deserved in the research. So we asked for a document that made the honest case for staying.
The Honest Case Found a Wrong Fact
The "stay on Rails" assessment went back and re-verified the earlier documents' claims against the original sources.
The Python recommendation had described the main Ruby OAuth library as stale, with almost two years between releases. It had the latest release date wrong by two years. The library had shipped a release two months earlier and had six releases in the prior two years. It was actively maintained.
That fact had helped justify the rewrite, and it was wrong. It sat in a document that read as thorough, and it went unchallenged until someone was asked to argue the other side.
The assessment also found the reverse. The existing Rails service already had a working, tested OAuth server, turned off behind a flag. It had never run in production. Staying wasn't free either.
Its conclusion wasn't "stay." It was:
Neither column is dominated. That's the honest answer.
The two paths carried different risks. Python meant standing up the platform's first customer-facing Python service, built on a library with a busy security record. Rails meant assembling several libraries with known weak spots and owning some custom security code indefinitely. The document laid out both and left the call to the people who'd live with it.
What Actually Decided It
With the advocacy stripped out, the decision came down to one question: which risk is this team best equipped to manage, seven weeks before go-live?
The team decided that same day to table the Python service until after go-live, and to keep running what it already knew. The new repository is archived now, and the work was scheduled to be backed out of the dev environment.
Five months later, the Rails service has had more than 200 commits of real work. A steady share of them extend the token-exchange code that was one of the three original reasons to leave. It's written in Ruby and maintained by us, just as the assessment predicted. That was the cost of staying, and we chose it knowingly.
What I Took From It
A confident recommendation isn't a finished one. The first document was careful, cited and wrong in ways that only showed up when someone argued against it. AI makes a polished recommendation cheap. Checking it is still work.
Make it re-check the facts the decision rests on. Pick the two or three facts a recommendation leans on hardest and have them checked against the original source, not the AI's summary of it. That's how the wrong release date surfaced.
More analysis isn't more certainty. AI made research so cheap that we produced four long documents in five days, and the verdict moved almost every time a new one landed. That's what over-thinking looks like when it's nearly free. Name the question you're actually deciding up front, or the analysis will keep finding new things to weigh.
Use AI as an advocate, not a judge. The most useful document in the whole process was the one we asked to argue against our plan. It found the wrong fact, and it found the real tradeoff. You get a different answer when you ask for the case against than when you ask whether the plan is sound.
AI lowers the cost of building, not the cost of owning. The team's "we can build the gaps ourselves with AI" argument was right about building. The assessment's answer was the part I keep coming back to: for security code, the expensive part isn't the first version. It's tracking every new vulnerability and spec change after that, forever. A library spreads that cost across all its users. Code you write yourself doesn't.
The deciding factor usually isn't technical. The research compared libraries, security records and data models in detail. What decided it was a go-live date, and which kind of risk the team would rather carry into it. That's a leadership call, and the best thing the analysis did was stop pretending it wasn't.
We didn't avoid the rewrite because the AI talked us out of it. We avoided it because we made the AI argue both sides, and then made the decision ourselves.
