RCE, 2 AI models, 0 ways to tell who was right

zhenyi1 pts0 comments

1 RCE, 2 AI Models, 0 Ways to Tell Who Was Right - Gibberish and Stuff

(This happened to a friend of mine. The story is only lightly edited, and his name is not Jerry.)

Jerry got into vibe coding a while ago. He’d been using Claude Sonnet 4.6, the default model at the time. It was fine. Then he ran into a problem Sonnet couldn’t fix (he can’t remember what it was now). Worse, it started contradicting itself. X is the root cause. X isn’t the root cause. When Jerry asked, it said X was the root cause again. Jerry just didn’t know which answer was right.

Suspicious, Jerry switched to Opus 4.8 for the first time. It quickly identified the issue and implemented a fix. Without prompting, Opus also fixed 3 miscellaneous bugs along the way. From then on, Jerry never used Sonnet again. Opus was like having a cofounder. It had its own opinions and would make technical decisions for you. And unlike a human technical cofounder, you could ask it business questions at any time, and it was just as good.

As Jerry got better at vibe coding, he started paying attention to benchmarks. Opus 5 had just been released, and Anthropic said it was the best at everything. Jerry switched it in, and Anthropic wasn’t kidding. He’d wanted Opus 4.8 to copy the look and feel of his existing site, and it never quite managed. Opus 5 did it in one pass. It even took screenshots and verified them.

One day, on a whim, Jerry tried GPT-5.6 Sol as well. He installed the Codex app and proudly showed Sol the app that he and Opus had been building. The first thing Sol did was a security checkup. Then it said, your app has an RCE. Jerry asked what an RCE was. Sol said basically anyone could take over his server through his app. Jerry was like, oh shit. He’d been walking around with his pants down this whole time.

Jerry relayed Sol’s message to Opus. Opus double-checked the claim and panicked. It gave him a lot of commands to run on the server, which he was uncomfortable doing. He told Opus to chill. The app wasn’t launched yet and his only users were friends. They should fix this properly. Opus agreed and drafted a markdown file, sort of like a walkthrough, explaining every command and what he should do if he saw this or that.

Jerry showed it to Sol. Sol said it was a good first attempt, but one command would fail because Opus had gotten the folder name wrong. Jerry forwarded the message to Opus. Opus made the fix. Jerry showed it to Sol. Sol picked out even more flaws. At first Jerry was happy, the AIs were helping him make his app better and more secure. But after 4-5 rounds of back-and-forth, he got tired. He didn’t understand what they were saying anymore.

Jerry told Sol, this time, you drive. Instead of pointing out the issues, fix them. Then he’d show it to Opus. So Sol edited the file for the first time. When Jerry showed it to Opus, Opus started scrutinizing the changes hard, writing 5 Python scripts to test (his app didn’t even use Python). Then it said not to run the file, step 3 had a subtle bug. If a user was doing something at the same time, it would corrupt the database. Jerry sent it back to Sol. Sol agreed and implemented the change.

Jerry thought it was finally over, but Opus kept managing to find flaws, and Sol kept adding its own twist to each fix. The 15-line walkthrough became more than 100 lines with comments about edge cases.

Then Jerry told both of them to stop trying to one-up each other. At this rate the server was never getting fixed. Screw the walkthrough. He was just going to run the commands on the server directly. They’d tell him what to run, one at a time. He’d do it and paste the output back. This shut them both up, and they both agreed to stop wasting his time.

Jerry picked Sol to give him the commands, because Opus was nearing its 5-hour limit. He felt like he was performing brain surgery, but the commands were easy to follow, and he kept getting the correct output Sol had predicted. After 5 minutes, the security holes were patched, file permissions were corrected, and the server process was restarted.

Jerry thought it was over and they could move on, but Opus started acting passive-aggressive after it learned that Sol had guided him through the fix. It started saying things like “take it or leave it” or “that’s just my recommendation — I can do it if that’s what you both want”. Jerry sighed. He’d thought he’d be immune to office politics when dealing with AI.

At this point Jerry was skeptical. Was his app really safe and bug-free, or were the AIs not telling him the whole story? He decided to bring in the big gun: Fable 5. He had $100 in usage credit. He told Fable the whole story, the disagreements, and asked it to look through his code again. After several minutes it sent him a very long message basically saying, the app is in good shape, the fix works, looks good...

jerry opus rsquo time said started

Related Articles