It's been a good two weeks now that I've been working full-on with Codex and I can say it, it's the best thing out there right now with OpenAI's latest models.
I also want to come back quickly on my very simple setup and see if the results are there.
Results of my setup
If you read this newsletter regularly, you know I change setups like I change shirts. Always looking for something complex, beautiful, well built, that lets me do everything I want.
I tested a ton of setups, apps, hackbots, whatever you want. And two weeks ago, I went back to the simplest setup possible: my server, ssh and tmux.
And it's been the best decision I made. It's so simple that there is zero friction and no annoying setup to do. I can switch from Claude to Codex whenever I want and back, I spin up tmux panes as soon as I have a new project or target, and off we go.
I found a lot of bugs these last few days (mostly with AI), and the latest models have completely changed how bug bounty works. It's amazing and scary at the same time.
But might as well make the most of it, and secure as much as possible while it's still possible.
There is also a lot of context switching when you work on several targets and projects at once. You jump from one conversation to another, from one project to another, and it's not great for the brain even if it's pretty stimulating, knowing it's almost unlimited.
I also fixed the problem I had of scrolling non stop while the agents work. I found something more stimulating and smarter to do, playing chess.
I'm completely bad at it, but the upside is that I have a huge margin for progress and so much to learn. And you can play games that are more or less long, and it makes you think during that time, so I feel less like I'm losing my brain. That's not bad at all. We'll see how it holds up over time.
Claude is useless
Claude is getting worse and worse in my opinion, even if the Opus 5 benchmarks say it's incredible, I don't get that feeling.
I find that it doesn't want to work. Maybe it's because I talk to it in French and we're lazy and we don't like working, but still.
The other day I gave it a precise goal, find me a certain number of vulnerabilities, and it stopped after 2, telling me "I don't feel like going further, this is already what I found, I won't continue".
While GPT keeps going without flinching. I have to fight with Claude to make it do what I want, it's way more stubborn in my opinion (it's crazy, we're personifying an LLM that is just a statistical machine, but anyway).
All that to say that for now, I'm using Claude less and less. Maybe they will come back stronger, but in my case, it's a breakup.
Especially since on the GPT side we have our friend Tibo burning OpenAI's whole budget by giving us resets every 2 days, which is more than pleasant for us.
Bug ownership
Another feeling I have is that the bugs the AI finds, I don't feel like I'm the one who found them (and it's true, honestly).
But that's the current meta and it's a strange one. When I report a bug, I'm still happy, but I don't tell myself that I found it. Sometimes I'm almost a bit ashamed to report it, almost like it was stealing.
When in the end what we do is still helping companies secure their products, so it's not so bad. We'll see how all this evolves, we're living through a strange period, and a pretty exciting one.
And I think there's no point being pessimistic about it anymore. Either way the tools and the models are here, so might as well use them to their full potential rather than being bitter and getting left behind in a few months.
On that note, I'm going back to enjoying the spa while my agents work for me, have a good end of the week everyone.
Comments