Astra was the first model that made me stop and go “wow.” With other releases, I could usually predict what they’d try next when a task got hard. Astra spun up multiple Playwright browsers to test and compare things, then recorded videos to show what it had done. I’d seen people talk about doing that online. I hadn’t asked Astra to do it.

It also overengineered some open-ended problems I gave it. And I can’t use it every day without upgrading my plan. That matters more to my recommendation than how impressive it was.

What I actually use

GPT 6 Sol on low is my default in Pi. So far it feels like a flat upgrade over 5.6 Sol for my work. The change I notice most is the prose it writes without any writing skills. Low is enough for most tasks. I only turn up the reasoning for really hard ones.

Fable is amazing for code I can trust enough to review and merge. I can give it a wildly broad task, and if it gets stuck, it comes back to me instead of burning tokens forever. I don’t have it on my current plan, though. Sol on low does enough that I rarely need to reach for the big guns.

Opus 5.5 is what I’ve started using for my longest-running tasks with lots of context. It feels fast, and the prose is way better. The “Claude speak” that bothered me before seems to be gone. Opus 5 was a flat no for me. 4.8 was alright, but Sol 5.6 beat it easily. I didn’t expect to come back to Opus, but here I am.

The plan changes the answer

My recommendations assume subsidized subscription plans, not pay-as-you-go API pricing. Running agents at API rates every day would be too expensive for me. Check what’s included in your plan before copying anyone’s model ranking. The expensive option needs to do something your cheaper default can’t.

I’ve updated my Recommended AI Tools page with what I use now. If you’re choosing a setup, start with a model you can afford to run often, then work out which tasks actually need something bigger.