When to reach for Opus

After I updated my model recommendations, I got some feedback from people still reaching for Sonnet or models like Luna. I get it. If I’m breaking a job down and checking every step, Sonnet can get me pretty much the same result as Opus most of the time. But why am I doing all that planning for it? That defeats the purpose of a fully agentic approach. What I typically try to do with Opus is give it a wider task: give it the goal, let it work through the branches, and come back to review what it did. ...

2026-09-28 · 2 min · Tyler Collins

Why I still use Sol on low

Astra was the first model that made me stop and go “wow.” With other releases, I could usually predict what they’d try next when a task got hard. Astra spun up multiple Playwright browsers to test and compare things, then recorded videos to show what it had done. I’d seen people talk about doing that online. I hadn’t asked Astra to do it. It also overengineered some open-ended problems I gave it. And I can’t use it every day without upgrading my plan. That matters more to my recommendation than how impressive it was. ...

2026-09-23 · 2 min · Tyler Collins

Before You Rebuild Your RAG Pipeline, Try an Agent

If your RAG system keeps giving incomplete answers, try going fully agentic. I’m not saying retrieval-augmented generation is useless. It can make a large collection of documents cheap and fast enough to query. But if you have some tokens to burn and your corpus isn’t massive, it may be better to let an agent search the original documents itself. I ran into this with a public-facing technical wiki. The articles are dense. They have long tables, code samples, configuration details, and links to many related pages. A useful answer might depend on a warning above a table, one row in the table, and an example farther down the page. ...

2026-09-15 · 4 min · Tyler Collins

Read the incident report before declaring AGI

I keep seeing people describe the OpenAI and Hugging Face security incident as evidence of AGI or how we are all doomed in the next year. The bots communicated with each other, escaped their sandboxes, and got loose on the internet. It sounds completely wild when you reduce it to a few headlines. Then you read what happened. Marius Horatau’s article, The Hugging Face Incident Is Not an AI Story, does a great job of reading the incident as a security engineer instead of treating “AI” as the explanation for everything. In my opinion, he’s got it exactly right. ...

2026-09-14 · 3 min · Tyler Collins

Check what your AI plan does with your data

There is a lot of drama around OpenAI’s claimed solution to the Navier–Stokes existence and smoothness problem. I’m not going to try to evaluate the proof or untangle the dispute around it. One sentence near the bottom of OpenAI’s announcement was very interesting to me: While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. Anyone using an AI tool for research stop and check their plan. ...

2026-09-09 · 3 min · Tyler Collins

Stop installing things your agent doesn't need

Stop adding things to Claude Code or Codex that you use once and then forget about. I like customizing my agents. Skills, plugins, extensions, whatever the particular tool calls them. But I want to combat the slow creep towards the dumb zone, and loading a bunch of unrelated stuff into every session works against that. Matt Pocock has a useful description of the smart zone and the dumb zone. Early in a session, the agent is at its best. The longer the conversation goes, the more mistakes start to pop up. This is where those early-days patterns of hallucinations really came from. The context window isn’t a promise that it’ll work well at all times. ...

2026-09-08 · 3 min · Tyler Collins

From a Throwaway QC Prototype to PyLossless

In March 2025, I had an idea for improving quality control in PyLossless. I wanted to bring back being able to review an EEG recording to select independent components, choose a time window, and immediately compare the raw signal against the result of removing those components. It had been done in the previous version in MATLAB, but I was less familiar with the inner working of PyLossless at the time. I was also a complete novice when it came to things like Qt 5, PyQt, and the mne-qt-browser. ...

2026-08-28 · 5 min · Tyler Collins

Writing More Code with AI Agents

I recently gave a SHARCNET General Interest Webinar called “Writing More Code with AI Agents.” More than 300 people registered. Attendance was excellent, there were lots of questions, and I’ve had a decent number of follow-up conversations over email. Pretty happy with how it turned out! The response also confirmed why I wanted to give the talk. People are constantly asking me agents and what they should be doing with them. They’re watching other researchers and developers move very quickly with these tools, and there’s a real fear of missing out. They want to try them, but they don’t necessarily know where to start or how much of the output they should trust. ...

2026-08-27 · 6 min · Tyler Collins

Revisiting Cookiecutter in the Age of Coding Agents

A few years ago, I used Cookiecutter to show how to get from import blah to pip install blah. This sounds simple, but it isn’t. A Python script can live almost anywhere, but a formal package needs the right structure to be built into a wheel. It needs dependency metadata, versions, releases, a licence, and some way to test that it still works. Cookiecutter gave us a standard template. Instead of remembering every file and setting and creating them manually, we answered a few questions and started with a working package. I gave a talk about that workflow in 2022. ...

2026-08-24 · 5 min · Tyler Collins