Imagine having a bunch of Claude agents doing time sensitive work and this happens and now a single human needs to pick up several 8k PRs authored by bots that can't talk to you now. Not sure how effective it would be to switch models using OpenRouter and the like given that, as the Earendil guys said, there is a lot of session detail that the provider doesn't share. The uptime is really going to cap their market penetration imo.
If you're not using Codex and Claude at the same time, you're missing out heaps. Even without outages, OpenAI models are great reviewers for Claude; and vice versa.
Other than trivial PRs; everything I do with Opus/Fable gets reviewed by Astra; and everything I do with Astra gets reviewed by Opus/Fable.
Using only models from a single vendor, is like testing your website/webapp only on Chrome.
There's far more than zero evidence, see <a href="https://arxiv.org/html/2402.08806v1" rel="nofollow">https://arxiv.org/html/2402.08806v1 ; or the industry-standard practice of using multiple model families as LLM judges; or even 1P implementations <a href="https://code.claude.com/docs/en/advisor" rel="nofollow">https://code.claude.com/docs/en/advisor (which misses most of the benefit; since you want different model families, with different pretrains and posttrains).
Eh that’s about medical diagnoses. Not directly transferable. For regular usage, multiple models just make you feel productive but I bet they aren’t any more productive than just one.
In my workflow, I still struggle to employ both providers in a way that makes sense and doesn't require me passing data or prompts between the two.
For me, commits are the atomic block; and a good commit should be self-contained anyway; where no additional context is necessary. If an agent can't figure out what a commit is supposed to do, with only the commit title/description and diff, it's a bad commit, and this has always been true in software engineering. I do not pass prompts between the two ordinarily.
After claude or codex finishes a commit, I switch console tabs and ask the other to review it; with the commit ID. I often find it helpful to inject a bit of human knowledge, and callout any areas of attention I see from a quick skim. (But that could just be me wanting to not abstract myself away from software engineering that much :)
Sometimes, I do ask Codex to read my ~/.claude/; and vice-versa. But, generally, I try to keep as much knowledge (e.g. investigations, reports, deep dives) inside the git tree as possible; so that is not necessary.
I don't use skills, but I do have ~/.codex/AGENTS.md and ~/.claude/CLAUDE.md. These are high-level instructions for what (1) I consider readable, maintainable code and patterns, and (2) workarounds for empirically observed model behavior IOdon't like, such as Astra being a bit of a "over-correct over-validation nit-picker". I keep these human-authored, and update regularly based on what I find annoying.
I do all of this before I submit a PR; but of course, for trivial stuff (e.g. CSS changes, copy/string changes, etc), I don't bother.
My practices and workflows do change over time. Back in the ~Opus 4.5 days I'd often define a rubric/criteria in a markdown file, iterate with AI to improve it, and that's the "spec". I've stopped doing that since GPT 5.6; partly because models have gotten a lot better at understanding high level intent from the context; and partly because nearly all models these days feel 'gradermaxxed' when working like that.
Finally, consumer $100/$200mo subs get you _so_ far, I get a lot of value from both. I used to have multiple Claude subs for a while, but trying to 'get full value' made me work on projects just for the sake of it; so 2x$200/mo is my cap :)
netdevphoenix · · focus · HN ↗
dannyw · · focus · HN ↗
Other than trivial PRs; everything I do with Opus/Fable gets reviewed by Astra; and everything I do with Astra gets reviewed by Opus/Fable.
Using only models from a single vendor, is like testing your website/webapp only on Chrome.
dingaling911 · · focus · HN ↗
It's just feelings.
dannyw · · focus · HN ↗
dingaling911 · · focus · HN ↗
pruzicka · · focus · HN ↗
Could you share a bit more about how you do it?
dannyw · · focus · HN ↗
After claude or codex finishes a commit, I switch console tabs and ask the other to review it; with the commit ID. I often find it helpful to inject a bit of human knowledge, and callout any areas of attention I see from a quick skim. (But that could just be me wanting to not abstract myself away from software engineering that much :)
Sometimes, I do ask Codex to read my ~/.claude/; and vice-versa. But, generally, I try to keep as much knowledge (e.g. investigations, reports, deep dives) inside the git tree as possible; so that is not necessary.
I don't use skills, but I do have ~/.codex/AGENTS.md and ~/.claude/CLAUDE.md. These are high-level instructions for what (1) I consider readable, maintainable code and patterns, and (2) workarounds for empirically observed model behavior IOdon't like, such as Astra being a bit of a "over-correct over-validation nit-picker". I keep these human-authored, and update regularly based on what I find annoying.
I do all of this before I submit a PR; but of course, for trivial stuff (e.g. CSS changes, copy/string changes, etc), I don't bother.
My practices and workflows do change over time. Back in the ~Opus 4.5 days I'd often define a rubric/criteria in a markdown file, iterate with AI to improve it, and that's the "spec". I've stopped doing that since GPT 5.6; partly because models have gotten a lot better at understanding high level intent from the context; and partly because nearly all models these days feel 'gradermaxxed' when working like that.
Finally, consumer $100/$200mo subs get you _so_ far, I get a lot of value from both. I used to have multiple Claude subs for a while, but trying to 'get full value' made me work on projects just for the sake of it; so 2x$200/mo is my cap :)
pruzicka · · focus · HN ↗