Show HN: Open-source model routing for coding agents at Astra-level performance
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Show HN: Open-source model routing for coding agents at Astra-level performance
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
xms17189 · · focus · HN ↗
[dead]
redrove · · focus · HN ↗
adchurch · · focus · HN ↗
itsmeduncan · · focus · HN ↗
[dead]
thefourthchime · · focus · HN ↗
aschla · · focus · HN ↗
Seems those are the direct competitors to this, and would be good to know whether it’s worthwhile to stand up a separate solution over those.
adchurch · · focus · HN ↗
adchurch · · focus · HN ↗
Conceptually very similar to Cursor's auto mode. The key distinctions are:
- We plug into any harness (e.g. Claude Code, Codex, OpenCode, Pi) - We aren't incentivized to route to our own model, we're incentivized to route to the best model whatever it may be
jamesforestwest · · focus · HN ↗
adchurch · · focus · HN ↗
The latter is an interesting question! In practice because the session continues, we can see it's going down the wrong path and escalate. Basically no decision we make when routing is entirely unsalvageable (but we do have a performance penalty for every incorrect decision we make so of course we try to avoid it).
aminsamir45 · · focus · HN ↗
adurnos · · focus · HN ↗
[dead]
1minusp · · focus · HN ↗
adchurch · · focus · HN ↗
rirze · · focus · HN ↗
adchurch · · focus · HN ↗
OpenRouter isn't strictly necessary but it does make it easier to not have to set up accounts/API keys with several different providers to get started.
ajspig1 · · focus · HN ↗
And for both opensource and closed source, does the router account for provider quality, or catch it when a provider degrades?
adchurch · · focus · HN ↗
Yes for both! We have some logic to put providers on cooldowns and/or deprioritize them.
YuechenLi · · focus · HN ↗
conception · · focus · HN ↗
YuechenLi · · focus · HN ↗
waximabbax · · focus · HN ↗
conception · · focus · HN ↗
gitowiec · · focus · HN ↗
adchurch · · focus · HN ↗
svnt · · focus · HN ↗
adchurch · · focus · HN ↗
svnt · · focus · HN ↗
adchurch · · focus · HN ↗
mrkn1 · · focus · HN ↗
[dead]
361994752 · · focus · HN ↗
adchurch · · focus · HN ↗
361994752 · · focus · HN ↗
ryanhecht · · focus · HN ↗
creativeSlumber · · focus · HN ↗
creativeSlumber · · focus · HN ↗
Is this correct? in practice you would only be deciding what the next turn is going to be. Because the turn after the next turn is determined by the outcome of the next turn. So shouldn't this number be 10 * 100?
rlprlprlp · · focus · HN ↗
adchurch · · focus · HN ↗
dackdel · · focus · HN ↗
adchurch · · focus · HN ↗
I’d argue it’s similar to looking at the Meta algorithm for serving ads and asking “what about astra” - it could come up with an answer but it’s not the right model for the task.
leej111 · · focus · HN ↗
[dead]
leanuard · · focus · HN ↗
[dead]
devmor · · focus · HN ↗
1) Claude Code's "advisor mode" (nominally, Sonnet 5.5 with a Fable advisor)
2) Copilot's "HydraFusion" model router/advisor combo.
Specifically I would like to see them compared on architecture planning (both human assisted and hands-off with a draft document) and code review, as these are what I have found the most significant improvement on with multi-model systems.
> our initial hypothesis that an ensemble of models can do better than any single model ever could.
To your hypothesis, anecdotally I find both of these offerings to be far superior to any single model for most tasks of any real complexity, and both to have general frustration/failure cases that single models do not. I would not be surprised that any ensemble approach that utilizes more than one single model meets this hypothesis.
adchurch · · focus · HN ↗
Also I’m very interested in the unique failure cases you’re referring to! What have you noticed?
devmor · · focus · HN ↗
> Also I’m very interested in the unique failure cases you’re referring to! What have you noticed?
Permissions issues would be the most common - models collaborating with eachother on a task seem to try to convince eachother they have either more or less permission to do things than they actually do. Especially when transitioning between creating a plan and executing the plan.
Another is deciding that there is a limit to the “loops” they are allowed to run to iterate on something. In many cases I have set an explicit goal, and come back to an agent stopped and reporting that it has hit the “maximum allowable loops of [insert arbitrary number that changes every time].”
Now that I think about it further, I believe the other examples I have also all fall into the models hallucinating the presence of control instructions, or attempting repeatedly to violate permission boundaries that a single model’s harness instructions would usually guide it away from re-attempting.
adchurch · · focus · HN ↗
Fair enough but what exactly were you thinking when you said:
> I would like to see this (and any model router, frankly) benchmarked against two things
And thanks for sharing about failure cases! Those do sound like strange harness-level things, honestly we haven't seen failure cases like that crop up in our own usage & testing.
parisbs · · focus · HN ↗
krzychoo87 · · focus · HN ↗
[dead]
capocasa · · focus · HN ↗
Interested in adding support to my own 3code- <a href="https://3code.capocasa.dev-" rel="nofollow">https://3code.capocasa.dev- which aims to reduce cost by using more compact system prompts and a clever compaction variant.
ipvolt · · focus · HN ↗
[dead]
knifelemon · · focus · HN ↗
[dead]
Igsta · · focus · HN ↗
[dead]
folayii · · focus · HN ↗
[dead]
shurshilov · · focus · HN ↗
[dead]
tarekabouzeid · · focus · HN ↗