Show HN: Open-source model routing for coding agents at Astra-level performance
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Show HN: Open-source model routing for coding agents at Astra-level performance
Unofficial Hacker News client; not affiliated with Y Combinator.
devmor · · focus · HN ↗
1) Claude Code's "advisor mode" (nominally, Sonnet 5.5 with a Fable advisor)
2) Copilot's "HydraFusion" model router/advisor combo.
Specifically I would like to see them compared on architecture planning (both human assisted and hands-off with a draft document) and code review, as these are what I have found the most significant improvement on with multi-model systems.
> our initial hypothesis that an ensemble of models can do better than any single model ever could.
To your hypothesis, anecdotally I find both of these offerings to be far superior to any single model for most tasks of any real complexity, and both to have general frustration/failure cases that single models do not. I would not be surprised that any ensemble approach that utilizes more than one single model meets this hypothesis.
Quite frequently I delegate review and restructuring loops to subagents acting as judges/advisors to tell the primary agent if it met the goal it was instructed to. For some workloads, I will even vet every tool call and user-facing output this way.
adchurch · · focus · HN ↗
Also I’m very interested in the unique failure cases you’re referring to! What have you noticed?
devmor · · focus · HN ↗
> Also I’m very interested in the unique failure cases you’re referring to! What have you noticed?
Permissions issues would be the most common - models collaborating with eachother on a task seem to try to convince eachother they have either more or less permission to do things than they actually do. Especially when transitioning between creating a plan and executing the plan.
Another is deciding that there is a limit to the “loops” they are allowed to run to iterate on something. In many cases I have set an explicit goal, and come back to an agent stopped and reporting that it has hit the “maximum allowable loops of [insert arbitrary number that changes every time].”
Now that I think about it further, I believe the other examples I have also all fall into the models hallucinating the presence of control instructions, or attempting repeatedly to violate permission boundaries that a single model’s harness instructions would usually guide it away from re-attempting.
adchurch · · focus · HN ↗
Fair enough but what exactly were you thinking when you said:
> I would like to see this (and any model router, frankly) benchmarked against two things
And thanks for sharing about failure cases! Those do sound like strange harness-level things, honestly we haven't seen failure cases like that crop up in our own usage & testing.