> Planning improves success at additional cost for weaker models but mainly reduces cost, with small decreases in success rate, for stronger models.
> Predefined tools raise success rates for models with weak bash control, whereas bash-only yields higher success at lower cost for bash-capable models, most clearly on shell-centric task types.
> context management extends execution trajectories without substantially altering agent behavior and is most beneficial under tight context budgets
> planning sustains the trajectories of models that abandon tasks too early and trims repeated verification in models that verify too long
> structured tools support models with limited shell proficiency, while bash-only enables capable models to combine multiple code modifications in a single tool call
Seems fairly intuitive to me, based on feeling. But also fairly kind of obvious; bash-only tooling has higher success for bash-capable models, compared to using predefined tools for models that aren't good at bash? Yeah... They all seem a bit "duh" to me. The final piece of the conclusion is agreeable regardless of how they arrived at it though:
> Harness design is thus a conditional systems problem in which each component should be selected for the target model, task type, and resource budget rather than adopted as a default.
I think lots of people treat the harness/model/prompts combo as interchangeable, but in my experience the quality and efficiently depends heavily on the combo of the harness/model, and using the harness + model made by the same lab, has vastly better experience compared to more "general purpose" (for the lack of a better term) harnesses. Most likely because they use their own traces when training future model iterations.
In other words, MCP was just a bunch of bullshit that maybe helped a little bit until the models got good at bash, and now it's basically useless.
bash scripts, famously the last word in software engineering. all these castles of sand we've built atop the beautiful, perfect, timeless Bourne Again SHell. all for naught. fools!
What an intensely wrong comment. The trajectory of all software is the opposite of what you say. Bash is the entry level, everything flees.
At one point the web was bash scripts glued together. Now everything is brought into a runtime. Then brought into virtualized containers to isolate and hide from every other aspect of the operating system.
Eventually the agents will be on a runtime as well. The existing ones just aren't good enough. Rather than forking like made as a way of doing everything, exfiltration into arbitrary executables will be something more tightly controlled.
Organizations need better in-product control and auditing of the operations of the agent. They need it to work across OSes and not depend on the state of the machine, and not conflict with what else the developer is trying to do with it.
embedding-shape · · focus · HN ↗
> Planning improves success at additional cost for weaker models but mainly reduces cost, with small decreases in success rate, for stronger models.
> Predefined tools raise success rates for models with weak bash control, whereas bash-only yields higher success at lower cost for bash-capable models, most clearly on shell-centric task types.
> context management extends execution trajectories without substantially altering agent behavior and is most beneficial under tight context budgets
> planning sustains the trajectories of models that abandon tasks too early and trims repeated verification in models that verify too long
> structured tools support models with limited shell proficiency, while bash-only enables capable models to combine multiple code modifications in a single tool call
Seems fairly intuitive to me, based on feeling. But also fairly kind of obvious; bash-only tooling has higher success for bash-capable models, compared to using predefined tools for models that aren't good at bash? Yeah... They all seem a bit "duh" to me. The final piece of the conclusion is agreeable regardless of how they arrived at it though:
> Harness design is thus a conditional systems problem in which each component should be selected for the target model, task type, and resource budget rather than adopted as a default.
I think lots of people treat the harness/model/prompts combo as interchangeable, but in my experience the quality and efficiently depends heavily on the combo of the harness/model, and using the harness + model made by the same lab, has vastly better experience compared to more "general purpose" (for the lack of a better term) harnesses. Most likely because they use their own traces when training future model iterations.
jimbokun · · focus · HN ↗
No. The conclusion is that:
bash-capable models + bash-only tools > bash-capable models + predefined tools
In other words, MCP was just a bunch of bullshit that maybe helped a little bit until the models got good at bash, and now it's basically useless.
themgt · · focus · HN ↗
zzbzq · · focus · HN ↗
At one point the web was bash scripts glued together. Now everything is brought into a runtime. Then brought into virtualized containers to isolate and hide from every other aspect of the operating system.
Eventually the agents will be on a runtime as well. The existing ones just aren't good enough. Rather than forking like made as a way of doing everything, exfiltration into arbitrary executables will be something more tightly controlled.
Organizations need better in-product control and auditing of the operations of the agent. They need it to work across OSes and not depend on the state of the machine, and not conflict with what else the developer is trying to do with it.