I think whichever one is used, there needs to be a way to enforce what's written.
If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json. Yet that constantly happens to me anyway.
There is a command in oh-my-pi called "/omfg <problem>". You explain what is wrong with agent's response, and it writes a hook to make sure that the problem doesn't happen again. It then re-runs your previous prompt to make sure that hook is triggered, and if not, it rewrites the hook to make your previous prompt trigger the hook. Then each next agent's response is checked by the hook, and if it is triggered, the agent receives feedback on what's wrong and what must be done differently.
Hooks should (in my opinion) be deterministic. As an example I’ve also noticed Claude writing Python scripts to extract fields from JSON when jq is available to it. They almost always contain the same patterns, and having seen this I’m going to write a hook which triggers on those and fails the turn telling it to use jq instead.
Because custom Python scripts means burning more tokens which costs more money. Reaching for an existing tool like `jq` means burning a LOT fewer tokens.
Imagine you have application code with a function to parse some json. But every time you decide to use it, you re-implement the entire function. This is despite it working the same way every time.
It's not like it's being rewritten for efficiency. "Just because".
And on top it's far more likely to be faulty than a well-designed tool specifically crafted for the job.
Like it may issue an incorrect jq invocation but that's surely cheaper to fix than rewriting the python script. And then it's a simple problem, but let's say you clearly see that "manually editing git files instead of using git" is worse, and most tasks are somewhere in-between.
I think it's kind of cute the way it writes Python scripts, but I've never seen it do that when the relevant native tool is on the PATH. It's like the most competent ever intern, on speed. No tool to convert SVG to PNG? No problem, I'll write a Python program to do that!
> I think it's kind of cute the way it writes Python scripts, but I've never seen it do that when the relevant native tool is on the PATH.
You haven't been paying attention then. I routinely see Claude and GPT models churning out python code to do stupid things like linting. Last week I even had a TypeScript project with prettier configured all over the place, including in a custom agent skill I added with the express purpose of getting the damned model to lint the code, being constantly prompted to run ad-hoc python code supposedly to format whitespaces. I even explicitly prompted one session to just use npm run lint, where I pointed out the exact line of code where prettier was invoked, and the session still churned python code to hande whitespaces.
My experience is the same, I can put that decision in the prompt, in a skill, in a hook, in agent.md, etc, it doesn't matter, after few iterations it starts again using useless python scripts to do anything from linting, to error checking, to parsing one liners of code, etc etc...
Treat it like any other software system: rules that must not be violated are enforced by static type-checking or a trusted runtime monitor. There’s no other option.
You absolutely can, that’s what your harness is for. You don’t need your environment to “reason” about things when deterministic tools exist - You have a really fancy hammer, but that doesn’t make everything a nail.
But what if we used the fancy hammer to change the shape of everything to be a nail? And what if we build the handle of the fancy hammer with a fancy hammer? With all of this we could build a very good fancy hammer manufacturing company.
To offer a possible example: What would the game Zork™ look like with an LLM? Assume we do not want to let players sweet-talk the system into letting them teleport to the end.
The LLM's job would be to channel "I perambulate in the direction of the Arctic circle" into go(north). You saved writing the grammar parser, but you still need to write the game world.
Yet the hammer seller continues to scream everything is a nail and their hammer will replace your entire job eventually. So are you telling me the hammer seller is lying or am I the one using it wrong?
> If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json.
I think there is a deeper problem emerging from this sort of behavior. Even when we bother to create agent skills with there own scripts that call tools like jq a specific way to achieve a goal, AI coding assistants and agents still go way out of their way to generate ad-hoc scripts to do the most absurdly stupid tasks such as parsing output in structured language, and even remove whitespaces from a markdown file. This means AI coding assistants and coding agents treat agent skills as mere suggestions of using a alternative option that more often than not the choose to ignore.
This has a very dangerous implication: your average user is trained to develop a pavlovian reflex to authorize agents to just execute their ad-hoc scripting code with our own permissions and credentials in our systems, which includes the ability to call anything over the internet.
You could try <a href="https://github.com/ioni-dev/mati" rel="nofollow">https://github.com/ioni-dev/mati the constant changes in reasoning effort in consumer models can break things and its more dangerous for devs that relay to much on agents.
spike021 · · focus · HN ↗
If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json. Yet that constantly happens to me anyway.
ACCount39 · · focus · HN ↗
ozim · · focus · HN ↗
gojogs · · focus · HN ↗
zahrevsky · · focus · HN ↗
aliasxneo · · focus · HN ↗
gf000 · · focus · HN ↗
jon-wood · · focus · HN ↗
PcChip · · focus · HN ↗
rmunn · · focus · HN ↗
spike021 · · focus · HN ↗
It's not like it's being rewritten for efficiency. "Just because".
gf000 · · focus · HN ↗
Like it may issue an incorrect jq invocation but that's surely cheaper to fix than rewriting the python script. And then it's a simple problem, but let's say you clearly see that "manually editing git files instead of using git" is worse, and most tasks are somewhere in-between.
amelius · · focus · HN ↗
girvo · · focus · HN ↗
dboreham · · focus · HN ↗
locknitpicker · · focus · HN ↗
You haven't been paying attention then. I routinely see Claude and GPT models churning out python code to do stupid things like linting. Last week I even had a TypeScript project with prettier configured all over the place, including in a custom agent skill I added with the express purpose of getting the damned model to lint the code, being constantly prompted to run ad-hoc python code supposedly to format whitespaces. I even explicitly prompted one session to just use npm run lint, where I pointed out the exact line of code where prettier was invoked, and the session still churned python code to hande whitespaces.
yulaow · · focus · HN ↗
jkhdigital · · focus · HN ↗
koolba · · focus · HN ↗
Otherwise you can get a python one liner that execs a different script engine.
devmor · · focus · HN ↗
TZubiri · · focus · HN ↗
Terr_ · · focus · HN ↗
The LLM's job would be to channel "I perambulate in the direction of the Arctic circle" into go(north). You saved writing the grammar parser, but you still need to write the game world.
altmanaltman · · focus · HN ↗
user_of_the_wek · · focus · HN ↗
sick_of_slop · · focus · HN ↗
[dead]
locknitpicker · · focus · HN ↗
I think there is a deeper problem emerging from this sort of behavior. Even when we bother to create agent skills with there own scripts that call tools like jq a specific way to achieve a goal, AI coding assistants and agents still go way out of their way to generate ad-hoc scripts to do the most absurdly stupid tasks such as parsing output in structured language, and even remove whitespaces from a markdown file. This means AI coding assistants and coding agents treat agent skills as mere suggestions of using a alternative option that more often than not the choose to ignore.
This has a very dangerous implication: your average user is trained to develop a pavlovian reflex to authorize agents to just execute their ad-hoc scripting code with our own permissions and credentials in our systems, which includes the ability to call anything over the internet.
rojaneerdev · · focus · HN ↗
[dead]
ozim · · focus · HN ↗
<a href="https://gist.github.com/cynthiateeters/6868ca26c059a3106cd93736d33015af" rel="nofollow">https://gist.github.com/cynthiateeters/6868ca26c059a3106cd93...
HisashiSpace · · focus · HN ↗
[dead]
fennect · · focus · HN ↗
I’m the author of the project.
triyambakam · · focus · HN ↗
joquarky · · focus · HN ↗
eivindmeyer · · focus · HN ↗
[dead]