I think whichever one is used, there needs to be a way to enforce what's written.
If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json. Yet that constantly happens to me anyway.
Treat it like any other software system: rules that must not be violated are enforced by static type-checking or a trusted runtime monitor. There’s no other option.
You absolutely can, that’s what your harness is for. You don’t need your environment to “reason” about things when deterministic tools exist - You have a really fancy hammer, but that doesn’t make everything a nail.
To offer a possible example: What would the game Zork™ look like with an LLM? Assume we do not want to let players sweet-talk the system into letting them teleport to the end.
The LLM's job would be to channel "I perambulate in the direction of the Arctic circle" into go(north). You saved writing the grammar parser, but you still need to write the game world.
spike021 · · focus · HN ↗
If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json. Yet that constantly happens to me anyway.
jkhdigital · · focus · HN ↗
koolba · · focus · HN ↗
Otherwise you can get a python one liner that execs a different script engine.
devmor · · focus · HN ↗
Terr_ · · focus · HN ↗
The LLM's job would be to channel "I perambulate in the direction of the Arctic circle" into go(north). You saved writing the grammar parser, but you still need to write the game world.