A coworker was trying to tell me that models perform better in their own agent harnesses. I use both `pi` and `omp` and I'm somewhat skeptical. I understand that the tool calls might be slightly different. But really how much impact on the model itself does the harness have?
I'm no academic, but I have read that harnesses dictate the output more than the models themselves. I'm not educated enough on the topic so I will defer to those smarter than me to chime in.
I think it's often the other way around. The differences in perceived coding productivity that many people attribute to claude vs codex is more the harness than the model it's running (if running comparable classes of models).
What is the harness doing that's not the model.
I've been assuming a harness is basically a set of tools and a TUI for passing text to the model and the model coming back with tool calls and user responses.
Are the tools really that complex and different between harnesses?
I have run into some models that seem to mind. Like for the life of me I couldn't get North Mini Coder to play nice on Pi, which was sad since it seemed good otherwise. It just couldn't grasp the tools.
I found this: <a href="https://arena.ai/blog/coding-agents-harness-tax" rel="nofollow">https://arena.ai/blog/coding-agents-harness-tax
nixpulvis · · focus · HN ↗
fishfasell · · focus · HN ↗
jupp0r · · focus · HN ↗
nixpulvis · · focus · HN ↗
I've been assuming a harness is basically a set of tools and a TUI for passing text to the model and the model coming back with tool calls and user responses.
Are the tools really that complex and different between harnesses?
Pxtl · · focus · HN ↗
nixpulvis · · focus · HN ↗
- read - write - edit - bash
Pxtl · · focus · HN ↗
nixpulvis · · focus · HN ↗
Seems to support my skepticism.
capocasa · · focus · HN ↗
Each family gets their own system prompt, based on the original harness one, but compacted and includes instructions for more brevity.