I am increasingly hesitant to use non-native harnesses - model providers are now starting to train their agents for use within the harness. An eval like terminal bench can only capture so much data. I don't want to have to assess each harness every model release to make sure it's working as well as it can.
johnmlussier · · focus · HN ↗
PromptAlo · · focus · HN ↗
[dead]