I am increasingly hesitant to use non-native harnesses - model providers are now starting to train their agents for use within the harness. An eval like terminal bench can only capture so much data. I don't want to have to assess each harness every model release to make sure it's working as well as it can.
Is your main concern privacy? I've been using the zcode of and on for a few months now and it's definitely improved over what it was back in the spring. I also use DeepSeek's harness. Not sure which I prefer at this point. Used to be I preferred DSH, but zcode has some features I like over DSH.
johnmlussier · · focus · HN ↗
crossroadsguy · · focus · HN ↗
UncleOxidant · · focus · HN ↗
mongrelion · · focus · HN ↗
UncleOxidant · · focus · HN ↗