I assume the bigger problem is the cache expiring while you wait for the tool call to complete. A bunch of hosted providers only keep the cache alive for ~5 minutes - if you sit there waiting for a 5 minute tool call, you get to pay to reload the entire context into cache
nylonstrung · · focus · HN ↗
It says that it removes tokens wasted while a model is waiting on synchronous tool calls. What tokens exactly is Pi using when waiting?
swiftcoder · · focus · HN ↗