I don't really understand how async tool calls translate to token savingsIt says that it removes tokens wasted while a model is waiting on synchronous tool calls. What tokens exactly is Pi using when waiting?
"More tool work per model turn" could reduce the number of cache reads (or even cache misses) and associated cost?
nylonstrung · · focus · HN ↗
It says that it removes tokens wasted while a model is waiting on synchronous tool calls. What tokens exactly is Pi using when waiting?
332451b · · focus · HN ↗