Context needs to pre-filled into a GPU memory in a node (usually 8xB300 or 8xH200) so there isn't any context or cache sharing between model families given their different parameter sizes, tokenizers, unlikely they are co-located in the same node.
Sub-agents not sharing context is a useful design-pattern when you want adversarial or independent reviews.
Cache reads could be shared between sub-agents, A single node(8GPU cluster) supports few hundred concurrent user sessions, that all share the same KV cache memory, so it is likely model providers do colocate your sub-agents in one node, it is more efficient , but may not be guaranteed so performance could vary; like we have with elastic compute and storage[1]
This can be cheaper depending on your coding flow i.e. cache hit % and the billing plan - cache reads are basically free or charged very little in subscription plans.
[1] Modern AWS does offer collocation at additional costs for compute but that is not the default and most other clouds do not offer it
heyjstn · · focus · HN ↗
- Fable 5.1 for planning/adversarial reviewer
- Opus 5.5 for well-scoped tasks break down
- Sonnet 5.5 for these well-scoped tasks implementation
I think the blocker might be how efficient the context is compacted and sending around between these agents
chrismustcode · · focus · HN ↗
Changing model would be cache busting spiking usage for no good reason when Opus can do it all.
Haiku 5.5 might fit well though depending on pricing.
SirMadam · · focus · HN ↗
manquer · · focus · HN ↗
Sub-agents not sharing context is a useful design-pattern when you want adversarial or independent reviews.
Cache reads could be shared between sub-agents, A single node(8GPU cluster) supports few hundred concurrent user sessions, that all share the same KV cache memory, so it is likely model providers do colocate your sub-agents in one node, it is more efficient , but may not be guaranteed so performance could vary; like we have with elastic compute and storage[1]
This can be cheaper depending on your coding flow i.e. cache hit % and the billing plan - cache reads are basically free or charged very little in subscription plans.
[1] Modern AWS does offer collocation at additional costs for compute but that is not the default and most other clouds do not offer it