‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. heyjstn · · focus · HN ↗
    Have anyone tried a workflow that:

    - Fable 5.1 for planning/adversarial reviewer

    - Opus 5.5 for well-scoped tasks break down

    - Sonnet 5.5 for these well-scoped tasks implementation

    I think the blocker might be how efficient the context is compacted and sending around between these agents

    1. chrismustcode · · focus · HN ↗
      You might as well use Opus for everything there.

      Changing model would be cache busting spiking usage for no good reason when Opus can do it all.

      Haiku 5.5 might fit well though depending on pricing.

      1. SirMadam · · focus · HN ↗
        Do subagents share context? If Opus delegates to a different Sonnet window, I don't believe this busts cache?
        1. enraged_camel · · focus · HN ↗
          Subagents don't share context. But that's why delegating implementation to a subagent doesn't work well except for things that are truly mechanical in nature: the subagent needs to independently reason about the task it is given, and then the output will also be reasoned about by the main agent. So you end up wasting time and tokens.
          1. mnicky · · focus · HN ↗
            On the contrary, subagents save context overall, when the task is sufficiently large.

            Also, my experience is that Fable 5.1 is very good at prompting/orchestrating Opus/Sonnet subagents when working on a larger task (e.g. 1-2M context window use only for the orchestrator itself).

          2. esafak · · focus · HN ↗
            If you use subagents your main agent won't need to compact as often, with the loss of information that entails.
        2. manquer · · focus · HN ↗
          Context needs to pre-filled into a GPU memory in a node (usually 8xB300 or 8xH200) so there isn't any context or cache sharing between model families given their different parameter sizes, tokenizers, unlikely they are co-located in the same node.

          Sub-agents not sharing context is a useful design-pattern when you want adversarial or independent reviews.

          Cache reads could be shared between sub-agents, A single node(8GPU cluster) supports few hundred concurrent user sessions, that all share the same KV cache memory, so it is likely model providers do colocate your sub-agents in one node, it is more efficient , but may not be guaranteed so performance could vary; like we have with elastic compute and storage[1]

          This can be cheaper depending on your coding flow i.e. cache hit % and the billing plan - cache reads are basically free or charged very little in subscription plans.

          [1] Modern AWS does offer collocation at additional costs for compute but that is not the default and most other clouds do not offer it

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.