‹ BackHN Continuity

Thread

Show HN: Lossless-memory – a personal AI memory that never summarizes

67 points · 34 comments · aru-labs

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. StilesCrisis · · focus · HN ↗
    Is the date really useful? Usually I want the AI to remember rules, like "always do X before committing Y" and a timestamp doesn't change anything there.
    1. drsopp · · focus · HN ↗
      I can't resist the opportunity to paste some of the text I wrote in the Anthropic AI interview last December on an idea I had about adding temporal awareness to Claude.

      Anthropic: Let's focus on AI with capabilities that feel within reach—not necessarily limited to exactly what exists today, but the kind of AI tools you can reasonably imagine existing in the near future. What would you want AI to help you with?

      Me: There is something fundamental missing in today's capabilities that is well within reach of the tech today. I thought I would make a Github repo with a prototype but too much in my life right now. If a human in Anthropic reads this, please let me know if you run with this idea: The missing thing is lack of the bot's initiative. It is always the human who takes initiative and writes something, then the bot makes a (perhaps) elaborate response and shuts up. So something seemingly very simple like telling the bot: "Write me a limerick about a swedish chef, wait 50 seconds, then write a recipe for a danish christmas dinner not involving duck". The bot will not be able to do this. It has no concept of time. No ability to take initiative. There are many ways to frame this but you get the point. I am confident there must be a way to implement this. I read some research on this but all articles I could find are (in my opinion) unnecessarily complicated. The current chat style is Query, Response. The bot waits forever for the next input. This is not natural. When you (A) chat with customer support (B) (a real human), the context is such that it is natural for both A and B to give many inputs in a row before the other party gives their input. And, crucially, if either A or B takes too long to give input, the other party will react to that. This applies to most conversations between humans. So a way to get this behavior from a bot is as follows: the moment either the human or the bot gives any input, the chat system tries to understand the current context, and immediately decides how long time it is natural to wait until the bot should say something by its own initiative, and what it is that it should say. Then the system has a simple timer in the background that waits this amount of time and then says the thing. Very simple. Would it work? I have no idea, but it is worth a shot. This should probably be a separate mode, Proactive Messaging. it could be linked to an agent dynamically. So for example, i have a conversation in Proactive Messaging asking claude to follow the lowest price of something and let me now if it goes below a certain value. Then a few weeks down the line, i get a popup from claude about this if it happens. There should also be a menu in this case with what tasks like this have been defined.

      1. cassianoleal · · focus · HN ↗
        > So for example, i have a conversation in Proactive Messaging asking claude to follow the lowest price of something and let me now if it goes below a certain value. Then a few weeks down the line, i get a popup from claude about this if it happens. There should also be a menu in this case with what tasks like this have been defined.

        I've been looking for a new cooker. I have a Hermes agent running on VM that does this. I told it to keep an eye on prices. It created a cron job that triggers the scrapes and messages me if anything interesting happened since the last time it checked.

      2. StilesCrisis · · focus · HN ↗
        Claude can definitely wait 50 seconds and then do a second thing. I see it all the time when running tools that never stop but iterate towards a goal. It can either run a bash `sleep 50` or use the ScheduleWakeup tool.
      3. altruios · · focus · HN ↗
        There is a KISS way to do this with our current models:

        use a time string that updates:

        every second (that it is not responding) erase a byte string at the end, then send a new byte string (that represents time from last response): bot either outputs a null terminating token, or if the number gets high enough, instruct/train it to take initiative then.

        1. Lerc · · focus · HN ↗
          That only really gets you a limited sense, certainly better than nothing. Though.

          The approach I would like to see would be more expensive token wise, but I don't think that can be avoided for fully interactive real time AI.

          Add another dimension to position encoding to allow for multiple simultaneous streams. Run a model on each stream with an additional "I must speak" output value. A simple adjudicator model (maybe just softmax would work) decides which token(s) are emitted, each tagged by the channel that emitted them, which in-turn goes to the additional position encoding dimension. Then reinforcement learn them all at once with attention being allowed to look at all streams.

          Then the model is always emitting tokens but while waiting it might just be emitting dum-dee-dum twiddle thumbs, on a channel that doesn't go out to the user.

          It needs a way to focus attention though because you need a much larger potential context.

          Is there any work on attention being limited to looking at a subset of embeddings defined by a parameterisable function because if a model could emit special tokens to change those parameters, it would potentially be able to scan it's own memory. Could be tricky to train but a training process that reduced the number of locations attended to over time might force it to compensate for the loss in scope by focusing.

    2. lamplight77 · · focus · HN ↗

      [dead]

    3. jv22222 · · focus · HN ↗
      Oh cool. I'm working on exactly that here!

      <a href="https:&#x2F;&#x2F;innerloop.works&#x2F;breadcrumb" rel="nofollow">https:&#x2F;&#x2F;innerloop.works&#x2F;breadcrumb

      1. StilesCrisis · · focus · HN ↗
        If you can make an agent reliably follow instructions, you&#x27;ve solved alignment completely. So I&#x27;d be particularly surprised if you can square that circle as a hobby project!
        1. jv22222 · · focus · HN ↗
          Its working pretty well to be honest! here’s some more background :)

          <a href="https:&#x2F;&#x2F;innerloop.works&#x2F;blog&#x2F;introducing-breadcrumb" rel="nofollow">https:&#x2F;&#x2F;innerloop.works&#x2F;blog&#x2F;introducing-breadcrumb

          1. hokapo · · focus · HN ↗
            Does it store all your screen re recordings indefinitely? I&#x27;d imagine it&#x27;s a lot of data. Or does it clean them up somehow and remove redundant or unnecessary ones?
            1. jv22222 · · focus · HN ↗
              It’s not in there yet, but there will be a lot of data management options. It just hasn’t existed long enough yet for it to cause the storage problem for anyone but that’s something I’m going to get too real quick.
      2. skinfaxi · · focus · HN ↗
        Am interested but don&#x27;t want to sign up for discord to try something out.
        1. jv22222 · · focus · HN ↗
          Fair enough I will try to unlock download in next few days.

          I guess I&#x27;ll do a Show HN too but I really wanted a few more users before that but yolo.

    4. aru-labs · · focus · HN ↗

      [dead]

    5. CharlieDigital · · focus · HN ↗
      The specific design of this system uses the raw memories and allows rebuilding the operational memory from the raw memory (the LLL &quot;Left Leg Layer&quot;).

      To do this requires that there is an ordering of which memory came last.

          &quot;always do X before committing Y&quot;
          &quot;always do X before committing Y except after Z&quot;
      
      Which of these is the current state? Without the date, it is not possible to rebuild the operational state of the rule from the raw records.
      1. StilesCrisis · · focus · HN ↗
        When I look at Claude&#x27;s memory markdown it tends to list dates for significant requests already, so it can untangle this sort of thing. And if it can&#x27;t it will just ask directly.
        1. CharlieDigital · · focus · HN ↗
          The point is to build systems that don&#x27;t require the agent to &quot;ask&quot;.

          The human becomes the bottleneck as systems become more agentic.

          OP&#x27;s system is first ingesting and storing the individual messages and then recompiling it into &quot;working memory&quot;.

    6. TimByte · · focus · HN ↗
      A timeline turns contradictory instructions into an orderly sequence of updates. Without it the model just gambles between the old guideline and the new one on every run
  2. alansaber · · focus · HN ↗
    looks like a nifty information retrieval approach. cool. temporal&#x2F;versioned chronology is absolutely useful.
  3. rbansal2 · · focus · HN ↗
    Have you thought about using llm-wiki? It&#x27;s pretty powerful and works well.
    1. demeyer1 · · focus · HN ↗
      I’ve been thinking about going that route.

      I’ve seen some good research where xx% of the time, all LLMs essentially just get lazy and won’t actually look things up.

      Been working on an informational retrieval system similar, but endless Nerfs combined with non deterministic and finicky behaviors make it hard to reliably depend on.

      How’s your luck been with LLm-wiki?

  4. VidiAI · · focus · HN ↗

    [dead]

  5. ls-a · · focus · HN ↗

    [dead]

  6. kia_ilands · · focus · HN ↗

    [dead]

  7. 0xbadcafebee · · focus · HN ↗
    It may seem useful at first, but eventually you&#x27;ll hit a wall where this doesn&#x27;t work, and you have to add another memory technique. Eventually you end up with a complex multi-layered system, because what people want by &quot;memory&quot; is actually 10 different things which all need their own solution.
  8. startup_zombie_ · · focus · HN ↗

    [dead]

  9. theresLand · · focus · HN ↗
    Seems like this would break the cache often. That would increase billing rates with certain providers and, for local models, take a while to generate responses, especially with long-running agentic sessions. Interesting idea though.
    1. MikhailTal · · focus · HN ↗
      No? from a quick skim it doesnt look like it goes into the system prompt everytime, you just use search&#x2F;grep over it. Pretty much most memory approaches relying on a big set of info where agent chooses what to &#x27;recall&#x27;. Its like any other tool
      1. messh · · focus · HN ↗
        If there are no summaries then when context is full messages need to get evicted. If doing so one by one then it would indeed destroy the cache. Of course... maybe the implementation evicted 50% of messages at once, I didnt verify in code
    2. TimByte · · focus · HN ↗
      Placing LLL inside the user message right ahead of the new query keeps the common conversation prefix fully intact
  10. cemaluresin · · focus · HN ↗
    The part I like most is the LLL rule: &quot;the AI reads it; the human writes it.&quot; I ended up at the same place from a different direction. I keep a small index that is loaded whole at the start of every session, and the only way to keep it useful was to cap it hard and make it hold pointers, never content. The moment the model was allowed to write into it, it filled with things that looked important and weren&#x27;t.

    Honest question about &quot;never summarize&quot;: how do you decide what goes into context on a turn when the time range is wide, say &quot;last month, about the budget&quot;? Chronological order is great for a day, but a month of raw lines won&#x27;t fit. Is that where the semantic fallback kicks in, or do you just cut at N results?

  11. schainks · · focus · HN ↗
    How is this different than <a href="https:&#x2F;&#x2F;github.com&#x2F;obra&#x2F;episodic-memory" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;obra&#x2F;episodic-memory?
  12. TimByte · · focus · HN ↗
    As long as the relative time parser is hardcoded for japanese, english sessions are stuck with manual ISO dates. Wiring up dateparser or duckling would take an evening , so leaving that on the roadmap is an odd choice
  13. syndiary · · focus · HN ↗
    I have two questions: how a user can inspect the original source when the AI’s derived memory changes? Is chronology&#x2F;versioning becoming more important than semantic summaries?
  14. firstline-dev · · focus · HN ↗
    Congrats on the launch. The &quot;never summarizes&quot; choice matches what I&#x27;ve seen from the other side: I build a tiny local-only context saver for devs with ADHD, and the failure mode of summaries is that they drop exactly the detail you need to restart — the half-finished function name, the thing you were about to check. Lossy compression is fine for recalling facts, but bad for resuming action. Resumption needs pointers, not prose. Curious: how do you handle stale memories when the codebase changes under them?
  15. tnspacetime · · focus · HN ↗
    This is sort of clever! It seems to me though you can almost design a small dsl and you don’t need ai at all.
  16. amit2403 · · focus · HN ↗
    The thing I did not expect, running a persistent multi agent system for about eight months, is that the summarizing is rarely what kills recall. What kills it is that everything you keep has equal weight forever. I had relationship scores between agents that only ever accumulated, and within a few weeks every pair was pinned at the maximum and the number stopped carrying any information at all. Adding decay and diminishing returns near the ceiling fixed more than any retrieval change did. So I would be curious whether never summarizing eventually produces the same flattening for you, just later and with more tokens.
  17. aru-labs · · focus · HN ↗
    Thank you for your questions. I have posted all the answers in the GitHub FAQ. Please take a look. Best regards. <a href="https:&#x2F;&#x2F;github.com&#x2F;aru-labs&#x2F;lossless-memory&#x2F;blob&#x2F;main&#x2F;docs&#x2F;faq.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;aru-labs&#x2F;lossless-memory&#x2F;blob&#x2F;main&#x2F;docs&#x2F;f...
  18. keparlak · · focus · HN ↗

    [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.