‹ BackHN Continuity

Thread

Building a RAG pipeline for semantic code search

31 points · 6 comments · saikatsg

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. duhhhhh1212 · · focus · HN ↗
    <a href="https:&#x2F;&#x2F;www.pangram.com&#x2F;history&#x2F;c901e80e-9cb7-46e4-bf10-7348afcab3cb?ucc=t8c6BkzVOh8" rel="nofollow">https:&#x2F;&#x2F;www.pangram.com&#x2F;history&#x2F;c901e80e-9cb7-46e4-bf10-7348...

    I don&#x27;t want to say don&#x27;t waste your time since the first half is human written. Questions for the authors: did y&#x27;all just get tired of writing and said &quot;fuck it let&#x27;s have the LLM finish the rest&quot;? Or did one of you use LLM to write the last half and the other used their own words?

    1. verdverm · · focus · HN ↗
      This looks like the only kind of comment you make lately (to a Ai writing detector)

      1. detectors are unreliable

      2. more people don&#x27;t care as long as the content is quality

      3. there is a spectrum of ai-human writing and how people make use of the tools, some good, some NS;NT

      <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47089907">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47089907

      1. duhhhhh1212 · · focus · HN ↗
        &gt; detectors are unreliable

        If you can prove pangram is unreliable I will stop using it. By prove, I mean show their false positives and false negative rates are made up.

        &gt; more people don’t care as long as the content is quality

        That’s your opinion. If you want to read AI generated content then there are wonderful sites that can produce that within seconds.

        1. hactually · · focus · HN ↗
          try feeding it content already in the corpus - niche religious texts are a good one where they&#x27;ve tried to special case out the mainstream things like the King James but missed other ones. They claim it&#x27;s 100% AI.

          Garbage app In my experience.

  2. simianwords · · focus · HN ↗
    Here we go again, the industry largely gave on up RAG. In fact I have hardly seen any case where grep doesn&#x27;t work as well as RAG.
    1. dang · · focus · HN ↗
      Can you please review <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;newsguidelines.html">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;newsguidelines.html and do a better job of staying within those guidelines? A bunch of your recent posts have been breaking them—not egregiously (which is good), but noticeably.
  3. keeda · · focus · HN ↗
    I think something like this would be key to improving the quality of coding agents. A very common issue (maybe the biggest one) observed by many people is that agents often produce a lot of duplicate and redundant code; multiple abstractions, methods, classes, data structures, etc. serving minor variations of the same purpose... sometimes within the same file!

    My theory is that this is due to a kind of &quot;tunnel vision&quot; these models have as they execute on a given task, because engineers new to a company do the same thing until they learn the &quot;lay of the land&quot; and figure out that similar problems have been solved elsewhere.

    In a past job my team owned the internal multi-repo codesearch tool, which was by far the most popular internal tool, and later another team added a similar semantic search capability. This was very exciting, but I left before I could see how well it worked out in real-life.

    Like, you&#x27;d do a keyword search and explore if you need some major piece of functionality that would require significant work, or whenever you encounter an abstraction whose code does not exist in your repo and you want to learn more about it. But when you&#x27;re in the flow and inventing smaller abstractions, like a class or utility method, you don&#x27;t necessarily think to search for it. Worse, even if you did, you could not search for it effectively because something similar may exist with slightly different naming or terminology or a typo that a keyword search would miss. Predictably, at scale you ended up with a dozen different implementations doing the same thing.

    Now however, you could automate this with agents. I suspect these days simply prompting an agent to look for any relevant code to reuse would actually work pretty well. But they would need to store the entire codebase in their context window (if it fits at all) to refer to it all the time, which would burn a ton of tokens AND reduce performance due to a heavily polluted context. Instead, something like this would be invaluable to provide as a tool &#x2F; MCP to the agent so that it could locate relevant, reusable code during its planning phase. (Or maybe a post-codegen linter-like check, which IIRC some people have tried, but why fix when you could prevent?)

    I&#x27;m not sure if the exact method in TFA would work best, though; maybe a pipeline that generates comments&#x2F;docs for each unit of functionality and then semantically indexes those, rather than structure-aware chunks of code itself?

    1. nharziro · · focus · HN ↗
      You should be reviewing code generated by llms for this kind of shit. Also scope your changes. All of this stuff is preventable if you&#x27;re reviewing your code and having your llm review your code. You can setup skills for the review agent to use to look for these things.
  4. kaycebasques · · focus · HN ↗
    Oh, wow. I wasn&#x27;t aware that binary quantization was being applied to embeddings! And the code search related use cases were very helpful for grokking when quantization is OK versus when more precision must be maintained. So they sound bearish on Matryoshka? Post makes it sound like Air Context never reduces dimensions.

    Thanks, authors Please follow through on the rest of the series. I learned a bunch.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.