‹ BackHN Continuity

Thread

OpenDLSS: A Vulkan Reimplementation of Nvidia's DLSS 5 Neural Rendering Network

277 points · 126 comments · sagacity

  1. Lerc · · focus · HN ↗
    I'm rather surprised that it doesn't take the z-buffer as an input. I would have thought that would have provided useful information, it's one of the more useful forms of contolnet.
    1. rcarmo · · focus · HN ↗
      The official one seems to do, as well as other info from the engine (I think remember their mentioning LOD/UV map hints in one of the public demos, or articles, a few months back--or it might have been an Unreal Engine podcast)
      1. strangecasts · · focus · HN ↗
        The technical report suggests it does not use the depth buffer: <a href="https:&#x2F;&#x2F;research.nvidia.com&#x2F;labs&#x2F;adlr&#x2F;files&#x2F;DLSS5_Report.pdf" rel="nofollow">https:&#x2F;&#x2F;research.nvidia.com&#x2F;labs&#x2F;adlr&#x2F;files&#x2F;DLSS5_Report.pdf

        &gt; The inference interface uses the engine-rendered RGB image as a dense, registered observation of visible scene appearance. It provides dense, pixel-aligned evidence for object support, occlusion boundaries, composition, and local material properties; engine motion vectors separately provide temporal correspondence.

        &gt; Existing image generative models commonly rely on text embeddings, exemplar images, or spatial control fields such as depth, edges, segmentation, and pose [...] These conditions are effective for general-purpose generation and editing, but they do not uniquely determine the object identities, materials, visibility relationships, lighting decisions, and pixel-aligned detail contained in an engine-rendered frame. DLSS 5 is therefore conditioned on the rendered frame itself.

        1. mikepurvis · · focus · HN ↗
          Obviously they tried it multiple ways have brought receipts, but nonetheless it seems surprising that it wouldn&#x27;t be of benefit to bring as much of that kind of metadata to the model as possible. You&#x27;d think depth and segmentation in particular would basically just be a straight shortcut without which the model spends its own time and effort re-deriving that stuff.

          I&#x27;d also be interested in how post-processing fits in with this. Like if you&#x27;ve got weather effects, film grain, tone mapping, etc, I would have thought the model would do better working on the image before those processes.

          1. jayd16 · · focus · HN ↗
            Hmm well the depth buffer only has accurate depth for opaque objects (and even that&#x27;s not really true). Things like hair wouldn&#x27;t be in there. Its not ground truth depth.
            1. Lerc · · focus · HN ↗
              It is data, It doesn&#x27;t have to be ground truth Depth. In fact the difference between what optically appears to be at one depth and the depth map depth is in-itself information. You have described a mechanism that the model can use to identify that what it is supposed to be enhancing is more likely to be hair.
              1. awill88 · · focus · HN ↗
                I guess in this example since the model wasn’t trained on that, just like they said.

                If that data is some kind of missing key, go make your own and stop nitpicking someone who made a great point as to why they didn’t want that oh so precious data

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.