‹ BackHN Continuity

Thread

OpenDLSS: A Vulkan Reimplementation of Nvidia's DLSS 5 Neural Rendering Network

277 points · 126 comments · sagacity

  1. Lerc · · focus · HN ↗
    I'm rather surprised that it doesn't take the z-buffer as an input. I would have thought that would have provided useful information, it's one of the more useful forms of contolnet.
    1. avaer · · focus · HN ↗
      Even relatively small RGB -> depth models are pretty good. Which kind of implies depth is well encoded in the RGB, and adding depth would not really reduce entropy, while costing bandwidth.
    2. rcarmo · · focus · HN ↗
      The official one seems to do, as well as other info from the engine (I think remember their mentioning LOD/UV map hints in one of the public demos, or articles, a few months back--or it might have been an Unreal Engine podcast)
      1. strangecasts · · focus · HN ↗
        The technical report suggests it does not use the depth buffer: <a href="https:&#x2F;&#x2F;research.nvidia.com&#x2F;labs&#x2F;adlr&#x2F;files&#x2F;DLSS5_Report.pdf" rel="nofollow">https:&#x2F;&#x2F;research.nvidia.com&#x2F;labs&#x2F;adlr&#x2F;files&#x2F;DLSS5_Report.pdf

        &gt; The inference interface uses the engine-rendered RGB image as a dense, registered observation of visible scene appearance. It provides dense, pixel-aligned evidence for object support, occlusion boundaries, composition, and local material properties; engine motion vectors separately provide temporal correspondence.

        &gt; Existing image generative models commonly rely on text embeddings, exemplar images, or spatial control fields such as depth, edges, segmentation, and pose [...] These conditions are effective for general-purpose generation and editing, but they do not uniquely determine the object identities, materials, visibility relationships, lighting decisions, and pixel-aligned detail contained in an engine-rendered frame. DLSS 5 is therefore conditioned on the rendered frame itself.

        1. mikepurvis · · focus · HN ↗
          Obviously they tried it multiple ways have brought receipts, but nonetheless it seems surprising that it wouldn&#x27;t be of benefit to bring as much of that kind of metadata to the model as possible. You&#x27;d think depth and segmentation in particular would basically just be a straight shortcut without which the model spends its own time and effort re-deriving that stuff.

          I&#x27;d also be interested in how post-processing fits in with this. Like if you&#x27;ve got weather effects, film grain, tone mapping, etc, I would have thought the model would do better working on the image before those processes.

          1. strangecasts · · focus · HN ↗
            I think it has more to do with what kind of data they have access to at runtime - IIRC DLSS upscaling has only required the previous frames and motion vectors, so requiring depth buffers would mean it was no longer a &quot;drop-in&quot; replacement

            &gt; I&#x27;d also be interested in how post-processing fits in with this.

            I think screenspace effects like film grain and tonemapping are excluded in the same way UI elements are rendered separately from the game.

            1. Stevvo · · focus · HN ↗
              That doesn&#x27;t make sense; every game has a depth buffer, but not every game has motion vectors.
              1. ChocolateGod · · focus · HN ↗
                Because the depth buffer isn&#x27;t always accurate or accessible.

                WoW goes to lengths to hide its depth buffer.

              2. strangecasts · · focus · HN ↗
                Oh, d&#x27;oh, no you are right, DLSS 2 on does pass the depth buffer to the model <a href="https:&#x2F;&#x2F;developer.download.nvidia.com&#x2F;video&#x2F;gputechconf&#x2F;gtc&#x2F;2020&#x2F;presentations&#x2F;s22698-dlss-image-reconstruction-for-real-time-rendering-with-deep-learning.pdf" rel="nofollow">https:&#x2F;&#x2F;developer.download.nvidia.com&#x2F;video&#x2F;gputechconf&#x2F;gtc&#x2F;... (but the neural rendering does not seem to rely on it?)
          2. jayd16 · · focus · HN ↗
            Hmm well the depth buffer only has accurate depth for opaque objects (and even that&#x27;s not really true). Things like hair wouldn&#x27;t be in there. Its not ground truth depth.
            1. Lerc · · focus · HN ↗
              It is data, It doesn&#x27;t have to be ground truth Depth. In fact the difference between what optically appears to be at one depth and the depth map depth is in-itself information. You have described a mechanism that the model can use to identify that what it is supposed to be enhancing is more likely to be hair.
              1. awill88 · · focus · HN ↗
                I guess in this example since the model wasn’t trained on that, just like they said.

                If that data is some kind of missing key, go make your own and stop nitpicking someone who made a great point as to why they didn’t want that oh so precious data

          3. bcatanzaro · · focus · HN ↗
            Integration cost is very important to DLSS.

            More information is better, but more games is also better.

      2. cubefox · · focus · HN ↗
        The Nvidia video presentation on DLSS 5 says that the model was only trained with various G-buffers as input (including the depth buffer) but during inference, the model only uses the rendered frame. As well as the previous rendered frame reprojected via motion vectors, if I understand correctly, likely to improve temporal stability.
    3. strangecasts · · focus · HN ↗
      I think this is mainly so it can use the existing hooks for DLSS upscaling without requiring changes to the renderer, AMD is working on a comparable method which uses adapter networks to slot normals and material properties from the renderer into the diffusion model: <a href="https:&#x2F;&#x2F;gpuopen.com&#x2F;learn&#x2F;temporally-stable-generative-illumination&#x2F;" rel="nofollow">https:&#x2F;&#x2F;gpuopen.com&#x2F;learn&#x2F;temporally-stable-generative-illum...
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.