SDF vs. MSDF vs. Slug: GPU Text Rendering
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
SDF vs. MSDF vs. Slug: GPU Text Rendering
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
bel8 · · focus · HN ↗
I use MSDF to render crisp text in my webgl hobby game. Hope to publish it with source code when I get the time.
__m · · focus · HN ↗
[dead]
Keyframe · · focus · HN ↗
dcrazy · · focus · HN ↗
jdanford · · focus · HN ↗
cute_boi · · focus · HN ↗
kevin_thibedeau · · focus · HN ↗
20k · · focus · HN ↗
nurumaik · · focus · HN ↗
JoshTriplett · · focus · HN ↗
poly2it · · focus · HN ↗
I think it's a mix of human and LLM writing.
nuxi · · focus · HN ↗
> Drawing it on a GPU, crisply, at any size, under any 3D transform, while the text changes every frame, is not.
> One small texture, resolution independent within reason, one cheap shader.
English is not my native language, so I may be wrong here.
tokenscoper · · focus · HN ↗
Specifically, these sentence structures: 1. ... so <blah> costs <blah>, not <blah>. 2. And <blah>, no <blah>. 3. ... fixed resolution ahead of time, which quietly ...
mbrumlow · · focus · HN ↗
Maybe what we are seeing is us internet dwellers being exposed to wiring styles of humanity we had not been exposed to ?
furyofantares · · focus · HN ↗
mbrock · · focus · HN ↗
There are tells in almost every sentence but some flagrant ones: “where every technique below makes its trade”, “every size you want crisp is another atlas”, “a bitmap has no idea it is being viewed in perspective”, “the fix that carried the industry for a decade”, “it lies about corners”, “built for artwork that moves”, “what separates them is technique, not resolution”, “text is where Slug earned its name”.
These are Claudisms, cloying mannered prose cliches that often anthropomorphize concepts, with an uncanny and annoying smugness.
The post is an AI generated advertisement for a product.
kristianp · · focus · HN ↗
> A letter is not a picture, it is a set of outlines: closed loops of straight lines and Bezier curves, filled according to a winding rule.
Also the way the comma is used to add an inscrutable end to the sentence: "filled according to a winding rule."
tripzilch · · focus · HN ↗
> reduces antialiased vector paths into unique triangle patches and rasterizes them through a massively parallel pipeline with pixel local storage
this is typical type of LLM sentence that kinda runs under the radar, sounds clever, but if you actually look at it, none of it makes sense
- what are "antialiased vector paths"? - why mention "unique" triangle patches? (surely it wouldn't render duplicate ones) - a "massively parallel pipeline"? like a GPU you mean? - "pixel local storage" sounds like an obnoxious marketing term if it's not defined in the article (which it isn't)
anyway I quickly skimmed the article after that but it was little substance
stuaxo · · focus · HN ↗
[dead]
Ithildin · · focus · HN ↗
tokenscoper · · focus · HN ↗
It does bias me negatively towards the author, and makes me a little sad that people are choosing to not put more thought and effort into their writing, but this is a trend that is here to stay.
Making a big fuss about this hasn't gone anywhere from what I can tell, and it's unlikely that the outcome will be any different going forward.
marssaxman · · focus · HN ↗
XenonofArcticus · · focus · HN ↗
Yes, LLMs were used to help write this article and there are tells. I can smell them too.
However, as several have picked out, it's not a one-shot LLM Leeeeroy-Jenkins writing process. I am a domain expert in these sorts of things, and the source of much of the information, concept and organization of the article.
I wrote a bunch of notes and material. I then used an LLM trained on my own styleguide from my writing to organize, expand and fill in places where it was mostly gathering info from other sources. Then I went through THAT and rewrote a bunch because it doesn't always follow directions and color within the lines. It took several days of iteration to put this together even with an LLM assisting.
So, there are parts that are entirely human. There are parts that I felt were less important for me to write personally that I let the LLM writing stand, with some edits.
It is a HUGE piece of work, and writing it all by hand would have taken more time than I could dedicate to it because I have to make a living too.
I should go back and de-LLM it some more, but legit our time is spent working on Slughorn to make it better. I welcome all suggestions for edits and errata on the article (or on Slughorn itself). Our goal isn't to trash other rendering methods. It's not a competition. The post is intending just to fairly compare all of them so people can decide what technique is best for their use, because Slug/Slughorn ISN'T the right choice always, and there's so many techniques with different tradeoffs that it can be hard to grasp where your own use-case falls and how to decide.
Also, Cubicool, the primary developer of the Slughorn code will probably chime in below to address some of the more code- and feature-centric responses.
(Finally, cheers to Eric L. for releasing the Slughorn patent that made this all possible.)
rezmason · · focus · HN ↗
The project I'm most known for is basically an MSDF shader with a bloom pass. It serves my needs 100%, though if I expand to support arbitrary text, I may reach for Slug.
The one issue with MSDF I want to raise is, it seems everybody uses the same msdfgen texture creation program from Viktor Chlumský's master's thesis 11 years ago. I wish there were other implementations. Who ever heard of a graphics technique that was only ever programmed once, and then used everywhere without substantial iteration? We need to de-XKCD-2347 MSDFs for everyone's sake, including and especially Chlumský.
torginus · · focus · HN ↗
<a href="https://web.archive.org/web/20120505013814/https://www.valvesoftware.com/publications/2007/SIGGRAPH2007_AlphaTestedMagnification.pdf" rel="nofollow">https://web.archive.org/web/20120505013814/https://www.valve...
The pdf details only simple SDF, but in the closing paragraphs, it mentions the weakness of the technique and mentions how it can be solved with multiple SDFs, but the exact technique wasn't showcased.
Which led to people trying to reverse engineering it, and making their implementation, for years (it's Valve after all). Including me. Not sure if what I came up with was exactly MSDF, but certainly there are a lot of implementations out there.
rezmason · · focus · HN ↗
I'd say what you and the others did was research, not implementation. Your results were solutions to the same problem, not variations of the same recipe. What you did is important but separate from what's concerned me.
A recipe of Chlumský's— the msdfgen utility— is widely used, but only has one producer. I'm just saying that that's a liability. Like, imagine if HarfBuzz was the only text shaper, and was maintained by one person.
tripzilch · · focus · HN ↗
while I agree with your general point, since you're asking, the table from Paul Bourke's marching cubes code comes to mind:
<a href="https://paulbourke.net/geometry/polygonise/" rel="nofollow">https://paulbourke.net/geometry/polygonise/
another one is the "fast hash" GLSL pseudo random number generator that everyone copies from Inigo Quilez:
despite those constants clearly being the result of some keyboard-bashing, they are copied far and wide :-)reactordev · · focus · HN ↗
flohofwoe · · focus · HN ↗
via WebGPU backend: <a href="https://floooh.github.io/sokol-webgpu/slug-sapp.html" rel="nofollow">https://floooh.github.io/sokol-webgpu/slug-sapp.html
via WebGL2 backend: <a href="https://floooh.github.io/sokol-html5/slug-sapp.html" rel="nofollow">https://floooh.github.io/sokol-html5/slug-sapp.html
There's quite a bit of helper code plus stb_truetype.h and stb_ds.h under the hood to parse TTF files and crunch the TTF curve data into the runtime format expected by the Slug shader (this stuff should better go into an offline asset pipeline tool):
<a href="https://github.com/floooh/sokol-samples/blob/master/libs/slugutil/slugutil.c" rel="nofollow">https://github.com/floooh/sokol-samples/blob/master/libs/slu...
...the actual text rendering code is also taking a couple of shortcuts, e.g. no kerning, no right-to-left, and also no text shaping.
There's also a new and complete text rendering stack by Mikko Mononen called Skribidi (AFAIK not based on Slug though):
<a href="https://github.com/memononen/Skribidi" rel="nofollow">https://github.com/memononen/Skribidi
...the list of external dependencies is a bit scary for a small self-contained sample though (Harfbuzz, SheenBidi, libunibreak, etc...), but that basically shows that proper international text rendering is really damn hard, even when trying to simplify the code as much as possible.
whizzter · · focus · HN ↗
Question though, you implemented the slug system? The article mentions root eligibility but doesn't expand on it, what's the point of it because slug doesn't seem too "magical" for only doing winding counts?
hncbw02z5a · · focus · HN ↗
mattdesl · · focus · HN ↗
It may be of interest to some game/graphics devs here...
[1] <a href="https://github.com/texel-org/windfoil-algorithm" rel="nofollow">https://github.com/texel-org/windfoil-algorithm
yeoyeo42 · · focus · HN ↗
the reason slug has two bands is precisely because of its anti-aliasing. you having only one implies you do the AA differently, or are less efficient with searches off the main direction.
for readers: the slug shader traces two rays, one horizontally, and one vertically, to find out if a pixel is inside a glyph or outside. one would be enough for pure inside outside, but for anti-aliasing purposes it's also helpful to know how far away you are from the closest edge. but if you have a horizontal ray running parallel to a horizontal glyph edge (not uncommon), the ray glyph intersection will return no value at all. so slug traces two rays and blends between them for anti-aliasing that always works in both cases.
it's still approximate. a true ground truth would just supersample and do many point-in-glyph tests within one real pixel - at edges some of them would be inside, and some of them would be outside, giving you a smooth value to display depending on the shape of the glyph within the edge pixel.
mattdesl · · focus · HN ↗
pavlov · · focus · HN ↗
It's nice that he gave the patent to public domain, but this is not how patents are supposed to work. You can't patent something two years after it was already published.
I'm guessing he actually filed for a patent before publishing, and the article should read: "Lengyel was granted a patent for it in 2019"
It's an AI-written article, so maybe it's not reasonable to expect it to be consistent on this level...
flohofwoe · · focus · HN ↗
<a href="https://terathon.com/blog/decade-slug.html" rel="nofollow">https://terathon.com/blog/decade-slug.html
The timeline basically checks out.
eviks · · focus · HN ↗
elengyel · · focus · HN ↗
eviks · · focus · HN ↗
eviks · · focus · HN ↗
pavlov · · focus · HN ↗
A sibling comment notes that he indeed filed a provisional patent in 2017 before going public. That’s the part that the article muddled. 2019 was when the patent was granted, not filed.
eviks · · focus · HN ↗
> provisional application can be filed up to 12 months following an inventor's public disclosure of the invention.
<a href="https://www.uspto.gov/patents/basics/apply/provisional-application" rel="nofollow">https://www.uspto.gov/patents/basics/apply/provisional-appli...
elengyel · · focus · HN ↗
elengyel · · focus · HN ↗
logdahl · · focus · HN ↗
flohofwoe · · focus · HN ↗
<a href="https://github.com/AlphaPixel/slughorn" rel="nofollow">https://github.com/AlphaPixel/slughorn
cubicool · · focus · HN ↗
XenonofArcticus · · focus · HN ↗
Slughorn is coded by hard-working humans. Sure, we use LLM tools for some investigation and development, but this is not an AI slop project. It's been in the works steadily since not long after Eric announced the patent release, grinding through improvements with every hour of the day we can dedicate to it.
Every important line of code here was written, tested and is understood by Cubicool, the lead programmer (there are some LLM-generated ancillary parts of the codebase, and I think git shows those).
logdahl · · focus · HN ↗
sirwhinesalot · · focus · HN ↗
<a href="https://rookandpossum.com/posts/scanline-sweeper/" rel="nofollow">https://rookandpossum.com/posts/scanline-sweeper/
Sean Barret (creator of the stb public domain libraries) independently invented a CPU-based implementation of the same idea, used in stb_truetype.
cubicool · · focus · HN ↗
I'll be keeping my eye on it, for sure. Like I mentioned elsewhere, slughorn has become WAY more than "just text." Glyph-centric usage is what made Slug popular, but I've been spending my free time lately trying to figure out how to make it work for vector graphics IN GENERAL (to varying degrees of success, based on the amount of time I can dedicate to it...)
seanw265 · · focus · HN ↗
I'm a bit confused because at some points it seems like the author is conflating "tessellation" and "Rive". Are they the same thing? As an uneducated reader, my understanding would be that Rive is an implementation of a renderer using the tessellation approach. But surely a generic tessellation approach could support perfect arbitrary transformations even if Rive doesn't?
Maybe there's something I'm missing.
Also, a nitpick: in the "head to head" section, the author highlights Slug's better performance in green for the entries that it wins. For the sake of fairness, shouldn't we highlight the winners in every category? Surely Rive's "low" memory usage beats Slug's "moderate"?
exDM69 · · focus · HN ↗
Yes, but at typical font sizes this generates a lot of tiny triangles (0 to few pixels) which perfoms badly on GPUs.
GuB-42 · · focus · HN ↗
Slug doesn't seem to support any of these, it is just "for a given point, am I in or am I out?", but it doesn't tell you by how much, which is great on really high resolution displays and large sizes, but you would lose the ability to do the kind of effects you can do with SDFs, and have to deal with antialiasing separately.
JoshTriplett · · focus · HN ↗
That said, there are interesting effects other than grayscale antialiasing, and I wonder how well slug could handle things like outlines (useful for subtitles over video, to make them readable on any background).
eviks · · focus · HN ↗
JoshTriplett · · focus · HN ↗
Not especially, unless you're running something like a high-refresh-rate gaming monitor that's still 1080p or less.
On most current phones, laptops, and monitors, I would expect the difference between grayscale antialiasing and monochrome rendering to be hard to notice.
mncharity · · focus · HN ↗
jasomill · · focus · HN ↗
exDM69 · · focus · HN ↗
It does give you an anti-aliased value between 0 and 1 that estimates how much of a pixel is being covered.
But this is a linear estimate based on horizontal and vertical distance to the Bezier curve. It does not look correct at long distances, which is why you shouldn't use it for outlines, drop shadows or the other cool things you can do with (M)SDF. A single pixel outline works fine but is not really legible with modern display resolutions (very thin lines).
jayd16 · · focus · HN ↗
MSDF is pretty much just the target texel in question plus the surrounding samples in a way the GPU can do entirely upfront before the math starts. (M)SDF glyphs also play nicely with mip mapping and I would think this Slug algorithm needs uncompressed data. Maybe that doesn't matter because you just don't scale the data ever?
flohofwoe · · focus · HN ↗
The pixel shader is indeed quite complex compared to SDF (this is using textures instead of storage buffers, because the example also needs to work on WebGL):
<a href="https://github.com/floooh/sokol-samples/blob/8afa83928ce1870efeb0d513e7c4dce4f5db7b3e/sapp/slug-sapp.glsl#L30-L169" rel="nofollow">https://github.com/floooh/sokol-samples/blob/8afa83928ce1870...
sophietaylor · · focus · HN ↗
[dead]
Const-me · · focus · HN ↗
“Chinese, Japanese, and Korean have tens of thousands of glyphs, and baking all of them at several sizes is a memory disaster” One possible solution is dynamic atlas built on CPU for visible glyphs only.
“The distance-field panels notch, where interpolating between stored samples no longer matches the true curve” Can’t it be fixed in the shader, using screen-space derivatives of the SDF? I think in theory, SDF value for pixel center combined with screen-space gradient vector of that number delivers enough data to compute partial coverage for the pixels on the edge.
exDM69 · · focus · HN ↗
At first glance, subdividing the Bezier curves sounds like a bad idea (more Beziers to rasterize) but it opens doors for some parallelism, and most Bezier curves that appear in fonts are monotonic in the first place (so the increase is very modest). This was inspired by this entertaining but not very serious video about font rasterization [0].
The first parallelism optimization is checking against the curve bounding box vs. a rectangular (in uv-space) region of pixels, and this can quickly determine if the Bezier needs to be evaluated in the first place. This can be done per GPU warp.
The second optimization works only for rectilinear transformation (no rotation, skew or perspective). Solving the quadratic equation involves a square root and a division (which alone are >30% of the computation), which can be computed for each row and column of pixels instead of for each pixel (2n instead of n^2).
Both optimizations rely on mathematical invariants of monotonicity, ie. the derivative of the Bezier curve must be non-zero. All Bezier curves can be robustly subdivided into monotonic sections using de Casteljau's algorithm.
My simple benchmarks compare favorably to Slug on the GPU and to "fast" rasterization algorithms on the CPU (which is an order of magnitude faster than "fancy" rasterization algorithms with hinting etc).
Unfortunately there are so many hobby projects and so little time. All I have is messy shaders that draw individual characters and a few benchmarks to see how quickly (and something similar for the CPU). Going from there to a complete text rendering system would be a lot of work. Writing a more detailed article with illustrative code examples is something I'd want to do but haven't gotten around to.
If you want to offer words of encouragement or geek out about rasterization algorithms, I welcome any input.
[0] <a href="https://www.youtube.com/watch?v=SO83KQuuZvg" rel="nofollow">https://www.youtube.com/watch?v=SO83KQuuZvg Sebastian Lague - Coding Adventures: Rendering text.
LoganDark · · focus · HN ↗
What a name!
Skipped the rest of the article cause it's AI.
XenonofArcticus · · focus · HN ↗
Take a seat over there.
And then go read my reply above about how AI was and was not used to write the article.
I'm looking forward to your competing article that does it better without using AI.
LoganDark · · focus · HN ↗
YuechenLi · · focus · HN ↗
The other thing is MSDF rendering is fairly cheap and can be done on the CPU quite easily without GPU shaders, the atlas generation/upload is the expensive part, but it's a one-time cost per character per font as the atlas is just a normal bitmap texture file, MSDF text have sharp edges at most zoom resolutions, and the small size text is better handled with simple CPU raster anyways.
I really failed to see significant benefit of using Slug over MSDF + raster fallback for small fonts, it's definitely more exact, but I'm not sure if the marginal resolution benefit is worth it over much more complicated GPU dependent rendering, so I'd really want to test it out myself when I have the time over taking the word of an obviously AI written article for it.
decoding · · focus · HN ↗
psyclyx · · focus · HN ↗
When the Slug patent was released to the public domain, I put together[1] Snail[2], a Slug implementation in Zig.
One thing I found was that it was tough to get small text to look good with some fonts. TrueType fonts often have bytecode that tweaks curve points to better fit the pixel grid at a particular size. Part of the pitch for Slug is that it doesn't require per-size glyph prep, so Slug text is just unhinted.
As monitors have gotten denser, the major font renderers have moved away from bytecode hinting, either toward auto-hinting (ignore the bytecode, look at the outline, and decide what to do) or toward no hinting at all.
Snail has GPU auto-hinting that tries to replicate a lot of that, so I could have hinted text without per-size prep. It precomputes knots at various glyph features, and the shader then moves the knots to stretch/squeeze parts of the glyph. It's not perfect (especially for serif fonts!), benefits from per-font tuning, and focuses on Latin glyphs (pretty much always draws CJK glyphs unhinted because they're too complex for the current shaders to handle). Also, cases where you want hinting and wouldn't be better served by just scissoring prepared bitmaps aren't all that common. It also includes a TrueType VM, for cases where the per-size cost is acceptable (DejaVu Mono looks decent with the auto-hinter, but it has phenomenal hinting bytecode).
It was a lot of fun to put together, I learned a lot, and as far as I know the auto-hinter is unique among Slug implementations. Slug is a very cool algorithm.
[1] through high-effort delegation to Claude [2] <a href="https://github.com/psyclyx/snail" rel="nofollow">https://github.com/psyclyx/snail - includes some diagrams (made with Snail!) that explain most of the prep/rendering of a glyph with Slug
elcritch · · focus · HN ↗
Same here. For scaling from medium to large fonts the MSDFs were good. However at smaller "normal" font sizes they just didn't tend to look good even with larger MSDF sizes.
Plus for font rendering the GPU overhead was too large to be worthwhile for normal GUI apps. Font atlas'es are pretty small, and most fonts for GUIs are statically sized.
However I did find MSDF are fantastic for rendering complex SVG paths with lots of benefits [1]! You can use a 64x64 SVG star and scale it up fullscreen with reasonable loss in quality [2] (or use 128x128 for almost lossless images).
I use them as the core of my GUI library to render complex SVG paths for icons and such and compose individual paths on the GPU.
You can use them to implement fair chunks of Lottie animation as well [3]. Though I never got enough into the Lottie work to go beyond basic demos, but you could easily scale MSDFs, add shadows, feathering, outlines, etc at 100+ FPS.
1: <a href="https://forum.nim-lang.org/t/14062#85314" rel="nofollow">https://forum.nim-lang.org/t/14062#85314
2: <a href="https://github.com/elcritch/figdraw#msdf-bitmap-based-sdf-rendering" rel="nofollow">https://github.com/elcritch/figdraw#msdf-bitmap-based-sdf-re...
3: <a href="https://github.com/elcritch/lotty" rel="nofollow">https://github.com/elcritch/lotty
vidarh · · focus · HN ↗
Frankly, I don't bother with hinting at all for my renderer (based off libschrift, which also does no hinting) for that reason. I was essential. It probably still looks marginally better. But the complexity just doesn't feel worth it any more.
musicale · · focus · HN ↗
<a href="https://news.ycombinator.com/item?id=49909610">https://news.ycombinator.com/item?id=49909610
theandrewbailey · · focus · HN ↗
zackmorris · · focus · HN ↗
<a href="https://gamedev.net/tutorials/programming/graphics/s-buffer-faq-r668" rel="nofollow">https://gamedev.net/tutorials/programming/graphics/s-buffer-...
<a href="https://en.wikipedia.org/wiki/Scanline_rendering" rel="nofollow">https://en.wikipedia.org/wiki/Scanline_rendering
<a href="https://mcejp.github.io/2021/02/17/s-buffers.html" rel="nofollow">https://mcejp.github.io/2021/02/17/s-buffers.html
It works great except it can have z-fighting issues, especially near t-junctions.
I'm wondering if there's a way to calculate scan lines along 2 or more axes, then choose the ones with the least uncertainty.
A similar issue happens with convex hull algorithms. If we wrap along the x axis, we might get a different triangle mesh than if we wrap along the y or z axis, because vertices may be coplanar along an axis. I've considered wrapping along all 3 axes, then choosing the mesh that matches in the most axes.
Is there a general solution for this? Some point clouds have sets of vertices that are coplanar regardless of which axis we use. I suspect that the number of axes needed might be proportional to the number of vertices. I wonder if we could use something like change of coordinates to wrap the mesh from the frame of reference of each vertex.
If we had that technique, we could apply it to Slug and render glyphs having minimal error, regardless of how they're transformed, if we're willing to sacrifice some performance.
cubicool · · focus · HN ↗
We're still nowhere near the levels of libraries like Noesis or Rive (which are fundamentally different approaches to the same idea), but hey--maybe one day. There's already an impressively large stack of stuff I've been able to pull of with just some clever shader tricks, and I'm coming up with more things I want to try every day.
Right now I'm working on adding completely GPU-driven "path stroking" to slughorn/osgSlug (which uses SSBO to store/update a `slughorn::Path` and then does some `gl_InstanceID` tricks to dynamically "emit" the necessary geometric bounds Slug needs to do its analytical fill/antialiasing), and after that I'll go back to finishing the full Porter-Duff compositing stack.
There are also a handful of UIs from games I want to try "cloning" (Deadspace's diegetic style, the radio scanner from the Arkham games, etc). I also keep this HUGE directory of UIs I see in shows/movies; every time I see some mocked-up thing and think to myself, "I wonder how they did that?", I usually snap a screenshot to poke at later.
Here are some "just for fun" screenshots of interesting stuff I haven't put up on Github yet:
<a href="https://slughorn.io/funsies-00.png" rel="nofollow">https://slughorn.io/funsies-00.png <- which one is 2D!? ): <a href="https://slughorn.io/funsies-01.png" rel="nofollow">https://slughorn.io/funsies-01.png <- decaling slughorn-based "shapes" on 3D meshes; you can zoom until the GPU runs out of float precision
There's a whole slew of cool changes I haven't fully updated the repo with; there just aren't enough hours in the day, and we still gotta pay bills. Anyway, I'll check back later if anyone has any questions/ideas/things they'd like to see... that kinda feedback dictates what I focus on next. :)
icemic · · focus · HN ↗
[dead]