ReBAR stands for "Resizable BAR (Base Address Register)" - it makes the CPU's access window into GPU memory resizable; instead of accessing VRAM in small chunks, the CPU can access much more at once, which can improve performance on some modern GPUs.
i literally don't understand how people move through life - do you expect every single link on this site (or any site) to be ELI5 for you specifically? do you not understand that some things require effort/homework on your part?
What is the purpose of your comments? You failed to answer his question, then attack him for asking the question. Consider the two possibilities: a) doing that homework is easy, b) it's hard. If it's easy, you could have just posted the relevant link and be done. If it's hard, then his question is warranted. Either way, your comments add nothing. One can only conclude that you don't actually know how to answer the question. The real question, though, is: why even bother replying?
Also, ReBAR can get complicated as you can see from other responses above. The articles on the web are also not really good as they tend to replicate marketing material and make vague statements about performance, but fail to explain how specifically performance is improved.
They wanted a basic explanation in response to that.
Or to put it differently, I would take the sentence "He asked for a purpose and a use-case not an explanation." and replace the word "not" with a comma.
BARs are part of the PCI spec, they're the way for firmware and OS to determine the size of the memory or i/o window the card controls as well as the way for the system firmware (usually) to pick what address range the card should use.
In the old days, they were a fixed size. If your GPU has 16 GB of ram and you want to access all of it via memory addresses, you'd need a 16 GB BAR ... but lots of (older) systems wouldn't be able to map in a large BAR because of a lack of address lines (or wahtever), so GPUs had stayed with 256MB for VRAM access because it was compatible. With a smaller BAR than the VRAM, you have to use some sort of windowing / paging setup. Resizable BAR lets the BAR start small so older systems will work, but grow larger with capable systems.
Much better than having a jumper to set the BAR to big or small, and you can skip VRAM window management.
It should come with a warning label. If you're in a scenario where VRAM is being maxed out and thrashing main memory then ReBAR will worsen the frequency and magnitude of the stutters. I ran into this with a 3070 (8 GB) recently.
I think you’re experiencing paging to main memory rather than anything ReBAR related. ReBAR itself reduces memory copy operations. It’s of course possible you’re experiencing some sort of BIOS bug.
It's entirely possible a game manages paging VRAM badly. But allowing the game more flexibility isn't a problem with a larger BAR, it's that the game is stupid and gets dumber the more VRAM you give it.
If a program runs slower when you give it more RAM, the problem isn't giving more RAM.
Resizable Base Address Register. In short, rather than the 256MB of mapped memory available to the CPU for any PCI device, ReBAR capable devices can map larger memory to the CPU’s addressable memory space. Without ReBAR or sufficient tricks (that Intel GPUs lack) you have to shuffle 256MB chunks around between GPU and CPU.
Using an Intel discrete GPU on a system that lacks ReBAR support results in somewhat lower performance, not a complete inability to use more than 256MB of VRAM.
Source? I don't know what happened for the Linux drivers, but the Windows drivers definitely booted and ran benchmarks without ReBAR when the Arc A380 launched, because there were lots of published comparisons of the performance impact on workloads that definitely didn't fit in 256MB.
Unfortunately I can't edit my earlier post, because I can't find it - might have been distorted rumour that reached me, then :/
Wasn't helped by my experience reacting very strongly to story of early Arc drivers being problematic because the devs were too used to GPU being just a ring-bus away ;-)
ReBar's commercial name is AMD Smart Access Memory, it allows a PCIe device such as a GPU to map more VRAM to the system at once, which improves performance by reducing access overhead. This generated much fanfare in the early 2020s after AMD officially supported it in the newly released AMD Zen 3 CPUs with RX6800 series GPUs. It was marketed as a new technology to boost GPU/gaming performance. What AMD did was just rebranding an obscure feature in the PCIe specification [1]. As shown by this project, it was actually supported by the PCIe controller since Sandy Bridge, just disabled in the firmware. For a decade nobody bothered to use it. Presumably, AMD saw an opportunity and enabled it, presumably after validating the hardware and fixing any driver compatibility problems.
[1] This is nothing new in the tech industry. Intel rebrands DVFS as SpeedStep, IOMMU as VT-d, AMD rebrands the NX bit as Enhanced Virus Protection, etc.
I feel part of the "Smart Access Memory" name is actually the driver features and paths to actually make use of it.
It's not like ReBar is a single toggle "Make Things Faster", but a different option in how it can map gpu memory to the cpu. The driver still needs to use it - and decide where it's use vs the "staging buffer" approach would actually be be benefitial.
It's hard to explain without going into a few low-level details of PCI Express, but let me try.
Most PCI devices expose some memory and/or I/O ports to the CPU. That memory (or I/O ports) is mapped to somewhere in the address space visible to the CPU. Besides the memory and I/O ports, all PCI devices also expose a separate set of configuration registers; among these registers, there are the Base Address Registers (BARs), which configure where the memory or I/O ports is mapped.
Here's an example output from "lspci -vv" for a GPU:
Region 0: Memory at 7c00000000 (64-bit, prefetchable) [size=8G]
Region 2: Memory at 7e00000000 (64-bit, prefetchable) [size=256M]
Region 4: I/O ports at f000 [size=256]
Region 5: Memory at fca00000 (32-bit, non-prefetchable) [size=1M]
Expansion ROM at fcb00000 [disabled] [size=128K]
Note that regions 0 and 2 are above the 4GB addressable by old 32-bit CPUs. To be compatible with these old CPUs, this card and many others like it allow the firmware (and/or the operating system) to choose not only where the memory is mapped, but also its size. We can see this in the same "lspci -vv" output for this GPU:
Capabilities: [200 v1] Physical Resizable BAR
BAR 0: current size: 8GB, supported: 256MB 512MB 1GB 2GB 4GB 8GB
BAR 2: current size: 256MB, supported: 2MB 4MB 8MB 16MB 32MB 64MB 128MB 256MB
Older systems which do not understand this extended capability will still treat these regions as fixed size, probably with the first size in this list (256MB for region 0, 2MB for region 2). Newer systems can tell the device to "resize" the BAR to a bigger size, which obviously needs the first region to be placed above the 4GB barrier since it's too big.
Why is this useful? This particular GPU has 8GB of VRAM; it's quite obvious that region 0 is a direct view into that VRAM. When using the maximum BAR size, the CPU can directly read and write anywhere into the VRAM; when using a smaller BAR, the CPU can only see a small window into the VRAM, and has to use less direct methods to access it.
(As an aside: go right now and do a "sudo lspci -vv" on your computer, if you see a Resizable BAR capability which isn't using the maximum size, you can probably gain a bit more speed for free by going into the BIOS and enabling "Resizable BAR" and/or "Above 4G decoding". If you can't find these options, well, AFAIU that's what this project is all about..)
> (As an aside: go right now and do a "sudo lspci -vv" on your computer, if you see a Resizable BAR capability which isn't using the maximum size, you can probably gain a bit more speed for free by going into the BIOS and enabling "Resizable BAR" and/or "Above 4G decoding". If you can't find these options, well, AFAIU that's what this project is all about..)
IIRC Linux doesn't need resizable bar enabled in the BIOS since the kernel will resize the bar if supported by the GPU, Windows however relies on the UEFI doing it which is where it being enabled in the BIOS is needed.
Since about 5 or 6 years ago, PCIe Video Cards advertise a capability known as Resizeable BAR. PCI Devices like GPUs require some memory to use for PCI MMIO, which is directly visible on the CPU Address Space. As a side note, this Address Space is shared with RAM, and anyone that was around when having 4 GiB RAM with a 32 Bit OS was common (Earlier than 2010 or so) knows that you only saw about 3.25 GiB RAM or so because of sharing the Address Space with PCI MMIO, and those fortunate enough to have used SLI usually saw even less than that, like 2.87 GiB RAM.
GPUs has been using a 256 MiB PCI MMIO window regardless of how much VRAM they actually have since... nearly forever? At least since PCIe is a thing, since I recall than AGP Aperture Size was seteable in era accurate BIOSes. PCIe 3.0 specification introduced a feature known as Resizeable BAR, where the PCI Device can tell a compatible Firmware how much MMIO it actually wants. GPUs uses that to tell a ReBAR capable UEFI Firmware that it wants more MMIO (Usually as big as the GPU VRAM), or uses legacy 256 MiB otherwise.
Just to make sure I understand: Is the resizable BAR/MMIO a RAM buffer for PCie packets? Or does it have some deeper integration with DMA or something? I assume it's not like memory mapped peripherals on an AXI bus which is why you need the buffer?
For historic reasons (e.g. 32 bit address spaces, plus the need to reserve the space for multiple pci peripherals) it has been a narrow, movable aperture.
Resizable BAR lets the size of the aperture be chosen (which is usually chosen to allow all of VRAM to fit in and be directly accessible).
No idea on the AXI Bus you're talking about, so can't make comparisons.
MMIO (Memory Mapped I/O) is essentially memory (Whenever RAM or ROM) from OTHER devices that is directly visible on the CPU Address Space. My understanding is that from the CPU side, MMIO is mostly transparent (Except for the massive increase in latency) because it gets used like if it was interacting with its own workspace with regular instructions like MOV.
What PCIe ReBAR changes is that before, you could only see a 256 MiB window onto the GPU VRAM, so there was an added overhead since the GPU may need to relocate things from inside that window somewhere else on its total VRAM (So yes, it may be interpreted as if what you see from the CPU side is just some kind of exchange buffer).
I believe the best way to describe how it operates is comparing it to EMS (Expanded Memory) from the DOS days since it also worked with a similar, if not the same idea. You could only see a portion of the total memory from what was installed on the EMS card (A 128 KiB window located on the upper part of the 1 MiB address space from the 8086 CPU), so you had to switch which Page (Region) of the memory was visible, adding a lot of overhead and most likely requiring an additional buffer in main RAM to move data from one Page to another. However, since I have no knowledge if the GPUs really work like that I can't confirm. I never knew whenever the 256 MiB is "fixed" (You always see the same Region) or if you can decide which section of the VRAM to make visible.
There is no buffer at all. BAR is "base address register" (hint, there is no "the BAR"). Each PCIe device / function can have up to 6 32-bit BARs or 3 64 bit BARs.
When you insert a GPU into a PCIe slot, the memory mapped regions in memory can't be put in a hard wired location because any arbitrary device can be inserted and it can provide an arbitrary amount of memory (yes the GPU provides its own memory to the CPU). A BAR reserves a memory mapped region in the CPU space that is backed by the PCIe device.
When the BAR is smaller than the memory of the inserted device, the CPU cannot communicate with all of the memory on the inserted device directly anmore. This means if you want to perform a write to a region in the GPU outside a BAR region you have to go through the BAR region anyway. It's not a RAM buffer for PCIe packets.
>I assume it's not like memory mapped peripherals on an AXI bus which is why you need the buffer?
GPUs have been using a 256MB window since they started coming with 256MB of VRAM, as it would otherwise be a pointless waste of address space to have a window larger than the actual amount of VRAM. Previously, they would have a window exactly equal to the size of VRAM (likely rounded up to the next power of 2.)
There would also be a block of memory mapped hardware registers also needing address space that you have to poke to make the GPU actually do GPU things, instead of just being an expensive way to add extra memory to a system, no?
And if my experience from embedded development is in any way transferable, they're probably fairly spread out and probably takes a fairly big chunk of address space too.
Yes, typically the memory mapped memory is in one (64-bit) bar, memory mapped registers in another (64-bit) bar, plus for compatability with vga, probably a 32-bit memory bar and an i/o bar. 64-bit bars take up two bar slots, so that fills all six slots in the PCI config.
bigwheels · · focus · HN ↗
(Posting this under the assumption others will also appreciate a bit of quick context.)
mathisfun123 · · focus · HN ↗
Literally the second sentence in the repo:
> This provides performance benefits and is even required for Intel Arc GPUs to function optimally.
qudat · · focus · HN ↗
What is a BAR let alone a resizable one? Readme just jumps in, which is fine, but I’m not sure why this is on HN or why I should care.
pizza234 · · focus · HN ↗
docenttx · · focus · HN ↗
mathisfun123 · · focus · HN ↗
yunnpp · · focus · HN ↗
Also, ReBAR can get complicated as you can see from other responses above. The articles on the web are also not really good as they tend to replicate marketing material and make vague statements about performance, but fail to explain how specifically performance is improved.
mathisfun123 · · focus · HN ↗
Same as his: to voice my frustration about something on the internet.
Dylan16807 · · focus · HN ↗
mathisfun123 · · focus · HN ↗
Dylan16807 · · focus · HN ↗
Wrong.
mathisfun123 · · focus · HN ↗
> I'm still clueless on what the purpose and use-case is for ReBAR
Are you seriously debating this?
Dylan16807 · · focus · HN ↗
Or to put it differently, I would take the sentence "He asked for a purpose and a use-case not an explanation." and replace the word "not" with a comma.
mathisfun123 · · focus · HN ↗
Dylan16807 · · focus · HN ↗
I agree with "He asked for a purpose and a use-case". (This is what your quote supports.)
I disagree with "not an explanation".
When asking for a purpose and use case, they were asking for an explanation.
toast0 · · focus · HN ↗
In the old days, they were a fixed size. If your GPU has 16 GB of ram and you want to access all of it via memory addresses, you'd need a 16 GB BAR ... but lots of (older) systems wouldn't be able to map in a large BAR because of a lack of address lines (or wahtever), so GPUs had stayed with 256MB for VRAM access because it was compatible. With a smaller BAR than the VRAM, you have to use some sort of windowing / paging setup. Resizable BAR lets the BAR start small so older systems will work, but grow larger with capable systems.
Much better than having a jumper to set the BAR to big or small, and you can skip VRAM window management.
wmf · · focus · HN ↗
willis936 · · focus · HN ↗
Moto7451 · · focus · HN ↗
jmalicki · · focus · HN ↗
If a program runs slower when you give it more RAM, the problem isn't giving more RAM.
Moto7451 · · focus · HN ↗
bpye · · focus · HN ↗
Isn't it just being able to shift the window of GPU memory visible to the CPU?
p_l · · focus · HN ↗
wtallis · · focus · HN ↗
p_l · · focus · HN ↗
wtallis · · focus · HN ↗
p_l · · focus · HN ↗
Wasn't helped by my experience reacting very strongly to story of early Arc drivers being problematic because the devs were too used to GPU being just a ring-bus away ;-)
segfaultbuserr · · focus · HN ↗
[1] This is nothing new in the tech industry. Intel rebrands DVFS as SpeedStep, IOMMU as VT-d, AMD rebrands the NX bit as Enhanced Virus Protection, etc.
bpye · · focus · HN ↗
kimixa · · focus · HN ↗
It's not like ReBar is a single toggle "Make Things Faster", but a different option in how it can map gpu memory to the cpu. The driver still needs to use it - and decide where it's use vs the "staging buffer" approach would actually be be benefitial.
account42 · · focus · HN ↗
cesarb · · focus · HN ↗
Most PCI devices expose some memory and/or I/O ports to the CPU. That memory (or I/O ports) is mapped to somewhere in the address space visible to the CPU. Besides the memory and I/O ports, all PCI devices also expose a separate set of configuration registers; among these registers, there are the Base Address Registers (BARs), which configure where the memory or I/O ports is mapped.
Here's an example output from "lspci -vv" for a GPU:
Note that regions 0 and 2 are above the 4GB addressable by old 32-bit CPUs. To be compatible with these old CPUs, this card and many others like it allow the firmware (and/or the operating system) to choose not only where the memory is mapped, but also its size. We can see this in the same "lspci -vv" output for this GPU: Older systems which do not understand this extended capability will still treat these regions as fixed size, probably with the first size in this list (256MB for region 0, 2MB for region 2). Newer systems can tell the device to "resize" the BAR to a bigger size, which obviously needs the first region to be placed above the 4GB barrier since it's too big.Why is this useful? This particular GPU has 8GB of VRAM; it's quite obvious that region 0 is a direct view into that VRAM. When using the maximum BAR size, the CPU can directly read and write anywhere into the VRAM; when using a smaller BAR, the CPU can only see a small window into the VRAM, and has to use less direct methods to access it.
(As an aside: go right now and do a "sudo lspci -vv" on your computer, if you see a Resizable BAR capability which isn't using the maximum size, you can probably gain a bit more speed for free by going into the BIOS and enabling "Resizable BAR" and/or "Above 4G decoding". If you can't find these options, well, AFAIU that's what this project is all about..)
ChocolateGod · · focus · HN ↗
IIRC Linux doesn't need resizable bar enabled in the BIOS since the kernel will resize the bar if supported by the GPU, Windows however relies on the UEFI doing it which is where it being enabled in the BIOS is needed.
zir_blazer · · focus · HN ↗
GPUs has been using a 256 MiB PCI MMIO window regardless of how much VRAM they actually have since... nearly forever? At least since PCIe is a thing, since I recall than AGP Aperture Size was seteable in era accurate BIOSes. PCIe 3.0 specification introduced a feature known as Resizeable BAR, where the PCI Device can tell a compatible Firmware how much MMIO it actually wants. GPUs uses that to tell a ReBAR capable UEFI Firmware that it wants more MMIO (Usually as big as the GPU VRAM), or uses legacy 256 MiB otherwise.
akiselev · · focus · HN ↗
mlyle · · focus · HN ↗
For historic reasons (e.g. 32 bit address spaces, plus the need to reserve the space for multiple pci peripherals) it has been a narrow, movable aperture.
Resizable BAR lets the size of the aperture be chosen (which is usually chosen to allow all of VRAM to fit in and be directly accessible).
zir_blazer · · focus · HN ↗
MMIO (Memory Mapped I/O) is essentially memory (Whenever RAM or ROM) from OTHER devices that is directly visible on the CPU Address Space. My understanding is that from the CPU side, MMIO is mostly transparent (Except for the massive increase in latency) because it gets used like if it was interacting with its own workspace with regular instructions like MOV.
What PCIe ReBAR changes is that before, you could only see a 256 MiB window onto the GPU VRAM, so there was an added overhead since the GPU may need to relocate things from inside that window somewhere else on its total VRAM (So yes, it may be interpreted as if what you see from the CPU side is just some kind of exchange buffer). I believe the best way to describe how it operates is comparing it to EMS (Expanded Memory) from the DOS days since it also worked with a similar, if not the same idea. You could only see a portion of the total memory from what was installed on the EMS card (A 128 KiB window located on the upper part of the 1 MiB address space from the 8086 CPU), so you had to switch which Page (Region) of the memory was visible, adding a lot of overhead and most likely requiring an additional buffer in main RAM to move data from one Page to another. However, since I have no knowledge if the GPUs really work like that I can't confirm. I never knew whenever the 256 MiB is "fixed" (You always see the same Region) or if you can decide which section of the VRAM to make visible.
amstan · · focus · HN ↗
[1] <a href="https://en.wikipedia.org/wiki/Advanced_eXtensible_Interface" rel="nofollow">https://en.wikipedia.org/wiki/Advanced_eXtensible_Interface
sedatk · · focus · HN ↗
imtringued · · focus · HN ↗
When you insert a GPU into a PCIe slot, the memory mapped regions in memory can't be put in a hard wired location because any arbitrary device can be inserted and it can provide an arbitrary amount of memory (yes the GPU provides its own memory to the CPU). A BAR reserves a memory mapped region in the CPU space that is backed by the PCIe device.
When the BAR is smaller than the memory of the inserted device, the CPU cannot communicate with all of the memory on the inserted device directly anmore. This means if you want to perform a write to a region in the GPU outside a BAR region you have to go through the BAR region anyway. It's not a RAM buffer for PCIe packets.
>I assume it's not like memory mapped peripherals on an AXI bus which is why you need the buffer?
The PCIe controller is an AXI peripheral...
userbinator · · focus · HN ↗
joha4270 · · focus · HN ↗
And if my experience from embedded development is in any way transferable, they're probably fairly spread out and probably takes a fairly big chunk of address space too.
toast0 · · focus · HN ↗