Resident Evil 4 (GameCube) – complete byte-identical decompilation to C/C++
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Resident Evil 4 (GameCube) – complete byte-identical decompilation to C/C++
Unofficial Hacker News client; not affiliated with Y Combinator.
wk_end · · focus · HN ↗
Made using the leaked debug build and its symbols, which shows how meaningful the work the game preservation community that acquires and distributes these things is.
OTOH this is pretty off-putting:
> Where the compiler needed a particular source shape to reproduce a register choice or a schedule and no natural spelling was found, the construct is marked with a // COMPILER-DIFF: comment (644 of them: dead tests, empty asm("") launders and anchors, register T x asm("rN") pins, padding statements).
To me the value of a decomp isn't reproducing the original bytes per se (we already have the original bytes after all) - it's about reconstructing the understanding of the original game as represented by human-readable source code; getting byte-for-byte is just an indication that you've gotten it right. Needing to add a bunch of slop to force the compiler to match the original output is actually just an indication that you've gotten it wrong - and it's a demonstration of the danger of Goodhart's law, especially as it applies to AI.
Pannoniae · · focus · HN ↗
mitxela · · focus · HN ↗
Pannoniae · · focus · HN ↗
"get the order of local variables in this function right, otherwise the register allocation doesn't match. Oh and there's 150 local variables just in this function, good luck trying them all"
"Find out the translation unit boundaries exactly (assume there's no pdb otherwise this is trivial) and after doing so, figure out the order they were compiled in, otherwise it won't match"
"brute force the compilation flags for the project and if you're done, also bruteforce it for the CRT or any other middleware which usually came prebuilt so it doesn't match the main game"
Should I continue;)
mitxela · · focus · HN ↗
Pannoniae · · focus · HN ↗
So these are mostly problems with more advanced compilers yk
dezgeg · · focus · HN ↗
On the other hand, I have heard even old MSVC is a nightmare. Various things are affected by hash table ordering (so the names of variables matter in some situations), stuff like order of #includes mattering, etc.
Pannoniae · · focus · HN ↗
toast0 · · focus · HN ↗
Maybe someone else comes and helps out here and there.
brandonpelfrey · · focus · HN ↗
The reason I’m finding this is much faster than a traditional decomp is that while it’d be nice for the bytes to match, finding the perfect blend of compiler version, compiler args, permitting variables etc to try and find the perfect register assignments, etc is all very time consuming. My ultimate goal is not a byte for byte match, that’s just one way to ensure correctness. I’ve found agents are much faster and effective at reading the original assembly and understanding what’s going on then writing semantically equivalent C.
Tiberium · · focus · HN ↗
ErroneousBosh · · focus · HN ↗
Heh. I was just looking at the RE family on Steam. Nothing over a fiver.
I guess they really did make it more convenient than piracy.
jchw · · focus · HN ↗
There is also value in decomps beyond just understanding and general interest. They can also be used to make more advanced mods, better translation patches, etc. The fidelity here matters, like having asm code accessing structures makes it hard to modify structures, but any fidelity improvement beyond pure asm is very welcome.
Of course I still prefer to try to recover the original code that caused the compiler to do what it did, but it's a really challenging problem sometimes. I've been working on decompiling code from old versions of MSVC for literally years now and you accumulate some knowledge of what things impact register allocation or the order of symbols but some of it comes from things that get fully erased from the source. Like for example, debug builds generally seem to retain symbols that aren't actually referenced anywhere, but those symbols only actually make it into an object file if they are. For functions that were only ever inlined and not actually referenced anywhere... They still wind up in the object files and thus in debug builds, despite nothing referencing them. They are also COMDAT any'd because they can appear in multiple objects legally, which means the exact object that winds up retaining it in the final linked executable is arbitrary (and the compilation flags of the object containing it, too - I bet that was fun for developers to debug.) This is incredibly useful but very challenging, needless to say. It may even be feasible to construct examples that would be legitimately infeasible to simply guess back to equivalent source, which I suspect is a major reason why until it was finally shown to be possible in larger scale projects many people wrote fully matching decomps off as a fool's errand..
Pannoniae · · focus · HN ↗
jchw · · focus · HN ↗
I haven't tried this, but I also suspect that once you have a lot of code fully matching, it might make it possible to ratchet your way up further into builds that you don't have debug information for, that may have more aggressive compilation options. I am not sure if you would manage to get /LTCG builds fully matching even with this advantage, but it's going to be the best shot at it. You're possibly 90% of the way there already.
Pannoniae · · focus · HN ↗
And yes if you have a pdb / an Od build then things are much easier, I was assuming arbitrary game i.e. release binaries.
The "knowledge laundering" approach you describe might help in reconstructing headers, class layouts and function names which is a godsend although I don't think it would be enough to get a match. Getting functionally equivalent code is muuuch easier (although there's the problem of "how do you verify that without running every function")
jchw · · focus · HN ↗
AI is pretty powerful for decompilation, especially because you can also just have an LLM go and start reverse engineering bits of the linker and compiler if you want. (I suspect this decomp is AI assisted if the Clauded out README is any indication.) Maybe future models will be able to come up with clever and novel ways to reduce the number of possibilities and converge faster on possible matching source codes. Or maybe not; I think Astra is the best LLMs have ever been at decompilation and yet I find LLMs frustrating and prone to getting deeply stuck in local maxima in my experimentation.
Pannoniae · · focus · HN ↗
StilesCrisis · · focus · HN ↗
void foo() { bar(); }
int foo(int x, float y) { return bar(x, y); }
This is because parameter-passing and return-value handling can sometimes be done in zero instructions in PPC, if the values are already in the right registers. So you need to continually go back and reassess old functions once you learn the signatures of newer ones!
jchw · · focus · HN ↗
There are plenty of examples where two very semantically different source codes can have the same output, which is really counter-intuitive when combined with the difficulty that often comes with trying to find a single source code that does match.
This is just another reason why debug information is such a godsend; having symbol names for C++ code will usually give you most of the function signature, and type information will give you the rest. That greatly enriches the disassembly, and the automated decompilation output, and no doubt constrains the number of possible source codes that could match both the output and the debug information, somewhat alleviating this issue.
StilesCrisis · · focus · HN ↗
It's a huge benefit that things like the OS and SDK are known and predictable. This got me a foothold. Otherwise I would have gotten nowhere.
I'm still impressed that this decomp has macros (fully lost in the assembly) and nearly-100% meaningful variable names and struct layouts. That's not easy!
jchw · · focus · HN ↗
StilesCrisis · · focus · HN ↗
jchw · · focus · HN ↗
StilesCrisis · · focus · HN ↗
mitxela · · focus · HN ↗
Sesse__ · · focus · HN ↗