The computed jump into a NOP sled is a nice trick for getting a fine-grained delay without spending registers on a counter. I'm curious how much contended memory got in the way: on the 48K, code or sample data sitting in the lower 16K gets stalled by the ULA during screen fetch, which seems like it would wreck cycle-exact pulse widths unless everything lives above 0x8000. Did you end up having to place things carefully for that, and did the 128K/+2/+3 contention differences show up in the hardware tests? It's also interesting that the first attempts mostly produced the 8kHz carrier, since that's squarely in the audible range where the PC's 16kHz-ish carrier was easier to ignore.
corbinvachal · · focus · HN ↗