monocasa 3 hours ago

It says in the rules

> Trapped/emulated/virtualized instructions may only time the trap, not the handler.

But I feel like that 12ms write to an ACPI IO port at current leaderboard position 8 is probably trapping to SMM and being handled there.

layer8 5 hours ago

Nop should be #1, because it is infinitely slow for what it does. ;)

  • jooops1 5 hours ago

    It increments rip by one.

    • EvanAnderson 24 minutes ago

      I thought I remembered reading somewhere re: the 8086 microcode disassembly that NOP, which is encoded as XCHG AX,AX actually does run the XCHG microcode and uses an internal scratchpad register to do the exchange.

      • JoeAltmaier 23 minutes ago

        There were several NOPs - XCHG BX,BX and so on. Those were taken later to be prefixes for new classes of opcodes.

    • fluoridation 2 hours ago

      No, that's done by the decoder. It actually does nothing.

      • russdill an hour ago

        I mean....there are several architectures out there which has a nop that is a jump forward. Kind of a tree forest issue imho

        • fluoridation a minute ago

          I don't know about other architectures in as much detail. I know x86 NOP does nothing.

  • mito88 4 hours ago

    Strategy: nop does nothing. It opens the leaderboard accordingly.

    Score: 1 cycles Time: 0 nanoseconds

    • layer8 4 hours ago

      It opens the leaderboard as #27, so in the last place.

TomatoCo 6 hours ago

This author also has other things like: A compiler that emits only `mov` instructions and another compiler that deliberately messes with the control flow so that, if disassembled, common debuggers will draw symbols like skulls or threats. https://github.com/xoreaxeaxeax/repsych

  • inigyou 4 hours ago

    He also bruteforced the entire opcode space to find undocumented instructions (sandsifter).

metadat 6 hours ago

It’s crazy how computers still seem to get perceivably slow every few years, given how many instructions can be executed in 1ms. Shameful, even..

What’s that law called about programmers wasting all the compute on abstraction?

  • AceJohnny2 21 minutes ago

    There are two aspects to compute performance: latency and throughput.

    There has been a ton of work to improve throughput, but that's often come at the cost of latency, because a great technique to improve throughput is batching, to avoid the per-task overhead cost.

    We've also added many layers of abstraction.

    A seminal example is Dan Luu's "computer latency" table (2017) https://danluu.com/input-lag/, which measures the time between a keypress and visible changes on the screen, and wherein an Apple IIe has latency 5x lower than a Lenovo X1 Carbon on Windows.

    From some perspective, you could say that the Lenovo on Windows example is an impressive feat of engineering, considering all the subsystems involved (USB, interrupt dispatcher, input layer, windowing system, double-buffering...)

    This throughput-over-latency tradeoff can be seen across the entire computer landscape design.

  • mwigdahl 5 hours ago

    Wirth's Law I believe.

    • inigyou 4 hours ago

      And remember to call him by name, not by value!

  • HappyPanacea 6 hours ago

    The new windows notepad is a disgrace

    • adamrezich 3 hours ago

      The new mspaint fucked, then unfucked, then refucked my decades-old muscle memory of Win+R mspaint Enter Ctrl+E 1 Tab 1 Enter Ctrl+V to open Paint, resize canvas to minimum, then paste from clipboard. When you press Ctrl+E now, the Units control is selected by default, for some completely asinine reason!!!

      • EvanAnderson 16 minutes ago

        You aren't alone in being frustrated by decades of muscle memory being disrespected by MSFT.

      • fluoridation 2 hours ago

        That's pretty cool. I've never needed to do that because Paint remembers the last canvas size set manually.

      • sitzkrieg 2 hours ago

        i’m on LTSC for this (and many other) reasons lol! don’t touch my mspaint and notepad

  • summarybot 5 hours ago

    The OS should do less not more

    • RiverCrochet 2 hours ago

      vi should be a kernel-level system call, tunable with sysctls. For agentic management.

  • inigyou 4 hours ago

    Andy and Bill's Law

  • LoganDark 6 hours ago

    Huh? A millisecond is an eternity!

    • m463 5 hours ago

      I remember reading once somewhere:

      If some app responds in 10ms or less, it is INTERACTIVE.

      makes you think.

      • Xirdus 4 hours ago

        It is literally impossible to respond to input in 10ms on most platforms, for various reasons. The USB input lag of 12-30ms and the 60Hz refresh rate of most monitors being just the first two.

        • m463 3 hours ago

          I stand corrected.

          I looked it up and it is .1 seconds (100ms)

          The basic advice regarding response times has been about the same for thirty years [Miller 1968; Card et al. 1991]:

          - 0.1 second is about the limit for having the user feel that the system is reacting instantaneously, meaning that no special feedback is necessary except to display the result.

          - 1.0 second is about the limit for the user's flow of thought to stay uninterrupted, even though the user will notice the delay. Normally, no special feedback is necessary during delays of more than 0.1 but less than 1.0 second, but the user does lose the feeling of operating directly on the data.

          - 10 seconds is about the limit for keeping the user's attention focused on the dialogue. For longer delays, users will want to perform other tasks while waiting for the computer to finish, so they should be given feedback indicating when the computer expects to be done. Feedback during the delay is especially important if the response time is likely to be highly variable, since users will then not know what to expect.

          from Jakob Nielsen:

          https://www.nngroup.com/articles/response-times-3-important-...

          less readable but the original paper:

          https://www.yusufarslan.net/sites/yusufarslan.net/files/uplo...

          • citelao 2 hours ago

            That’s the prevailing statistic, but your original number isn’t that wrong either:

            Humans can perceive much smaller latencies.

            If you look at the Card & Miller reference, at least some humans can perceive differences in ~50ms vs 100ms latencies when typing (in my limited testing, it’s likely you can!). There’s some newer research I don’t have handy that I believe found error rates decreased and NSAT improved until around at least 30ms (if not 20ms).

            On that note, humans can definitely distinguish 60hz vs 120hz reliably (about 8ms faster per frame).

            Even faster: with a reference (eg when dragging on a touchscreen), humans can distinguish down to at least 1ms vs 10ms of latency: https://m.youtube.com/watch?v=vOvQCPLkPt4

            And you can probably distinguish metronomes that are off by about 1-2ms. Much smaller for other things (like metronomes that slightly slower or faster than one another).

            This is a special interest of mine XD

        • xboxnolifes 3 hours ago

          60Hz monitors definitely prevent it, but I'm pretty sure USB lag is far less than 12-30ms. My USB mouse can make a round-trip to a remote server faster than that.

        • sitzkrieg 2 hours ago

          this is also why 60 hz refresh rate is all but dead outside of console gaming (not to mention 1000hz poll rate devices being the norm)

      • faresahmed 2 hours ago

        Sorry for nitpicking but the logical inverse of this statement is the stronger and more appropriate version: If some app does not respond in 10ms or less, it is not interactive.

codeshaunted 6 hours ago

what im seeing from this chart is that we should be using the nop instruction for everything

  • bee_rider 5 hours ago

    Well the best code is no code. Nop could be second best though.

    • inigyou 4 hours ago

      Instructions unclear. Set the NX bit to ensure no code, and got a general protection fault.

markus_zhang 3 hours ago

Does that mean Chris Domas is ready for his next adventure?

michalsustr 5 hours ago

Very cool! Also, huh interesting. I’ve used rdtsc to measure cycle diffs but had no idea its execution takes that long. Is that common across architectures?

  • pbsd 3 hours ago

    The cycle count for RDTSC is ~25 cycles on Skylake-era microarchitectures. The 49 number shown in the OP seems off.

baddash 2 hours ago

just curious, how much do these actually discover useful practices or pitfalls, on top of just being for fun?

vardump 6 hours ago

A great resource for any performance deoptimization.

spoocecow 5 hours ago

Oh wow, glad to see Chris Domas active online again!

IshKebab 4 hours ago

Using MMIO is cheating and makes the results very boring.

It would be much more interesting to know the results if you're only allowed to use main memory.

achierius 6 hours ago

It'd be really interesting to see whether the winning (losing?) instructions/strategies would be different on other architectures. At least right now the top spot (`fxrstor64` on MMIO, starve PCIe) seems relatively architecture-independent, but maybe something about MMIO ordering rules on e.g. POWER would be different enough to change that -- or perhaps open up new avenues?

I wonder what the actual limit on this `fxrstor64` is right now. If you can stall the PCIe bus for that long, then why not indefinitely? Certainly there's no forward progress guarantee here.

arn3n 6 hours ago

There’s definitely strategies here; A lot of the floating point operations use subnormals, and a lot of the worst instructions are slowed down by really, really fucking with MMIO.