Only need an exact sleep for pre-frame sleep, since the frame is scheduled
to finish just in time.
Otherwise waking slightly late is harmless, the video thread handles the
presentation timing.
In the worst case, input would get polled up to 0.5ms later than originally
expected which isn't going to hurt anything. If you have that little frame
budget, disable optimal frame pacing.
With lm set, every normal facing away from a light gets clamped to zero,
so unlike elsewhere, the saturation branches are frequently mispredicted
here. Clamp unconditionally and derive the flag from the comparison.
With the helpers inlined, FLAG is kept in a register, which led to the
compiler turning the IR/colour saturation checks into conditional moves.
That is slower than a predicted branch when saturation is rare, and made
DPCS slower than it was before inlining.
The error bit is always clear when it is computed, since FLAG is either
reset at the start of the instruction, or masked when the guest writes
it. So it can be ORed in, instead of being cleared and reinserted.
The cross product can be computed with half as many multiplies. None of
the intermediates can overflow, and only the sum was being range checked,
so MAC0 and the flags are unchanged.
The recompilers resolve the handler once per instruction, so return a
variant with the sf/lm bits baked in. Shift amounts and IR saturation
limits become constants, rather than being decoded for each execution.
The triple variants were making up to nine calls per instruction, with
FLAG and IR1-3 going through memory between each of them. PGXP's part of
RTPS is moved out of line instead, so the integer path stays compact.
Sign extending to 44 bits is a no-op unless the value is out of range,
so move it to the overflow path, and test both bounds with a single
comparison. Shortens the dependency chain of each matrix row.
Covers IR/colour saturation, 44-bit MAC overflow wrapping, NCLIP and ORGB
against reference formulas, and hashes the register state after random
instructions through both the interpreter and recompiler entry points.
Use three packed multiply-adds for the 20-tap upsampling FIR and widen
the shared coefficients for downsampling. Preserve the final shift and clamp.
The absolute coefficient sum is 31142, keeping every s32 sum within range.
Validated output, persistent state, and RAM against the original across
8,192 frames using both SSE4.1 and SSE2; compiled with clang-cl /O2.
Bypass the read-only upsampling FIR for channels with zero output volume.
Keep updating feedback and resampling history so re-enabling output is exact.
Validated output and persistent state against the original implementation
across 8,192 frames with independent output-volume and enable transitions.
Decode shifted nibbles directly when both predictor coefficients are zero,
preserving interpolation history, predictor history, and block flags.
Validated against the original decoder across 262,144 blocks and all
header filters/shifts; compiled with clang-cl /O2.
Exit generated blocks when the interpreter raises an exception, preserving its handler PC across all backends and implementing the missing RISC-V and LoongArch fallback paths.