Skip to content

X87 FPU Support

opencode-agent[bot] edited this page Sep 27, 2026 · 1 revision

X87 FPU Support

JNode's x87 floating-point support: the global control word, JLS-correct float→int conversion emitted by the JIT, and the StrictMath argument-reduction hang.

Overview

Every x87 floating-point operation JNode emits — whether from a JIT-compiled guest method, from BootImageBuilder running on the host JVM, or from the cpu.asm kernel initializer — goes through the same x87 FPU. That gives one place to enforce Java's strictfp semantics and one place to break when the FPU's native behaviour diverges from the JLS.

Three things live here:

  1. The global control word, initialized once by init_fpu in core/src/native/x86/cpu.asm, set to double precision (PC=10) and round-to-nearest-even (RC=00).
  2. Float→int conversion helpers, X86CompilerHelper.emitF2I / emitF2L, which wrap a single FISTP in a temporary truncate control word and patch the x87 "integer indefinite" sentinel into the JLS-mandated result.
  3. Guest-side StrictMath, where a missing recompute = false in the fdlibm argument-reduction kernel livelocked the whole VM.

The Control Word

fninit (DB E3) loads CW = 0x037F: exception masks set, PC=11 (64-bit extended significand), RC=00 (nearest-even). JNode then rewrites PC and RC for Java semantics:

; core/src/native/x86/cpu.asm:10-24
init_fpu:
        fninit
        ; Setup x87 for Java strictfp: double precision (PC=10),
        ; round to nearest even (RC=00). fninit gives CW=0x037F
        ; (PC=11 extended, RC=00); clear RC and PC bit 8, keep PC bit 9.
        lea ASP,[ASP-SLOT_SIZE]
        fstcw [ASP]
        and word [ASP], 0xF2FF        ; clear bits 11:8 (RC + PC bit 8)
        or  word [ASP], 0x0200        ; set PC bit 9  -> PC = 10
        fldcw [ASP]
        lea ASP,[ASP+SLOT_SIZE]
        mov AAX,cr0
        or eax,CR0_MP                 ; Enable monitoring FPU

init_fpu is called once from core/src/native/x86/kernel.asm:181. The identical init_fpu text is duplicated as a JNasm test fixture in builder/src/test/org/jnode/jnasm/jnode32.asm:1166-1178 and is kept in sync so the JNasm and native paths cannot drift.

Control Word Bit Layout

Bits Field Values
5:0 Exception masks —
7 IC denormal handling
9:8 PC (precision control) 00=single, 10=double (53-bit), 11=extended (64-bit)
11:10 RC (rounding control) 00=nearest-even, 01=down, 10=up, 11=toward zero
12 TC (truncate) never set by JNode
Configuration Value PC RC
fninit default 0x037F 11 extended 00 nearest
Pre-#642 (buggy) 0x0F7F (0x037F | 0x0C00) 11 extended 11 truncate
Current kernel value 0x027F ((0x037F & 0xF2FF) | 0x0200) 10 double 00 nearest
Temporary, inside emitF2I/emitF2L saved CW | 0x0C00 10 double 11 truncate

Why Double Precision

strictfp requires every intermediate result to be an IEEE-754 binary64 double with a 53-bit significand. x87 defaults to PC=11 (80-bit extended, 64-bit significand), which lets a JVM observe fraction bits Java forbids. The canonical StrictMath.rint idiom depends on this:

double t52 = (double) (1L << 52);
(t52 + 2.3) - t52     // must be exactly 2.0

With a 53-bit significand, ulp(2^52) == 1.0, so adding a fraction and subtracting 2^52 back leaves only the round-to-nearest result. With a 64-bit significand, ulp in [2^52, 2^53) is 2^-11, so 2^52 + 2.3 is not rounded and rint returns 2.2998046875. core/src/test/org/jnode/test/bugs/RintTest.java documents this and pins rint(2.3) == 2.0.

RC must be 00 because rint is defined as ties-to-even: rint(2.5) == 2.0, rint(3.5) == 4.0, rint(4.5) == 4.0, rint(-2.5) == -2.0. The pre-#642 truncate mode failed every one of those.

Math.round(double) is not implemented in-tree — it lives in classlib.jar — but it is pure double arithmetic (floor(a+0.5) for a >= 0, ceil(a-0.5) otherwise) and so it was equally broken under extended precision.

The control word survives thread switches: core/src/native/x86/vm-ints.asm:182 fnsave [ABX] and :311 frstor [ABX] save and restore the full 108-byte FPU state, control word included.

Float→Int Conversion

The JLS Rule

JLS §5.1.3 (narrowing primitive conversion, double/float → int/long):

  • NaN → 0
  • too-large positive, including +Infinity → Integer.MAX_VALUE / Long.MAX_VALUE
  • too-large negative, including -Infinity → Integer.MIN_VALUE / Long.MIN_VALUE
  • otherwise the fractional part is discarded — rounding toward zero

The x87 FISTP instruction implements none of the special cases: out-of-range and NaN conversions all yield the integer indefinite value 0x80000000 (32-bit) or 0x8000000000000000 (64-bit), and the rounding direction comes from the control word's RC field.

The Helpers

core/src/core/org/jnode/vm/x86/compiler/X86CompilerHelper.java (1219 lines) adds two static methods at the end of the class:

Method Line Purpose
emitF2I(X86Assembler os, GPR destReg, int destDisp) :988 float/double → int, JLS-correct
emitF2L(X86Assembler os, GPR destReg, int destDisp) :1108 float/double → long, JLS-correct

They are static (no X86CompilerHelper instance needed) so the L2 code generator can call them, and they take an arbitrary base register plus displacement. A file-scope private static int fpuConvertLabelCounter (:972) supplies globally unique label names, because Label is a VmAddress subclass that compares by its string and the helpers can be invoked many times inside one compiled method.

Emitted Sequence (emitF2I, lines 1026-1091)

os.writeLEA(sp, sp, -reserve);              // :1026
os.writeMOV(BITS32, sp, 8, scratch);        // :1028  save EAX
os.writeFSTP64(sp, 0);                      // :1032  DD /3  store the double for bit inspection
os.writeFLD64(sp, 0);                       // :1033  DD /0  reload (also forces a double rounding)
os.writeFSTCW(sp, cwSave);                  // :1036  9B D9 /7  save the global CW
os.writeMOV(BITS32, scratch, sp, cwSave);
os.writeOR(scratch, 0x0C00);                // :1038  RC = 11b = round toward zero
os.writeMOV(BITS32, sp, cwTrunc, scratch);
os.writeFLDCW(sp, cwTrunc);                 // :1040  D9 /5
os.writeFISTP32(destRegAfter, destDispAfter);// :1041  DB /3
os.writeFLDCW(sp, cwSave);                  // :1043  restore round-to-nearest

Scratch frame layout (32-bit: reserve = 20, cwSave = 12, cwTrunc = 16; 64-bit: reserve = 24, cwSave = 16, cwTrunc = 20):

Offset Contents
sp+0 double low dword
sp+4 double high dword
sp+8 saved EAX (4 B) / RAX (8 B)
sp+cwSave FNSTCW word + 2 pad
sp+cwTrunc FLDCW word + 2 pad

The truncate control word is applied by OR-ing 0x0C00 into a copy of the saved CW, so the live 0x027F is reloaded immediately afterwards. Cost per conversion: two FWAIT-prefixed FNSTCW, two FLDCW, three ALU ops.

Sentinel Classification and Immediate Patching

The critical trick is that the negative cases need no store at all, because the x87 indefinite value is Integer.MIN_VALUE / Long.MIN_VALUE:

os.writeMOV(BITS32, scratch, destRegAfter, destDispAfter);
os.writeCMP_Const(scratch, 0x80000000);   // :1046  EAX == Integer.MIN_VALUE?
os.writeJCC(done, X86Constants.JNE);      // :1047  no -> in-range value, keep the FISTP result
os.writeMOV(BITS32, scratch, sp, 4);      // :1049  high dword of the double
os.writeAND(scratch, 0x7FF00000);         // :1050  25 000000F0 7F
os.writeCMP_Const(scratch, 0x7FF00000);   // :1051  exponent == 0x7FF ?
os.writeJCC(ovf, X86Constants.JNE);       // :1052  finite source that overflowed -> ovf
os.writeMOV(BITS32, scratch, sp, 0);      // :1054  low dword
os.writeTEST(scratch, scratch);
os.writeJCC(nan, X86Constants.JNE);       // :1056  low mantissa bits set -> NaN -> 0
os.setObjectRef(inf);
os.writeMOV(BITS32, scratch, sp, 4);
os.writeAND(scratch, 0x000FFFFF);         // :1060  top 20 mantissa bits
os.writeTEST(scratch, scratch);
os.writeJCC(nan, X86Constants.JNE);       // :1062  not an infinity -> it is a NaN after all
os.writeMOV(BITS32, scratch, sp, 4);
os.writeTEST(scratch, 0x80000000);        // :1065  sign bit
os.writeJCC(done, X86Constants.JNE);      // :1066  -Inf -> LEAVE 0x80000000 in place
os.writeMOV_Const(scratch, 0x7FFFFFFF);   // :1068  B8 7FFFFFFF
os.writeMOV(BITS32, destRegAfter, destDispAfter, scratch);   // :1069  +Inf -> MAX_VALUE
os.writeJMP(done);
// nan:
os.writeXOR(scratch, scratch);            // :1073  35 00000000  -> 0
// ovf:
os.writeTEST(scratch, 0x80000000);
os.writeJCC(done, X86Constants.JNE);      // :1080  negative overflow -> keep 0x80000000
os.writeMOV_Const(scratch, 0x7FFFFFFF);   // :1082  positive overflow -> MAX_VALUE

The canonical Double.NaN (0x7FF8000000000000) has a zero low dword, so it fails the low-mantissa NaN test at :1056, reaches the inf label, and is caught by the 0x000FFFFF mantissa test at :1062 (since 0x800000 != 0). Both NaN shapes are handled by the same path.

emitF2L tests the 64-bit indefinite sentinel as two 32-bit compares (:1165-1168) and writes Long.MAX_VALUE as two halves (:1189-1192, :1206-1209): 0xFFFFFFFF into the low dword, 0x7FFFFFFF into the high dword.

Both helpers end with os.setObjectRef(done) plus restoration of EAX/RAX and writeLEA(sp, sp, reserve) (:1085-1091 / :1211-1217).

SP-Relative Destination Fixup

// :1010-1013
final boolean destIsSP = (destReg.getNr() == 4);   // ESP and RSP both have nr 4
if (destIsSP) { destRegAfter = sp; destDispAfter = destDisp + reserve; }

Because the helper does lea sp,[sp-reserve], an SP-relative destination must be re-based to sp+reserve, otherwise the FISTP would store over the saved double and the saved EAX at sp+8. This is exactly the path the L2 register-destination F2I sequences use.

Labels

final String uid = "f2i_" + (++fpuConvertLabelCounter) + "_";
final Label done = new Label(uid + "done");
final Label nan  = new Label(uid + "nan");
final Label inf  = new Label(uid + "inf");
final Label ovf  = new Label(uid + "ovf");

The monotone counter replaced an earlier (os.getLength() % 256) uniquifier. Labels are bound with os.setObjectRef(Label) and consumed by writeJCC(Label, cc), which picks jcc rel8 when the target is already resolved and in byte range, else 0F 8x rel32.

Call Sites

Every raw FISTP in the L1/L2 generators has been replaced:

File Line Site
l1a/IntItem.java :94-96 popFromFPU → emitF2I
l1a/LongItem.java :110-112 popFromFPU → emitF2L
l1a/X86BytecodeVisitor.java :4443 store-to-local of an FPUSTACK int
l1a/X86BytecodeVisitor.java :499 store-to-local of an FPUSTACK long
l1b/IntItem.java :95 popFromFPU → emitF2I
l1b/LongItem.java :105 popFromFPU → emitF2L
l1b/X86BytecodeVisitor.java :5419 / :491 store-to-local int / long
l2/GenericX86CodeGenerator.java :433, 512, 592, 676 case F2I: (reg→reg, mem→reg, mem→mem, mem→mem with different disps)

Bytecodes funnel in through X86BytecodeVisitor.visit_d2i/d2l/f2i/f2l (l1a/X86BytecodeVisitor.java:1478, 1486, 1780, 1787) → fpCompiler.convert(fromType, toType) → an IntItem/LongItem whose popFromFPU is the patched method.

Example L2 register-destination site (GenericX86CodeGenerator.java:430-434):

case F2I:
    os.writePUSH((GPR) rhsReg);
    os.writeFLD32(X86Register.ESP, 0);
    X86CompilerHelper.emitF2I(os, X86Register.ESP, 0);   // exercises the SP displacement shift
    os.writePOP((GPR) lhsReg);

The F2L, D2I and D2L cases in the same switch still throw new IllegalArgumentException("Unknown operation: …") on master — D2I/D2L reach the helpers through IntItem/LongItem.popFromFPU instead.

StrictMath Argument Reduction Hang

core/src/openjdk/vm/java/lang/NativeStrictMath.java is JNode's fdlibm-derived java.lang.StrictMath. It is explicitly excluded from the JNode header-fix target (all/build.xml:830) because it keeps its GNU Classpath headers. An identical copy lives at core/src/test/org/jnode/test/StrictMathTest.java so the functions can be exercised on a normal JVM.

The bug was a single missing line in remPiOver2(double[] x, double[] y, int e0, int nx):

  do {
+     recompute = false;
      // Distill q[] into iq[] reversingly.
      for (i = 0, j = jz, z = q[jz]; j > 0; i++, j--) {

The multi-precision argument-reduction kernel declares boolean recompute = false; (:505), sets it to true at :600 when the 24-bit chunk distillation of x * 2/pi collapses to all-zero iq[jk..jz-1] with a zero fractional remainder, and loops with while (recompute) (:603-604). Pre-fix the flag was set but never cleared, so the first time the recomputation branch fired the method never returned. On the second pass the freshly added terms make j != 0 so the branch is not taken again — but the flag is still true, so the loop spins at 100% CPU forever with a fixed jz, re-distilling the same arrays. A pure livelock: no exception, no allocation growth.

remPiOver2(double, double[]) (:418) is only entered for |x| > TWO_20 * PI/2 ≈ 1.647e6 (small args take the polynomial path, medium args the PIO2_1/2/3 split at :433-461). sin (:794-814), cos (:819-839) and tan (:844) all call it, and java.lang.Math delegates to StrictMath, so any guest Math.sin(x) with a triggering x hung the entire VM. Because StrictMath is a VM-side class, the spin happens in Java code inside the kernel's address space — the machine appears completely dead, with no timer threads progressing.

Measured trigger (extracted from StrictMathTest.java and run under Zulu JDK 8 with a trip counter): for exact powers of two the condition holds when e0 ∈ {23,24,25,26}, i.e. |x| = 2^46 … 2^49. sin(2^46) loops pre-fix and returns 0.9994730524837995 post-fix. 2^21 … 2^45 and 2^50 upwards do not loop.

Note this is not the OpenJDK haveTan/takeTan static cache — grepping the tree for haveTan|takeTan returns nothing. JNode's fdlibm port is stateless apart from the loop-local recompute.

Instruction Encoding Reference

Instruction Encoding Assembler method
FLD m64fp DD /0 X86BinaryAssembler.writeFLD64:1836
FSTCW m16int 9B D9 /… (FWAIT-prefixed FNSTCW) writeFSTCW:1942
FLDCW m2byte D9 /5 writeFLDCW:1860
FISTP m32int DB /3 writeFISTP32:1792
FISTP m64int DF /7 writeFISTP64:1802
FLD m32fp D9 /0 writeFLD32:1811
FSTP m64fp DD /3 writeFSTP64:1974
FNINIT DB E3 writeFNINIT:1893

FISTP m64int is legal in 32-bit protected mode (32-bit addressing, 64-bit data), so emitF2L is used unchanged by the 32-bit JIT.

X86CompilerHelper is a dual-mode emitter: the JIT always uses X86BinaryAssembler (raw bytes, AbstractX86Compiler.java:78-84), while X86TextAssembler is used only by disassemble() (:154). Both implement the abstract writeF2x primitives in X86Assembler (writeFISTP32:1071, writeFISTP64:1079, writeFLDCW:1125, writeFSTCW:1187), so the new helpers work unchanged in both modes. Text output prints fistp dword [...] / fistp qword [...] / fldcw / fstcw word [...].

Tests

Test Location Covers
ConversionTest core/src/test/org/jnode/test/bugs/ConversionTest.java 19 assertions on (int)/(long) of NaN, ±Infinity, ±1.5, ±0.0, 1e20, -1e20, Long.MAX_VALUE/MIN_VALUE
RintTest core/src/test/org/jnode/test/bugs/RintTest.java rint(2.3)=2.0, rint(2.7)=3.0, ties-to-even 2.5/3.5/4.5/-2.5, pass-through of ±Inf, NaN, ±0.0, and the literal (t52 + 2.3) - t52 trick
StrictMathTest core/src/test/org/jnode/test/StrictMathTest.java In-VM sin/cos/tan behaviour, including the large-argument path that used to livelock

All three are main()-style guest programs in the org.jnode.test.bugs / org.jnode.test style (like bug778001.java), not JUnit.

Gotchas & Non-Obvious Behavior

  • Changing the global control word breaks emitF2I/emitF2L. The helpers OR 0x0C00 into a copy of the live CW to get truncation. If a thread or a future change leaves RC at 11 globally, the OR is a no-op and the conversion silently becomes double-rounding; if RC is set to something other than 00/11, (int) 3.7 comes out wrong.
  • Round-to-nearest is load-bearing for more than rint. Any library code that relies on a 53-bit significand (StrictMath, Math.round, transcendentals) is wrong under extended precision. Fix it in the CW, not per-instruction.
  • 0x80000000 is both "the answer" and "no answer". The negative-overflow and -Infinity paths deliberately leave the x87 indefinite value in place instead of storing a constant. Do not "simplify" them into a MOV.
  • Labels must be globally unique, not method-unique. Label.equals compares the string, so two f2i_1_done labels in one compiled method silently bind to the wrong jump target.
  • The helpers are synchronized-free but not reentrant across threads — they only touch sp-relative scratch, so they are safe, but they must never be given a destination that aliases the scratch frame.
  • emitF2L on a 32-bit JIT is fine. FISTP m64int with 32-bit addressing is legal in protected mode.
  • The control word is thread-shared via fnsave/frstor, not per-thread. vm-ints.asm saves the whole 108-byte state, so the truncate window inside emitF2I cannot leak across a context switch — but a debug agent that flips the CW behind the JIT's back will.
  • A sin hang is a VM hang, not an application hang. StrictMath lives in the VM address space, so a livelocked remPiOver2 freezes timer threads and the machine appears dead.
  • NativeStrictMath.java is excluded from header-fix. Editing it does not require touching license headers; editing files under core/src/openjdk/ generally does.

Related Pages

Clone this wiki locally