Repository navigation
X87 FPU Support
JNode's x87 floating-point support: the global control word, JLS-correct float→int conversion emitted by the JIT, and the
StrictMathargument-reduction hang.
Every x87 floating-point operation JNode emits — whether from a JIT-compiled guest method, from BootImageBuilder running on the host JVM, or from the cpu.asm kernel initializer — goes through the same x87 FPU. That gives one place to enforce Java's strictfp semantics and one place to break when the FPU's native behaviour diverges from the JLS.
Three things live here:
-
The global control word, initialized once by
init_fpuincore/src/native/x86/cpu.asm, set to double precision (PC=10) and round-to-nearest-even (RC=00). -
Float→int conversion helpers,
X86CompilerHelper.emitF2I/emitF2L, which wrap a singleFISTPin a temporary truncate control word and patch the x87 "integer indefinite" sentinel into the JLS-mandated result. -
Guest-side
StrictMath, where a missingrecompute = falsein the fdlibm argument-reduction kernel livelocked the whole VM.
fninit (DB E3) loads CW = 0x037F: exception masks set, PC=11 (64-bit extended significand), RC=00 (nearest-even). JNode then rewrites PC and RC for Java semantics:
; core/src/native/x86/cpu.asm:10-24
init_fpu:
fninit
; Setup x87 for Java strictfp: double precision (PC=10),
; round to nearest even (RC=00). fninit gives CW=0x037F
; (PC=11 extended, RC=00); clear RC and PC bit 8, keep PC bit 9.
lea ASP,[ASP-SLOT_SIZE]
fstcw [ASP]
and word [ASP], 0xF2FF ; clear bits 11:8 (RC + PC bit 8)
or word [ASP], 0x0200 ; set PC bit 9 -> PC = 10
fldcw [ASP]
lea ASP,[ASP+SLOT_SIZE]
mov AAX,cr0
or eax,CR0_MP ; Enable monitoring FPUinit_fpu is called once from core/src/native/x86/kernel.asm:181. The identical init_fpu text is duplicated as a JNasm test fixture in builder/src/test/org/jnode/jnasm/jnode32.asm:1166-1178 and is kept in sync so the JNasm and native paths cannot drift.
| Bits | Field | Values |
|---|---|---|
5:0 |
Exception masks | — |
7 |
IC | denormal handling |
9:8 |
PC (precision control) |
00=single, 10=double (53-bit), 11=extended (64-bit) |
11:10 |
RC (rounding control) |
00=nearest-even, 01=down, 10=up, 11=toward zero |
12 |
TC (truncate) | never set by JNode |
| Configuration | Value | PC | RC |
|---|---|---|---|
fninit default |
0x037F |
11 extended |
00 nearest |
| Pre-#642 (buggy) |
0x0F7F (0x037F | 0x0C00) |
11 extended |
11 truncate |
| Current kernel value |
0x027F ((0x037F & 0xF2FF) | 0x0200) |
10 double |
00 nearest |
Temporary, inside emitF2I/emitF2L
|
saved CW | 0x0C00
|
10 double |
11 truncate |
strictfp requires every intermediate result to be an IEEE-754 binary64 double with a 53-bit significand. x87 defaults to PC=11 (80-bit extended, 64-bit significand), which lets a JVM observe fraction bits Java forbids. The canonical StrictMath.rint idiom depends on this:
double t52 = (double) (1L << 52);
(t52 + 2.3) - t52 // must be exactly 2.0With a 53-bit significand, ulp(2^52) == 1.0, so adding a fraction and subtracting 2^52 back leaves only the round-to-nearest result. With a 64-bit significand, ulp in [2^52, 2^53) is 2^-11, so 2^52 + 2.3 is not rounded and rint returns 2.2998046875. core/src/test/org/jnode/test/bugs/RintTest.java documents this and pins rint(2.3) == 2.0.
RC must be 00 because rint is defined as ties-to-even: rint(2.5) == 2.0, rint(3.5) == 4.0, rint(4.5) == 4.0, rint(-2.5) == -2.0. The pre-#642 truncate mode failed every one of those.
Math.round(double) is not implemented in-tree — it lives in classlib.jar — but it is pure double arithmetic (floor(a+0.5) for a >= 0, ceil(a-0.5) otherwise) and so it was equally broken under extended precision.
The control word survives thread switches: core/src/native/x86/vm-ints.asm:182 fnsave [ABX] and :311 frstor [ABX] save and restore the full 108-byte FPU state, control word included.
JLS §5.1.3 (narrowing primitive conversion, double/float → int/long):
- NaN →
0 - too-large positive, including
+Infinity→Integer.MAX_VALUE/Long.MAX_VALUE - too-large negative, including
-Infinity→Integer.MIN_VALUE/Long.MIN_VALUE - otherwise the fractional part is discarded — rounding toward zero
The x87 FISTP instruction implements none of the special cases: out-of-range and NaN conversions all yield the integer indefinite value 0x80000000 (32-bit) or 0x8000000000000000 (64-bit), and the rounding direction comes from the control word's RC field.
core/src/core/org/jnode/vm/x86/compiler/X86CompilerHelper.java (1219 lines) adds two static methods at the end of the class:
| Method | Line | Purpose |
|---|---|---|
emitF2I(X86Assembler os, GPR destReg, int destDisp) |
:988 |
float/double → int, JLS-correct |
emitF2L(X86Assembler os, GPR destReg, int destDisp) |
:1108 |
float/double → long, JLS-correct |
They are static (no X86CompilerHelper instance needed) so the L2 code generator can call them, and they take an arbitrary base register plus displacement. A file-scope private static int fpuConvertLabelCounter (:972) supplies globally unique label names, because Label is a VmAddress subclass that compares by its string and the helpers can be invoked many times inside one compiled method.
os.writeLEA(sp, sp, -reserve); // :1026
os.writeMOV(BITS32, sp, 8, scratch); // :1028 save EAX
os.writeFSTP64(sp, 0); // :1032 DD /3 store the double for bit inspection
os.writeFLD64(sp, 0); // :1033 DD /0 reload (also forces a double rounding)
os.writeFSTCW(sp, cwSave); // :1036 9B D9 /7 save the global CW
os.writeMOV(BITS32, scratch, sp, cwSave);
os.writeOR(scratch, 0x0C00); // :1038 RC = 11b = round toward zero
os.writeMOV(BITS32, sp, cwTrunc, scratch);
os.writeFLDCW(sp, cwTrunc); // :1040 D9 /5
os.writeFISTP32(destRegAfter, destDispAfter);// :1041 DB /3
os.writeFLDCW(sp, cwSave); // :1043 restore round-to-nearestScratch frame layout (32-bit: reserve = 20, cwSave = 12, cwTrunc = 16; 64-bit: reserve = 24, cwSave = 16, cwTrunc = 20):
| Offset | Contents |
|---|---|
sp+0 |
double low dword |
sp+4 |
double high dword |
sp+8 |
saved EAX (4 B) / RAX (8 B) |
sp+cwSave |
FNSTCW word + 2 pad |
sp+cwTrunc |
FLDCW word + 2 pad |
The truncate control word is applied by OR-ing 0x0C00 into a copy of the saved CW, so the live 0x027F is reloaded immediately afterwards. Cost per conversion: two FWAIT-prefixed FNSTCW, two FLDCW, three ALU ops.
The critical trick is that the negative cases need no store at all, because the x87 indefinite value is Integer.MIN_VALUE / Long.MIN_VALUE:
os.writeMOV(BITS32, scratch, destRegAfter, destDispAfter);
os.writeCMP_Const(scratch, 0x80000000); // :1046 EAX == Integer.MIN_VALUE?
os.writeJCC(done, X86Constants.JNE); // :1047 no -> in-range value, keep the FISTP result
os.writeMOV(BITS32, scratch, sp, 4); // :1049 high dword of the double
os.writeAND(scratch, 0x7FF00000); // :1050 25 000000F0 7F
os.writeCMP_Const(scratch, 0x7FF00000); // :1051 exponent == 0x7FF ?
os.writeJCC(ovf, X86Constants.JNE); // :1052 finite source that overflowed -> ovf
os.writeMOV(BITS32, scratch, sp, 0); // :1054 low dword
os.writeTEST(scratch, scratch);
os.writeJCC(nan, X86Constants.JNE); // :1056 low mantissa bits set -> NaN -> 0
os.setObjectRef(inf);
os.writeMOV(BITS32, scratch, sp, 4);
os.writeAND(scratch, 0x000FFFFF); // :1060 top 20 mantissa bits
os.writeTEST(scratch, scratch);
os.writeJCC(nan, X86Constants.JNE); // :1062 not an infinity -> it is a NaN after all
os.writeMOV(BITS32, scratch, sp, 4);
os.writeTEST(scratch, 0x80000000); // :1065 sign bit
os.writeJCC(done, X86Constants.JNE); // :1066 -Inf -> LEAVE 0x80000000 in place
os.writeMOV_Const(scratch, 0x7FFFFFFF); // :1068 B8 7FFFFFFF
os.writeMOV(BITS32, destRegAfter, destDispAfter, scratch); // :1069 +Inf -> MAX_VALUE
os.writeJMP(done);
// nan:
os.writeXOR(scratch, scratch); // :1073 35 00000000 -> 0
// ovf:
os.writeTEST(scratch, 0x80000000);
os.writeJCC(done, X86Constants.JNE); // :1080 negative overflow -> keep 0x80000000
os.writeMOV_Const(scratch, 0x7FFFFFFF); // :1082 positive overflow -> MAX_VALUEThe canonical Double.NaN (0x7FF8000000000000) has a zero low dword, so it fails the low-mantissa NaN test at :1056, reaches the inf label, and is caught by the 0x000FFFFF mantissa test at :1062 (since 0x800000 != 0). Both NaN shapes are handled by the same path.
emitF2L tests the 64-bit indefinite sentinel as two 32-bit compares (:1165-1168) and writes Long.MAX_VALUE as two halves (:1189-1192, :1206-1209): 0xFFFFFFFF into the low dword, 0x7FFFFFFF into the high dword.
Both helpers end with os.setObjectRef(done) plus restoration of EAX/RAX and writeLEA(sp, sp, reserve) (:1085-1091 / :1211-1217).
// :1010-1013
final boolean destIsSP = (destReg.getNr() == 4); // ESP and RSP both have nr 4
if (destIsSP) { destRegAfter = sp; destDispAfter = destDisp + reserve; }Because the helper does lea sp,[sp-reserve], an SP-relative destination must be re-based to sp+reserve, otherwise the FISTP would store over the saved double and the saved EAX at sp+8. This is exactly the path the L2 register-destination F2I sequences use.
final String uid = "f2i_" + (++fpuConvertLabelCounter) + "_";
final Label done = new Label(uid + "done");
final Label nan = new Label(uid + "nan");
final Label inf = new Label(uid + "inf");
final Label ovf = new Label(uid + "ovf");The monotone counter replaced an earlier (os.getLength() % 256) uniquifier. Labels are bound with os.setObjectRef(Label) and consumed by writeJCC(Label, cc), which picks jcc rel8 when the target is already resolved and in byte range, else 0F 8x rel32.
Every raw FISTP in the L1/L2 generators has been replaced:
| File | Line | Site |
|---|---|---|
l1a/IntItem.java |
:94-96 |
popFromFPU → emitF2I
|
l1a/LongItem.java |
:110-112 |
popFromFPU → emitF2L
|
l1a/X86BytecodeVisitor.java |
:4443 |
store-to-local of an FPUSTACK int |
l1a/X86BytecodeVisitor.java |
:499 |
store-to-local of an FPUSTACK long |
l1b/IntItem.java |
:95 |
popFromFPU → emitF2I
|
l1b/LongItem.java |
:105 |
popFromFPU → emitF2L
|
l1b/X86BytecodeVisitor.java |
:5419 / :491
|
store-to-local int / long |
l2/GenericX86CodeGenerator.java |
:433, 512, 592, 676 |
case F2I: (reg→reg, mem→reg, mem→mem, mem→mem with different disps) |
Bytecodes funnel in through X86BytecodeVisitor.visit_d2i/d2l/f2i/f2l (l1a/X86BytecodeVisitor.java:1478, 1486, 1780, 1787) → fpCompiler.convert(fromType, toType) → an IntItem/LongItem whose popFromFPU is the patched method.
Example L2 register-destination site (GenericX86CodeGenerator.java:430-434):
case F2I:
os.writePUSH((GPR) rhsReg);
os.writeFLD32(X86Register.ESP, 0);
X86CompilerHelper.emitF2I(os, X86Register.ESP, 0); // exercises the SP displacement shift
os.writePOP((GPR) lhsReg);The F2L, D2I and D2L cases in the same switch still throw new IllegalArgumentException("Unknown operation: …") on master — D2I/D2L reach the helpers through IntItem/LongItem.popFromFPU instead.
core/src/openjdk/vm/java/lang/NativeStrictMath.java is JNode's fdlibm-derived java.lang.StrictMath. It is explicitly excluded from the JNode header-fix target (all/build.xml:830) because it keeps its GNU Classpath headers. An identical copy lives at core/src/test/org/jnode/test/StrictMathTest.java so the functions can be exercised on a normal JVM.
The bug was a single missing line in remPiOver2(double[] x, double[] y, int e0, int nx):
do {
+ recompute = false;
// Distill q[] into iq[] reversingly.
for (i = 0, j = jz, z = q[jz]; j > 0; i++, j--) {The multi-precision argument-reduction kernel declares boolean recompute = false; (:505), sets it to true at :600 when the 24-bit chunk distillation of x * 2/pi collapses to all-zero iq[jk..jz-1] with a zero fractional remainder, and loops with while (recompute) (:603-604). Pre-fix the flag was set but never cleared, so the first time the recomputation branch fired the method never returned. On the second pass the freshly added terms make j != 0 so the branch is not taken again — but the flag is still true, so the loop spins at 100% CPU forever with a fixed jz, re-distilling the same arrays. A pure livelock: no exception, no allocation growth.
remPiOver2(double, double[]) (:418) is only entered for |x| > TWO_20 * PI/2 ≈ 1.647e6 (small args take the polynomial path, medium args the PIO2_1/2/3 split at :433-461). sin (:794-814), cos (:819-839) and tan (:844) all call it, and java.lang.Math delegates to StrictMath, so any guest Math.sin(x) with a triggering x hung the entire VM. Because StrictMath is a VM-side class, the spin happens in Java code inside the kernel's address space — the machine appears completely dead, with no timer threads progressing.
Measured trigger (extracted from StrictMathTest.java and run under Zulu JDK 8 with a trip counter): for exact powers of two the condition holds when e0 ∈ {23,24,25,26}, i.e. |x| = 2^46 … 2^49. sin(2^46) loops pre-fix and returns 0.9994730524837995 post-fix. 2^21 … 2^45 and 2^50 upwards do not loop.
Note this is not the OpenJDK haveTan/takeTan static cache — grepping the tree for haveTan|takeTan returns nothing. JNode's fdlibm port is stateless apart from the loop-local recompute.
| Instruction | Encoding | Assembler method |
|---|---|---|
FLD m64fp |
DD /0 |
X86BinaryAssembler.writeFLD64:1836 |
FSTCW m16int |
9B D9 /… (FWAIT-prefixed FNSTCW) |
writeFSTCW:1942 |
FLDCW m2byte |
D9 /5 |
writeFLDCW:1860 |
FISTP m32int |
DB /3 |
writeFISTP32:1792 |
FISTP m64int |
DF /7 |
writeFISTP64:1802 |
FLD m32fp |
D9 /0 |
writeFLD32:1811 |
FSTP m64fp |
DD /3 |
writeFSTP64:1974 |
FNINIT |
DB E3 |
writeFNINIT:1893 |
FISTP m64int is legal in 32-bit protected mode (32-bit addressing, 64-bit data), so emitF2L is used unchanged by the 32-bit JIT.
X86CompilerHelper is a dual-mode emitter: the JIT always uses X86BinaryAssembler (raw bytes, AbstractX86Compiler.java:78-84), while X86TextAssembler is used only by disassemble() (:154). Both implement the abstract writeF2x primitives in X86Assembler (writeFISTP32:1071, writeFISTP64:1079, writeFLDCW:1125, writeFSTCW:1187), so the new helpers work unchanged in both modes. Text output prints fistp dword [...] / fistp qword [...] / fldcw / fstcw word [...].
| Test | Location | Covers |
|---|---|---|
ConversionTest |
core/src/test/org/jnode/test/bugs/ConversionTest.java |
19 assertions on (int)/(long) of NaN, ±Infinity, ±1.5, ±0.0, 1e20, -1e20, Long.MAX_VALUE/MIN_VALUE
|
RintTest |
core/src/test/org/jnode/test/bugs/RintTest.java |
rint(2.3)=2.0, rint(2.7)=3.0, ties-to-even 2.5/3.5/4.5/-2.5, pass-through of ±Inf, NaN, ±0.0, and the literal (t52 + 2.3) - t52 trick |
StrictMathTest |
core/src/test/org/jnode/test/StrictMathTest.java |
In-VM sin/cos/tan behaviour, including the large-argument path that used to livelock |
All three are main()-style guest programs in the org.jnode.test.bugs / org.jnode.test style (like bug778001.java), not JUnit.
-
Changing the global control word breaks
emitF2I/emitF2L. The helpers OR0x0C00into a copy of the live CW to get truncation. If a thread or a future change leaves RC at11globally, theORis a no-op and the conversion silently becomes double-rounding; if RC is set to something other than00/11,(int) 3.7comes out wrong. -
Round-to-nearest is load-bearing for more than
rint. Any library code that relies on a 53-bit significand (StrictMath,Math.round, transcendentals) is wrong under extended precision. Fix it in the CW, not per-instruction. -
0x80000000is both "the answer" and "no answer". The negative-overflow and-Infinitypaths deliberately leave the x87 indefinite value in place instead of storing a constant. Do not "simplify" them into aMOV. -
Labels must be globally unique, not method-unique.
Label.equalscompares the string, so twof2i_1_donelabels in one compiled method silently bind to the wrong jump target. -
The helpers are
synchronized-free but not reentrant across threads — they only touchsp-relative scratch, so they are safe, but they must never be given a destination that aliases the scratch frame. -
emitF2Lon a 32-bit JIT is fine.FISTP m64intwith 32-bit addressing is legal in protected mode. -
The control word is thread-shared via
fnsave/frstor, not per-thread.vm-ints.asmsaves the whole 108-byte state, so the truncate window insideemitF2Icannot leak across a context switch — but a debug agent that flips the CW behind the JIT's back will. -
A
sinhang is a VM hang, not an application hang.StrictMathlives in the VM address space, so a livelockedremPiOver2freezes timer threads and the machine appears dead. -
NativeStrictMath.javais excluded fromheader-fix. Editing it does not require touching license headers; editing files undercore/src/openjdk/generally does.
- JIT-Compilers — The L1/L2 compiler pipeline that emits these sequences
-
L1-Compiler-Deep-Dive —
IntItem/LongItempopFromFPU, FPU stack item management -
L2-Compiler-Deep-Dive —
GenericX86CodeGeneratorF2I/F2Lcases and the SP-displacement pattern -
JNAsm-Instruction-Encoding —
X86BinaryAssemblervsX86TextAssemblerand the abstractwriteF2xcontract -
Assembly-Files — Role of
cpu.asmandvm-ints.asm -
OpenJDK-Patches — Where
NativeStrictMath.javasits in the class-library override scheme -
Testing — How the
org.jnode.test.bugsguest programs are run