Fixing OS/2 Warp 4.52 running under QEMU (TCG)
Sat Jun 29 2024
Browsing through the QEMU issue tracker on GitLab, my curiosity was piqued by the issue reported at https://gitlab.com/qemu-project/qemu/-/work_items/2198 that OS/2 Warp was unable to boot under QEMU. Many many years ago, I remember several PC Magazines here in the UK running editions with an extra OS/2 Warp CD-ROM included on the front cover so I was vaguely familiar with the OS, and given Paolo’s recent rework of the i386 target decoder which fixed a lot of bugs, I was surprised that this didn’t work. So, time to dust off my debugger and investigate!
The original GitLab issue made reference to a thread on osworld.com indicating that the cause was due to an issue with QEMU’s implementation of the 16-bit SGDT instruction, so I started comparing the logic between QEMU’s TCG accelerator and KVM for this instruction. The official Intel documentation seemed to be inconsistent across CPU releases which had me scratching my head, until I was pointed towards the article at https://www.os2museum.com/wp/sgdtsidt-fiction-and-reality/ containing a technical description of the problem. In short, the Intel documentation far from reflects the behaviour of the SGDT instruction on real CPUs, so if you’ve used the Intel documentation as a reference then you’re doing it wrong.
The fix seemed easy enough: fix up the size of the SGDT and SIDT instructions to write all 32-bits to memory which I implemented in a patch and sent upstream. The patch was merged into QEMU git, but the original reporter noted that it was still not possible to boot OS/2 Warp 4.52 in QEMU. Clearly something else was still amiss, so I requested a copy of the test image so that I could investigate further.
Initially I performed a number of searches related to OS/2 and QEMU which returned a number of articles, most of which
contained historical instructions (as far back as 2008) describing how to get OS/2 Warp booting successfully in QEMU. Based
upon the most recent article I could find reporting that it was working fine, I could determine the approximate QEMU release
on that date and build it to confirm if this was a regression. Despite this I was still unable to boot the test image
successfully which had me scratching my head a little… until I realised that all the working examples were using the
-accel=kvm command line option to enable KVM. A local test confirmed that when using KVM the test image would boot
fine, whereas when using TCG (QEMU’s internal emulator) the test image would fail with a number of exceptions upon boot.
The good news was that this almost certainly confirmed that it not an issue with the device emulation, but more specifically a problem with QEMU’s x86 CPU emulation. And even better by changing a simple command line I was able to switch between the working and non-working cases to try and figure out what the difference was when using the emulated x86 CPU. So now it was time to find and fix the bug.
My normal strategy for debugging CPU issues is to run QEMU with -d in_asm and then use this to determine the address of
the faulting function. I then place a breakpoint at an address close to the fault, and then start to work back until I
find the point at which the working case and the non-working case diverge. Normally this works well, but in this case I
could see that the breakpoint was getting hit constantly during boot. When this happens the only solution is to start
working backwards from the TCG trace output to find each caller of the offending routine, and then test each caller in
turn to see if it takes the faulty path.
This part took me several weeks to figure out, but in the end I had a very fiddly recipe I could use with gdbstub to set a breakpoint at address A, then set a breakpoint at address B, continue C times and then single step D times at which point the until-then identical results of booting OS/2 Warp under KVM and TCG would diverge. Phew!
At this point I started scratching my head, since from all I could see I was simply stepping over a normal instruction and then TCG would generate a fault, whilst KVM would continue as normal. Inspecting the x86 registers between TCG and KVM showed a number of differences with the CS, FS and GS segment registers which is impossible, unless of course an interrupt handler is being executed. Adding some extra debugging in QEMU showed that indeed when stepping over the instruction in question, we were generating a memory fault for the address. By default QEMU does not step into interrupt routines, unless there is an explicit breakpoint set there, but from the logging I was able to figure out that an address fault was happening and place a breakpoint within it.
Stepping through the fault handler showed that it was working identically between TCG and KVM until I spotted one important detail upon entry to the fault handler:
KVM (good):
(gdb) x/64xw 0xffd70780 + $rsp
0xffd75f74: 0x000c006f 0x8d1e0017 0x00043da5 0x8d060000
0xffd75f84: 0x00002246 0x000013bd 0x0000000f 0x4200006f
0xffd75f94: 0x0000971e 0x00000017 0x00005790 0x00000000
0xffd75fa4: 0xfa27b1d8 0x00000000 0x00000000 0x0000000e
0xffd75fb4: 0x00000000 0x00000000 0x10000000 0x7000ffff
0xffd75fc4: 0xf60082dc 0x0780487f 0xff1096d7 0x00000000
0xffd75fd4: 0x00000000 0x00300030 0x0000f204 0x00000000
0xffd75fe4: 0x00008087 0x00000000 0x00000000 0x00000000
0xffd75ff4: 0x00000000 0x00000000 0x00000000 Cannot access memory at address 0xffd76000
-> into page fault handler
Breakpoint 3, 0x00000000fff1ddc8 in ?? ()
(gdb) x/64xw 0xffd70780 + $rsp
0xffd75f64: 0x0000006c 0x00008c0c 0x00000148 0x00012302
0xffd75f74: 0x000c006f 0x8d1e0017 0x00043da5 0x8d060000
0xffd75f84: 0x00002246 0x000013bd 0x0000000f 0x4200006f
0xffd75f94: 0x0000971e 0x00000017 0x00005790 0x00000000
0xffd75fa4: 0xfa27b1d8 0x00000000 0x00000000 0x0000000e
0xffd75fb4: 0x00000000 0x00000000 0x10000000 0x7000ffff
0xffd75fc4: 0xf60082dc 0x0780487f 0xff1096d7 0x00000000
0xffd75fd4: 0x00000000 0x00300030 0x0000f204 0x00000000
0xffd75fe4: 0x00008087 0x00000000 0x00000000 0x00000000
0xffd75ff4: 0x00000000 0x00000000 0x00000000 Cannot access memory at address 0xffd76000
TCG (bad):
(gdb) x/64xw 0xffd70780 + $rsp
0xffd75f74: 0x0000006f 0x00000017 0x00043da5 0x8d060000
0xffd75f84: 0x00002246 0x000013bd 0x0000000f 0x4200006f
0xffd75f94: 0x0000971e 0x00000017 0x000057a4 0x00000000
0xffd75fa4: 0xfa27b1d8 0x00000000 0x00000000 0x0000000e
0xffd75fb4: 0x00000000 0x00000000 0x10000000 0x7000ffff
0xffd75fc4: 0xf60082dc 0x0780487f 0xff1096d7 0x00000000
0xffd75fd4: 0x00000000 0x00300030 0x0000f204 0x00000000
0xffd75fe4: 0x00008087 0x00000000 0x00000000 0x00000000
0xffd75ff4: 0x00000000 0x00000000 0x00000000 Cannot access memory at address 0xffd76000
-> into page fault handler
0xffd75f68: 0x0000006c 0x00008c0c 0x00000148 0x00002202 <--- 0x6f has been overwritten!
0xffd75f78: 0x00000017 0x00043da5 0x8d060000 0x00002246
0xffd75f88: 0x000013bd 0x0000000f 0x4200006f 0x0000971e
0xffd75f98: 0x00000017 0x000057a4 0x00000000 0xfa27b1d8
0xffd75fa8: 0x00000000 0x00000000 0x0000000e 0x00000000
0xffd75fb8: 0x00000000 0x10000000 0x7000ffff 0xf60082dc
0xffd75fc8: 0x0780487f 0xff1096d7 0x00000000 0x00000000
0xffd75fd8: 0x00300030 0x0000f204 0x00000000 0x00008087
0xffd75fe8: 0x00000000 0x00000000 0x00000000 0x00000000
0xffd75ff8: 0x00000000 0x00000000 Cannot access memory at address 0xffd76000
Yes! Stepping into the page fault handler was overwriting the topmost item of the stack, so whilst the fault handler appeared to be working fine, upon its return the stack contents were corrupted causing a panic very quickly after returning back to the currently executing code. With this knowledge I was able to trace through the QEMU’s memory fault handling, and determined that the problem was caused by the stack pointer being adjusted by the POP instruction before checking the segment register contents: if it faulted (which it would on OS/2 since it used a fault-on-demand implementation) then after resolving the fault address, the stack would be corrupted causing the OS to crash soon afterwards.
As luck would have it, the fix was very simple: ensure the writeback (and hence segment validation) occurs before
the POP instruction updates the registers to ensures that the SP remains correct if a fault is taken. With a fix for
this added to QEMU, I then tried to reboot OS/2 Warp once again hoping that it would work. But whilst the process got
further, the image was still unable to boot with a different exception being shown upon boot.
Once again I followed the same process as above, checking the output of -d in_asm at the point where the problem occurs
and then working backwards to find the place where the behaviour diverges between TCG and KVM. Fortunately the routine
in question diverged the first time it was called, and it was fairly easy to see where things were going wrong:
KVM (good):
Breakpoint 3, 0x00000000fff66a25 in ?? ()
(gdb) i r
rax 0x0 0
rbx 0xfff3d0b8 4294168760
rcx 0xfa27b7c0 4196906944
rdx 0xf9809ea4 4185956004
rsi 0xfa27b7c0 4196906944
rdi 0x0 0
rbp 0x504c 0x504c
rsp 0x504c 0x504c
r8 0x0 0
r9 0x0 0
r10 0x0 0
r11 0x0 0
r12 0x0 0
r13 0x0 0
r14 0x0 0
r15 0x0 0
rip 0xfff66a25 0xfff66a25
eflags 0x2246 [ IOPL=2 IF ZF PF ]
cs 0x168 360
ss 0x30 48
ds 0x160 352
es 0x160 352
fs 0x0 0
gs 0x0 0
fs_base 0x0 0
gs_base 0x0 0
k_gs_base 0x0 0
cr0 0x80010013 [ PG WP ET MP PE ]
cr2 0xbefe4 782308
cr3 0x1ff000 [ PDBR=511 PCID=0 ]
cr4 0x200 [ OSFXSR ]
cr8 0x0 0
efer 0x0 [ ]
TCG (bad):
Breakpoint 2, 0x00000000fff66a25 in ?? ()
(gdb) i r
rax 0x73 115
rbx 0xfff3d0b8 4294168760
rcx 0xfa27b7c0 4196906944
rdx 0xffed0053 4293722195
rsi 0xfa27b7c0 4196906944
rdi 0x0 0
rbp 0x5504c 0x5504c <--- value is different
rsp 0x504c 0x504c
r8 0x0 0
r9 0x0 0
r10 0x0 0
r11 0x0 0
r12 0x0 0
r13 0x0 0
r14 0x0 0
r15 0x0 0
rip 0xfff66a25 0xfff66a25
eflags 0x2206 [ IOPL=2 IF PF ]
cs 0x168 360
ss 0x30 48
ds 0x160 352
es 0x160 352
fs 0x0 0
gs 0x0 0
fs_base 0x0 0
gs_base 0x0 0
k_gs_base 0x0 0
cr0 0x80010013 [ PG WP ET MP PE ]
cr2 0xf9859f34 4186283828
cr3 0x1ff000 [ PDBR=511 PCID=0 ]
cr4 0x200 [ OSFXSR ]
cr8 0x0 0
efer 0x0 [ ]
Here it seems that executing an ENTER instruction left extra “junk” in the top 16-bits of RBP (the frame pointer
register) causing an exception when returning back from the current stack frame. Once again the fix itself was
really simple, and with that bug fixed… OS/2 Warp 4.52 was now able to boot under QEMU’s TCG emulator!
The final series was sent upstream and applied in time for the QEMU 9.1 release. In summary I would say that fixing this proved to be a hard challenge in that the main bug occurred upon entry to a memory fault handler and corrupted the stack, meaning that the exception generated appeared some time after the actual bug. However with a systematic approach and a working example from KVM, it was possible to systematically work through the code generated by TCG to determine the underlying issue and fix it.
