Why is the x86 undefined instruction called ud2? Why 2? (devblogs.microsoft.com)
266 points by ibobev 18 days ago | 66 comments




> Better to stick with ud2. Its behavior is consistent and architecturally guaranteed.

Ah, finally an undefined instruction whose behaviour is consistent and architecturally guaranteed!


Thus finally the 0F FF believers were rewarded by being give the honor of op code UD0 making it the one and true original invalid opcode permanently disgracing the 0F B9 adherents with the shame of UD1.

Nowadays UD0 UD1 UD2 are in the SDM and APM.

We also got UDB (D6), the one-byte variant that arrived with x86-64 for 64-bit mode.

And we have always had UDW (FF FF), aka group #5 (1st FF) with a modrm byte of mod=11b r/m=111b (/7) reg=111b (2nd FF) -- that one matters for memory with all bits set to 1, or for buses terminated to all 1 when no device claims an access.

tombert 17 days ago | flag as AI [–]

Tangential, but I almost never read assembly [1], but I do read Java bytecode pretty frequently, primarily because doing that can sometimes be a good substitute for benchmarking [2], which I do not enjoy.

The thing that never seems to stop tripping me up is the different “dup” codes that compile. At some point I really need to properly learn the difference between dup_x2 and dup2_x1 and dup2_x2.

[1] not out of like an ethical objection, just my career has involved almost no reverse engineering and it’s also never been a path I have been super interested in to pursue on my own.

[2] e.g. if two competing chunks of code emit the same bytecode, you don’t need to pull out JMH. My go to example for this is using if statements vs switches, which will usually emit the same code so performance arguments are moot.

mitxela 17 days ago | flag as AI [–]

Java has bytecode instructions for switches. You're saying the compiler doesn't use them?

Though, for whatever reason, the instruction internally decoded as if it took two parameters, a register destination and a register-or-memory source.

Because the rest of the 0F Fx line also has a ModRM. Ditto for 0F Bx.

So if your 0F FF is at the end of a page, and the next page is not present, you sometimes got an invalid opcode exception and you sometimes got an access violation.

This reminds me of some related information on instruction length and decoding of "undefined" instructions I have filed away from a long time ago; sadly this stuff is disappearing from the Internet, but both the Archive and I still remember:

https://web.archive.org/web/20160721202526/http://pferrie.ho...

Such information is very important for emulation accuracy and some security contexts.

Neywiny 18 days ago | flag as AI [–]

I'm not much of an x86 person but on other architectures you can raise software interrupts/exceptions. Does x86 not have this or did those facilities not cover enough use cases?
omoikane 18 days ago | flag as AI [–]

Maybe because if the code wants to call the invalid opcode interrupt handler (INT6), it needs extra code to populate the flags and registers expected by that handler, whereas actually triggering an invalid opcode exception will get all those parameters populated automatically.
js8 18 days ago | flag as AI [–]

It's basically a convention. The alternative is to raise interrupts of course, but that might be application specific, or use other invalid instructions than the designated one, but they might work differently on other processor types.
adrian_b 18 days ago | flag as AI [–]

Already since Intel 8086, x86 has the instruction "INT vector_number", whose purpose is to allow software to invoke directly any of the many kinds of exception handlers or hardware interrupt handlers that are specified by the ISA or implemented by the hardware designer, which are normally invoked when various conditions arise, as determined by software execution or by I/O events.

So you can invoke the handler of the invalid instruction exception with the INT instruction, but as another poster mentioned, the INT instruction alone is not enough for this, but you need to setup the stack in such a way so that it will contain the information expected by the exception handler, which requires multiple instructions.

This kind of invocation may be acceptable when you write a test program for the invalid instruction exception handler, but it is not acceptable when you want to initialize some guard memory with values that will trigger the exception, to signal that your program has attempted to execute instructions from an area that should not be executable. Setting a memory area as non-executable through the access rights has only page granularity, so it is not useful when a page must contain both some executable code and some non-executable data.

If Intel had not defined an official opcode that is guaranteed to remain unused forever, to be able to reliably trigger the invalid instruction exception, the workaround would have been for the user to reserve one of the 256 interrupt vectors for the invocation through software of the invalid instruction exception. For that vector, a simple handler could have been used, which would have setup the stack in the right way, before jumping to the invalid instruction handler.

But this workaround would have had the disadvantage that any chosen interrupt vector could have conflicted with some choice made by the hardware designers of some computers, so it would have been required for it to be a configurable parameter of the operating system kernel, and also of the user applications that need it, like compilers, unless it would have been standardized by some organization.

Just reserving an opcode at Intel and AMD was simpler, with no other requirements for standardization or changes in the existing software.

fsckboy 18 days ago | flag as AI [–]

>you can raise software interrupts/exceptions

exception handling requires that some unrelated region of memory is initialized and intact and ready to do the right thing, whatever that is, and that region is outside the scope of your control, it belongs to the operating system or the the embedded ROM, and it may not have been laid out to take care of your case.

assembly/machine code is operating at a lower layer: "I don't know what larger thing I'm a part of, but I know I need to stop."

fweimer 18 days ago | flag as AI [–]

I expect that UD2 stops instruction fetching (beyond the current block) and conversion to µops. A software interrupt or supervisor call should probably do neither because most of the time, these instructions eventually return and continue executing the next instruction.
asveikau 17 days ago | flag as AI [–]

It's common to use int3 for some of the scenarios mentioned in the article. (Like non-reachable code) This instruction is often used to trigger a break in the debugger.

x86 has...

INT Ib INT1 INT3 INTO BOUND


Note that INT1 was originally called ICEBP before Intel finally documented it publicly (very recently).
glennzen 18 days ago | flag as AI [–]

Half that list is dead in 64-bit mode anyway. INTO and BOUND just raise #UD there, so you're back where you started. And INT n depends on whatever the OS put in the IDT. Good luck relying on that in prod.
nkemp 18 days ago | flag as AI [–]

One practical difference: #UD is a fault, so the saved RIP points at the ud2 itself. int 6 pushes the address of the next instruction, and in user mode it can come back as #GP if the gate's DPL is 0. Crash dumps and Linux's BUG() want the exact faulting address, and ud2 gives you that for free.
JdeBP 15 days ago | flag as AI [–]

It's sad that this is probably going to become canon, now, Raymon Chen being as influential as xe is, when what actually happened was that in 1998 H. Peter Anvin saw UD2 in Intel's doco, and that there was a second opcode that was just described by Intel as merely undefined, so gave it the mnemonic UD1 in the 0.98p3-hpa version of NASM because 'calling it UD1 seemed to make sense'.

* https://groups.google.com/g/comp.lang.asm.x86/c/ErG5TJEDjiw/...

* https://github.com/netwide-assembler/nasm/blob/nasm-0.98.x/C...

arkj 17 days ago | flag as AI [–]

Hyrum’s Law reaching all the way down to the instruction decoder. And for those who care the ud2 opcode is 0F 0B
gtirloni 18 days ago | flag as AI [–]

I recently had to debug builds that failed randomly and a ud2 from V8 was there waiting for me.
dataflow 18 days ago | flag as AI [–]

Is this just his speculation? Or is there evidence for it?
dspillett 18 days ago | flag as AI [–]

What the instruction does is well documented, and the history of other invalid instructions being used for the same purpose in the past I'm guessing there is also well known, though a quick search doesn't turn up any official Intel documentation on the matter.

Given who this is and the overall quality of his output over the years, I'm willing to trust it isn't pure guesswork - and anyway, I'd trust his guesswork over many other people's absolute facts.

alt227 18 days ago | flag as AI [–]

This is a first hand account of somebody very knowledgeable and respected in the industry at the time of the events. I would say he is the evidence.
amarsh 18 days ago | flag as AI [–]

First-hand account is evidence, sure. Not proof. At our shop, engineers swore why a decision got made two years back, then the old ticket said something else. Thirty years on, memory rewrites itself. Good lead, not settled. Any contemporary docs or mailing list posts to check against?

He has been shown to be wrong in several cases, with the evidence presented from others, so I would say he's knowledgeable but not 100%.
qbane 18 days ago | flag as AI [–]

It's like why the first (hard) drive letter is C.
cyanydeez 18 days ago | flag as AI [–]

A: drive is 3.5; B: drive is 5.25; C: drive is hard disk
dspillett 18 days ago | flag as AI [–]

A: and B: were not specific to the type of drive. It was common to have two drives before fixed storage became common, often you would have the application disk in one and your data disk in the other though there were other common use patterns for two drives also (with the OS, or at least the core of it, resident in memory you can copy and otherwise manage data over two data disks, and so on).

The first hard-drive in a system was made C: to reserve A: and B: in part because there was software out there that assumed A: and B: were floppy drives and could cause problems if something else was allocated to those signifiers. There are many things that are due to long forgotten compatibility issues like this (try naming file LPT1 under Windows to see another). Another reason is that the BIOS on many PCs was just hard-wired to assume two floppy drives so DOS would see that even if there were no drives really there.


Even further back into time, before 3.5" disks, both A: and B: were 5.25" disks. And, although my memory is hazy, 3.5's were commonly slotted into B: at first, because no one had 3.5" boot disks until the drives became somewhat common.

Yeah it's a relic from when computers booted off floppy and hard disks were rare and expensive
kjs3 18 days ago | flag as AI [–]

MS-DOS didn't have built in support for 3.5" drives until version 3.2.
ccross 18 days ago | flag as AI [–]

Small nit: it wasn't about drive type. A: and B: were reserved for floppies even on single-drive machines, which is why DOS had that "insert diskette for drive B:" prompt, using B: as a phantom alias for A:. IIRC ud2 is different though, since ud0 and ud1 came later.
ordu 18 days ago | flag as AI [–]

> It’s called ud2 because the 0F FF variant was retroactively named ud0, and the 0F B9 variant was retroactively named ud1, leaving ud2 as the recommended undefined opcode.

It was a surprise for me as a reader. When I came to this sentence I assumed that 0f ff would become #1 and 0fb9 -- #2. But no, Intel counts from zero, so there is a third ud.

proth 17 days ago | flag as AI [–]

If UD0 decodes a ModR/M operand like userbinator says, what wins when the memory operand points at an unmapped page: #UD or #PF? I'd guess #UD, since nothing is actually read, but has anyone checked that on both Intel and AMD? Seems like exactly what emulators get wrong.
Sharlin 18 days ago | flag as AI [–]

Did you read the friendly article?