The contents of this document are provided in connection with Advanced Micro
Devices, Inc. (“AMD”) products. AMD makes no representations or warranties with
respect to the accuracy or completeness of the contents of this publication and
reserves the right to make changes to specifications and product descriptions at
any time without notice. The information contained herein may be of a preliminary
or advance nature and is subject to change without notice. No license, whether
express, implied, arising by estoppel or otherwise, to any intellectual property rights
is granted by this publication. Except as set forth in AMD’s Standard Terms and
Conditions of Sale, AMD assumes no liability whatsoever, and disclaims any
express or implied warranty, relating to its products including, but not limited to, the
implied warranty of merchantability, fitness for a particular purpose, or infringement
of any intellectual property right.
AMD’s products are not designed, intended, authorized or warranted for use as
components in systems intended for surgical implant into the body, or in other applications intended to support or sustain life, or in any other application in which the
failure of AMD’s product could create a situation where personal injury, death, or
severe property or environmental damage may occur. AMD reserves the right to
discontinue or make changes to its products at any time without notice.
Trademarks
AMD, the AMD arrow logo, AMD Athlon, and AMD Opteron, and combinations thereof, AMD Virtualization and 3DNow!
are trademarks, and AMD-K6 is a registered trademark of Advanced Micro Devices, Inc.
MMX is a trademark and Pentium is a registered trademark of Intel Corporation.
Windows NT is a registered trademark of Microsoft Corporation.
HyperTransport is a licensed trademark of the HyperTransport Technology Consortium.
Other product names used in this publication are for identification purposes only and may be trademarks of their
respective companies.
Added information on ”Speculative Caching of Address Translations,”
”Caching of Upper Level Translation Table Entries,” ”Use of Cached Entries
When Reporting a Page Fault Exception,” ”Use of Cached Entries When
Reporting a Page Fault Exception,” ”Handling of D-Bit Updates,” ”Invalidation
September
2007
July 20073.13
3.14
of Cached Upper-level Entries by INVLPG” and ”Handling of PDPT Entries
in PAE Mode” to section 5.5.2, ”TLB Management” on page 140.
Added 15.20.7, ”Interrupt Masking in Local APIC” on page 395.
Added 16.3.5, ”Extended APIC Control Register” on page 431; clarified the
use of the ICR DS bit in 16.5, ”Interprocessor Interrupts (IPI)” on page 438.
Added minor clarifications and corrected typographical and formatting
errors.
Added 5.3.5, ”1-Gbyte Page Translation” on page 133.
Added 7.2, ”Multiprocessor Memory Access Ordering” on page 164
Added divide-by-zero exception to Table 8-8, “Simultaneous Interrupt
Priorities”‚ on page 226.
Added information on ”CPU Watchdog Timer Register” and ”Machine-Check
Miscellaneous-Error Information Registers (MCi_MISCj)” to Chapter 9.
Added SSE4A support to Chapter 11, ”128-Bit, 64-Bit, and x87
Programming” on page 289.
Added Monitor and MWAIT intercept information to section 15.8, ”Instruction
Intercepts” on page 378 and reorganized intercept information; clarified
15.15.1, ”TLB Flush” on page 390.
Added Monitor and MWAIT intercepts to tables B-1, ”VMCB Layout, Control
Area” on page 471 and C-1, ”SVM Intercept Codes” on page 477.
Added Chapter 16, ”Advanced Programmable Interrupt Controller (APIC)”
on page 425, Chapter 17, ”OS-Visible Workaround Information” on page
453, Chapter 18, ”Hardware P-State Control” on page 457.
Corrected Table 8-6, “General-Protection Exception Conditions”‚ on
page 219. Added SSE3 information. Clarified and corrected information on
the CPUID instruction and feature identification. Added information on the
RDTSCP instruction. Clarified information about MTRRs and PATs in
multiprocessing systems.
Revision Historyxxv
Page 28
AMD64 Technology24593—Rev. 3.14—September 2007
DateRevisionDescription
September
2003
April 20033.08
September
2002
3.09Corrected numerous minor typographical errors.
3.07
Clarified terms in section on FXSAVE/FXSTOR. Corrected several minor
errors of omission. Documentation of CR0.NW bit has been corrected.
Several register diagrams and figure labels have been corrected.
Description of shared cache lines has been clarified in 7.3, ”Memory
Coherency and Protocol” on page 167.
Made numerous small grammatical changes and factual clarifications.
Added Revision History.
xxviRevision History
Page 29
24593—Rev. 3.14—September 2007AMD64 Technology
Preface
About This Book
This book is part of a multivolume work entitled the AMD64 Architecture Pr ogrammer’s Manual. This
table lists each volume and its order number.
TitleOrder No.
Volume 1: Application Programming24592
Volume 2: System Programming24593
Volume 3: General-Purpose and System Instructions24594
Volume 4: 128-Bit Media Instructions26568
Volume 5: 64-Bit Media and x87 Floating-Point Instructions26569
Audience
This volume (Volume 2) is intended for programmers writing operating systems, loaders, linkers,
device drivers, or system utilities. It assumes an understanding of AMD64 architecture applicationlevel programming as described in Volume 1.
This volume describes the AMD64 architecture’s resources and functions that are managed by system
software, including operating-mode control, memory management, interrupts and exceptions, task and
state-change management, system-management mode (including power management), multiprocessor support, debugging, and processor initialization.
Application-programming topics are described in Volume 1. Details about each instruction are
described in volumes 3, 4, and 5.
Organization
This volume begins with an overview of system programming and differences between the x86 and
AMD64 architectures. This is followed by chapters that describe the following details of system
programming:
•System Resour ces—The system registers and processor ID (CPUID) functions.
•Segmented Virtual Memory—The segmented-memory models supported by the architecture and
their associated data structures and protection checks.
•Page Translation and Protection—The page-translation functions supported by the architecture
and their associated data structures and protection checks.
Prefacexxvii
Page 30
AMD64 Technology24593—Rev. 3.14—September 2007
•System-Management Instructions—The instructions used to manage system functions.
•Memory System—The memory-system hierarchy and its resources and protocols, including
memory-characterization, caching, and buffering functions.
•Exceptions and Interrupts—Details about the types and causes of exceptions and interrupts, and
the methods of transferring control during these events.
•Machine-Check Mechanism—The resources and functions that support detection and handling of
machine-check errors.
•System-Management Mode—The resources and functions that support system-management mode
(SMM), including power-management functions.
•128-Bit, 64-Bit, and x87 Programming—The resources and functions that support use (by
application software) and state-saving (by the operation system) of the 128-bit media, 64-bit
media, and x87 floating-point instructions.
•Multiple-Pr ocessor Management—The features of the instruction set and the system resources and
functions that support multiprocessing environments.
•Debug and Performance Resources—The system resources and functions that support software
debugging and performance monitoring.
•Legacy Task Management—Support for the legacy hardware multitasking functions, including
register resources and data structures.
•Processor Initialization and Long-Mode Activation—The methods by which system software
initializes and changes operating modes.
•Mixing Code Across Operating Modes—Things to remember when running programs in different
operating modes.
•Secure Virtual Machine—The system resources that support virtualization development and
deployment.
There are appendices describing details of model-specific registers (MSRs) and machine-check
implementations. Definitions assumed throughout this volume are listed below. The index at the end of
this volume cross-references topics within the volume. For other topics relating to the AMD64
architecture, see the tables of contents and indexes of the other volumes.
xxviiiPreface
Page 31
24593—Rev. 3.14—September 2007AMD64 Technology
Definitions
Some of the following definitions assume a knowledge of the legacy x86 architecture. See “Related
Documents” on page xxxix for descriptions of the legacy x86 architecture.
Terms and Notation
1011b
A binary value—in this example, a 4-bit value.
F0EAh
A hexadecimal value—in this example a 2-byte value.
[1,2)
A range that includes the left-most value (in this case, 1) but excludes the right-most value (in this
case, 2).
7–4
A bit range, from bit 7 to 4, inclusive. The high-order bit is shown first.
128-bit media instructions
Instructions that use the 128-bit XMM registers. These are a combination of the SSE and SSE2
instruction sets.
64-bit media instructions
Instructions that use the 64-bit MMX registers. These are primarily a combination of MMX and
3DNow!™ instruction sets, with some additional instructions from the SSE and SSE2 instruction
sets.
16-bit mode
Legacy mode or compatibility mode in which a 16-bit address size is active. See legacy mode and
compatibility mode.
32-bit mode
Legacy mode or compatibility mode in which a 32-bit address size is active. See legacy mode and
compatibility mode.
64-bit mode
A submode of long mode. In 64-bit mode, the default address size is 64 bits and new features, such
as register extensions, are supported for system and application software.
#GP(0)
Notation indicating a general-protection exception (#GP) with error code of 0.
Prefacexxix
Page 32
AMD64 Technology24593—Rev. 3.14—September 2007
absolute
Said of a displacement that references the base of a code segment rather than an instruction pointer.
Contrast with relative.
ASID
Address space identifier.
biased exponent
The sum of a floating-point value’s exponent and a constant bias for a particular floating-point data
type. The bias makes the range of the biased exponent always positive, which allows reciprocation
without overflow.
byte
Eight bits.
clear
To write a bit value of 0. Compare set.
compatibility mode
A submode of long mode. In compatibility mode, the default address size is 32 bits, and legacy 16-
bit and 32-bit applications run without modification.
commit
To irreversibly write, in program order, an instruction’s result to software-visible storage, such as a
register (including flags), the data cache, an internal write buffer, or memory.
CPL
Current privilege level.
CR0–CR4
A register range, from register CR0 through CR4, inclusive, with the low-order register first.
CR0.PE = 1
Notation indicating that the PE bit of the CR0 register has a value of 1.
direct
Referencing a memory location whose address is included in the instruction’s syntax as an
immediate operand. The address may be an absolute or relative address. Compare indirect.
dirty data
Data held in the processor’s caches or internal buffers that is more recent than the copy held in
main memory.
displacement
A signed value that is added to the base of a segment (absolute addressing) or an instruction pointer
(relative addressing). Same as offset.
xxxPreface
Page 33
24593—Rev. 3.14—September 2007AMD64 Technology
doubleword
Two words, or four bytes, or 32 bits.
double quadword
Eight words, or 16 bytes, or 128 bits. Also called octword.
DS:rSI
The contents of a memory location whose segment address is in the DS register and whose offset
relative to that segment is in the rSI register.
EFER.LME = 0
Notation indicating that the LME bit of the EFER register has a value of 0.
effective address size
The address size for the current instruction after accounting for the default address size and any
address-size override prefix.
effective operand size
The operand size for the current instruction after accounting for the default operand size and any
operand-size override prefix.
element
See vector.
exception
An abnormal condition that occurs as the result of executing an instruction. The processor’s
response to an exception depends on the type of the exception. For all exceptions except 128-bit
media SIMD floating-point exceptions and x87 floating-point exceptions, control is transferred to
the handler (or service routine) for that exception, as defined by the exception’s vector. For
floating-point exceptions defined by the IEEE 754 standard, there are both masked and unmasked
responses. When unmasked, the exception handler is called, and when masked, a default response
is provided instead of calling the handler.
FF /0
Notation indicating that FF is the first byte of an opcode, and a subopcode in the ModR/M byte has
a value of 0.
flush
An often ambiguous term meaning (1) writeback, if modified, and invalidate, as in “flush the cache
line,” or (2) invalidate, as in “flush the pipeline,” or (3) change a value, as in “flush to zero.”
GDT
Global descriptor table.
Prefacexxxi
Page 34
AMD64 Technology24593—Rev. 3.14—September 2007
GIF
Global interrupt flag.
IDT
Interrupt descriptor table.
IGN
Ignore. Field is ignored.
indirect
Referencing a memory location whose address is in a register or other memory location. The
address may be an absolute or relative address. Compare direct.
IRB
The virtual-8086 mode interrupt-redirection bitmap.
IST
The long-mode interrupt-stack table.
IVT
The real-address mode interrupt-vector table.
LDT
Local descriptor table.
legacy x86
The legacy x86 architecture. See “Related Documents” on page xxxix for descriptions of the
legacy x86 architecture.
legacy mode
An operating mode of the AMD64 architecture in which existing 16-bit and 32-bit applications and
operating systems run without modification. A processor implementation of the AMD64
architecture can run in either long mode or legacy mode. Legacy mode has three submodes, real
mode, pr otected mode, and virtual-8086 mode.
long mode
An operating mode unique to the AMD64 architecture. A processor implementation of the
AMD64 architecture can run in either long mode or legacy mode. Long mode has two submodes,
64-bit mode and compatibility mode.
lsb
Least-significant bit.
LSB
Least-significant byte.
xxxiiPreface
Page 35
24593—Rev. 3.14—September 2007AMD64 Technology
main memory
Physical memory, such as RAM and ROM (but not cache memory) that is installed in a particular
computer system.
mask
(1) A control bit that prevents the occurrence of a floating-point exception from invoking an
exception-handling routine. (2) A field of bits used for a control purpose.
MBZ
Must be zero. If software attempts to set an MBZ bit to 1, a general-protection exception (#GP)
occurs.
memory
Unless otherwise specified, main memory.
ModRM
A byte following an instruction opcode that specifies address calculation based on mode (Mod),
register (R), and memory (M) variables.
moffset
A 16, 32, or 64-bit offset that specifies a memory operand directly, without using a ModRM or SIB
byte.
msb
Most-significant bit.
MSB
Most-significant byte.
multimedia instructions
A combination of 128-bit media instructions and 64-bit media instructions.
octword
Same as double quadword.
offset
Same as displacement.
overflow
The condition in which a floating-point number is larger in magnitude than the largest, finite,
positive or negative number that can be represented in the data-type format being used.
packed
See vector.
Prefacexxxiii
Page 36
AMD64 Technology24593—Rev. 3.14—September 2007
PAE
Physical-address extensions.
physical memory
Actual memory, consisting of main memory and cache.
probe
A check for an address in a processor’s caches or internal buffers. External probes originate
outside the processor, and internal pr obes originate within the processor.
protected mode
A submode of legacy mode.
quadword
Four words, or eight bytes, or 64 bits.
RAZ
Read as zero (0), regardless of what is written.
real-addr ess mode
See real mode.
real mode
A short name for real-addr ess mode, a submode of legacy mode.
relative
Referencing with a displacement (also called offset) from an instruction pointer rather than the
base of a code segment. Contrast with absolute.
reserved
Fields marked as reserved may be used at some future time.
To preserve compatibility with future processors, reserved fields require special handling when
read or written by software.
Reserved fields may be further qualified as MBZ, RAZ, SBZ or IGN (see definitions).
Software must not depend on the state of a reserved field, nor upon the ability of such fields to
return to a previously written state.
If a reserved field is not marked with one of the above qualifiers, software must not change the state
of that field; it must reload that field with the same values returned from a prior read.
REX
An instruction prefix that specifies a 64-bit operand size and provides access to additional
registers.
RIP-relative addr essing
Addressing relative to the 64-bit RIP instruction pointer.
xxxivPreface
Page 37
24593—Rev. 3.14—September 2007AMD64 Technology
SBZ
Should be zero. An attempt by software to set an SBZ bit to 1 results in undefined behavior.
set
To write a bit value of 1. Compare clear.
SIB
A byte following an instruction opcode that specifies address calculation based on scale (S), index
(I), and base (B).
SIMD
Single instruction, multiple data. See vector.
SSE
Streaming SIMD extensions instruction set. See 128-bit media instructions and 64-bit media
instructions.
SSE2
Extensions to the SSE instruction set. See 128-bit media instructions and 64-bit media
instructions.
SSE3
Further extensions to the SSE instruction set. See 128-bit media instructions.
sticky bit
A bit that is set or cleared by hardware and that remains in that state until explicitly changed by
software.
TOP
The x87 top-of-stack pointer.
TSS
Task-state segment.
underflow
The condition in which a floating-point number is smaller in magnitude than the smallest nonzero,
positive or negative number that can be represented in the data-type format being used.
vector
(1) A set of integer or floating-point values, called elements, that are packed into a single operand.
Most of the 128-bit and 64-bit media instructions use vectors as operands. Vectors are also called
packed or SIMD (single-instruction multiple-data) operands.
(2) An index into an interrupt descriptor table (IDT), used to access exception handlers. Compare
exception.
Prefacexxxv
Page 38
AMD64 Technology24593—Rev. 3.14—September 2007
virtual-8086 mode
A submode of legacy mode.
VMCB
Virtual machine control block.
VMM
Virtual machine monitor.
word
Two bytes, or 16 bits.
x86
See legacy x86.
Registers
In the following list of registers, the names are used to refer either to a given register or to the contents
of that register:
AH–DH
The high 8-bit AH, BH, CH, and DH registers. Compare AL–DL.
AL–DL
The low 8-bit AL, BL, CL, and DL registers. Compare AH–DH.
AL–r15B
The low 8-bit AL, BL, CL, DL, SIL, DIL, BPL, SPL, and R8B–R15B registers, available in 64-bit
mode.
BP
Base pointer register.
CRn
Control register number n.
CS
Code segment register.
eAX–eSP
The 16-bit AX, BX, CX, DX, DI, SI, BP, and SP registers or the 32-bit EAX, EBX, ECX, EDX,
EDI, ESI, EBP, and ESP registers. Compare rAX–rSP.
EFER
Extended features enable register.
xxxviPreface
Page 39
24593—Rev. 3.14—September 2007AMD64 Technology
eFLAGS
16-bit or 32-bit flags register. Compare rFLAGS.
EFLAGS
32-bit (extended) flags register.
eIP
16-bit or 32-bit instruction-pointer register. Compare rIP.
EIP
32-bit (extended) instruction-pointer register.
FLAGS
16-bit flags register.
GDTR
Global descriptor table register.
GPRs
General-purpose registers. For the 16-bit data size, these are AX, BX, CX, DX, DI, SI, BP, and SP.
For the 32-bit data size, these are EAX, EBX, ECX, EDX, EDI, ESI, EBP, and ESP. For the 64-bit
data size, these include RAX, RBX, RCX, RDX, RDI, RSI, RBP, RSP, and R8–R15.
IDTR
Interrupt descriptor table register.
IP
16-bit instruction-pointer register.
LDTR
Local descriptor table register.
MSR
Model-specific register.
r8–r15
The 8-bit R8B–R15B registers, or the 16-bit R8W–R15W registers, or the 32-bit R8D–R15D
registers, or the 64-bit R8–R15 registers.
rAX–rSP
The 16-bit AX, BX, CX, DX, DI, SI, BP, and SP registers, or the 32-bit EAX, EBX, ECX, EDX,
EDI, ESI, EBP, and ESP registers, or the 64-bit RAX, RBX, RCX, RDX, RDI, RSI, RBP, and RSP
registers. Replace the placeholder r with nothing for 16-bit size, “E” for 32-bit size, or “R” for 64-
bit size.
Prefacexxxvii
Page 40
AMD64 Technology24593—Rev. 3.14—September 2007
RAX
64-bit version of the EAX register.
RBP
64-bit version of the EBP register.
RBX
64-bit version of the EBX register.
RCX
64-bit version of the ECX register.
RDI
64-bit version of the EDI register.
RDX
64-bit version of the EDX register.
rFLAGS
16-bit, 32-bit, or 64-bit flags register. Compare RFLAGS.
RFLAGS
64-bit flags register. Compare rFLAGS.
rIP
16-bit, 32-bit, or 64-bit instruction-pointer register. Compare RIP.
RIP
64-bit instruction-pointer register.
RSI
64-bit version of the ESI register.
RSP
64-bit version of the ESP register.
SP
Stack pointer register.
SS
Stack segment register.
TPR
Task priority register (CR8), a new register introduced in the AMD64 architecture to speed
interrupt management.
xxxviiiPreface
Page 41
24593—Rev. 3.14—September 2007AMD64 Technology
TR
Task register.
Endian Order
The x86 and AMD64 architectures address memory using little-endian byte-ordering. Multibyte
values are stored with their least-significant byte at the lowest byte address, and they are illustrated
with their least significant byte at the right side. Strings are illustrated in reverse order, because the
addresses of their bytes increase from right to left.
Related Documents
•Peter Abel, IBM PC Assembly Language and Pr ogramming , Prentice-Hall, Englewood Cliffs, NJ,
1995.
•Rakesh Agarwal, 80x86 Architecture & Programming: Volume II, Prentice-Hall, Englewood
Cliffs, NJ, 1991.
•AMD data sheets and application notes for particular hardware implementations of the AMD64
architecture.
•AMD, AMD-K6™ MMX™Enhanced Processor Multimedia Technology, Sunnyvale, CA, 2000.
•AMD, 3DNow!™ Technology Manual, Sunnyvale, CA, 2000.
•AMD, AMD Extensions to the 3DNow!™ and MMX™ Instruction Sets, Sunnyvale, CA, 2000.
•AMD, SYSCALL and SYSRET Instruction Specification Application Note, Sunnyvale, CA, 1998.
•Don Anderson and Tom Shanley, Pentium Processor System Architecture, Addison-Wesley, New
York, 1995.
•Nabajyoti Barkakati and Randall Hyde, Microsoft Macr o Assembler Bible, Sams, Carmel, Indiana,
1992.
•Barry B. Brey, 8086/8088, 80286, 80386, and 80486 Assembly Language Programming,
Macmillan Publishing Co., New York, 1994.
•Barry B. Brey, Programming the 80286, 80386, 80486, and Pentium Based Personal Computer,
Prentice-Hall, Englewood Cliffs, NJ, 1995.
•Ralf Brown and Jim Kyle, PC Interrupts, Addison-Wesley, New York, 1994.
•Penn Brumm and Don Brumm, 80386/80486 Assembly Language Programming, Windcrest
McGraw-Hill, 1993.
•Geoff Chappell, DOS Internals, Addison-Wesley, New York, 1994.
•Chips and Technologies, Inc. Super386 DX Programmer’s Reference Manual, Chips and
Technologies, Inc., San Jose, 1992.
•John Crawford and Patrick Gelsinger, Programming the 80386 , Sybex, San Francisco, 1987.
•Walter A. Triebel, The 80386DX Microprocessor, Prentice-Hall, Englewood Cliffs, NJ, 1992.
•John Wharton, The Complete x86, MicroDesign Resources, Sebastopol, California, 1994.
•Web sites and newsgroups:
-www.amd.com
-news.comp.arch
-news.comp.lang.asm.x86
-news.intel.microprocessors
-news.microsoft
Prefacexli
Page 44
AMD64 Technology24593—Rev. 3.14—September 2007
xliiPreface
Page 45
24593—Rev. 3.14—September 2007AMD64 Technology
1System-Programming Overview
This entire volume is intended for system-software developers—programmers writing operating
systems, loaders, linkers, device drivers, or utilities that require access to system resources. These
system resources are generally available only to software running at the highest-privilege level
(CPL=0), also referred to as privileged software. Privilege levels and their interactions are fully
described in “Segment-Protection Overview” on page 93.
This chapter introduces the basic features and capabilities of the AMD64 architecture that are available
to system-software developers. The concepts include:
•The supported address forms and how memory is organized.
•How memory-management hardware makes use of the various address forms to access memory.
•The processor operating modes, and how the memory-management hardware supports each of
those modes.
•The system-control registers used to manage system resources.
•The interrupt and exception mechanism, and how it is used to interrupt program execution and to
report errors.
•Additional, miscellaneous features available to system software, including support for hardware
Many of the legacy features and capabilities are enhanced by the AMD64 architecture to support 64bit operating systems and applications, while providing backward-compatibility with existing
software.
1.1Memory Model
The AMD64 architecture memory model is designed to allow system software to manage application
software and associated data in a secure fashion. The memory model is backward-compatible with the
legacy memory model. Hardware-translation mechanisms are provided to map addresses between
virtual-memory space and physical-memory space. The translation mechanisms allow system
software to relocate applications and data transparently, either anywhere in physical-memory space, or
in areas on the system hard drive managed by the operating system.
In long mode, the AMD64 architecture implements a flat-memory model. In legacy mode, the
architecture implements all legacy memory models.
System-Programming Overview1
Page 46
AMD64 Technology24593—Rev. 3.14—September 2007
1.1.1 Memory Addressing
The AMD64 architecture supports address relocation. To do this, several types of addresses are needed
to completely describe memory organization. Specifically, four types of addresses are defined by the
AMD64 architecture:
•Logical addresses
•Effective addresses, or segment offsets, which are a portion of the logical address.
•Linear (virtual) addresses
•Physical addresses
Logical Addresses. A logical address is a reference into a segmented-address space. It is comprised
of the segment selector and the effective address. Notationally, a logical address is represented as
Logical Address = Segment Selector : Offset
The segment selector specifies an entry in either the global or local descriptor table. The specified
descriptor-table entry describes the segment location in virtual-address space, its size, and other
characteristics. The effective address is used as an offset into the segment specified by the selector.
Logical addresses are often referred to as far pointers. Far pointers are used in software addressing
when the segment reference must be explicit (i.e., a reference to a segment outside the current
segment).
Effective Addresses. The offset into a memory segment is referred to as an effective address (see
“Segmentation” on page 5 for a description of segmented memory). Effective addresses are formed by
adding together elements comprising a base value, a scaled-index value, and a displacement value. The
effective-address computation is represented by the equation
Effective Address = Base + (Scale x Index) + Displacement
The elements of an effective-address computation are defined as follows:
•Base—A value stored in any general-purpose register.
•Scale—A positive value of 1, 2, 4, or 8.
•Index—A two’s-complement value stored in any general-purpose register.
•Displacement—An 8-bit, 16-bit, or 32-bit two’s-complement value encoded as part of the
instruction.
Effective addresses are often referred to as near pointers. A near pointer is used when the segment
selector is known implicitly or when the flat-memory model is used.
Long mode defines a 64-bit effective-address length. If a processor implementation does not support
the full 64-bit virtual-address space, the effective address must be in canonical form (see “Canonical
Address Form” on page 4).
2System-Programming Overview
Page 47
24593—Rev. 3.14—September 2007AMD64 Technology
Linear (Virtual) Addresses. The segment-selector portion of a logical address specifies a segment-
descriptor entry in either the global or local descriptor table. The specified segment-descriptor entry
contains the segment-base address, which is the starting location of the segment in linear-address
space. A linear address is formed by adding the segment-base address to the effective address
(segment offset), which creates a reference to any byte location within the supported linear-address
space. Linear addresses are often referred to as virtual addresses, and both terms are used
interchangeably throughout this document.
Linear Address = Segment Base Address + Effective Address
When the flat-memory model is used—as in 64-bit mode—a segment-base address is treated as 0. In
this case, the linear address is identical to the effective address. In long mode, linear addresses must be
in canonical address form, as described in “Canonical Address Form” on page 4.
Physical Addresses. A physical address is a reference into the physical-address space, typically
main memory. Physical addresses are translated from virtual addresses using page-translation
mechanisms. See “Paging” on page 7 for information on how the paging mechanism is used for
virtual-address to physical-address translation. When the paging mechanism is not enabled, the virtual
(linear) address is used as the physical address.
1.1.2 Memory Organization
The AMD64 architecture organizes memory into virtual memory and physical memory. Virtual-
memory and physical-memory spaces can be (and usually are) different in size. Generally, the virtualaddress space is much larger than physical-address memory. System software relocates applications
and data between physical memory and the system hard disk to make it appear that much more
memory is available than really exists. System software then uses the hardware memory-management
mechanisms to map the larger virtual-address space into the smaller physical-address space.
Virtual Memory. Software uses virtual addresses to access locations within the virtual-memory
space. System software is responsible for managing the relocation of applications and data in virtualmemory space using segment-memory management. System software is also responsible for mapping
virtual memory to physical memory through the use of page translation. The AMD64 architecture
supports different virtual-memory sizes using the following address-translation modes:
•Pr otected Mode—This mode supports 4 gigabytes of virtual-address space using 32-bit virtual
addresses.
•Long Mode—This mode supports 16 exabytes of virtual-address space using 64-bit virtual
addresses.
System-Programming Overview3
Page 48
AMD64 Technology24593—Rev. 3.14—September 2007
Physical Memory. Physical addresses are used to directly access main memory. For a particular
computer system, the size of the available physical-address space is equal to the amount of main
memory installed in the system. The maximum amount of physical memory accessible depends on the
processor implementation and on the address-translation mode. The AMD64 architecture supports
varying physical-memory sizes using the following address-translation modes:
•Real-Addr ess Mode—This mode, also called r eal mode, supports 1 megabyte of physical-address
space using 20-bit physical addresses. This address-translation mode is described in “Real
Addressing” on page 10. Real mode is available only from legacy mode (see “Legacy Modes” on
page 14).
•Legacy Pr otected Mode—This mode supports several different address-space sizes, depending on
the translation mechanism used and whether extensions to those mechanisms are enabled.
Legacy protected mode supports 4 gigabytes of physical-address space using 32-bit physical
addresses. Both segment translation (see “Segmentation” on page 5) and page translation (see
“Paging” on page 7) can be used to access the physical address space, when the processor is
running in legacy protected mode.
When the physical-address size extensions are enabled (see “Physical-Address Extensions (PAE)
Bit” on page 119), the page-translation mechanism can be extended to support 52-bit physical
addresses. 52-bit physical addresses allow up to 4 petabytes of physical-address space to be
supported. (Currently, the AMD64 architecture supports 40-bit addresses in this mode, allowing up
to 1 terabyte of physical-address space to be supported.
•Long Mode—This mode is unique to the AMD64 architecture. This mode supports up to 4
petabytes of physical-address space using 52-bit physical addresses. Long mode requires the use of
page-translation and the physical-address size extensions (PAE).
1.1.3 Canonical Address Form
Long mode defines 64 bits of virtual-address space, but processor implementations can support less.
Although some processor implementations do not use all 64 bits of the virtual address, they all check
bits 63 through the most-significant implemented bit to see if those bits are all zeros or all ones. An
address that complies with this property is in canonical address form . In most cases, a virtual-memory
reference that is not in canonical form causes a general-protection exception (#GP) to occur. However,
implied stack references where the stack address is not in canonical form causes a stack exception
(#SS) to occur. Implied stack references include all push and pop instructions, and any instruction
using RSP or RBP as a base register.
By checking canonical-address form, the AMD64 architecture prevents software from exploiting
unused high bits of pointers for other purposes. Software complying with canonical-address form on a
specific processor implementation can run unchanged on long-mode implementations supporting
larger virtual-address spaces.
4System-Programming Overview
Page 49
24593—Rev. 3.14—September 2007AMD64 Technology
1.2Memory Management
Memory management consists of the methods by which addresses generated by software are translated
by segmentation and/or paging into addresses in physical memory. Memory management is not visible
to application software. It is handled by the system software and processor hardware.
1.2.1 Segmentation
Segmentation was originally created as a method by which system software could isolate software
processes (tasks), and the data used by those processes, from one another in an effort to increase the
reliability of systems running multiple processes simultaneously.
The AMD64 architecture is designed to support all forms of legacy segmentation. However, most
modern system software does not use the segmentation features available in the legacy x86
architecture. Instead, system software typically handles program and data isolation using page-level
protection. For this reason, the AMD64 architecture dispenses with multiple segments in 64-bit mode
and, instead, uses a flat-memory model. The elimination of segmentation allows new 64-bit system
software to be coded more simply, and it supports more efficient management of multi-processing than
is possible in the legacy x86 architecture.
Segmentation is, however, used in compatibility mode and legacy mode. Here, segmentation is a form
of base memory-addressing that allows software and data to be relocated in virtual-address space off of
an arbitrary base address. Software and data can be relocated in virtual-address space using one or
more variable-sized memory segments. The legacy x86 architecture provides several methods of
restricting access to segments from other segments so that software and data can be protected from
interfering with each other.
In compatibility and legacy modes, up to 16,383 unique segments can be defined. The base-address
value, segment size (called a limit), protection, and other attributes for each segment are contained in a
data structure called a segment descriptor. Collections of segment descriptors are held in descriptor
tables. Specific segment descriptors are referenced or selected from the descriptor table using a
segment selector register. Six segment-selector registers are available, providing access to as many as
six segments at a time.
Figure 1-1 on page 6 shows an example of segmented memory. Segmentation is described in
Chapter 4, “Segmented Virtual Memory.”
System-Programming Overview5
Page 50
AMD64 Technology24593—Rev. 3.14—September 2007
Virtual Address
Space
Effective Address
Descriptor Table
Selectors
Virtual Address
CS
DS
ES
FS
GS
SS
Limit
Base
Segment
Limit
Base
Segment
513-201.eps
Figure 1-1.Segmented-Memory Model
Flat Segmentation. One special case of segmented memory is the flat-memory model. In the legacy
flat-memory model, all segment-base addresses have a value of 0, and the segment limits are fixed at
4 Gbytes. Segmentation cannot be disabled but use of the flat-memory model effectively disables
segment translation. The result is a virtual address that equals the effective address. Figure 1-2 on
page 7 shows an example of the flat-memory model.
Software running in 64-bit mode automatically uses the flat-memory model. In 64-bit mode, the
segment base is treated as if it were 0, and the segment limit is ignored. This allows an effective
addresses to access the full virtual-address space supported by the processor.
6System-Programming Overview
Page 51
24593—Rev. 3.14—September 2007AMD64 Technology
Virtual Address
Space
Effective Address
Virtual Address
Flat Segment
513-202.eps
Figure 1-2.Flat Memory Model
1.2.2 Paging
Paging allows software and data to be relocated in physical-address space using fixed-size blocks
called physical pages. The legacy x86 architecture supports three different physical-page sizes of
4 Kbytes, 2 Mbytes, and 4 Mbytes. As with segment translation, access to physical pages by lesserprivileged software can be restricted.
Page translation uses a hierarchical data structure called a page-translation table to translate virtual
pages into physical-pages. The number of levels in the translation-table hierarchy can be as few as one
or as many as four, depending on the physical-page size and processor operating mode. Translation
tables are aligned on 4-Kbyte boundaries. Physical pages must be aligned on 4-Kbyte, 2-Mbyte, or 4Mbyte boundaries, depending on the physical-page size.
Each table in the translation hierarchy is indexed by a portion of the virtual-address bits. The entry
referenced by the table index contains a pointer to the base address of the next-lower-level table in the
translation hierarchy. In the case of the lowest-level table, its entry points to the physical-page base
address. The physical page is then indexed by the least-significant bits of the virtual address to yield
the physical address.
Figure 1-3 on page 8 shows an example of paged memory with three levels in the translation-table
hierarchy. Paging is described in Chapter 5, “Page Translation and Protection.”
System-Programming Overview7
Page 52
AMD64 Technology24593—Rev. 3.14—September 2007
Physical Address
Virtual Address
Table 3Table 2Table 1
Space
Physical Address
Page Translation Tables
Physical Page
Page Table Base Address
513-203.eps
Figure 1-3.Paged Memory Model
Software running in long mode is required to have page translation enabled.
1.2.3 Mixing Segmentation and Paging
Memory-management software can combine the use of segmented memory and paged memory.
Because segmentation cannot be disabled, paged-memory management requires some minimum
initialization of the segmentation resources. Paging can be completely disabled, so segmentedmemory management does not require initialization of the paging resources.
Segments can range in size from a single byte to 4 Gbytes in length. It is therefore possible to map
multiple segments to a single physical page and to map multiple physical pages to a single segment.
Alignment between segment and physical-page boundaries is not required, but memory-management
software is simplified when segment and physical-page boundaries are aligned.
8System-Programming Overview
Page 53
24593—Rev. 3.14—September 2007AMD64 Technology
The simplest, most efficient method of memory management is the flat-memory model. In the flatmemory model, all segment base addresses have a value of 0 and the segment limits are fixed at 4
Gbytes. The segmentation mechanism is still used each time a memory reference is made, but because
virtual addresses are identical to effective addresses in this model, the segmentation mechanism is
effectively ignored. Translation of virtual (or effective) addresses to physical addresses takes place
using the paging mechanism only.
Because 64-bit mode disables segmentation, it uses a flat, paged-memory model for memory
management. The 4 Gbyte segment limit is ignored in 64-bit mode. Figure 1-4 shows an example of
this model.
Effective Address
Virtual Address
Space
Virtual Address
Physical Address
Space
Physical Address
Page Translation Tables
Page Frame
Flat Segment
Page Table Base Address
513-204.eps
Figure 1-4.64-Bit Flat, Paged-Memory Model
System-Programming Overview9
Page 54
AMD64 Technology24593—Rev. 3.14—September 2007
1.2.4 Real Addressing
Real addressing is a legacy-mode form of address translation used in real mode. This simplified form
of address translation is backward compatible with 8086-processor effective-to-physical address
translation. In this mode, 16-bit effective addresses are mapped to 20-bit physical addresses, providing
a 1-Mbyte physical-address space.
Segment selectors are used in real-address translation, but not as an index into a descriptor table.
Instead, the 16-bit segment-selector value is shifted left by 4 bits to form a 20-bit segment-base
address. The 16-bit effective address is added to this 20-bit segment base address to yield a 20-bit
physical address. If the sum of the segment base and effective address carries over into bit 20, that bit
can be optionally truncated to mimic the 20-bit address wrapping of the 8086 processor by using the
A20M# input signal to mask the A20 address bit.
Real-address translation supports a 1-Mbyte physical-address space using up to 64K segments aligned
on 16-byte boundaries. Each segment is exactly 64K bytes long. Figure 1-5 shows an example of realaddress translation.
015
Effective Address
0000Effective Address0000Selector
+
Physical Address
Selectors
CS
DS
ES
FS
GS
SS
019019
019
513-205.eps
Figure 1-5.Real-Address Memory Model
10System-Programming Overview
Page 55
24593—Rev. 3.14—September 2007AMD64 Technology
1.3Operating Modes
The legacy x86 architecture provides four operating modes or environments that support varying
forms of memory management, virtual-memory and physical-memory sizes, and protection:
•Real Mode.
•Protected Mode.
•Virtual-8086 Mode.
•System Management Mode.
The AMD64 architecture supports all these legacy modes, and it adds a new operating mode called
long mode. Table 1-1 shows the differences between long mode and legacy mode. Software can move
between all supported operating modes as shown in Figure 1-6 on page 12. Each operating mode is
described in the following sections.
Table 1-1.Operating Modes
System
Mode
64-Bit
Long
Mode
Legacy
Mode
Note:
1. Defaults can be overridden in most modes using an instruction prefix or system control bit.
2. Register extensions includes eight new GPRs and eight new XMM registers (also called SSE registers).
3. Long mode supports only x86 protected mode. It does not support x86 real mode or virtual-8086 mode.
Mode
3
Compatibility
Mode
Protected
Mode
Virtual-8086
Mode
Real Mode
Software
Required
New
64-bit OS
Legacy
32-bit OS
Legacy
16-bit OS
Application
Recompile
Required
yes64
no
no
Defaults
Address
Size
(bits)
32
1616
3232
1616
161632
1
Operand
Size
(bits)
32
Register
Extensions
yes64
no32
no
Maximum
2
GPR
Width
(bits)
32
System-Programming Overview11
Page 56
AMD64 Technology24593—Rev. 3.14—September 2007
Long Mode
RSMSMI#
System
Management
Mode
64-bit
Mode
CS.L=0
EFER.LME=1, CR4.PAE=1
then CR0.PG=1
RSM
CR0.PE=1
Reset
Reset
CS.L=1
CS.L=0
Protected
Mode
Real
Mode
Compatibility
CR0.PG=0
then EFER.LME=0
SMI#
EFLAGS.VM=0
EFLAGS.VM=1
CR0.PE=0
SMI#
Mode
RSM
Reset
RSM
SMI#
RSM
SMI#
Reset
Virtual
8086
Mode
513-206.eps
Figure 1-6.Operating Modes of the AMD64 Architecture
1.3.1 Long Mode
Long mode consists of two submodes: 64-bit mode and compatibility mode. 64-bit mode supports
several new features, including the ability to address 64-bit virtual-address space. Compatibility mode
provides binary compatibility with existing 16-bit and 32-bit applications when running on 64-bit
system software.
Throughout this document, references to long mode refer collectively to both 64-bit mode and
compatibility mode. If a function is specific to either 64-bit mode or compatibility mode, then those
specific names are used instead of the name long mode.
Before enabling and activating long mode, system software must first enable protected mode. The
process of enabling and activating long mode is described in Chapter 14, “Processor Initialization and
12System-Programming Overview
Page 57
24593—Rev. 3.14—September 2007AMD64 Technology
Long Mode Activation.” Long mode features are described throughout this document, where
applicable.
1.3.2 64-Bit Mode
64-bit mode, a submode of long mode, provides support for 64-bit system software and applications by
adding the following new features:
•64-bit virtual addresses (processor implementations can have fewer).
•Register extensions through a new instruction prefix (REX):
•Flat-segment address space with single code, data, and stack space.
The mode is enabled by the system software on an individual code-segment basis. Although code
segments are used to enable and disable 64-bit mode, the legacy segmentation mechanism is largely
disabled. Page translation is required for memory management purposes. Because 64-bit mode
supports a 64-bit virtual-address space, it requires 64-bit system software and development tools.
In 64-bit mode, the default address size is 64 bits, and the default operand size is 32 bits. The defaults
can be overridden on an instruction-by-instruction basis using instruction prefixes. A new REX prefix
is introduced for specifying a 64-bit operand size and the new registers.
1.3.3 Compatibility Mode
Compatibility mode, a submode of long mode, allows system software to implement binary
compatibility with existing 16-bit and 32-bit x86 applications. It allows these applications to run,
without recompilation, under 64-bit system software in long mode, as shown in Table 1-1 on page 11.
In compatibility mode, applications can only access the first 4 Gbytes of virtual-address space.
Standard x86 instruction prefixes toggle between 16-bit and 32-bit address and operand sizes.
Compatibility mode, like 64-bit mode, is enabled by system software on an individual code-segment
basis. Unlike 64-bit mode, however, segmentation functions the same as in the legacy-x86
architecture, using 16-bit or 32-bit protected-mode semantics. From an application viewpoint,
compatibility mode looks like a legacy protected-mode environment. From a system-software
viewpoint, the long-mode mechanisms are used for address translation, interrupt and exception
handling, and system data-structures.
System-Programming Overview13
Page 58
AMD64 Technology24593—Rev. 3.14—September 2007
1.3.4 Legacy Modes
Legacy mode consists of three submodes: real mode, protected mode, and virtual-8086 mode.
Protected mode can be either paged or unpaged. Legacy mode preserves binary compatibility not only
with existing x86 16-bit and 32-bit applications but also with existing x86 16-bit and 32-bit system
software.
Real Mode. In this mode, also called real-address mode, the processor supports a physical-memory
space of 1 Mbyte and operand sizes of 16 bits (default) or 32 bits (with instruction prefixes). Interrupt
handling and address generation are nearly identical to the 80286 processor's real mode. Paging is not
supported. All software runs at privilege level 0.
Real mode is entered after reset or processor power-up. The mode is not supported when the processor
is operating in long mode because long mode requires that paged protected mode be enabled.
Protected Mode. In this mode, the processor supports virtual-memory and physical-memory spaces
of 4 Gbytes and operand sizes of 16 or 32 bits. All segment translation, segment protection, and
hardware multitasking functions are available. System software can use segmentation to relocate
effective addresses in virtual-address space. If paging is not enabled, virtual addresses are equal to
physical addresses. Paging can be optionally enabled to allow translation of virtual addresses to
physical addresses and to use the page-based memory-protection mechanisms.
In protected mode, software runs at privilege levels 0, 1, 2, or 3. Typically, application software runs at
privilege level 3, the system software runs at privilege levels 0 and 1, and privilege level 2 is available
to system software for other uses. The 16-bit version of this mode was first introduced in the 80286
processor.
Virtual-8086 Mode. Virtual-8086 mode allows system software to run 16-bit real-mode software on a
virtualized-8086 processor. In this mode, software written for the 8086, 8088, 80186, or 80188
processor can run as a privilege-level-3 task under protected mode. The processor supports a virtualmemory space of 1 Mbytes and operand sizes of 16 bits (default) or 32 bits (with instruction prefixes),
and it uses real-mode address translation.
Virtual-8086 mode is enabled by setting the virtual-machine bit in the EFLAGS register
(EFLAGS.VM). EFLAGS.VM can only be set or cleared when the EFLAGS register is loaded from
the TSS as a result of a task switch, or by executing an IRET instruction from privileged software. The
POPF instruction cannot be used to set or clear the EFLAGS.VM bit.
Virtual-8086 mode is not supported when the processor is operating in long mode. When long mode is
enabled, any attempt to enable virtual-8086 mode is silently ignored.
14System-Programming Overview
Page 59
24593—Rev. 3.14—September 2007AMD64 Technology
1.3.5 System Management Mode (SMM)
System management mode (SMM) is an operating mode designed for system-control activities that are
typically transparent to conventional system software. Power management is one popular use for
system management mode. SMM is primarily targeted for use by the basic input-output system (BIOS)
and specialized low-level device drivers. The code and data for SMM are stored in the SMM memory
area, which is isolated from main memory by the SMM output signal.
SMM is entered by way of a system management interrupt (SMI). Upon recognizing an SMI, the
processor enters SMM and switches to a separate address space where the SMM handler is located and
executes. In SMM, the processor supports real-mode addressing with 4 Gbyte segment limits and
default operand, address, and stack sizes of 16 bits (prefixes can be used to override these defaults).
1.4System Registers
Figure 1-7 on page 16 shows the system registers defined for the AMD64 architecture. System
software uses these registers to, among other things, manage the processor operating environment,
define system resource characteristics, and to monitor software execution. With the exception of the
RFLAGS register, system registers can be read and written only from privileged software.
Except for the descriptor-table registers and task register, the AMD64 architecture defines all system
registers to be 64 bits wide. The descriptor table and task registers are defined by the AMD64
architecture to include 64-bit base-address fields, in addition to their other fields.
As shown in Figure 1-7 on page 16, the system registers include:
•Contr ol Registers—These registers are used to control system operation and some system features.
See “System-Control Registers” on page 41 for details.
•System-Flags Register—The RFLAGS register contains system-status flags and masks. It is also
used to enable virtual-8086 mode and to control application access to I/O devices and interrupts.
See “RFLAGS Register” on page 50 for details.
•Descriptor-Table Registers—These registers contain the location and size of descriptor tables
stored in memory. Descriptor tables hold segmentation data structures used in protected mode. See
“Descriptor Tables” on page 71 for details.
•Task Register—The task register contains the location and size in memory of the task-state
segment. The hardware-multitasking mechanism uses the task-state segment to hold state
information for a given task. The TSS also holds other data, such as the inner-level stack pointers
used when changing to a higher privilege level. See “Task Register” on page 311 for details.
•Debug Registers—Debug registers are used to control the software-debug mechanism, and to
report information back to a debug utility or application. See “Debug Registers” on page 328 for
details.
System-Programming Overview15
Page 60
AMD64 Technology24593—Rev. 3.14—September 2007
Control Registers
CR0
CR2
CR3
CR4
CR8
System-Flags Register
RFLAGS
Debug Registers
DR0
DR1
DR2
DR3
DR6
DR7
Descriptor-Table Registers
GDTR
IDTR
LDTR
Extended-Feature-Enable Register
EFER
System-Configuration Register
SYSCFG
System-Linkage Registers
STAR
LSTAR
CSTAR
SFMASK
FS.base
GS.base
KernelGSbase
SYSENTER_CS
SYSENTER_ESP
SYSENTER_EIP
Debug-Extension Registers
DebugCtlMSR
LastBranchFromIP
LastBranchToIP
LastIntFromIP
LastIntToIP
Memory-Typing Registers
MTRRcap
MTRRdefType
MTRRphysBasen
MTRRphysMaskn
MTRRfixn
PAT
TOP_MEM
TOP_MEM2
Performance-Monitoring Registers
TSC
PerfEvtSeln
PerfCtrn
Machine-Check Registers
MCG_CAP
MCG_STAT
MCG_CTL
MCi_CTL
MCi_STATUS
MCi_ADDR
MCi_MISC
Task Register
TR
Model-Specific Registers
513-260.eps
Figure 1-7.System Registers
Also defined as system registers are a number of model-specific registers included in the AMD64
architectural definition, and shown in Figure 1-7:
•Extended-Featur e-Enable Register—The EFER register is used to enable and report status on
special features not controlled by the CRn control registers. In particular, EFER is used to control
activation of long mode. See “Extended Feature Enable Register (EFER)” on page 54 for more
information.
16System-Programming Overview
Page 61
24593—Rev. 3.14—September 2007AMD64 Technology
•System-Configuration Register—The SYSCFG register is used to enable and configure system-
bus features. See “System Configuration Register (SYSCFG)” on page 57 for more information.
•System-Linkage Registers—These registers are used by system-linkage instructions to specify
operating-system entry points, stack locations, and pointers into system-data structures. See “Fast
System Call and Return” on page 149 for details.
•Memory-Typing Registers—Memory-typing registers can be used to characterize (type) system
memory. Typing memory gives system software control over how instructions and data are cached,
and how memory reads and writes are ordered. See “MTRRs” on page 184 for details.
•Debug-Extension Registers—These registers control additional software-debug reporting features.
See “Debug Registers” on page 328 for details.
•Performance-Monitoring Registers—Performance-monitoring registers are used to count
processor and system events, or the duration of events. See “Performance Optimization” on
page 341 for more information.
•Machine-Check Registers—The machine-check registers control the response of the processor to
non-recoverable failures. They are also used to report information on such failures back to system
utilities designed to respond to such failures. See “Machine Check MSRs” on page 256 for more
information.
1.5System-Data Structures
Figure 1-8 on page 18 shows the system-data structures defined for the AMD64 architecture. Systemdata structures are created and maintained by system software for use by the processor when running
in protected mode. A processor running in protected mode uses these data structures to manage
memory and protection, and to store program-state information when an interrupt or task switch
occurs.
System-Programming Overview17
Page 62
AMD64 Technology24593—Rev. 3.14—September 2007
Segment Descriptors (Contained in Descriptor Tables)
Code
Stack
Data
Descriptor Tables
Global-Descriptor Table
Descriptor
Descriptor
. . .
Descriptor
Page-Translation Tables
Page-Map Level-4
Gate
Task-State Segment
Local-Descriptor Table
Interrupt-Descriptor Table
Gate Descriptor
Gate Descriptor
. . .
Gate Descriptor
Task-State Segment
Local-Descriptor Table
Descriptor
Descriptor
. . .
Descriptor
Page TablePage DirectoryPage-Directory Pointer
513-261.eps
Figure 1-8.System-Data Structures
As shown in Figure 1-8, the system-data structures include:
•Descriptors—A descriptor provides information about a segment to the processor, such as its
location, size and privilege level. A special type of descriptor, called a gate, is used to provide a
code selector and entry point for a software routine. Any number of descriptors can be defined, but
system software must at a minimum create a descriptor for the currently executing code segment
and stack segment. See “Legacy Segment Descriptors” on page 77, and “Long-Mode Segment
Descriptors” on page 86 for complete information on descriptors.
•Descriptor Tables—As the name implies, descriptor tables hold descriptors. The global-descriptor
table holds descriptors available to all programs, while a local-descriptor table holds descriptors
used by a single program. The interrupt-descriptor table holds only gate descriptors used by
18System-Programming Overview
Page 63
24593—Rev. 3.14—September 2007AMD64 Technology
interrupt handlers. System software must initialize the global-descriptor and interrupt-descriptor
tables, while use of the local-descriptor table is optional. See “Descriptor Tables” on page 71 for
more information.
•Task-S tate Segment—The task-state segment is a special segment for holding processor-state
information for a specific program, or task. It also contains the stack pointers used when switching
to more-privileged programs. The hardware multitasking mechanism uses the state information in
the segment when suspending and resuming a task. Calls and interrupts that switch stacks cause the
stack pointers to be read from the task-state segment. System software must create at least one
task-state segment, even if hardware multitasking is not used. See “Legacy Task-State Segment”
on page 313, and “64-Bit Task State Segment” on page 317 for details.
•Page-Translation Tables—Use of page translation is optional in protected mode, but it is required
in long mode. A four-level page-translation data structure is provided to allow long-mode
operating systems to translate a 64-bit virtual-address space into a 52-bit physical-address space.
Legacy protected mode can use two- or three-level page-translation data structures. See “Page
Translation Overview” on page 115 for more information on page translation.
1.6Interrupts
The AMD64 architecture provides a mechanism for the processor to automatically suspend (interrupt)
software execution and transfer control to an interrupt handler when an interrupt or exception occurs.
An interrupt handler is privileged software designed to identify and respond to the cause of an interrupt
or exception, and return control back to the interrupted software. Interrupts can be caused when
system hardware signals an interrupt condition using one of the external-interrupt signals on the
processor. Interrupts can also be caused by software that executes an interrupt instruction. Exceptions
occur when the processor detects an abnormal condition as a result of executing an instruction. The
term “interrupts” as used throughout this volume includes both interrupts and exceptions when the
distinction is unnecessary.
System software not only sets up the interrupt handlers, but it must also create and initialize the data
structures the processor uses to execute an interrupt handler when an interrupt occurs. The data
structures include the code-segment descriptors for the interrupt-handler software and any datasegment descriptors for data and stack accesses. Interrupt-gate descriptors must also be supplied.
Interrupt gates point to interrupt-handler code-segment descriptors, and the entry point in an interrupt
handler. Interrupt gates are stored in the interrupt-descriptor table. The code-segment and datasegment descriptors are stored in the global-descriptor table and, optionally, the local-descriptor table.
When an interrupt occurs, the processor uses the interrupt vector to find the appropriate interrupt gate
in the interrupt-descriptor table. The gate points to the interrupt-handler code segment and entry point,
and the processor transfers control to that location. Before invoking the interrupt handler, the processor
saves information required to return to the interrupted program. For details on how the processor
transfers control to interrupt handlers, see “Legacy Protected-Mode Interrupt Control Transfers” on
page 231, and “Long-Mode Interrupt Control Transfers” on page 241.
System-Programming Overview19
Page 64
AMD64 Technology24593—Rev. 3.14—September 2007
Table 1-2 shows the supported interrupts and exceptions, ordered by their vector number. Refer to
“Vectors” on page 208 for a complete description of each interrupt, and a description of the interrupt
mechanism.
Table 1-2.Interrupts and Exceptions
VectorDescription
0Integer Divide-by-Zero Exception
1Debug Exception
2Non-Maskable-Interrupt
3Breakpoint Exception (INT 3)
4Overflow Exception (INTO instruction)
5Bound-Range Exception (BOUND instruction)
6Invalid-Opcode Exception
7Device-Not-Available Exception
8Double-Fault Exception
9
10Invalid-TSS Exception
11Segment-Not-Present Exception
12Stack Exception
13General-Protection Exception
14Page-Fault Exception
15(Reserved)
16x87 Floating-Point Exception
17Alignment-Check Exception
18Machine-Check Exception
19SIMD Floating-Point Exception
0-255Interrupt Instructions
AnyHardware Maskable Interrupts
Coprocessor-Segment-Overrun Exception (reserved in
AMD64)
1.7Additional System-Programming Facilities
1.7.1 Hardware Multitasking
A task is any program that the processor can execute, suspend, and later resume executing at the point
of suspension. During the time a task is suspended, other tasks are allowed to execute. Each task has its
own execution space, consisting of a code segment, data segments, and a stack segment for each
privilege level. Tasks can also have their own virtual-memory environment managed by the pagetranslation mechanism. The state information defining this execution space is stored in the task-state
segment (TSS) maintained for each task.
20System-Programming Overview
Page 65
24593—Rev. 3.14—September 2007AMD64 Technology
Support for hardware multitasking is provided by implementations of the AMD64 architecture when
software is running in legacy mode. Hardware multitasking provides automated mechanisms for
switching tasks, saving the execution state of the suspended task, and restoring the execution state of
the resumed task. When hardware multitasking is used to switch tasks, the processor takes the
following actions:
•The processor automatically suspends execution of the task, allowing any executing instructions to
complete and save their results.
•The execution state of a task is saved in the task TSS.
•The execution state of a new task is loaded into the processor from its TSS.
•The processor begins executing the new task at the location specified in the new task TSS.
Use of hardware-multitasking features is optional in legacy mode. Generally, modern operating
systems do not use the hardware-multitasking features, and instead perform task management entirely
in software. Long mode does not support hardware multitasking at all.
Whether hardware multitasking is used or not, system software must create and initialize at least one
task-state segment data-structure. This requirement holds for both long-mode and legacy-mode
software. The single task-state segment holds critical pieces of the task execution environment and is
referenced during certain control transfers.
Detailed information on hardware multitasking is available in Chapter 12, “Task Management,” along
with a full description of the requirements that must be met in initializing a task-state segment when
hardware multitasking is not used.
1.7.2 Machine Check
Implementations of the AMD64 architecture support the machine-check exception. This exception is
useful in system applications with stringent requirements for reliability, availability, and serviceability.
The exception allows specialized system-software utilities to report hardware errors that are generally
severe and non-recoverable. Providing the capability to report such errors can allow complex system
problems to be pinpointed rapidly.
The machine-check exception is described in Chapter 9, “Machine Check Mechanism.” Much of the
error-reporting capabilities is implementation dependent. For more information, developers of
machine-check error-reporting software should also refer to the BIOS writer’s guide for a specific
implementation.
1.7.3 Software Debugging
A software-debugging mechanism is provided in hardware to help software developers quickly isolate
programming errors. This capability can be used to debug system software and application software
alike. Only privileged software can access the debugging facilities. Generally, software-debug support
is provided by a privileged application program rather than by the operating system itself.
The facilities supported by the AMD64 architecture allow debugging software to perform the
following:
System-Programming Overview21
Page 66
AMD64 Technology24593—Rev. 3.14—September 2007
•Set breakpoints on specific instructions within a program.
•Set breakpoints on an instruction-address match.
•Set breakpoints on a data-address match.
•Set breakpoints on specific I/O-port addresses.
•Set breakpoints to occur on task switches when hardware multitasking is used.
•Single step an application instruction-by-instruction.
•Single step only branches and interrupts.
•Record a history of branches and interrupts taken by a program.
The debugging facilities are fully described in “Software-Debug Resources” on page 327. Some
processors provide additional, implementation-specific debug support. For more information, refer to
the BIOS writer’s guide for the specific implementation.
1.7.4 Performance Monitoring
For many software developers, the ability to identify and eliminate performance bottlenecks from a
program is nearly as important as quickly isolating programming errors. Implementations of the
AMD64 architecture provide hardware performance-monitoring resources that can be used by special
software applications to identify such bottlenecks. Non-privileged software can access the
performance monitoring facilities, but only if privileged software grants that access.
The performance-monitoring facilities allow the counting of events, or the duration of events.
Performance-analysis software can use the data to calculate the frequency of certain events, or the time
spent performing specific activities. That information can be used to suggest areas for improvement
and the types of optimizations that are helpful.
The performance-monitoring facilities are fully described in “Performance Optimization” on
page 341. The specific events that can be monitored are generally implementation specific. For more
information, refer to the BIOS writer’s guide for the specific implementation.
22System-Programming Overview
Page 67
24593—Rev. 3.14—September 2007AMD64 Technology
2x86 and AMD64 Architecture Differences
The AMD64 architecture is designed to provide full binary compatibility with all previous AMD
implementations of the x86 architecture. This chapter summarizes the new features and architectural
enhancements introduced by the AMD64 architecture, and compares those features and enhancements
with previous AMD x86 processors. Most of the new capabilities introduced by the AMD64
architecture are available only in long mode (64-bit mode, compatibility mode, or both). However,
some of the new capabilities are also available in legacy mode, and are mentioned where appropriate.
The material throughout this chapter assumes the reader has a solid understanding of the x86
architecture. For those who are unfamiliar with the x86 architecture, please read the remainder of this
volume before reading this chapter.
2.1Operating Modes
See “Operating Modes” on page 11 for a complete description of the operating modes supported by the
AMD64 architecture.
2.1.1 Long Mode
The AMD64 architecture introduces long mode and its two sub-modes: 64-bit mode and compatibility
mode.
64-Bit Mode. 64-bit mode provides full support for 64-bit system software and applications. The new
features introduced in support of 64-bit mode are summarized throughout this chapter. To use 64-bit
mode, a 64-bit operating system and tool chain are required.
Compatibility Mode. Compatibility mode allows 64-bit operating systems to implement binary
compatibility with existing 16-bit and 32-bit x86 applications. It allows these applications to run,
without recompilation, under control of a 64-bit operating system in long mode. The architectural
enhancements introduced by the AMD64 architecture that support compatibility mode are
summarized throughout this chapter.
Unsupported Modes. Long mode does not support the following two operating modes:
•Virtual-8086 Mode—The virtual-8086 mode bit (EFLAGS.VM) is ignored when the processor is
running in long mode. When long mode is enabled, any attempt to enable virtual-8086 mode is
silently ignored. System software must leave long mode in order to use virtual-8086 mode.
•Real Mode—Real mode is not supported when the processor is operating in long mode because
long mode requires that protected mode be enabled.
2.1.2 Legacy Mode
The AMD64 architecture supports a pure x86 legacy mode, which preserves binary compatibility not
only with existing 16-bit and 32-bit applications but also with existing 16-bit and 32-bit operating
x86 and AMD64 Architecture Differences23
Page 68
AMD64 Technology24593—Rev. 3.14—September 2007
systems. Legacy mode supports real mode, protected mode, and virtual-8086 mode. A reset always
places the processor in legacy mode (real mode), and the processor continues to run in legacy mode
until system software activates long mode. New features added by the AMD64 architecture that are
supported in legacy mode are summarized in this chapter.
2.1.3 System-Management Mode
The AMD64 architecture supports system-management mode (SMM). SMM can be entered from both
long mode and legacy mode, and SMM can return directly to either mode. The following differences
exist between the support of SMM in the AMD64 architecture and the SMM support found in previous
processor generations:
•The SMRAM state-save area format is changed to hold the 64-bit processor state. This state-save
area format is used regardless of whether SMM is entered from long mode or legacy mode.
•The auto-halt restart and I/O-instruction restart entries in the SMRAM state-save area are one byte
instead of two bytes.
•The initial processor state upon entering SMM is expanded to reflect the 64-bit nature of the
processor.
•New conditions exist that can cause a processor shutdown while exiting SMM.
•SMRAM caching considerations are modified because the legacy FLUSH# external signal
(writeback, if modified, and invalidate) is not supported on implementations of the AMD64
architecture.
See Chapter 10, “System-Management Mode,” for more information on the SMM differences.
2.2Memory Model
The AMD64 architecture provides enhancements to the legacy memory model to support very large
physical-memory and virtual-memory spaces while in long mode. Some of this expanded support for
physical memory is available in legacy mode.
2.2.1 Memory Addressing
Virtual-Memory Addressing. Virtual-memory support is expanded to 64 address bits in long mode.
This allows up to 16 exabytes of virtual-address space to be accessed. The virtual-address space
supported in legacy mode is unchanged.
Physical-Memory Addressing. Physical-memory support is expanded to 52 address bits in long
mode and legacy mode. This allows up to 4 petabytes of physical memory to be accessed. The
expanded physical-memory support is achieved by using paging and the page-size extensions.
Implementations can support fewer than 52 physical-address bits. The first implementation of the
AMD64 architecture, for example, supports 40-bit physical addressing in both long mode and legacy
mode.
24x86 and AMD64 Architecture Differences
Page 69
24593—Rev. 3.14—September 2007AMD64 Technology
Effective Addressing. The effective-address length is expanded to 64 bits in long mode. An
effective-address calculation uses 64-bit base and index registers, and sign-extends 8-bit and 32-bit
displacements to 64 bits. In legacy mode, effective addresses remain 32 bits long.
2.2.2 Page Translation
The AMD64 architecture defines an expanded page-translation mechanism supporting translation of a
64-bit virtual address to a 52-bit physical address. See “Long-Mode Page Translation” on page 128 for
detailed information on the enhancements to page translation in the AMD64 architecture. The
enhancements are summarized below.
Physical-Address Extensions (PAE). The AMD64 architecture requires physical-address
extensions to be enabled (CR4.PAE=1) before long mode is entered. When PAE is enabled, all paging
data-structures are 64 bits, allowing references into the full 52-bit physical-address space supported by
the architecture.
Page-Size Extensions (PSE). Page-size extensions (CR4.PSE) are ignored in long mode. Long
mode does not support the 4-Mbyte page size enabled by page-size extensions. Long mode does,
however, support 4-Kbyte and 2-Mbyte page sizes.
Paging Data Structures. The AMD64 architecture extends the page-translation data structures in
support of long mode. The extensions are:
•Page-map level-4 (PML4)—Long mode defines a new page-translation data structure, the PML4
table. The PML4 table sits at the top of the page-translation hierarchy and references PDP tables.
•Page-directory pointer (PDP)—The PDP tables in long mode are expanded from 4 entries to 512
entries each.
•Page-dir ectory pointer entry (PDPE)—Previously undefined fields within the legacy-mode PDPE
are defined by the AMD64 architecture.
CR3 Register. The CR3 register is expanded to 64 bits for use in long-mode page translation. When
long mode is active, the CR3 register references the base address of the PML4 table. In legacy mode,
the upper 32 bits of CR3 are masked by the processor to support legacy page translation. CR3
references the PDP base-address when physical-address extensions are enabled, or the page-directory
table base-address when physical-address extensions are disabled.
Legacy-Mode Enhancements. Legacy-mode software can take advantage of the enhancements
made to the physical-address extension (PAE) support and page-size extension (PSE) support. The
four-level page translation mechanism introduced by long mode is not available to legacy-mode
software.
•PAE—When physical-address extensions are enabled (CR4.PAE=1), the AMD64 architecture
allows legacy-mode software to load up to 52-bit (maximum size) physical addresses into the PDE
and PTE. (Addresses are expanded to the maximum physical address size supported by the
implementation.)
x86 and AMD64 Architecture Differences25
Page 70
AMD64 Technology24593—Rev. 3.14—September 2007
•PSE—The use of page-size extensions allows legacy mode software to define 4-Mbyte pages using
the 32-bit page-translation tables. When page-size extensions are enabled (CR4.PSE=1), the
AMD64 architecture enhances the 4-Mbyte PDE to support 40 physical-address bits.
See “Legacy-Mode Page Translation” on page 120 for more information on these enhancements.
2.2.3 Segmentation
In long mode, the effects of segmentation depend on whether the processor is running in compatibility
mode or 64-bit mode:
•In compatibility mode, segmentation functions just as it does in legacy mode, using legacy 16-bit
or 32-bit protected mode semantics.
•64-bit mode requires a flat-memory model for creating a flat 64-bit virtual-address space. Much of
the segmentation capability present in legacy mode and compatibility mode is disabled when the
processor is running in 64-bit mode.
The differences in the segmentation model as defined by the AMD64 architecture are summarized in
the following sections. See Chapter 4, “Segmented Virtual Memory,” for a thorough description of
these differences.
Descriptor-Table Registers. In long mode, the base-address portion of the descriptor-table registers
(GDTR, IDTR, LDTR, and TR) are expanded to 64 bits. The full 64-bit base address can only be
loaded by software when the processor is running in 64-bit mode (using the LGDT, LIDT, LLDT, and
LTR instructions, respectively). However, the full 64-bit base address is used by a processor running in
compatibility mode (in addition to 64-bit mode) when making a reference into a descriptor table.
A processor running in legacy mode can only load the low 32 bits of the base address, and the high 32
bits are ignored when references are made to the descriptor tables.
Code-Segment Descriptors. The AMD64 architecture defines a new code-segment descriptor
attribute, L (long). In compatibility mode, the processor treats code-segment descriptors as it does in
legacy mode, with the exception that the processor recognizes the L attribute. If a code descriptor with
L=1 is loaded in compatibility mode, the processor leaves compatibility mode and enters 64-bit mode.
In legacy mode, the L attribute is reserved.
The following differences exist for code-segment descriptors in 64-bit mode only:
•The CS base-address field is ignored by the processor.
•The CS limit field is ignored by the processor.
•Only the L (long), D (default size), and DPL (descriptor-privilege level) fields are used by the
processor in 64-bit mode. All remaining attributes are ignored.
Data-Segment Descriptors. The following differences exist for data-segment descriptors in 64-bit
mode only:
•The DS, ES, and SS descriptor base-address fields are ignored by the processor.
26x86 and AMD64 Architecture Differences
Page 71
24593—Rev. 3.14—September 2007AMD64 Technology
•The FS and GS descriptor base-address fields are expanded to 64 bits and used in effective-address
calculations. The 64 bits of base address are mapped to model-specific registers (MSRs), and can
only be loaded using the WRMSR instruction.
•The limit fields and attribute fields of all data-segment descriptors (DS, ES, FS, GS, and SS) are
ignored by the processor.
In compatibility mode, the processor treats data-segment descriptors as it does in legacy mode.
Compatibility mode ignores the high 32 bits of base address in the FS and GS segment descriptors
when calculating an effective address.
System-Segment Descriptors. In 64-bit mode only, The LDT and TSS system-segment descriptor
formats are expanded by 64 bits, allowing them to hold 64-bit base addresses. LLDT and LTR
instructions can be used to load these descriptors into the LDTR and TR registers, respectively, from
64-bit mode.
In compatibility mode and legacy mode, the formats of the LDT and TSS system-segment descriptors
are unchanged. Also, unlike code-segment and data-segment descriptors, system-segment descriptor
limits ar e checked by the processor in long mode.
Some legacy mode LDT and TSS type-field encodings are illegal in long mode (both compatibility
mode and 64-bit mode), and others are redefined to new types. See “System Descriptors” on page 88
for additional information.
Gate Descriptors. The following differences exist between gate descriptors in long mode (both
compatibility mode and 64-bit mode) and in legacy mode:
•In long mode, all 32-bit gate descriptors are redefined as 64-bit gate descriptors, and are expanded
to hold 64-bit offsets. The length of a gate descriptor in long mode is therefore 128 bits (16 bytes),
versus the 64 bits (8 bytes) in legacy mode.
•Some type-field encodings are illegal in long mode, and others are redefined to new types. See
“Gate Descriptors” on page 90 for additional information.
•The interrupt-gate and trap-gate descriptors define a new field, called the interrupt-stack table
(IST) field.
2.3Protection Checks
The AMD64 architecture makes the following changes to the protection mechanism in long mode:
•The page-protection-check mechanism is expanded in long mode to include the U/S and R/W
protection bits stored in the PML4 entries and PDP entries.
•Several system-segment types and gate-descriptor types that are legal in legacy mode are illegal in
long mode (compatibility mode and 64-bit mode) and fail type checks when used in long mode.
•Segment-limit checks are disabled in 64-bit mode for the CS, DS, ES, FS, GS, and SS segments.
Segment-limit checks remain enabled for the LDT, GDT, IDT and TSS system segments.
All segment-limit checks are performed in compatibility mode.
x86 and AMD64 Architecture Differences27
Page 72
AMD64 Technology24593—Rev. 3.14—September 2007
•Code and data segments used in 64-bit mode are treated as both readable and writable.
See “Page-Protection Checks” on page 142 and “Segment-Protection Overview” on page 93 for
detailed information on the protection-check changes.
2.4Registers
The AMD64 architecture adds additional registers to the architecture, and in many cases expands the
size of existing registers to 64 bits. The 80-bit floating-point stack registers and their overlaid 64-bit
MMX™ registers are not modified by the AMD64 architecture.
2.4.1 General-Purpose Registers
In 64-bit mode, the general-purpose registers (GPRs) are 64 bits wide, and eight additional GPRs are
available. The GPRs are: RAX, RBX, RCX, RDX, RDI, RSI, RBP, RSP, and the new R8–R15
registers. To access the full 64-bit operand size, or the new R8–R15 registers, an instruction must
include a new REX instruction-prefix byte (see “REX Prefixes” on page 29 for a summary of this
prefix).
In compatibility and legacy modes, the GPRs consist only of the eight legacy 32-bit registers. All
legacy rules apply for determining operand size.
2.4.2 128-Bit Media Registers
In 64-bit mode, eight additional 128-bit XMM registers are available, XMM8–XMM15. A REX
instruction prefix is used to access these registers. In compatibility and legacy modes, the XMM
registers consist of the eight 128-bit legacy registers, XMM0–XMM7.
2.4.3 Flags Register
The flags register is expanded to 64 bits, and is called RFLAGS. All 64 bits can be accessed in 64-bit
mode, but the upper 32 bits are reserved and always read back as zeros. Compatibility mode and legacy
mode can read and write only the lower-32 bits of RFLAGS (the legacy EFLAGS).
2.4.4 Instruction Pointer
In long mode, the instruction pointer is extended to 64 bits, to support 64-bit code offsets. This 64-bit
instruction pointer is called RIP.
2.4.5 Stack Pointer
In 64-bit mode, the size of the stack pointer, RSP, is always 64 bits. The stack size is not controlled by
a bit in the SS descriptor, as it is in compatibility or legacy mode, nor can it be overridden by an
instruction prefix. Address-size overrides are ignored for implicit stack references.
28x86 and AMD64 Architecture Differences
Page 73
24593—Rev. 3.14—September 2007AMD64 Technology
2.4.6 Control Registers
The AMD64 architecture defines several enhancements to the control registers (CRn). In long mode,
all control registers are expanded to 64 bits, although the entire 64 bits can be read and written only
from 64-bit mode. A new control register, the task-priority register (CR8 or TPR) is added, and can be
read and written from 64-bit mode. Last, the function of the page-enable bit (CR0.PG) is expanded.
When long mode is enabled, the PG bit is used to activate and deactivate long mode.
2.4.7 Debug Registers
In long mode, all debug registers are expanded to 64 bits, although the entire 64 bits can be read and
written only from 64-bit mode. Expanded register encodings for the decode registers allow up to eight
new registers to be defined (DR8–DR15), although presently those registers are not supported by the
AMD64 architecture.
2.4.8 Extended Feature Register (EFER)
The EFER is expanded by the AMD64 architecture to include a long-mode-enable bit (LME), and a
long-mode-active bit (LMA). These new bits can be accessed from legacy mode and long mode.
2.4.9 Memory Type Range Registers (MTRRs)
The legacy MTRRs are architecturally defined as 64 bits, and can accommodate the maximum 52-bit
physical address allowed by the AMD64 architecture. From both long mode and legacy mode,
implementations of the AMD64 architecture reference the entire 52-bit physical-address value stored
in the MTRRs. Long mode and legacy mode system software can update all 64 bits of the MTRRs to
manage the expanded physical-address space.
2.4.10 Other Model-Specific Registers (MSRs)
Several other MSRs have fields holding physical addresses. Examples include the APIC-base register
and top-of-memory register. Generally, any model-specific register that contains a physical address is
defined architecturally to be 64 bits wide, and can accommodate the maximum physical-address size
defined by the AMD64 architecture. When physical addresses are read from MSRs by the processor,
the entire value is read regardless of the operating mode. In legacy implementations, the high-order
MSR bits are reserved, and software must write those values with zeros. In legacy mode on AMD64
architecture implementations, software can read and write all supported high-order MSR bits.
2.5Instruction Set
2.5.1 REX Prefixes
REX prefixes are a new family of instruction-prefix bytes used in 64-bit mode to:
•Specify the new GPRs and XMM registers.
•Specify a 64-bit operand size.
x86 and AMD64 Architecture Differences29
Page 74
AMD64 Technology24593—Rev. 3.14—September 2007
•Specify additional control registers. One additional control register, CR8, is defined in 64-bit
mode.
•Specify additional debug registers (although none are currently defined).
Not all instructions require a REX prefix. The prefix is necessary only if an instruction references one
of the extended registers or uses a 64-bit operand. If a REX prefix is used when it has no meaning, it is
ignored.
Default 64-Bit Operand Size. In 64-bit mode, two groups of instructions have a default operand size
of 64 bits and thus do not need a REX prefix for this operand size:
•Near branches.
•All instructions, except far branches, that implicitly reference the RSP. See “Instructions that
Reference RSP” on page 31 for additional information.
2.5.2 Segment-Override Prefixes in 64-Bit Mode
In 64-bit mode, the DS, ES, SS, and CS segment-override prefixes have no effect. These four prefixes
are no longer treated as segment-override prefixes in the context of multiple-prefix rules. Instead, they
are treated as null prefixes.
The FS and GS segment-override prefixes are treated as segment-override prefixes in 64-bit mode. Use
of the FS and GS prefixes cause their respective segment bases to be added to the effective address
calculation. See “FS and GS Registers in 64-Bit Mode” on page 70 for additional information on using
these segment registers.
2.5.3 Operands and Results
The AMD64 architecture provides support for using 64-bit operands and generating 64-bit results
when operating in 64-bit mode. See “Operands” in Volume 1 for details.
Operand-Size Overrides. In 64-bit mode, the default operand size is 32 bits. A REX prefix can be
used to specify a 64-bit operand size. Software uses a legacy operand-size (66h) prefix to toggle to 16bit operand size. The REX prefix takes precedence over the legacy operand-size prefix.
Zero Extension of Results. In 64-bit mode, when performing 32-bit operations with a GPR
destination, the processor zero-extends the 32-bit result into the full 64-bit destination. Both 8-bit and
16-bit operations on GPRs preserve all unwritten upper bits of the destination GPR. This is consistent
with legacy 16-bit and 32-bit semantics for partial-width results.
2.5.4 Address Calculations
The AMD64 architecture modifies aspects of effective-address calculation to support 64-bit mode.
These changes are summarized in the following sections. See “Memory Addressing” in Volume 1 for
details.
30x86 and AMD64 Architecture Differences
Page 75
24593—Rev. 3.14—September 2007AMD64 Technology
Address-Size Overrides. In 64-bit mode, the default-address size is 64 bits. The address size can be
overridden to 32 bits by using the address-size prefix (67h). 16-bit addresses are not supported in 64bit mode. In compatibility mode and legacy mode, address-size overrides function the same as in x86
legacy architecture.
Displacements and Immediates. Generally, displacement and immediate values in 64-bit mode are
not extended to 64 bits. They are still limited to 32 bits and are sign extended during effective-address
calculations. In 64-bit mode, however, support is provided for some 64-bit displacement and
immediate forms of the MOV instruction.
Zero Extending 16-Bit and 32-Bit Addresses. All 16-bit and 32-bit address calculations are zero-
extended in long mode to form 64-bit addresses. Address calculations are first truncated to the
effective-address size of the current mode (64-bit mode or compatibility mode), as overridden by any
address-size prefix. The result is then zero-extended to the full 64-bit address width.
RIP-Relative Addressing. A new addressing form, RIP-relative (instruction-pointer relative)
addressing, is implemented in 64-bit mode. The effective address is formed by adding the
displacement to the 64-bit RIP of the next instruction.
2.5.5 Instructions that Reference RSP
With the exception of far branches, all instructions that implicitly reference the 64-bit stack pointer,
RSP, default to a 64-bit operand size in 64-bit mode (see Table 2-1 for a listing). Pushes and pops of
32-bit stack values are not possible in 64-bit mode with these instructions, but they can be overridden
to 16 bits.
Table 2-1.Instructions That Reference RSP
Mnemonic
ENTERC8Create Procedure Stack Frame
LEAVEC9Delete Procedure Stack Frame
POP reg/mem8F/0Pop Stack (register or memory)
POP reg58-5FPop Stack (register)
POP FS0F A1Pop Stack into FS Segment Register
POP GS0F A9Pop Stack into GS Segment Register
POPF, POPFD, POPFQ9DPop to rFLAGS Word, Doubleword, or Quadword
PUSH reg/memFF/6Push onto Stack (register or memory)
PUSH reg50-57Push onto Stack (register)
PUSH FS0F A0Push FS Segment Register onto Stack
PUSH GS0F A8Push GS Segment Register onto Stack
PUSHF, PUSHFD, PUSHFQ9CPush rFLAGS Word, Doubleword, or Quadword onto Stack
Opcode
(hex)
Description
x86 and AMD64 Architecture Differences31
Page 76
AMD64 Technology24593—Rev. 3.14—September 2007
2.5.6 Branches
The AMD64 architecture expands two branching mechanisms to accommodate branches in the full 64bit virtual-address space:
•In 64-bit mode, near-branch semantics are redefined.
•In both 64-bit and compatibility modes, a 64-bit call-gate descriptor is defined for far calls.
In addition, enhancements are made to the legacy SYSCALL and SYSRET instructions.
Near Branches. In 64-bit mode, the operand size for all near branches defaults to 64 bits (see
Table 2-2 for a listing). Therefore, these instructions update the full 64-bit RIP without the need for a
REX operand-size prefix. The following aspects of near branches default to 64 bits:
•Truncation of the instruction pointer.
•Size of a stack pop or stack push, resulting from a CALL or RET.
•Size of a stack-pointer increment or decrement, resulting from a CALL or RET.
•Size of operand fetched by indirect-branch operand size.
The operand size for near branches can be overridden to 16 bits in 64-bit mode.
Table 2-2.64-Bit Mode Near Branches, Default 64-Bit Operand Size
Mnemonic
CALLE8, FF/2Call Procedure Near
JccmanyJump Conditional Near
JMPE9, EB, FF/4Jump Near
LOOPE2Loop
LOOPccE0, E1Loop Conditional
RETC3, C2Return From Call (near)
Opcode
(hex)
Description
The address size of near branches is not forced in 64-bit mode. Such addresses are 64 bits by default,
but they can be overridden to 32 bits by a prefix.
The size of the displacement field for relative branches is still limited to 32 bits.
Far Branches Through Long-Mode Call Gates. Long mode redefines the 32-bit call-gate
descriptor type as a 64-bit call-gate descriptor and expands the call-gate descriptor size to hold a 64-bit
offset. The long-mode call-gate descriptor allows far branches to reference any location in the
supported virtual-address space. In long mode, the call-gate mechanism is changed as follows:
•In long mode, CALL and JMP instructions that reference call-gates must reference 64-bit call
gates.
•A 64-bit call-gate descriptor must reference a 64-bit code-segment.
32x86 and AMD64 Architecture Differences
Page 77
24593—Rev. 3.14—September 2007AMD64 Technology
•When a control transfer is made through a 64-bit call gate, the 64-bit target address is read from the
64-bit call-gate descriptor. The base address in the target code-segment descriptor is ignored.
Stack Switching. Automatic stack switching is also modified when a control transfer occurs through
a call gate in long mode:
•The target-stack pointer read from the TSS is a 64-bit RSP value.
•The SS register is loaded with a null selector. Setting the new SS selector to null allows nested
control transfers in 64-bit mode to be handled properly. The SS.RPL value is updated to remain
consistent with the newly loaded CPL value.
•The size of pushes onto the new stack is modified to accommodate the 64-bit RIP and RSP values.
•Automatic parameter copying is not supported in long mode.
Far Returns. In long mode, far returns can load a null SS selector from the stack under the following
conditions:
•The target operating mode is 64-bit mode.
•The target CPL<3.
Allowing RET to load SS with a null selector under these conditions makes it possible for the
processor to unnest far CALLs (and interrupts) in long mode.
Task Gates. Control transfers through task gates are not supported in long mode.
Branches to 64-Bit Offsets. Because immediate values are generally limited to 32 bits, the only way
a full 64-bit absolute RIP can be specified in 64-bit mode is with an indirect branch. For this reason,
direct forms of far branches are eliminated from the instruction set in 64-bit mode.
SYSCALL and SYSRET Instructions. The AMD64 architecture expands the function of the legacy
SYSCALL and SYSRET instructions in long mode. In addition, two new STAR registers, LSTAR and
CSTAR, are provided to hold the 64-bit target RIP for the instructions when they are executed in long
mode. The legacy STAR register is not expanded in long mode. See “SYSCALL and SYSRET” on
page 150 for additional information.
SWAPGS Instruction. The AMD64 architecture provides the SWAPGS instruction as a fast method
for system software to load a pointer to system data-structures. SWAPGS is valid only in 64-bit mode.
An undefined-opcode exception (#UD) occurs if software attempts to execute SWAPGS in legacy
mode or compatibility mode. See “SWAPGS Instruction” on page 152 for additional information.
SYSENTER and SYSEXIT Instructions. The SYSENTER and SYSEXIT instructions are invalid in
long mode, and result in an invalid opcode exception (#UD) if software attempts to use them. Software
should use the SYSCALL and SYSRET instructions when running in long mode. See “SYSENTER
and SYSEXIT (Legacy Mode Only)” on page 152 for additional information.
x86 and AMD64 Architecture Differences33
Page 78
AMD64 Technology24593—Rev. 3.14—September 2007
2.5.7 NOP Instruction
The legacy x86 architecture commonly uses opcode 90h as a one-byte NOP. In 64-bit mode, the
processor treats opcode 90h specially in order to preserve this NOP definition. This is necessary
because opcode 90h is actually the XCHG EAX, EAX instruction in the legacy architecture. Without
special handling in 64-bit mode, the instruction would not be a true no-operation. Therefore, in 64-bit
mode the processor treats opcode 90h (the legacy XCHG EAX, EAX instruction) as a true NOP,
regardless of a REX operand-size prefix.
This special handling does not apply to the two-byte ModRM form of the XCHG instruction. Unless a
64-bit operand size is specified using a REX prefix byte, using the two-byte form of XCHG to
exchange a register with itself does not result in a no-operation, because the default operation size is 32
bits in 64-bit mode.
2.5.8 Single-Byte INC and DEC Instructions
In 64-bit mode, the legacy encodings for the 16 single-byte INC and DEC instructions (one for each of
the eight GPRs) are used to encode the REX prefix values. The functionality of these INC and DEC
instructions is still available, however, using the ModRM forms of those instructions (opcodes FF /0
and FF /1). See “Single-Byte INC and DEC Instructions in 64-Bit Mode” in Volume 3 for additional
information.
2.5.9 MOVSXD Instruction
MOVSXD is a new instruction in 64-bit mode (the legacy ARPL instruction opcode, 63h, is reassigned
as the MOVSXD opcode). It reads a fixed-size 32-bit source operand from a register or memory and (if
a REX prefix is used with the instruction) sign-extends the value to 64 bits. MOVSXD is analogous to
the MOVSX instruction, which sign-extends a byte to a word or a word to a doubleword, depending on
the effective operand size. See “General-Purpose Instruction Reference” in Volume 3 for additional
information.
2.5.10 Invalid Instructions
Table 2-3 lists instructions that are illegal in 64-bit mode. Table 2-4 on page 35 lists instructions that
are invalid in long mode (both compatibility mode and 64-bit mode). Attempted use of these
instructions causes an invalid-opcode exception (#UD) to occur.
Table 2-3.Invalid Instructions in 64-Bit Mode
Mnemonic
AAA37ASCII Adjust After Addition
AADD5ASCII Adjust Before Division
AAMD4ASCII Adjust After Multiply
AAS3FASCII Adjust After Subtraction
BOUND62Check Array Bounds
Opcode
(hex)
Description
34x86 and AMD64 Architecture Differences
Page 79
24593—Rev. 3.14—September 2007AMD64 Technology
Table 2-3.Invalid Instructions in 64-Bit Mode (continued)
Mnemonic
CALL (far)9AProcedure Call Far (absolute)
DAA27Decimal Adjust after Addition
DAS2FDecimal Adjust after Subtraction
INTOCEInterrupt to Overflow Vector
JMP (far)EAJump Far (absolute)
LDSC5Load DS Segment Register
LESC4Load ES Segment Register
POP DS1FPop Stack into DS Segment
POP ES07Pop Stack into ES Segment
POP SS17Pop Stack into SS Segment
POPA, POPAD61Pop All to GPR Words or Doublewords
PUSH CS0EPush CS Segment Selector onto Stack
PUSH DS1EPush DS Segment Selector onto Stack
PUSH ES06Push ES Segment Selector onto Stack
PUSH SS16Push SS Segment Selector onto Stack
PUSHA,
PUSHAD
Redundant Grp1
(undocumented)
SALC
(undocumented)
Opcode
(hex)
60
82
D6Set AL According to CF
Push All GPR Words or Doublewords onto
Stack
Redundant encoding of group1 Eb,Ib
opcodes
Description
Table 2-4.Invalid Instructions in Long Mode
Mnemonic
SYSENTER0F 34System Call
SYSEXIT0F 35System Return
Opcode
(hex)
Description
Table 2-5 on page 36 lists the instructions that are no longer valid in 64-bit mode because their
opcodes have been reassigned. The reassigned opcodes are used in 64-bit mode as REX instruction
prefixes.
x86 and AMD64 Architecture Differences35
Page 80
AMD64 Technology24593—Rev. 3.14—September 2007
Table 2-5.Reassigned Instructions in 64-Bit Mode
Mnemonic
ARPL63
DEC and INC40-4F
Opcode
(hex)
Description
Opcode for MOVSXD instruction in 64-bit
mode. In all other modes, this the Adjust
Requestor Privilege Level instruction opcode.
Decrement by 1, Increment by 1. Two-byte
versions of DEC and INC are still valid.
2.5.11 FXSAVE and FXRSTOR Instructions
The FXSAVE and FXRSTOR instructions are used to save and restore the entire 128-bit media, 64-bit
media, and x87 instruction-set environment during a context switch. The AMD64 architecture
modifies the memory format used by these instructions in order to save and restore the full 64-bit
instruction and data pointers, as well as the XMM8–XMM15 registers. Selection of the 32-bit legacy
format or the expanded 64-bit format is accomplished by using the corresponding operand size with
the FXSAVE and FXRSTOR instructions. When 64-bit software executes an FXSAVE and FXRSTOR
with a 32-bit operand size (no operand-size override) the 32-bit legacy format is used. When 64-bit
software executes an FXSAVE and FXRSTOR with a 64-bit operand size, the 64-bit format is used.
If the fast-FXSAVE/FXRSTOR (FFXSR) feature is enabled in EFER, FXSAVE and FXRSTOR do not
save or restore the XMM0-XMM15 registers when executed in 64-bit mode at CPL 0. The x87
environment and MXCSR are saved whether fast-FXSAVE/FXRSTOR is enabled or not. Software
can use CPUID to determine whether the fast-FXSAVE/FXRSTOR feature is available (CPUID
function 8000_0001h, EDX bit 25). The fast-FXSAVE/FXRSTOR feature has no effect on
FXSAVE/FXRSTOR in non 64-bit mode or when CPL > 0.
2.6Interrupts and Exceptions
When a processor is running in long mode, an interrupt or exception causes the processor to enter 64bit mode. All long-mode interrupt handlers must be implemented as 64-bit code. The AMD64
architecture expands the legacy interrupt-processing and exception-processing mechanism to support
handling of interrupts by 64-bit operating systems and applications. The changes are summarized in
the following sections. See “Long-Mode Interrupt Control Transfers” on page 241 for detailed
information on these changes.
2.6.1 Interrupt Descriptor Table
The long-mode interrupt-descriptor table (IDT) must contain 64-bit mode interrupt-gate or trap-gate
descriptors for all interrupts or exceptions that can occur while the processor is running in long mode.
Task gates cannot be used in the long-mode IDT, because control transfers through task gates are not
supported in long mode. In long mode, the IDT index is formed by scaling the interrupt vector by 16.
In legacy protected mode, the IDT is indexed by scaling the interrupt vector by eight.
36x86 and AMD64 Architecture Differences
Page 81
24593—Rev. 3.14—September 2007AMD64 Technology
2.6.2 Stack Frame Pushes
In legacy mode, the size of an IDT entry (16 bits or 32 bits) determines the size of interrupt-stackframe pushes, and SS:eSP is pushed only on a CPL change. In long mode, the size of interrupt stackframe pushes is fixed at eight bytes, because interrupts are handled in 64-bit mode. Long mode
interrupts also cause SS:RSP to be pushed unconditionally, rather than pushing only on a CPL change.
2.6.3 Stack Switching
Legacy mode provides a mechanism to automatically switch stack frames in response to an interrupt.
In long mode, a slightly modified version of the legacy stack-switching mechanism is implemented,
and an alternative stack-switching mechanism—called the interrupt stack table (IST)—is supported.
Long-Mode Stack Switches. When stacks are switched as part of a long-mode privilege-level
change resulting from an interrupt, the following occurs:
•The target-stack pointer read from the TSS is a 64-bit RSP value.
•The SS register is loaded with a null selector. Setting the new SS selector to null allows nested
control transfers in 64-bit mode to be handled properly. The SS.RPL value is cleared to 0.
•The old SS and RSP are saved on the new stack.
Interrupt Stack Table. In long mode, a new interrupt stack table (IST) mechanism is available as an
alternative to the modified legacy stack-switching mechanism. The IST mechanism unconditionally
switches stacks when it is enabled. It can be enabled for individual interrupt vectors using a field in the
IDT entry. This allows mixing interrupt vectors that use the modified legacy mechanism with vectors
that use the IST mechanism. The IST pointers are stored in the long-mode TSS. The IST mechanism is
only available when long mode is enabled.
2.6.4 IRET Instruction
In compatibility mode, IRET pops SS:eSP off the stack only if there is a CPL change. This allows
legacy applications to run properly in compatibility mode when using the IRET instruction.
In 64-bit mode, IRET unconditionally pops SS:eSP off of the interrupt stack frame, even if the CPL
does not change. This is done because the original interrupt always pushes SS:RSP. Because interrupt
stack-frame pushes are always eight bytes in long mode, an IRET from a long-mode interrupt handler
(64-bit code) must pop eight-byte items off the stack. This is accomplished by preceding the IRET
with a 64-bit REX operand-size prefix.
In long mode, an IRET can load a null SS selector from the stack under the following conditions:
•The target operating mode is 64-bit mode.
•The target CPL<3.
Allowing IRET to load SS with a null selector under these conditions makes it possible for the
processor to unnest interrupts (and far CALLs) in long mode.
x86 and AMD64 Architecture Differences37
Page 82
AMD64 Technology24593—Rev. 3.14—September 2007
2.6.5 Task-Priority Register (CR8)
The AMD64 architecture allows software to define up to 15 external interrupt-priority classes. Priority
classes are numbered from 1 to 15, with priority-class 1 being the lowest and priority-class 15 the
highest.
A new control register (CR8) is introduced by the AMD64 architecture for managing priority classes.
This register, also called the task-priority register (TPR), uses the four low-order bits for specifying a
task priority. How external interrupts are organized into these priority classes is implementation
dependent. See “External Interrupt Priorities” on page 228 for information on this feature.
2.6.6 New Exception Conditions
The AMD64 architecture defines a number of new conditions that can cause an exception to occur
when the processor is running in long mode. Many of the conditions occur when software attempts to
use an address that is not in canonical form. See “Vectors” on page 208 for information on the new
exception conditions that can occur in long mode.
2.7Hardware Task Switching
The legacy hardware task-switch mechanism is disabled when the processor is running in long mode.
However, long mode requires system software to create data structures for a single task—the longmode task.
•TSS Descriptors—A new TSS-descriptor type, the 64-bit TSS type, is defined for use in long
mode. It is the only valid TSS type that can be used in long mode, and it must be loaded into the TR
by executing the LTR instruction in 64-bit mode. See “TSS Descriptor” on page 310 for additional
information.
•Task Gates—Because the legacy task-switch mechanism is not supported in long mode, software
cannot use task gates in long mode. Any attempt to transfer control to another task through a task
gate causes a general-protection exception (#GP) to occur.
•Task-State Segment—A 64-bit task state segment (TSS) is defined for use in long mode. This new
TSS format contains 64-bit stack pointers (RSP) for privilege levels 0–2, interrupt-stack-table
(IST) pointers, and the I/O-map base address. See “64-Bit Task State Segment” on page 317 for
additional information.
2.8Long-Mode vs. Legacy-Mode Differences
Table 2-6 on page 39 summarizes several major system-programming differences between 64-bit
mode and legacy protected mode. The third column indicates whether the difference also applies to
compatibility mode. “Differences Between Long Mode and Legacy Mode” in Volume 3 summarizes
the application-programming model differences.
38x86 and AMD64 Architecture Differences
Page 83
24593—Rev. 3.14—September 2007AMD64 Technology
Table 2-6.Differences Between Long Mode and Legacy Mode
Applies To
Subject64-Bit Mode Difference
x86 ModesReal and virtual-8086 modes not supported
Task SwitchingTask switching not supported
64-bit virtual addresses
Addressing
Loaded Segment (Usage
during memory reference)
Exception and Interrupt
Handling
Call Gates
System-Descriptor
Registers
System-Descriptor Table
Entries and PseudoDescriptors
4-level paging structures
PAE must always be enabled
CS, DS, ES, SS segment bases are ignored
CS, DS, ES, FS, GS, SS segment limits are ignored
DS, ES, FS, GS attribute are ignored
CS, DS, ES, SS Segment prefixes are ignored
All pushes are 8 bytes
IDT entries are expanded to 16 bytes
SS is not changed for stack switch
SS:RSP is pushed unconditionally
All pushes are 8 bytes
16-bit call gates are illegal
32-bit call gate type is redefined as 64-bit call gate and is
expanded to 16 bytes
SS is not changed for stack switch
GDT, IDT, LDT, TR base registers expanded to 64 bits
LGDT and LIDT use expanded 10-byte pseudo-descriptors
LLDT and LTR use expanded 16-byte table entries
Compatibility
Mode?
Ye s
Ye s
No
Ye s
No
Ye s
Ye s
Ye s
No
x86 and AMD64 Architecture Differences39
Page 84
AMD64 Technology24593—Rev. 3.14—September 2007
40x86 and AMD64 Architecture Differences
Page 85
24593—Rev. 3.14—September 2007AMD64 Technology
3System Resources
The operating system manages the software-execution environment and general system operation
through the use of system resources. These resources consist of system registers (control registers and
model-specific registers) and system-data structures (memory-management and protection tables).
The system-control registers are described in detail in this chapter; many of the features they control
are described elsewhere in this volume. The model-specific registers supported by the AMD64
architecture are introduced in this chapter.
Because of their complexity, system-data structures are described in separate chapters. Refer to the
following chapters for detailed information on these data structures:
•Descriptors and descriptor tables are described in “Segmentation Data Structures and Registers”
on page 65.
•Page-translation tables are described in “Legacy-Mode Page Translation” on page 120 and “Long-
Mode Page Translation” on page 128.
•The task-state segment is described in “Legacy Task-State Segment” on page 313 and “64-Bit Task
State Segment” on page 317.
Not all processor implementations are required to support all possible features. The last section in this
chapter addresses processor-feature identification. System software uses the capabilities described in
that section to determine which features are supported so that the appropriate service routines are
loaded.
3.1System-Control Registers
The registers that control the AMD64 architecture operating environment include:
•CR0—Provides operating-mode controls and some processor-feature controls.
•CR2—This register is used by the page-translation mechanism. It is loaded by the processor with
the page-fault virtual address when a page-fault exception occurs.
•CR3—This register is also used by the page-translation mechanism. It contains the base address of
the highest-level page-translation table, and also contains cache controls for the specified table.
•CR4—This register contains additional controls for various operating-mode features.
•CR8—This new register, accessible in 64-bit mode using the REX prefix, is introduced by the
AMD64 architecture. CR8 is used to prioritize external interrupts and is referred to as the taskpriority register (TPR).
•RFLAGS—This register contains processor-status and processor-control fields. The status and
control fields are used primarily in the management of virtual-8086 mode, hardware multitasking,
and interrupts.
System Resources41
Page 86
AMD64 Technology24593—Rev. 3.14—September 2007
•EFER—This model-specific register contains status and controls for additional features not
managed by the CR0 and CR4 registers. Included in this register are the long-mode enable and
activation controls introduced by the AMD64 architecture.
Control registers CR1, CR5–CR7, and CR9–CR15 are reserved.
In legacy mode, all control registers and RFLAGS are 32 bits. The EFER register is 64 bits in all
modes. The AMD64 architecture expands all 32-bit system-control registers to 64 bits. In 64-bit mode,
the MOV CRn instructions read or write all 64 bits of these registers (operand-size prefixes are
ignored). In compatibility and legacy modes, control-register writes fill the low 32 bits with data and
the high 32 bits with zeros, and control-register reads return only the low 32 bits.
In 64-bit mode, the high 32 bits of CR0 and CR4 are reserved and must be written with zeros. Writing
a 1 to any of the high 32 bits results in a general-protection exception, #GP(0). All 64 bits of CR2 are
writable. However, the MOV CRn instructions do not check that addresses written to CR2 are within
the virtual-address limitations of the processor implementation.
All CR3 bits are writable, except for unimplemented physical address bits, which must be cleared to 0.
The upper 32 bits of RFLAGS are always read as zero by the processor. Attempts to load the upper 32
bits of RFLAGS with anything other than zero are ignored by the processor.
3.1.1 CR0 Register
The CR0 register is shown in Figure 3-1 on page 43. The legacy CR0 register is identical to the low 32
bits of the register shown in Figure 3-1 on page 43 (CR0 bits 31–0).
42System Resources
Page 87
24593—Rev. 3.14—September 2007AMD64 Technology
6332
Reserved, MBZ
3130292819181716156543210
A
PGCDN
W
BitsMnemonicDescriptionR/W
63–32ReservedReserved, Must be Zero
31PGPagingR/W
30CDCache DisableR/W
29NWNot WritethroughR/W
28–19ReservedReserved
18AMAlignment MaskR/W
17ReservedReserved
16WPWrite ProtectR/W
15-6ReservedReserved
5NENumeric ErrorR/W
4ETExtension TypeR
3TSTask SwitchedR/W
2EMEmulationR/W
1MPMonitor CoprocessorR/W
0PEProtection EnabledR/W
Reserved
W
R
M
P
Reserved
NEETTSEMMPP
E
Figure 3-1.Control Register 0 (CR0)
The functions of the CR0 control bits are (unless otherwise noted, all bits are read/write):
Protected-Mode Enable (PE) Bit. Bit 0. Software enables protected mode by setting PE to 1, and
disables protected mode by clearing PE to 0. When the processor is running in protected mode,
segment-protection mechanisms are enabled.
See “Segment-Protection Overview” on page 93 for information on the segment-protection
mechanisms.
Monitor Coprocessor (MP) Bit. Bit 1. Software uses the MP bit with the task-switched control bit
(CR0.TS) to control whether execution of the WAIT/FWAIT instruction causes a device-not-available
exception (#NM) to occur, as follows:
•If both the monitor-coprocessor and task-switched bits are set (CR0.MP=1 and CR0.TS=1), then
executing the WAIT/FWAIT instruction causes a device-not-available exception (#NM).
•If either the monitor-coprocessor or task-switched bits are clear (CR0.MP=0 or CR0.TS=0), then
executing the WAIT/FWAIT instruction proceeds normally.
System Resources43
Page 88
AMD64 Technology24593—Rev. 3.14—September 2007
Software typically should set MP to 1 if the processor implementation supports x87 instructions. This
allows the CR0.TS bit to completely control when the x87-instruction context is saved as a result of a
task switch.
Emulate Coprocessor (EM) Bit. Bit 2. Software forces all x87 instructions to cause a device-not-
available exception (#NM) by setting EM to 1. Likewise, setting EM to 1 forces an invalid-opcode
exception (#UD) when an attempt is made to execute any of the 64-bit or 128-bit media instructions.
The exception handlers can emulate these instruction types if desired. Setting the EM bit to 1 does not
cause an #NM exception when the WAIT/FWAIT instruction is executed.
Task Switched (TS) Bit. Bit 3. When an attempt is made to execute an x87 or media instruction while
TS=1, a device-not-available exception (#NM) occurs. Software can use this mechanism—sometimes
referred to as “lazy context-switching”—to save the unit contexts before executing the next instruction
of those types. As a result, the x87 and media instruction-unit contexts are saved only when necessary
as a result of a task switch.
When a hardware task switch occurs, TS is automatically set to 1. System software that implements
software task-switching rather than using the hardware task-switch mechanism can still use the TS bit
to control x87 and media instruction-unit context saves. In this case, the task-management software
uses a MOV CR0 instruction to explicitly set the TS bit to 1 during a task switch. Software can clear
the TS bit by either executing the CLTS instruction or by writing to the CR0 register directly. Longmode system software can use this approach even though the hardware task-switch mechanism is not
supported in long mode.
The CR0.MP bit controls whether the WAIT/FWAIT instruction causes an #NM exception when
TS=1.
Extension Type (ET) Bit. Bit 4, read-only. In some early x86 processors, software set ET to 1 to
indicate support of the 387DX math-coprocessor instruction set. This bit is now reserved and forced to
1 by the processor. Software cannot clear this bit to 0.
Numeric Error (NE) Bit. Bit 5. Clearing the NE bit to 0 disables internal control of x87 floating-point
exceptions and enables external control. When NE is cleared to 0, the IGNNE# input signal controls
whether x87 floating-point exceptions are ignored:
•When IGNNE# is 1, x87 floating-point exceptions are ignored.
•When IGNNE# is 0, x87 floating-point exceptions are reported by setting the FERR# input signal
to 1. External logic can use the FERR# signal as an external interrupt.
When NE is set to 1, internal control over x87 floating-point exception reporting is enabled and the
external reporting mechanism is disabled. It is recommended that software set NE to 1. This enables
optimal performance in handling x87 floating-point exceptions.
Write Protect (WP) Bit. Bit 16. Read-only pages are protected from supervisor-level writes when the
WP bit is set to 1. When WP is cleared to 0, supervisor software can write into read-only pages.
See “Page-Protection Checks” on page 142 for information on the page-protection mechanism.
44System Resources
Page 89
24593—Rev. 3.14—September 2007AMD64 Technology
Alignment Mask (AM) Bit. Bit 18. Software enables automatic alignment checking by setting the
AM bit to 1 when eFLAGS.AC=1. Alignment checking can be disabled by clearing either AM or
eFLAGS.AC to 0. When automatic alignment checking is enabled and CPL=3, a memory reference to
an unaligned operand causes an alignment-check exception (#AC).
Not Writethrough (NW) Bit. Bit 29. Ignored. This bit can be set to 1 or cleared to 0, but its value is
ignored. The NW bit exists only for legacy purposes.
Cache Disable (CD) Bit. Bit 30. When CD is cleared to 0, the internal caches are enabled. When CD
is set to 1, no new data or instructions are brought into the internal caches. However, the processor still
accesses the internal caches when CD=1 under the following situations:
•Reads that hit in an internal cache cause the data to be read from the internal cache that reported the
hit.
•Writes that hit in an internal cache cause the cache line that reported the hit to be written back to
memory and invalidated in the cache.
Cache misses do not affect the internal caches when CD=1. Software can prevent cache access by
writing back and invalidating the caches before setting CD to 1 (this avoids caching the instructions
that set CD to 1).
Setting CD to 1 also causes the processor to ignore the page-level cache-control bits (PWT and PCD)
when paging is enabled. These bits are located in the page-translation tables and CR3 register. See
“Page-Level Writethrough (PWT) Bit” on page 137 and “Page-Level Cache Disable (PCD) Bit” on
page 137 for information on page-level cache control.
See “Memory Caches” on page 176 for information on the internal caches.
Paging Enable (PG) Bit. Bit 31. Software enables page translation by setting PG to 1, and disables
page translation by clearing PG to 0. Page translation cannot be enabled unless the processor is in
protected mode (CR0.PE=1). If software attempts to set PG to 1 when PE is cleared to 0, the processor
causes a general-protection exception (#GP).
See “Page Translation Overview” on page 115 for information on the page-translation mechanism.
Reserved Bits. Bits 28–19, 17, 15–6, and 63–32. When writing the CR0 register, software should set
the values of reserved bits to the values found during the previous CR0 read. No attempt should be
made to change reserved bits, and software should never rely on the values of reserved bits. In long
mode, bits 63–32 are reserved and must be written with zero, otherwise a #GP occurs.
3.1.2 CR2 and CR3 Registers
The CR2 (page-fault linear address) register, shown in Figure 3-2 on page 46 and Figure 3-3 on
page 46, and the CR3 (page-translation-table base address) register, shown in Figure 3-4 and
Figure 3-5 on page 46, and Figure 3-6 on page 46, are used only by the page-translation mechanism.
System Resources45
Page 90
AMD64 Technology24593—Rev. 3.14—September 2007
310
Page-Fault Virtual Address
Figure 3-2.Control Register 2 (CR2)—Legacy-Mode
6332
Page-Fault Virtual Address
310
Page-Fault Virtual Address
Figure 3-3.Control Register 2 (CR2)—Long Mode
See “CR2 Register” on page 220 for a description of the CR2 register.
The CR3 register is used to point to the base address of the highest-level page-translation table.
(This is an architectural limit. A given implementation may support fewer bits.)
Page-Map Level-4 Table Base Address
3112 1154320
P
P
Page-Map Level-4 Table Base AddressReserved
C
D
Reserved
W
T
Figure 3-6.Control Register 3 (CR3)—Long Mode
46System Resources
Page 91
24593—Rev. 3.14—September 2007AMD64 Technology
The legacy CR3 register is described in “CR3 Register” on page 120, and the long-mode CR3 register
is described in “CR3” on page 128.
3.1.3 CR4 Register
The CR4 register is shown in Figure 3-7. In legacy mode, the CR4 register is identical to the low 32
bits of the register (CR4 bits 31–0). The features controlled by the bits in the CR4 register are modelspecific extensions. Except for the performance-counter extensions (PCE) feature, software can use
the CPUID instruction to verify that each feature is supported before using that feature.
6332
Reserved, MBZ
311110 9 876543210
Reserved, MBZ
O
S
X
OSF
XSR
P
P
M
P
P
T
P
C
G
C
E
E
E
D
A
S
S
E
E
E
D
V
V
M
I
E
BitsMnemonicDescriptionR/W
63–11 ReservedReserved, Must be Zero
10OSXMMEXCPTOperating System Unmasked Exception SupportR/W
9OSFXSROperating System FXSAVE/FXRSTOR SupportR/W
8PCEPerformance-Monitoring Counter EnableR/W
7PGEPage-Global EnableR/W
6MCEMachine Check EnableR/W
5PAEPhysical-Address ExtensionR/W
4PSEPage Size ExtensionsR/W
3DEDebugging ExtensionsR/W
2TSDTime Stamp DisableR/W
1PVIProtected-Mode Virtual InterruptsR/W
0VMEVirtual-8086 Mode ExtensionsR/W
Figure 3-7.Control Register 4 (CR4)
The function of the CR4 control bits are (all bits are read/write):
Virtual-8086 Mode Extensions (VME) Bit. Bit 0. Setting VME to 1 enables hardware-supported
performance enhancements for software running in virtual-8086 mode. Clearing VME to 0 disables
this support. The enhancements enabled when VME=1 include:
•Virtualized, maskable, external-interrupt control and notification using the VIF and VIP bits in the
rFLAGS register. Virtualizing affects the operation of several instructions that manipulate the
rFLAGS.IF bit.
•Selective intercept of software interrupts (INTn instructions) using the interrupt-redirection bitmap
in the TSS.
System Resources47
Page 92
AMD64 Technology24593—Rev. 3.14—September 2007
Protected-Mode Virtual Interrupts (PVI) Bit. Bit 1. Setting PVI to 1 enables support for protected-
mode virtual interrupts. Clearing PVI to 0 disables this support. When PVI=1, hardware support of two
bits in the rFLAGS register, VIF and VIP, is enabled.
Only the STI and CLI instructions are affected by enabling PVI. Unlike the case when CR0.VME=1,
the interrupt-redirection bitmap in the TSS cannot be used for selective INTn interception.
PVI enhancements are also supported in long mode. See “Virtual Interrupts” on page 247 for more
information on using PVI.
Time-Stamp Disable (TSD) Bit. Bit 2. The TSD bit allows software to control the privilege level at
which the time-stamp counter can be read. When TSD is cleared to 0, software running at any privilege
level can read the time-stamp counter using the RDTSC or RDTSCP instructions. When TSD is set to
1, only software running at privilege-level 0 can execute the RDTSC or RDTSCP instructions.
Debugging Extensions (DE) Bit. Bit 3. Setting the DE bit to 1 enables the I/O breakpoint capability
and enforces treatment of the DR4 and DR5 registers as reserved. Software that accesses DR4 or DR5
when DE=1 causes a invalid opcode exception (#UD).
When the DE bit is cleared to 0, I/O breakpoint capabilities are disabled. Software references to the
DR4 and DR5 registers are aliased to the DR6 and DR7 registers, respectively.
Page-Size Extensions (PSE) Bit. Bit 4. Setting PSE to 1 enables the use of 4-Mbyte physical pages.
With PSE=1, the physical-page size is selected between 4 Kbytes and 4 Mbytes using the pagedirectory entry page-size field (PS). Clearing PSE to 0 disables the use of 4-Mbyte physical pages and
restricts all physical pages to 4 Kbytes.
The PSE bit has no effect when physical-address extensions are enabled (CR4.PAE=1). Because long
mode requires CR4.PAE=1, the PSE bit is ignored when the processor is running in long mode.
See “4-Mbyte Page Translation” on page 123 for more information on 4-Mbyte page translation.
Physical-Address Extension (PAE) Bit. Bit 5. Setting PAE to 1 enables the use of physical-address
extensions and 2-Mbyte physical pages. Clearing PAE to 0 disables these features.
With PAE=1, the page-translation data structures are expanded from 32 bits to 64 bits, allowing the
translation of up to 52-bit physical addresses. Also, the physical-page size is selectable between
4 Kbytes and 2 Mbytes using the page-directory-entry page-size field (PS). Long mode requires PAE
to be enabled in order to use the 64-bit page-translation data structures to translate 64-bit virtual
addresses to 52-bit physical addresses.
See “PAE Paging” on page 124 for more information on physical-address extensions.
Machine-Check Enable (MCE) Bit. Bit 6. Setting MCE to 1 enables the machine-check exception
mechanism. Clearing this bit to 0 disables the mechanism. When enabled, a machine-check exception
(#MC) occurs when an uncorrectable machine-check error is encountered.
48System Resources
Page 93
24593—Rev. 3.14—September 2007AMD64 Technology
Regardless of whether machine-check exceptions are enabled, the processor records enabled-errors
when they occur. Error-reporting is performed by the machine-check error-reporting register banks.
Each bank includes a control register for enabling error reporting and a status register for capturing
errors. Correctable machine-check errors are also reported, but they do not cause a machine-check
exception.
See Chapter 9, “Machine Check Mechanism,” for a description of the machine-check mechanism, the
registers used, and the types of errors captured by the mechanism.
Page-Global Enable (PGE) Bit. Bit 7. When page translation is enabled, system-software
performance can often be improved by making some page translations global to all tasks and
procedures. Setting PGE to 1 enables the global-page mechanism. Clearing this bit to 0 disables the
mechanism.
When PGE is enabled, system software can set the global-page (G) bit in the lowest level of the pagetranslation hierarchy to 1, indicating that the page translation is global. Page translations marked as
global are not invalidated in the TLB when the page-translation-table base address (CR3) is updated.
When the G bit is cleared, the page translation is not global. All supported physical-page sizes also
support the global-page mechanism. See “Global Pages” on page 140 for information on using the
global-page mechanism.
Performance-Monitoring Counter Enable (PCE) Bit. Bit 8. Setting PCE to 1 allows software
running at any privilege level to use the RDPMC instruction. Software uses the RDPMC instruction to
read the four performance-monitoring MSRs, PerfCTR[3:0]. Clearing PCE to 0 allows only the mostprivileged software (CPL=0) to use the RDPMC instruction.
FXSAVE/FXRSTOR Support (OSFXSR) Bit. Bit 9. System software must set the OSFXSR bit to 1
to enable use of the 128-bit media instructions. When this bit is set to 1, it also indicates that system
software uses the FXSAVE and FXRSTOR instructions to save and restore the processor state for the
x87, 64-bit media, and 128-bit media instructions.
Clearing the OSFXSR bit to 0 indicates that 128-bit media instructions cannot be used. Attempts to use
those instructions while this bit is clear result in an invalid-opcode exception (#UD). Software can
continue to use the FXSAVE/FXRSTOR instructions for saving and restoring the processor state for
the x87 and 64-bit media instructions.
Unmasked Exception Support (OSXMMEXCPT) Bit. Bit 10. System software must set the
OSXMMEXCPT bit to 1 when it supports the SIMD floating-point exception (#XF) for handling of
unmasked 128-bit media floating-point errors. Clearing the OSXMMEXCPT bit to 0 indicates the
#XF handler is not supported. When OSXMMEXCPT=0, unmasked 128-bit media floating-point
exceptions cause an invalid-opcode exception (#UD). See “SIMD Floating-Point Exception Causes”
in Volume 1 for more information on 128-bit media unmasked floating-point exceptions.
3.1.4 CR1 and CR5–CR7 Registers
Control registers CR1, CR5–CR7, and CR9–CR15 are reserved. Attempts by software to use these
registers result in an undefined-opcode exception (#UD).
System Resources49
Page 94
AMD64 Technology24593—Rev. 3.14—September 2007
3.1.5 64-Bit-Mode Extended Control Registers
In 64-bit mode, additional encodings for control registers are available. The REX.R bit, in a REX
prefix, is used to modify the ModRM reg field when that field encodes a control register, as shown in
“REX Prefixes” in Volume 3. These additional encodings enable the processor to address CR8–CR15.
One additional control register, CR8, is defined in 64-bit mode for all hardware implementations, as
described in “CR8 (Task Priority Register, TPR),” below. Access to the CR9–CR15 registers is
implementation-dependent. Any attempt to access an unimplemented register results in an invalidopcode exception (#UD).
3.1.6 CR8 (Task Priority Register, TPR)
The AMD64 architecture introduces a new control register, CR8, defined as the task priority register
(TPR). The register is accessible in 64-bit mode using the REX prefix. See “External Interrupt
Priorities” on page 228 for a description of the TPR and how system software can use the TPR for
controlling external interrupts.
3.1.7 RFLAGS Register
The RFLAGS register contains two different types of information:
•Control bits provide system-software controls and directional information for string operations.
Some of these bits can have privilege-level restrictions.
•Status bits provide information resulting from logical and arithmetic operations. These are written
by the processor and can be read by software running at any privilege level.
Figure 3-8 on page 51 shows the format of the RFLAGS register. The legacy EFLAGS register is
identical to the low 32 bits of the register shown in Figure 3-8 (RFLAGS bits 31–0). The term rFLAGS
is used to refer to the 16-bit, 32-bit, or 64-bit flags register, depending on context.
50System Resources
Page 95
24593—Rev. 3.14—September 2007AMD64 Technology
6332
Reserved, RAZ
31222120191817161514131211109876543210
V
V
Reserved, RAZ
BitsMnemonicDescriptionR/W
63–22ReservedReserved, Read as Zero
21IDID FlagR/W
20VIPVirtual Interrupt PendingR/W
19VIFVirtual Interrupt FlagR/W
18ACAlignment CheckR/W
17VMVirtual-8086 ModeR/W
16RFResume FlagR/W
15ReservedReserved, Read as Zero
14NTNested TaskR/W
13-12IOPLI/O Privilege LevelR/W
11OFOverflow FlagR/W
10DFDirection FlagR/W
9IFInterrupt FlagR/W
8TFTrap FlagR/W
7SFSign FlagR/W
6ZFZero FlagR/W
5ReservedReserved, Read as Zero
4AFAuxiliary FlagR/W
3ReservedReserved, Read as Zero
2PFParity FlagR/W
1ReservedReserved, Read as One
0CFCarry FlagR/W
I
D
P
ACVMR
I
I
F
F
N
0
T
OFDFIFTFSFZ
IOPL
A
0
F
F
P
0
F
C
1
F
Figure 3-8.RFLAGS Register
The functions of the RFLAGS control and status bits used by application software are described in
“Flags Register” in Volume 1. The functions of RFLAGS system bits are (unless otherwise noted, all
bits are read/write):
Trap Flag (TF) Bit. Bit 8. Software sets the TF bit to 1 to enable single-step mode during software
debug. Clearing this bit to 0 disables single-step mode.
When single-step mode is enabled, a debug exception (#DB) occurs after each instruction completes
execution. Single stepping begins with the instruction following the instruction that sets TF. Single
stepping is disabled (TF=0) when the #DB exception occurs or when any exception or interrupt occurs.
System Resources51
Page 96
AMD64 Technology24593—Rev. 3.14—September 2007
See “Single Stepping” on page 339 for information on using the single-step mode during debugging.
Interrupt Flag (IF) Bit. Bit 9. Software sets the IF bit to 1 to enable maskable interrupts. Clearing this
bit to 0 causes the processor to ignore maskable interrupts. The state of the IF bit does not affect the
response of a processor to non-maskable interrupts, software-interrupt instructions, or exceptions.
The ability to modify the IF bit depends on several factors:
•The current privilege-level (CPL)
•The I/O privilege level (RFLAGS.IOPL)
•Whether or not virtual-8086 mode extensions are enabled (CR4.VME=1)
•Whether or not protected-mode virtual interrupts are enabled (CR4.PVI=1)
See “Masking External Interrupts” on page 207 for information on interrupt masking. See “Accessing
the RFLAGs Register” on page 154 for information on the specific instructions used to modify the IF
bit.
I/O Privilege Level Field (IOPL) Field. Bits 13–12. The IOPL field specifies the privilege level
required to execute I/O address-space instructions (i.e., instructions that address the I/O space rather
than memory-mapped I/O, such as IN, OUT, INS, OUTS, etc.). For software to execute these
instructions, the current privilege-level (CPL) must be equal to or higher than (lower numerical value
than) the privilege specified by IOPL (CPL <= IOPL). If the CPL is lower than (higher numerical value
than) that specified by the IOPL (CPL > IOPL), the processor causes a general-protection exception
(#GP) when software attempts to execute an I/O instruction. See “Protected-Mode I/O” in Volume 1
for information on how IOPL controls access to address-space I/O.
Virtual-8086 mode uses IOPL to control virtual interrupts and the IF bit when virtual-8086 mode
extensions are enabled (CR4.VME=1). The protected-mode virtual-interrupt mechanism (PVI) also
uses IOPL to control virtual interrupts and the IF bit when PVI is enabled (CR4.PVI=1). See “Virtual
Interrupts” on page 247 for information on how IOPL is used by the virtual interrupt mechanism.
Nested Task (NT) Bit. Bit 14, IRET reads the NT bit to determine whether the current task is nested
within another task. When NT is set to 1, the current task is nested within another task. When NT is
cleared to 0, the current task is at the top level (not nested).
The processor sets the NT bit during a task switch resulting from a CALL, interrupt, or exception
through a task gate. When an IRET is executed from legacy mode while the NT bit is set, a task switch
occurs. See “Task Switches Using Task Gates” on page 323 for information on switching tasks using
task gates, and “Nesting Tasks” on page 325 for information on task nesting.
Resume Flag (RF) Bit. Bit 16. The RF bit allows an instruction to be restarted following an
instruction breakpoint resulting in a debug exception (#DB). This bit prevents multiple debug
exceptions from occurring on the same instruction.
52System Resources
Page 97
24593—Rev. 3.14—September 2007AMD64 Technology
The processor clears the RF bit after every instruction is successfully executed, except when the
instruction is:
•An IRET that sets the RF bit.
•JMP, CALL, or INTn through a task gate.
In both of the above cases, RF is not cleared to 0 until the next instruction successfully executes.
When an exception occurs (or when a string instruction is interrupted), the processor normally sets
RF=1 in the rFLAGS image saved on the interrupt stack. However, when a #DB exception occurs as a
result of an instruction breakpoint, the processor clears the RF bit to 0 in the interrupt-stack rFLAGS
image.
For instruction restart to work properly following an instruction breakpoint, the #DB exception
handler must set RF to 1 in the interrupt-stack rFLAGS image. When an IRET is later executed to
return to the instruction that caused the instruction-breakpoint #DB exception, the set RF bit (RF=1) is
loaded from the interrupt-stack rFLAGS image. RF is not cleared by the processor until the instruction
causing the #DB exception successfully executes.
Virtual-8086 Mode (VM) Bit. Bit 17. Software sets the VM bit to 1 to enable virtual-8086 mode.
Software clears the VM bit to 0 to disable virtual-8086 mode. System software can only change this bit
using a task switch or an IRET. It cannot modify the bit using the POPFD instruction.
Alignment Check (AC) Bit. Bit 18. Software enables automatic alignment checking by setting the
AC bit to 1 when CR0.AM=1. Alignment checking can be disabled by clearing either AC or CR0.AM
to 0. When automatic alignment checking is enabled and the current privilege-level (CPL) is 3 (least
privileged), a memory reference to an unaligned operand causes an alignment-check exception (#AC).
Virtual Interrupt (VIF) Bit. Bit 19. The VIF bit is a virtual image of the RFLAGS.IF bit. It is enabled
when either virtual-8086 mode extensions are enabled (CR4.VME=1) or protected-mode virtual
interrupts are enabled (CR4.PVI=1), and the RFLAGS.IOPL field is less than 3. When enabled,
instructions that ordinarily would modify the IF bit actually modify the VIF bit with no effect on the
RFLAGS.IF bit.
System software that supports virtual-8086 mode should enable the VIF bit using CR4.VME. This
allows 8086 software to execute instructions that can set and clear the RFLAGS.IF bit without causing
an exception. With VIF enabled in virtual-8086 mode, those instructions set and clear the VIF bit
instead, giving the appearance to the 8086 software that it is modifying the RFLAGS.IF bit. System
software reads the VIF bit to determine whether or not to take the action desired by the 8086 software
(enabling or disabling interrupts by setting or clearing the RFLAGS.IF bit).
In long mode, the use of the VIF bit is supported when CR4.PVI=1. See “Virtual Interrupts” on
page 247 for more information on virtual interrupts.
Virtual Interrupt Pending (VIP) Bit. Bit 20. The VIP bit is provided as an extension to both virtual-
8086 mode and protected mode. It is used by system software to indicate that an external, maskable
interrupt is pending (awaiting) execution by either a virtual-8086 mode or protected-mode interrupt-
System Resources53
Page 98
AMD64 Technology24593—Rev. 3.14—September 2007
service routine. Software must enable virtual-8086 mode extensions (CR4.VME=1) or protectedmode virtual interrupts (CR4.PVI=1) before using VIP.
VIP is normally set to 1 by a protected-mode interrupt-service routine that was entered from virtual8086 mode as a result of an external, maskable interrupt. Before returning to the virtual-8086 mode
application, the service routine sets VIP to 1 if EFLAGS.VIF=1. When the virtual-8086 mode
application attempts to enable interrupts by clearing EFLAGS.VIF to 0 while VIP=1, a generalprotection exception (#GP) occurs. The #GP service routine can then decide whether to allow the
virtual-8086 mode service routine to handle the pending external, maskable interrupt. (EFLAGS is
specifically referred to in this case because virtual-8086 mode is supported only from legacy mode.)
In long mode, the use of the VIP bit is supported when CR4.PVI=1. See “Virtual Interrupts” on
page 247 for more information on virtual-8086 mode interrupts and the VIP bit.
Processor Feature Identification (ID) Bit. Bit 21. The ability of software to modify this bit
indicates that the processor implementation supports the CPUID instruction. See “Processor Feature
Identification” on page 61 for more information on the CPUID instruction.
3.1.8 Extended Feature Enable Register (EFER)
The extended-feature-enable register (EFER) contains control bits that enable additional processor
features not controlled by the legacy control registers. The EFER is a model-specific register (MSR)
with an address of C000_0080h (see “Model-Specific Registers (MSRs)” on page 56 for more
information on MSRs). It can be read and written only by privileged software. Figure 3-9 on page 55
shows the format of the EFER register.
54System Resources
Page 99
24593—Rev. 3.14—September 2007AMD64 Technology
6332
Reserved, MBZ
3115 14 13 12 11 10 98710
F
S
F
Reserved, MBZ
BitsMnemonicDescriptionR/W
63–15 Reserved, MBZReserved, Must be Zero
14FFXSRFast FXSAVE/FXRSTORR/W
13Reserved, MBZReserved, Must be Zero
12SVMESecure Virtual Machine EnableR/W
11NXENo-Execute EnableR/W
10LMALong Mode ActiveR
9Reserved, MBZReserved, Must be Zero
8LMELong Mode EnableR/W
7-1Reserved, RAZReserved, Read as Zero
0SCESystem Call ExtensionsR/W
The function of the EFER bits are (unless otherwise noted, all bits are read/write):
System-Call Extension (SCE) Bit. Bit 0. Setting this bit to 1 enables the SYSCALL and SYSRET
instructions. Application software can use these instructions for low-latency system calls and returns
in a non-segmented (flat) address space. See “Fast System Call and Return” on page 149 for additional
information.
Long Mode Enable (LME) Bit. Bit 8. Setting this bit to 1 enables the processor to activate long mode.
Long mode is not activated until software enables paging some time later. When paging is enabled
after LME is set to 1, the processor sets the EFER.LMA bit to 1, indicating that long mode is not only
enabled but also active. See Chapter 14, “Processor Initialization and Long Mode Activation,” for
more information on activating long mode.
Long Mode Active (LMA) Bit. Bit 10, read-only. This bit indicates that long mode is active. The
processor sets LMA to 1 when both long mode and paging have been enabled by system software. See
Chapter 14, “Processor Initialization and Long Mode Activation,” for more information on activating
long mode.
When LMA=1, the processor is running either in compatibility mode or 64-bit mode, depending on the
value of the L bit in a code-segment descriptor, as shown in Figure 1-6 on page 12.
System Resources55
Page 100
AMD64 Technology24593—Rev. 3.14—September 2007
When LMA=0, the processor is running in legacy mode. In this mode, the processor behaves like a
standard 32-bit x86 processor, with none of the new 64-bit features enabled.
No-Execute Enable (NXE) Bit. Bit 11. Setting this bit to 1 enables the no-execute page-protection
feature. The feature is disabled when this bit is cleared to 0. See “No Execute (NX) Bit” on page 143
for more information.
Before setting NXE, system software should verify the processor supports the feature by examining
the extended-feature flags returned by the CPUID instruction. For more information, see the CPUIDSpecification, order# 25481.
Secure Virtual Machine Enable (SVME) Bit. Bit 12. Enables the SVM extensions. When this bit is
zero, the SVM instructions cause #UD exceptions. EFER.SVME defaults to a reset value of zero. The
effect of turning off EFER.SVME while a guest is running is undefined; therefore, the VMM should
always prevent guests from writing EFER. SVM extensions can be disabled by setting
VM_CR.SVME_DISABLE. For more information, see descriptions of LOCK and SMVE_DISABLE
bits in Section 15.28.1, “VM_CR MSR (C001_0114h),” on page 420.
Fast FXSAVE/FXRSTOR (FFXSR) Bit. Bit 14. Setting this bit to 1 enables the FXSAVE and
FXRSTOR instructions to execute faster in 64-bit mode at CPL 0. This is accomplished by not saving
or restoring the XMM registers (XMM0-XMM15). The FFXSR bit has no effect when the
FXSAVE/FXRSTOR instructions are executed in non 64-bit mode, or when CPL > 0. The FFXSR bit
does not affect the save/restore of the legacy x87 floating-point state, or the save/restore of MXCSR.
Before setting FFXSR, system software should verify whether this feature is supported by examining
the CPUID extended feature flags returned by the CPUID instruction. For more information, see
"Function 8000_0001h: Processor Signature and AMD Features" in Volume 3.
3.2Model-Specific Registers (MSRs)
Processor implementations provide model-specific registers (MSRs) for software control over the
unique features supported by that implementation. Software reads and writes MSRs using the
privileged RDMSR and WRMSR instructions. Implementations of the AMD64 architecture can
contain a mixture of two basic MSR types:
•Legacy MSRs. The AMD family of processors often share model-specific features with other x86
processor implementations. Where possible, AMD implementations use the same MSRs for the
same functions. For example, the memory-typing and debug-extension MSRs are implemented on
many AMD and non-AMD processors.
•AMD model-specific MSRs. There are many MSRs common to the AMD family of processors but
not to legacy x86 processors. Where possible, AMD implementations use the same AMD-specific
MSRs for the same functions.
Every model-specific register, as the name implies, is not necessarily implemented by all members of
the AMD family of processors. Appendix A, “MSR Cross-Reference,” lists MSR-address ranges
currently used by various AMD and other x86 processors.
56System Resources
Loading...
+ hidden pages
You need points to download manuals.
1 point = 1 manual.
You can buy points or you can get point for every manual you upload.