(“AMD”) products. AMD makes no representations or warranties with respect to the accuracy or
completeness of the contents of this publication and reserves the right to make changes to
specifications and product descriptions at any time without notice. No license, whether express,
implied, arising by estoppel or otherwise, to any intellectual property rights is granted by this
publication. Except as set forth in AMD’s Standard Terms and Conditions of Sale, AMD assumes
no liability whatsoever, and disclaims any express or implied warranty, relating to its products
including, but not limited to, the implied warranty of merchantability, fitness for a particular purpose, or infringement of any intellectual property right.
AMD’s products are not designed, intended, authorized or warranted for use as components in
systems intended for surgical implant into the body, or in other applications intended to support
or sustain life, or in any other application in which the failure of AMD’s product could create a
situation where personal injury, death, or severe property or environmental damage may occur.
AMD reserves the right to discontinue or make changes to its products at any time without
notice.
Trademarks
AMD, the AMD arrow logo, AMD Athlon, and AMD Opteron, and combinations thereof, and 3DNow! are trademarks, and AMD-K6 is a
registered trademark of Advanced Micro Devices, Inc.
MMX is a trademark and Pentium is a registered trademark of Intel Corporation.
Windows NT is a registered trademark of Microsoft Corporation.
Other product names used in this publication are for identification purposes only and may be trademarks of their respective companies.
February 20053.10Clarified “Self-Modifying Code” on page 123. Made several patches to index
references. Added general descriptions of SSE3 instructions to Chapter 4. Added
description of the CMPXCHG16B instruction to Chapter 3. Corrected minor
typographical errors. Elaborated explanation of PREFETCHlevel instructions.
September 20033.09Corrected several factual errors.
September, 20023.07Corrected minor organizational problems in sections dealing with ‘Prefetch’
instructions in chapters 3, 4, and 5. Clarified the general description of the
operation of certain 128-bit media instructions in chapter 1. Corrected a factual
error in the description of the FNINIT/FINIT instructions in chapter 6. Corrected
operand descriptions for the CMOVcc instructions in chapter 3. Added Revision
History. Corrected marketing denotations.
: xvii
Revision Historyxvii
Page 18
AMD64 Technology24592—Rev. 3.10—March 2005
xviiiRevision History
Page 19
24592—Rev. 3.10—March 2005AMD64 Technology
Preface
About This Book
This book is part of a multivolume work entitled the AMD64
Architecture Programmer’s Manual. This table lists each volume
and its order number.
TitleOrder No.
Volume 1, Application Programming24592
Volume 2, System Programming24593
Volume 3, General-Purpose and System Instructions24594
Volume 4, 128-Bit Media Instructions26568
Volume 5, 64-Bit Media and x87 Floating-Point Instructions26569
Audience
Contact Information
This volume (Volume 1) is intended for programmers writing
application programs, compilers, or assemblers. It assumes
prior experience in microprocessor programming, although it
does not assume prior experience with the legacy x86 or
AMD64 microprocessor architecture.
This volume describes the AMD64 architecture’s resources and
functions that are accessible to application software, including
memory, registers, instructions, operands, I/O facilities, and
application-software aspects of control transfers (including
interrupts and exceptions) and performance optimization.
System-programming topics—including the use of instructions
running at a current privilege level (CPL) of 0 (mostprivileged)—are described in Volume 2. Details about each
instruction are described in volumes 3, 4, and 5.
To submit questions or comments concerning this document,
contact our technical documentation staff at
[email protected].
Prefacexix
Page 20
AMD64 Technology24592—Rev. 3.10—March 2005
Organization
This volume begins with an overview of the architecture and its
memory organization and is followed by chapters that describe
the four application-programming models available in the
AMD64 architecture:
General-Purpose Programming—This model uses the integer
general-purpose registers (GPRs). The chapter describing it
also describes the basic application environment for
exceptions, control transfers, I/O, and memory optimization
that applies to all other application-programming models.
128-bit Media Programming—This model uses the 128-bit
XMM registers and supports integer and floating-point
operations on vector (packed) and scalar data types.
64-bit Media Programming—This model uses the 64-bit
MMX™ registers and supports integer and floating-point
operations on vector (packed) and scalar data types.
x87 Floating-Point Programming—This model uses the 80-bit
x87 registers and supports floating-point operations on
scalar data types.
Definitions assumed throughout this volume are listed below.
The index at the end of this volume cross-references topics
within the volume. For other topics relating to the AMD64
architecture, see the tables of contents and indexes of the other
volumes.
Definitions
Some of the following definitions assume a knowledge of the
legacy x86 architecture. See “Related Documents” on page xxxi
for further information about the legacy x86 architecture.
Terms and Notation1011b
A binary value—in this example, a 4-bit value.
F0EAh
A hexadecimal value—in this example a 2-byte value.
[1,2)
A range that includes the left-most value (in this case, 1) but
excludes the right-most value (in this case, 2).
xxPreface
Page 21
24592—Rev. 3.10—March 2005AMD64 Technology
7–4
A bit range, from bit 7 to 4, inclusive. The high-order bit is
shown first.
128-bit media instructions
Instructions that use the 128-bit XMM registers. These are a
combination of the SSE and SSE2 instruction sets.
64-bit media instructions
Instructions that use the 64-bit MMX registers. These are
primarily a combination of MMX and 3DNow!™ instruction
sets, with some additional instructions from the SSE and
SSE2 instruction sets.
16-bit mode
Legacy mode or compatibility mode in which a 16-bit
address size is active. See legacy mode and compatibility
mode.
32-bit mode
Legacy mode or compatibility mode in which a 32-bit
address size is active. See legacy mode and compatibility
mode.
64-bit mode
A submode of long mode. In 64-bit mode, the default address
size is 64 bits and new features, such as register extensions,
are supported for system and application software.
#GP(0)
Notation indicating a general-protection exception (#GP)
with error code of 0.
absolute
Said of a displacement that references the base of a code
segment rather than an instruction pointer. Contrast with
relative.
biased exponent
The sum of a floating-point value’s exponent and a constant
bias for a particular floating-point data type. The bias makes
the range of the biased exponent always positive, which
allows reciprocation without overflow.
Prefacexxi
Page 22
AMD64 Technology24592—Rev. 3.10—March 2005
byte
Eight bits.
clear
To write a bit value of 0. Compare set.
compatibility mode
A submode of long mode. In compatibility mode, the default
address size is 32 bits, and legacy 16-bit and 32-bit
applications run without modification.
commit
To irreversibly write, in program order, an instruction’s
result to software-visible storage, such as a register
(including flags), the data cache, an internal write buffer, or
memory.
CPL
Current privilege level.
CR0–CR4
A register range, from register CR0 through CR4, inclusive,
with the low-order register first.
CR0.PE = 1
Notation indicating that the PE bit of the CR0 register has a
value of 1.
direct
Referencing a memory location whose address is included in
the instruction’s syntax as an immediate operand. The
address may be an absolute or relative address. Compare
indirect.
dirty data
Data held in the processor’s caches or internal buffers that is
more recent than the copy held in main memory.
displacement
A signed value that is added to the base of a segment
(absolute addressing) or an instruction pointer (relative
addressing). Same as offset.
doubleword
Two words, or four bytes, or 32 bits.
xxiiPreface
Page 23
24592—Rev. 3.10—March 2005AMD64 Technology
double quadword
Eight words, or 16 bytes, or 128 bits. Also called octword.
DS:rSI
The contents of a memory location whose segment address is
in the DS register and whose offset relative to that segment
is in the rSI register.
EFER.LME = 0
Notation indicating that the LME bit of the EFER register
has a value of 0.
effective address size
The address size for the current instruction after accounting
for the default address size and any address-size override
prefix.
effective operand size
The operand size for the current instruction after
accounting for the default operand size and any operandsize override prefix.
element
See vector.
exception
An abnormal condition that occurs as the result of executing
an instruction. The processor’s response to an exception
depends on the type of the exception. For all exceptions
except 128-bit media SIMD floating-point exceptions and
x87 floating-point exceptions, control is transferred to the
handler (or service routine) for that exception, as defined by
the exception’s vector. For floating-point exceptions defined
by the IEEE 754 standard, there are both masked and
unmasked responses. When unmasked, the exception
handler is called, and when masked, a default response is
provided instead of calling the handler.
FF /0
Notation indicating that FF is the first byte of an opcode,
and a subopcode in the ModR/M byte has a value of 0.
flush
An often ambiguous term meaning (1) writeback, if
modified, and invalidate, as in “flush the cache line,” or (2)
Prefacexxiii
Page 24
AMD64 Technology24592—Rev. 3.10—March 2005
invalidate, as in “flush the pipeline,” or (3) change a value,
as in “flush to zero.”
GDT
Global descriptor table.
IDT
Interrupt descriptor table.
IGN
Ignore. Field is ignored.
indirect
Referencing a memory location whose address is in a
register or other memory location. The address may be an
absolute or relative address. Compare direct.
IRB
The virtual-8086 mode interrupt-redirection bitmap.
IST
The long-mode interrupt-stack table.
IVT
The real-address mode interrupt-vector table.
LDT
Local descriptor table.
legacy x86
The legacy x86 architecture. See “Related Documents” on
page xxxi for descriptions of the legacy x86 architecture.
legacy mode
An operating mode of the AMD64 architecture in which
existing 16-bit and 32-bit applications and operating systems
run without modification. A processor implementation of
the AMD64 architecture can run in either long mode or legacy
mode. Legacy mode has three submodes, real mode, protected
mode, and virtual-8086 mode.
long mode
An operating mode unique to the AMD64 architecture. A
processor implementation of the AMD64 architecture can
run in either long mode or legacy mode. Long mode has two
submodes, 64-bit mode and compatibility mode.
xxivPreface
Page 25
24592—Rev. 3.10—March 2005AMD64 Technology
lsb
Least-significant bit.
LSB
Least-significant byte.
main memory
Physical memory, such as RAM and ROM (but not cache
memory) that is installed in a particular computer system.
mask
(1) A control bit that prevents the occurrence of a floatingpoint exception from invoking an exception-handling
routine. (2) A field of bits used for a control purpose.
MBZ
Must be zero. If software attempts to set an MBZ bit to 1, a
general-protection exception (#GP) occurs.
memory
Unless otherwise specified, main memory.
ModRM
A byte following an instruction opcode that specifies
address calculation based on mode (Mod), register (R), and
memory (M) variables.
moffset
A 16, 32, or 64-bit offset that specifies a memory operand
directly, without using a ModRM or SIB byte.
msb
Most-significant bit.
MSB
Most-significant byte.
multimedia instructions
A combination of 128-bit media instructions and 64-bit media
instructions.
octword
Same as double quadword.
offset
Same as displacement.
Prefacexxv
Page 26
AMD64 Technology24592—Rev. 3.10—March 2005
overflow
The condition in which a floating-point number is larger in
magnitude than the largest, finite, positive or negative
number that can be represented in the data-type format
being used.
packed
See vector.
PAE
Physical-address extensions.
physical memory
Actual memory, consisting of main memory and cache.
probe
A check for an address in a processor’s caches or internal
buffers. External probes originate outside the processor, and
internal probes originate within the processor.
protected mode
A submode of legacy mode.
quadword
Four words, or eight bytes, or 64 bits.
RAZ
Read as zero (0), regardless of what is written.
real-address mode
See real mode.
real mode
A short name for real-address mode, a submode of legacy
mode.
relative
Referencing with a displacement (also called offset) from an
instruction pointer rather than the base of a code segment.
Contrast with absolute.
reserved
Fields marked as reserved may be used at some future time.
xxviPreface
Page 27
24592—Rev. 3.10—March 2005AMD64 Technology
To preserve compatibility with future processors, reserved
fields require special handling when read or written by
software.
Reserved fields may be further qualified as MBZ, RAZ, SBZ
or IGN (see definitions).
Software must not depend on the state of a reserved field,
nor upon the ability of such fields to return to a previously
written state.
If a reserved field is not marked with one of the above
qualifiers, software must not change the state of that field; it
must reload that field with the same values returned from a
prior read.
REX
An instruction prefix that specifies a 64-bit operand size and
provides access to additional registers.
RIP-relative addressing
Addressing relative to the 64-bit RIP instruction pointer.
set
To write a bit value of 1. Compare clear.
SIB
A byte following an instruction opcode that specifies
address calculation based on scale (S), index (I), and base
(B).
SIMD
Single instruction, multiple data. See vector.
SSE
Streaming SIMD extensions instruction set. See 128-bit
media instructions and 64-bit media instructions.
SSE2
Extensions to the SSE instruction set. See 128-bit media
instructions and 64-bit media instructions.
SSE3
Further extensions to the SSE instruction set. See 128-bit
media instructions.
Prefacexxvii
Page 28
AMD64 Technology24592—Rev. 3.10—March 2005
sticky bit
A bit that is set or cleared by hardware and that remains in
that state until explicitly changed by software.
TOP
The x87 top-of-stack pointer.
TPR
Task-priority register (CR8).
TSS
Task-state segment.
underflow
The condition in which a floating-point number is smaller in
magnitude than the smallest nonzero, positive or negative
number that can be represented in the data-type format
being used.
vector
(1) A set of integer or floating-point values, called elements,
that are packed into a single operand. Most of the 128-bit
and 64-bit media instructions use vectors as operands.
Vectors are also called packed or SIMD (single-instruction
multiple-data) operands.
(2) An index into an interrupt descriptor table (IDT), used to
access exception handlers. Compare exception.
virtual-8086 mode
A submode of legacy mode.
word
Two bytes, or 16 bits.
x86
See legacy x86.
RegistersIn the following list of registers, the names are used to refer
either to a given register or to the contents of that register:
AH–DH
The high 8-bit AH, BH, CH, and DH registers. Compare
AL–DL.
xxviiiPreface
Page 29
24592—Rev. 3.10—March 2005AMD64 Technology
AL–DL
The low 8-bit AL, BL, CL, and DL registers. Compare AH–DH.
AL–r15B
The low 8-bit AL, BL, CL, DL, SIL, DIL, BPL, SPL, and
R8B–R15B registers, available in 64-bit mode.
BP
Base pointer register.
CRn
Control register number n.
CS
Code segment register.
eAX–eSP
The 16-bit AX, BX, CX, DX, DI, SI, BP, and SP registers or the
32-bit EAX, EBX, ECX, EDX, EDI, ESI, EBP, and ESP
registers. Compare rAX–rSP.
EFER
Extended features enable register.
eFLAGS
16-bit or 32-bit flags register. Compare rFLAGS.
EFLAGS
32-bit (extended) flags register.
eIP
16-bit or 32-bit instruction-pointer register. Compare rIP.
EIP
32-bit (extended) instruction-pointer register.
FLAGS
16-bit flags register.
GDTR
Global descriptor table register.
GPRs
General-purpose registers. For the 16-bit data size, these are
AX, BX, CX, DX, DI, SI, BP, and SP. For the 32-bit data size,
these are EAX, EBX, ECX, EDX, EDI, ESI, EBP, and ESP. For
Prefacexxix
Page 30
AMD64 Technology24592—Rev. 3.10—March 2005
the 64-bit data size, these include RAX, RBX, RCX, RDX,
RDI, RSI, RBP, RSP, and R8–R15.
IDTR
Interrupt descriptor table register.
IP
16-bit instruction-pointer register.
LDTR
Local descriptor table register.
MSR
Model-specific register.
r8–r15
The 8-bit R8B–R15B registers, or the 16-bit R8W–R15W
registers, or the 32-bit R8D–R15D registers, or the 64-bit
R8–R15 registers.
rAX–rSP
The 16-bit AX, BX, CX, DX, DI, SI, BP, and SP registers, or
the 32-bit EAX, EBX, ECX, EDX, EDI, ESI, EBP, and ESP
registers, or the 64-bit RAX, RBX, RCX, RDX, RDI, RSI,
RBP, and RSP registers. Replace the placeholder r with
nothing for 16-bit size, “E” for 32-bit size, or “R” for 64-bit
size.
RAX
64-bit version of the EAX register.
RBP
64-bit version of the EBP register.
RBX
64-bit version of the EBX register.
RCX
64-bit version of the ECX register.
RDI
64-bit version of the EDI register.
RDX
64-bit version of the EDX register.
xxxPreface
Page 31
24592—Rev. 3.10—March 2005AMD64 Technology
rFLAGS
16-bit, 32-bit, or 64-bit flags register. Compare RFLAGS.
RFLAGS
64-bit flags register. Compare rFLAGS.
rIP
16-bit, 32-bit, or 64-bit instruction-pointer register. Compare
RIP.
RIP
64-bit instruction-pointer register.
RSI
64-bit version of the ESI register.
RSP
64-bit version of the ESP register.
SP
Stack pointer register.
SS
Stack segment register.
TPR
Task priority register, a new register introduced in the
AMD64 architecture to speed interrupt management.
TR
Task register.
Endian OrderThe x86 and AMD64 architectures address memory using little-
endian byte-ordering. Multibyte values are stored with their
least-significant byte at the lowest byte address, and they are
illustrated with their least significant byte at the right side.
Strings are illustrated in reverse order, because the addresses of
their bytes increase from right to left.
Related Documents
Peter Abel, IBM PC Assembly Language and Programming,
Walter A. Triebel, The 80386DX Microprocessor, Prentice-
Hall, Englewood Cliffs, NJ, 1992.
John Wharton, The Complete x86, MicroDesign Resources,
Sebastopol, California, 1994.
Web sites and newsgroups:
-www.amd.com
-news.comp.arch
-news.comp.lang.asm.x86
-news.intel.microprocessors
-news.microsoft
xxxivPreface
Page 35
24592—Rev. 3.10—March 2005AMD64 Technology
1Overview of the AMD64 Architecture
1.1Introduction
The AMD64 architecture is a simple yet powerful 64-bit,
backward-compatible extension of the industry-standard
(legacy) x86 architecture. It adds 64-bit addressing and expands
register resources to support higher performance for
recompiled 64-bit programs, while supporting legacy 16-bit and
32-bit applications and operating systems without modification
or recompilation. It is the architectural basis on which new
processors can provide seamless, high-performance support for
both the vast body of existing software and new 64-bit software
required for higher-performance applications.
The need for a 64-bit x86 architecture is driven by applications
that address large amounts of virtual and physical memory,
such as high-performance servers, database management
systems, and CAD tools. These applications benefit from both
64-bit addresses and an increased number of registers. The
small number of registers available in the legacy x86
architecture limits performance in computation-intensive
applications. Increasing the number of registers provides a
performance boost to many such applications.
1.1.1 New FeaturesThe AMD64 architecture introduces these new features:
Register Extensions (see Figure 1-1 on page 2):
-8 new general-purpose registers (GPRs).
-All 16 GPRs are 64 bits wide.
-8 new 128-bit XMM registers.
-Uniform byte-register addressing for all GPRs.
-A new instruction prefix (REX) accesses the extended
registers.
1.1.2 RegistersTable 1-2 on page 4 compares the register and stack resources
available to application software, by operating mode. The left
set of columns shows the legacy x86 resources, which are
available in the AMD64 architecture’s legacy and compatibility
modes. The right set of columns shows the comparable
resources in 64-bit mode. Gray shading indicates differences
between the modes. These register differences (not including
stack-width difference) represent the register extensions shown
in Figure 1-1.
Chapter 1: Overview of the AMD64 Architecture3
Chapter 1: Overview of the AMD64 Architecture3
Page 38
AMD64 Technology24592—Rev. 3.10—March 2005
Table 1-2.Application Registers and Stack, by Operating Mode
Register
or Stack
General-Purpose
Registers (GPRs)
2
Legacy and Compatibility Modes
NameNumberSize (bits)NameNumberSize (bits)
EAX, EBX, ECX,
EDX, EBP, ESI,
832
EDI, ESP
128-Bit XMM RegistersXMM0–XMM78128
64-Bit MMX RegistersMMX0–MMX7
x87 RegistersFPR0–FPR7
Instruction Pointer
2
Flags
2
EIP132RIP164
EFLAGS132RFLAGS164
3
3
864MMX0–MMX73864
880FPR0–FPR73880
64-Bit Mode
RAX, RBX, RCX,
RDX, RBP, RSI,
RDI, RSP, R8–R15
XMM0–XMM1516128
1
1664
Stack—16 or 32—
Note:
1. Gray-shaded entries indicate differences between the modes. These differences (except stack-width difference) are the AMD64
architecture’s register extensions.
2. This list of GPRs shows only the 32-bit registers. The 16-bit and 8-bit mappings of the 32-bit registers are also accessible, as
described in “Registers” on page 27.
3. The MMX0–MMX7 registers are mapped onto the FPR0–FPR7 physical registers, as shown in Figure 1-1. The x87 stack registers,
ST(0)–ST(7), are the logical mappings of the FPR0–FPR7 physical registers.
64
As Table 1-2 shows, the legacy x86 architecture (called legacy
mode in the AMD64 architecture) supports eight GPRs. In
reality, however, the general use of at least four registers (EBP,
ESI, EDI, and ESP) is compromised because they serve special
purposes when executing many instructions. The AMD64
architecture’s addition of eight new GPRs—and the increased
width of these registers from 32 bits to 64 bits—allows
compilers to substantially improve software performance.
Compilers have more flexibility in using registers to hold
variables. Compilers can also minimize memory traffic—and
thus boost performance—by localizing work within the GPRs.
1.1.3 Instruction SetThe AMD64 architecture supports the full legacy x86
instruction set, and it adds a few new instructions to support
long mode (see Table 1-1 for a summary of operating modes).
The application-programming instructions are organized and
described in the following subsets:
General-Purpose Instructions—These are the basic x86
integer instructions used in virtually all programs. Most of
4Chapter 1: Overview of the AMD64 Architecture
Page 39
24592—Rev. 3.10—March 2005AMD64 Technology
these instructions load, store, or operate on data located in
the general-purpose registers (GPRs) or memory. Some of
the instructions alter sequential program flow program by
branching to other program locations.
128-Bit Media Instructions—These are the streaming SIMD
extension (SSE and SSE2) instructions that load, store, or
operate on data located primarily in the 128-bit XMM
registers. They perform integer and floating-point
operations on vector (packed) and scalar data types.
Because the vector instructions can independently and
simultaneously perform a single operation on multiple sets
of data, they are called single-instruction, multiple-data
(SIMD) instructions. They are useful for high-performance
media and scientific applications that operate on blocks of
data.
64-Bit Media Instructions—These are the multimedia
extension (MMX™ technology) and AMD 3DNow!™
technology instructions. They load, store, or operate on data
located primarily on the 64-bit MMX registers. Like their
128-bit counterparts, described above, they perform integer
and floating-point operations on vector (packed) and scalar
data types. Thus, they are also SIMD instructions and are
useful in media applications that operate on blocks of data.
1.1.4 Media
Instructions
x87 Floating-Point Instructions—These are the floating-point
instructions used in legacy x87 applications. They load,
store, or operate on data located in the x87 registers.
Some of these application-programming instructions bridge two
or more of the above subsets. For example, there are
instructions that move data between the general-purpose
registers and the XMM or MMX registers, and many of the
integer vector (packed) instructions can operate on either
XMM or MMX registers, although not simultaneously. If
instructions bridge two or more subsets, their descriptions are
repeated in all subsets to which they apply.
Media applications—such as image processing, music synthesis,
speech recognition, full-motion video, and 3D graphics
rendering—share certain characteristics:
They process large amounts of data.
They often perform the same sequence of operations
repeatedly across the data.
Chapter 1: Overview of the AMD64 Architecture5
Chapter 1: Overview of the AMD64 Architecture5
Page 40
AMD64 Technology24592—Rev. 3.10—March 2005
The data are often represented as small quantities, such as 8
bits for pixel values, 16 bits for audio samples, and 32 bits
for object coordinates in floating-point format.
The 128-bit and 64-bit media instructions are designed to
accelerate these applications. The instructions use a form of
vector (or packed) parallel processing known as singleinstruction, multiple data (SIMD) processing. This vector
technology has the following characteristics:
A single register can hold multiple independent pieces of
data. For example, a single 128-bit XMM register can hold 16
8-bit integer data elements, or four 32-bit single-precision
floating-point data elements.
The vector instructions can operate on all data elements in a
register, independently and simultaneously. For example, a
PADDB instruction operating on byte elements of two vector
operands in 128-bit XMM registers performs 16
simultaneous additions and returns 16 independent results
in a single operation.
1.1.5 Floating-Point
Instructions
128-bit and 64-bit media instructions take SIMD vector
technology a step further by including special instructions that
perform operations commonly found in media applications. For
example, a graphics application that adds the brightness values
of two pixels must prevent the add operation from wrapping
around to a small value if the result overflows the destination
register, because an overflow result can produce unexpected
effects such as a dark pixel where a bright one is expected. The
128-bit and 64-bit media instructions include saturatingarithmetic instructions to simplify this type of operation. A
result that otherwise would wrap around due to overflow or
underflow is instead forced to saturate at the largest or smallest
value that can be represented in the destination register.
The AMD64 architecture provides three floating-point
instruction subsets, using three distinct register sets:
128-Bit Media Instructions support 32-bit single-precision
and 64-bit double-precision floating-point operations, in
addition to integer operations. Operations on both vector
data and scalar data are supported, with a dedicated
floating-point exception-reporting mechanism. These
floating-point operations comply with the IEEE-754
standard.
6Chapter 1: Overview of the AMD64 Architecture
Page 41
24592—Rev. 3.10—March 2005AMD64 Technology
64-Bit Media Instructions (the subset of 3DNow! technology
instructions) support single-precision floating-point
operations. Operations on both vector data and scalar data
are supported, but these instructions do not support
floating-point exception reporting.
x87 Floating-Point Instructions support single-precision,
double-precision, and 80-bit extended-precision floatingpoint operations. Only scalar data are supported, with a
dedicated floating-point exception-reporting mechanism.
The x87 floating-point instructions contain special
instructions for performing trigonometric and logarithmic
transcendental operations. The single-precision and doubleprecision floating-point operations comply with the IEEE754 standard.
Maximum floating-point performance can be achieved using
the 128-bit media instructions. One of these vector instructions
can support up to four single-precision (or two doubleprecision) operations in parallel. In 64-bit mode, the AMD64
architecture doubles the number of legacy XMM registers from
8 to 16.
Applications gain additional benefits using the 64-bit media
and x87 instructions. The separate register sets supported by
these instructions relieve pressure on the XMM registers
available to the 128-bit media instructions. This provides
application programs with three distinct sets of floating-point
registers. In addition, certain high-end implementations of the
AMD64 architecture may support 128-bit media, 64-bit media,
and x87 instructions with separate execution units.
1.2Modes of Operation
Table 1-1 on page 3 summarizes the modes of operation
supported by the AMD64 architecture. In most cases, the
default address and operand sizes can be overridden with
instruction prefixes. The register extensions shown in the
second-from-right column of Table 1-1 are those illustrated in
Figure 1-1 on page 2.
1.2.1 Long ModeLong mode is an extension of legacy protected mode. Long
mode consists of two submodes: 64-bit mode and compatibilitymode. 64-bit mode supports all of the new features and register
extensions of the AMD64 architecture. Compatibility mode
Chapter 1: Overview of the AMD64 Architecture7
Chapter 1: Overview of the AMD64 Architecture7
Page 42
AMD64 Technology24592—Rev. 3.10—March 2005
supports binary compatibility with existing 16-bit and 32-bit
applications. Long mode does not support legacy real mode or
legacy virtual-8086 mode, and it does not support hardware task
switching.
Throughout this document, references to long mode refer to
both 64-bit mode and compatibility mode. If a function is specific
to either of these submodes, then the name of the specific
submode is used instead of the name long mode.
1.2.2 64-Bit Mode64-bit mode—a submode of long mode—supports the full range
of 64-bit virtual-addressing and register-extension features.
This mode is enabled by the operating system on an individual
code-segment basis. Because 64-bit mode supports a 64-bit
virtual-address space, it requires a new 64-bit operating system
and tool chain. Existing application binaries can run without
recompilation in compatibility mode, under an operating
system that runs in 64-bit mode, or the applications can also be
recompiled to run in 64-bit mode.
Addressing features include a 64-bit instruction pointer (RIP)
and a new RIP-relative data-addressing mode. This mode
accommodates modern operating systems by supporting only a
flat address space, with single code, data, and stack space.
Register Extensions. 64-bit mode implements register extensions
through a new group of instruction prefixes, called REX
prefixes. These extensions add eight GPRs (R8–R15), widen all
GPRs to 64 bits, and add eight 128-bit XMM registers
(XMM8–XMM15).
The REX instruction prefixes also provide a new byte-register
capability that makes the low byte of any of the sixteen GPRs
available for byte operations. This results in a uniform set of
byte, word, doubleword, and quadword registers that is better
suited to compiler register-allocation.
64-Bit Addresses and Operands. In 64-bit mode, the default virtualaddress size is 64 bits (implementations can have fewer). The
default operand size for most instructions is 32 bits. For most
instructions, these defaults can be overridden on an
instruction-by-instruction basis using instruction prefixes. REX
prefixes specify the 64-bit operand size and new registers.
RIP-Relative Data Addressing. 64-bit mode supports data addressing
relative to the 64-bit instruction pointer (RIP). The legacy x86
8Chapter 1: Overview of the AMD64 Architecture
Page 43
24592—Rev. 3.10—March 2005AMD64 Technology
architecture supports IP-relative addressing only in controltransfer instructions. RIP-relative addressing improves the
efficiency of position-independent code and code that
addresses global data.
Opcodes. A few instruction opcodes and prefix bytes are
redefined to allow register extensions and 64-bit addressing.
These differences are described in “General-Purpose
Instructions in 64-Bit Mode” in Volume 3 and “Differences
Between Long Mode and Legacy Mode” in Volume 3.
1.2.3 Compatibility
Mode
Compatibility mode—the second submode of long mode—
allows 64-bit operating systems to run existing 16-bit and 32-bit
x86 applications. These legacy applications run in compatibility
mode without recompilation.
Applications running in compatibility mode use 32-bit or 16-bit
addressing and can access the first 4GB of virtual-address
space. Legacy x86 instruction prefixes toggle between 16-bit
and 32-bit address and operand sizes.
As with 64-bit mode, compatibility mode is enabled by the
operating system on an individual code-segment basis. Unlike
64-bit mode, however, x86 segmentation functions the same as
in the legacy x86 architecture, using 16-bit or 32-bit protectedmode semantics. From the application viewpoint, compatibility
mode looks like the legacy x86 protected-mode environment.
From the operating-system viewpoint, however, address
translation, interrupt and exception handling, and system data
structures use the 64-bit long-mode mechanisms.
1.2.4 Legacy ModeLegacy mode preserves binary compatibility not only with
existing 16-bit and 32-bit applications but also with existing 16bit and 32-bit operating systems. Legacy mode consists of the
following three submodes:
Protected Mode—Protected mode supports 16-bit and 32-bit
programs with memory segmentation, optional paging, and
privilege-checking. Programs running in protected mode can
access up to 4GB of memory space.
mode programs running as tasks under protected mode. It
uses a simple form of memory segmentation, optional
paging, and limited protection-checking. Programs running
in virtual-8086 mode can access up to 1MB of memory space.
Chapter 1: Overview of the AMD64 Architecture9
Chapter 1: Overview of the AMD64 Architecture9
Page 44
AMD64 Technology24592—Rev. 3.10—March 2005
Real Mode—Real mode supports 16-bit programs using
simple register-based memory segmentation. It does not
support paging or protection-checking. Programs running in
real mode can access up to 1MB of memory space.
Legacy mode is compatible with existing 32-bit processor
implementations of the x86 architecture. Processors that
implement the AMD64 architecture boot in legacy real mode,
just like processors that implement the legacy x86 architecture.
Throughout this document, references to legacy mode refer to
all three submodes—protected mode, virtual-8086 mode, and realmode. If a function is specific to either of these submodes, then
the name of the specific submode is used instead of the name
legacy mode.
10Chapter 1: Overview of the AMD64 Architecture
Page 45
24592—Rev. 3.10—March 2005AMD64 Technology
2Memory Model
This chapter describes the memory characteristics that apply to
application software in the various operating modes of the
AMD64 architecture. These characteristics apply to all
instructions in the architecture. Several additional system-level
details about memory and cache management are described in
Vol um e 2 .
2.1Memory Organization
2.1.1 Virtual MemoryVirtual memory consists of the entire address space available to
programs. It is a large linear-address space that is translated by
a combination of hardware and operating-system software to a
smaller physical-address space, parts of which are located in
memory and parts on disk or other external storage media.
Figure 2-1 on page 12 shows how the virtual-memory space is
treated in the two submodes of long mode:
64-bit mode—This mode uses a flat segmentation model of
virtual memory. The 64-bit virtual-memory space is treated
as a single, flat (unsegmented) address space. Program
addresses access locations that can be anywhere in the
linear 64-bit address space. The operating system can use
separate selectors for code, stack, and data segments for
memory-protection purposes, but the base address of all
these segments is always 0. (For an exception to this general
rule, see “FS and GS as Base of Address Calculation” on
page 20.)
Compatibility mode—This mode uses a protected, multi-
segment model of virtual memory, just as in legacy
protected mode. The 32-bit virtual-memory space is treated
as a segmented set of address spaces for code, stack, and
data segments, each with its own base address and
protection parameters. A segmented space is specified by
adding a segment selector to an address.
Chapter 2: Memory Model11
Chapter 2: Memory Model11
Page 46
AMD64 Technology24592—Rev. 3.10—March 2005
64-Bit Mode
(Flat Segmentation Model)
264 - 1
Legacy and Compatibility Mode
(Multi-Segment Model)
232 - 1
Code Segment (CS) Base
code
stack
data
0
513-107.eps
Base Address for
All Segments
Stack Segment (SS) Base
Data Segment (DS) Base
0
Figure 2-1.Virtual-Memory Segmentation
Segmented memory has been used as a method by which
operating systems could isolate programs, and the data used by
programs, from each other in an effort to increase the reliability
of systems running multiple programs simultaneously. However,
most modern operating systems do not use the segmentation
features available in the legacy x86 architecture. Instead, these
operating systems handle segmentation functions entirely in
software. For this reason, the AMD64 architecture dispenses
with most of the legacy segmentation functions in 64-bit mode.
This allows new 64-bit operating systems to be coded more
simply, and it supports more efficient management of multiprogramming environments than is possible in the legacy x86
architecture.
2.1.2 Segment
Registers
Segment registers hold the selectors used to access memory
segments. Figure 2-2 on page 13 shows the application-visible
portion of the segment registers. In legacy and compatibility
modes, all segment registers are accessible to software. In 64bit mode, only the CS, FS, and GS segments are recognized by
12Chapter 2: Memory Model
Page 47
24592—Rev. 3.10—March 2005AMD64 Technology
the processor, and software can use the FS and GS segmentbase registers as base registers for address calculation, as
described in “FS and GS as Base of Address Calculation” on
page 20. For references to the DS, ES, or SS segments in 64-bit
mode, the processor assumes that the base for each of these
segments is zero, neither their segment limit nor attributes are
checked, and the processor simply checks that all such
addresses are in canonical form, as described in “64-bit
Canonical Addresses” on page 18.
2.1.3 Physical
Memory
Legacy Mode and
Compatibility Mode
CS
DS
ES
FS
GS
SS
150
64-Bit
Mode
CS
(Attributes only)
ignored
ignored
FS
(Base only)
GS
(Base only)
ignored
150
513-312.eps
Figure 2-2.Segment Registers
For details on segmentation and the segment registers, see
“Segmented Virtual Memory” in Volume 2.
Physical memory is the installed memory (excluding cache
memory) in a particular computer system that can be accessed
through the processor’s bus interface. The maximum size of the
physical memory space is determined by the number of address
bits on the bus interface. In a virtual-memory system, the large
virtual-address space (also called linear-address space) is
translated to a smaller physical-address space by a combination
of segmentation and paging hardware and software.
Segmentation is illustrated in Figure 2-1 on page 12. Paging is a
mechanism for translating linear (virtual) addresses into fixedsize blocks called pages, which the operating system can move,
as needed, between memory and external storage media
Chapter 2: Memory Model13
Chapter 2: Memory Model13
Page 48
AMD64 Technology24592—Rev. 3.10—March 2005
(typically disk). The AMD64 architecture supports an expanded
version of the legacy x86 paging mechanism, one that is able to
translate the full 64-bit virtual-address space into the physicaladdress space supported by the particular implementation.
2.1.4 Memory
Management
630
Memory management consists of the methods by which
addresses generated by programs are translated via
segmentation and/or paging into addresses in physical memory.
Memory management is not visible to application programs. It
is handled by the operating system and processor hardware.
The following description gives a very brief overview of these
functions. Details are given in “System-Management
Instructions” in Volume 2.
Long-Mode Memory Management. Figure 2-3 shows the flow, from
top to bottom, of memory management functions performed in
the two submodes of long mode.
64-Bit Mode
Virtual (Linear) Address
Compatibility Mode
031015
Effective AddressSelector
Segmentation
0313263
Virtual Address0
Paging
051
Physical Address
Paging
051
Physical Address
513-184.eps
Figure 2-3.Long-Mode Memory Management
In 64-bit mode, programs generate virtual (linear) addresses
that can be up to 64 bits in size. The virtual addresses are
14Chapter 2: Memory Model
Page 49
24592—Rev. 3.10—March 2005AMD64 Technology
passed to the long-mode paging function, which generates
physical addresses that can be up to 52 bits in size. (Specific
implementations of the architecture can support fewer virtualaddress and physical-address sizes.)
In compatibility mode, legacy 16-bit and 32-bit applications run
using legacy x86 protected-mode segmentation semantics. The
16-bit or 32-bit effective addresses generated by programs are
combined with their segments to produce 32-bit virtual (linear)
addresses that are zero-extended to a maximum of 64 bits. The
paging that follows is the same long-mode paging function used
in 64-bit mode. It translates the virtual addresses into physical
addresses. The combination of segment selector and effective
address is also called a logical address or far pointer. The virtualaddress is also called the linear address.
Legacy-Mode Memory Management. Figure 2-4 shows the memorymanagement functions performed in the three submodes of
legacy mode.
Protected Mode
031015
Effective Address (EA)Selector
Segmentation
031
Linear Address
Paging
031
Physical Address (PA)
Figure 2-4.Legacy-Mode Memory Management
Virtual-8086 Mode
015
Selector
Segmentation
Linear Address
Paging
Physical Address (PA)
EA
Real Mode
015
019
031
015
Selector
Segmentation
Linear Address
19031
0
015
EA
019
PA
513-185.eps
Chapter 2: Memory Model15
Chapter 2: Memory Model15
Page 50
AMD64 Technology24592—Rev. 3.10—March 2005
The memory-management functions differ, depending on the
submode, as follows:
Protected Mode—Protected mode supports 16-bit and 32-bit
programs with table-based memory segmentation, paging,
and privilege-checking. The segmentation function takes 32bit effective addresses and 16-bit segment selectors and
produces 32-bit linear addresses into one of 16K memory
segments, each of which can be up to 4GB in size. Paging is
optional. The 32-bit physical addresses are either produced
by the paging function or the linear addresses are used
without modification as physical addresses.
programs running as tasks under protected mode. 20-bit
linear addresses are formed in the same way as in real mode,
but they can optionally be translated through the paging
function to form 32-bit physical addresses that access up to
4GB of memory space.
Real Mode—Real mode supports 16-bit programs using
register-based shift-and-add segmentation, but it does not
support paging. Sixteen-bit effective addresses are zeroextended and added to a 16-bit segment-base address that is
left-shifted four bits, producing a 20-bit linear address. The
linear address is zero-extended to a 32-bit physical address
that can access up to 1MB of memory space.
2.2Memory Addressing
2.2.1 Byte OrderingInstructions and data are stored in memory in little-endian byte
order. Little-endian ordering places the least-significant byte of
the instruction or data item at the lowest memory address and
the most-significant byte at the highest memory address.
Figure 2-5 on page 17 shows a generalization of little-endian
memory and register images of a quadword data type. The leastsignificant byte is at the lowest address in memory and at the
right-most byte location of the register image.
16Chapter 2: Memory Model
Page 51
24592—Rev. 3.10—March 2005AMD64 Technology
Quadword in Memory
High (most-significant)
Quadword in General-Purpose Register
Figure 2-5.Byte Ordering
byte 7
byte 6
byte 5
byte 4
byte 3
byte 2
byte 1
byte 0
07h
06h
05h
04h
03h
02h
01h
00h
High (most-significant)
Low (least-significant)
Low (least-significant)
byte 0byte 1byte 2byte 3byte 4byte 5byte 6byte 7
063
513-116.eps
Figure 2-6 on page 18 shows the memory image of a 10-byte
instruction. Instructions are byte data types. They are read
from memory one byte at a time, starting with the leastsignificant byte (lowest address). For example, the following
instruction specifies the 64-bit instruction MOV RAX,
1122334455667788 instruction that consists of the following ten
bytes:
48 B8 8877665544332211
48 is a REX instruction prefix that specifies a 64-bit operand
size, B8 is the opcode that—together with the REX prefix—
specifies the 64-bit RAX destination register, and
8877665544332211 is the 8-byte immediate value to be moved,
where 88 represents the eighth (least-significant) byte and 11
represents the first (most-significant) byte. In memory, the REX
prefix byte (48) would be stored at the lowest address, and the
first immediate byte (11) would be stored at the highest
instruction address.
Chapter 2: Memory Model17
Chapter 2: Memory Model17
Page 52
AMD64 Technology24592—Rev. 3.10—March 2005
2.2.2 64-bit Canonical
Addresses
11
22
33
44
55
66
77
88
B8
48
09h
08h
07h
06h
05h
04h
03h
02h
01h
00h
High (most-significant)
Low (least-significant)
513-186.eps
Figure 2-6.Example of 10-Byte Instruction in Memory
Long mode defines 64 bits of virtual address, but
implementations of the AMD64 architecture may support fewer
bits of virtual address. Although implementations might not
use all 64 bits of the virtual address, they check bits 63 through
the most-significant implemented bit to see if those bits are all
zeros or all ones. An address that complies with this property is
said to be in canonical address form. If a virtual-memory
reference is not in canonical form, the implementation causes a
general-protection exception or stack fault.
2.2.3 Effective
Addresses
Programs provide effective addresses to the hardware prior to
segmentation and paging translations. Long-mode effective
addresses are a maximum of 64 bits wide, as shown in Figure 2-3
on page 14. Programs running in compatibility mode generate
(by default) 32-bit effective addresses, which the hardware zeroextends to 64 bits. Legacy-mode effective addresses, with no
address-size override, are 32 or 16 bits wide, as shown in
Figure 2-4. These sizes can be overridden with an address-size
instruction prefix, as described in “Instruction Prefixes” on
page 87.
There are five methods for generating effective addresses,
depending on the specific instruction encoding:
18Chapter 2: Memory Model
Page 53
24592—Rev. 3.10—March 2005AMD64 Technology
Absolute Addresses—These addresses are given as
displacements (or offsets) from the base address of a data
segment. They point directly to a memory location in the
data segment.
Instruction-Relative Addresses—These addresses are given as
displacements (or offsets) from the current instruction
pointer (IP), also called the program counter (PC). They are
generated by control-transfer instructions. A displacement
in the instruction encoding, or one read from memory, serves
as an offset from the address that follows the transfer. See
“RIP-Relative Addressing” on page 22 for details about RIPrelative addressing in 64-bit mode.
ModR/M Addressing—These addresses are calculated using a
scale, index, base, and displacement. Instruction encodings
contain two bytes—MODR/M and optional SIB (scale, index,
base) and a variable length displacement—that specify the
variables for the calculation. The base and index values are
contained in general-purpose registers specified by the SIB
byte. The scale and displacement values are specified
directly in the instruction encoding. Figure 2-7 shows the
components of a complex-address calculation. The resultant
effective address is added to the data-segment base address
to form a linear address, as described in “Segmented Virtual
Memory” in Volume 2. “Instruction Formats” in Volume 3
gives further details on specifying this form of address. The
encoding of instructions specifies how the address is
calculated.
Stack Addresses—PUSH, POP, CALL, RET, IRET, and INT
instructions implicitly use the stack pointer, which contains
the address of the procedure stack. See “Stack Operation”
on page 23 for details about the size of the stack pointer.
addresses using the rDI and rSI registers, as described in
“Implicit Uses of GPRs” on page 34.
In 64-bit mode, with no address-size override, the size of
effective-address calculations is 64 bits. An effective-address
calculation uses 64-bit base and index registers and signextends displacements to 64 bits. Due to the flat address space
in 64-bit mode, virtual addresses are equal to effective
addresses. (For an exception to this general rule, see “FS and
GS as Base of Address Calculation” on page 20.)
Long-Mode Zero-Extension of 16-Bit and 32-Bit Addresses. In long mode,
all 16-bit and 32-bit address calculations are zero-extended to
form 64-bit addresses. Address calculations are first truncated
to the effective-address size of the current mode (64-bit mode or
compatibility mode), as overridden by any address-size prefix.
The result is then zero-extended to the full 64-bit address width.
Because of this, 16-bit and 32-bit applications running in
compatibility mode can access only the low 4GB of the longmode virtual-address space. Likewise, a 32-bit address
generated in 64-bit mode can access only the low 4GB of the
long-mode virtual-address space.
Displacements and Immediates. In general, the maximum size of
address displacements and immediate operands is 32 bits. They
can be 8, 16, or 32 bits in size, depending on the instruction or,
for displacements, the effective address size. In 64-bit mode,
displacements are sign-extended to 64 bits during use, but their
actual size (for value representation) remains a maximum of 32
bits. The same is true for immediates in 64-bit mode, when the
operand size is 64 bits. However, support is provided in 64-bit
mode for some 64-bit displacement and immediate forms of the
MOV instruction.
FS and GS as Base of Address Calculation. In 64-bit mode, the FS and
GS segment-base registers (unlike the DS, ES, and SS segmentbase registers) can be used as non-zero data-segment base
registers for address calculations, as described in “Segmented
Virtual Memory” in Volume 2. 64-bit mode assumes all other
20Chapter 2: Memory Model
Page 55
24592—Rev. 3.10—March 2005AMD64 Technology
data-segment registers (DS, ES, and SS) have a base address of
0.
2.2.4 Address-Size
Prefix
The default address size of an instruction is determined by the
default-size (D) bit and long-mode (L) bit in the current codesegment descriptor (for details, see “Segmented Virtual
Memory” in Volume 2). Application software can override the
default address size in any operating mode by using the 67h
address-size instruction prefix byte. The address-size prefix
allows mixing 32-bit and 64-bit addresses on an instruction-byinstruction basis.
Table 2-1 shows the effects of using the address-size prefix in all
operating modes. In 64-bit mode, the default address size is 64
bits. The address size can be overridden to 32 bits. 16-bit
addresses are not supported in 64-bit mode. In compatibility
and legacy modes, the address-size prefix works the same as in
the legacy x86 architecture.
Table 2-1.Address-Size Prefixes
Address-
Size Prefix
1
(67h)
Required?
Operating Mode
Default
Address Size
(Bits)
Effective
Address Size
(Bits)
64no
64-Bit Mode
Long Mode
Compatibility Mode
Legacy Mode
(Protected, Virtual-8086, or Real
Mode)
Note:
1. “No’ indicates that the default address size is used.
Chapter 2: Memory Model21
Chapter 2: Memory Model21
64
32yes
32no
32
16ye s
32yes
16
16no
32no
32
16ye s
32yes
16
16no
Page 56
AMD64 Technology24592—Rev. 3.10—March 2005
2.2.5 RIP-Relative
Addressing
RIP-relative addressing—that is, addressing relative to the 64bit instruction pointer (also called program counter)—is
available in 64-bit mode. The effective address is formed by
adding the displacement to the 64-bit RIP of the next
instruction.
In the legacy x86 architecture, addressing relative to the
instruction pointer (IP or EIP) is available only in controltransfer instructions. In the 64-bit mode, any instruction that
uses ModRM addressing (see “ModRM and SIB Bytes” in
Volume 3) can use RIP-relative addressing. The feature is
particularly useful for addressing data in position-independent
code and for code that addresses global data.
Programs usually have many references to data, especially
global data, that are not register-based. To load such a program,
the loader typically selects a location for the program in
memory and then adjusts the program’s references to global
data based on the load location. RIP-relative addressing of data
makes this adjustment unnecessary.
Range of RIP-Relative Addressing. Without RIP-relative addressing,
instructions encoded with a ModRM byte address memory
relative to zero. With RIP-relative addressing, instructions with
a ModRM byte can address memory relative to the 64-bit RIP
using a signed 32-bit displacement. This provides an offset
range of ±2GB from the RIP.
Effect of Address-Size Prefix on RIP-relative Addressing. RIP-relative
addressing is enabled by 64-bit mode, not by a 64-bit addresssize. Conversely, use of the address-size prefix does not disable
RIP-relative addressing. The effect of the address-size prefix is
to truncate and zero-extend the computed effective address to
32 bits, like any other addressing mode.
Encoding. For details on instruction encoding of RIP-relative
addressing, see in “RIP-Relative Addressing” in Volume 3.
2.3Pointers
Pointers are variables that contain addresses rather than data.
They are used by instructions to reference memory. Instructions
access data using near and far pointers. Stack pointers locate
the current stack.
22Chapter 2: Memory Model
Page 57
24592—Rev. 3.10—March 2005AMD64 Technology
2.3.1 Near and Far
Pointers
Near pointers contain only an effective address, which is used
as an offset into the current segment. Far pointers contain both
an effective address and a segment selector that specifies one
of several segments. Figure 2-8 illustrates the two types of
pointers.
In 64-bit mode, the AMD64 architecture supports only the flatmemory model in which there is only one data segment, so the
effective address is used as the virtual (linear) address and far
pointers are not needed. In compatibility mode and legacy
protected mode, the AMD64 architecture supports multiple
memory segments, so effective addresses can be combined with
segment selectors to form far pointers, and the terms logicaladdress (segment selector and effective address) and far pointer
are synonyms. Near pointers can also be used in compatibility
mode and legacy mode.
2.4Stack Operation
A stack is a portion of a stack segment in memory that is used to
link procedures. Software conventions typically define stacks
using a stack frame, which consists of two registers—a stack-frame base pointer (rBP) and a stack pointer (rSP)—as shown in
Figure 2-9 on page 24. These stack pointers can be either near
pointers or far pointers.
The stack-segment (SS) register, points to the base address of
the current stack segment. The stack pointers contain offsets
from the base address of the current stack segment. All
instructions that address memory using the rBP or rSP registers
cause the processor to access the current stack segment.
Chapter 2: Memory Model23
Chapter 2: Memory Model23
Page 58
AMD64 Technology24592—Rev. 3.10—March 2005
Stack Frame Before Procedure CallStack Frame After Procedure Call
Stack-Frame Base Pointer (rBP)
and Stack Pointer (rSP)
Stack-Segment (SS) Base Address
Figure 2-9.Stack Pointer Mechanism
In typical APIs, the stack-frame base pointer and the stack
pointer point to the same location before a procedure call (the
top-of-stack of the prior stack frame). After data is pushed onto
the stack, the stack-frame base pointer remains where it was
and the stack pointer advances downward to the address below
the pushed data, where it becomes the new top-of-stack.
In legacy and compatibility modes, the default stack pointer
size is 16 bits (SP) or 32 bits (ESP), depending on the defaultsize (B) bit in the stack-segment descriptor, and multiple stacks
can be maintained in separate stack segments. In 64-bit mode,
stack pointers are always 64 bits wide (RSP).
Stack-Frame Base Pointer (rBP)
Stack Pointer (rSP)
Stack-Segment (SS) Base Address
passed data
513-110.eps
Further application-programming details on the stack
mechanism are described in “Control Transfers” on page 94.
System-programming details on the stack segments are
described in “Segmented Virtual Memory” in Volume 2.
2.5Instruction Pointer
The instruction pointer is used in conjunction with the codesegment (CS) register to locate the next instruction in memory.
The instruction-pointer register contains the displacement
(offset)—from the base address of the current CS segment, or
from address 0 in 64-bit mode—to the next instruction to be
executed. The pointer is incremented sequentially, except for
branch instructions, as described in “Control Transfers” on
page 94.
24Chapter 2: Memory Model
Page 59
24592—Rev. 3.10—March 2005AMD64 Technology
In legacy and compatibility modes, the instruction pointer is a
16-bit (IP) or 32-bit (EIP) register. In 64-bit mode, the
instruction pointer is extended to a 64-bit (RIP) register to
support 64-bit offsets. The case-sensitive acronym, rIP, is used to
refer to any of these three instruction-pointer sizes, depending
on the software context.
Figure 2-10 shows the relationship between RIP, EIP, and IP.
The 64-bit RIP can be used for RIP-relative addressing, as
described in “RIP-Relative Addressing” on page 22.
IP
EIP
RIP
6331032
513-140.eps
rIP
Figure 2-10.Instruction Pointer (rIP) Register
The contents of the rIP are not directly readable by software.
However, the rIP is pushed onto the stack by a call instruction.
The memory model described in this chapter is used by all of
the programming environments that make up the AMD64
architecture. The next four chapters of this volume describe the
application programming environments, which include:
General-purpose programming (Chapter 3 on page 27).
128-bit media programming (Chapter 4 on page 131).
64-bit media programming (Chapter 5 on page 237).
x87 floating-point programming (Chapter 6 on page 293).
Chapter 2: Memory Model25
Chapter 2: Memory Model25
Page 60
AMD64 Technology24592—Rev. 3.10—March 2005
26Chapter 2: Memory Model
Page 61
24592—Rev. 3.10—March 2005AMD64 Technology
3General-Purpose Programming
The general-purpose programming model includes the generalpurpose registers (GPRs), integer instructions and operands
that use the GPRs, program-flow control methods, memory
optimization methods, and I/O. This programming model
includes the original x86 integer-programming architecture,
plus 64-bit extensions and a few additional instructions. Only
the application-programming instructions and resources are
described in this chapter. Integer instructions typically used in
system programming, including all of the privileged
instructions, are described in Volume 2, along with other
system-programming topics.
The general-purpose programming model is used to some extent
by almost all programs, including programs consisting primarily
of 128-bit media instructions, 64-bit media instructions, x87
floating-point instructions, or system instructions. For this
reason, an understanding of the general-purpose programming
model is essential for any programming work using the AMD64
instruction set architecture.
3.1Registers
Figure 3-1 on page 28 shows an overview of the registers used in
general-purpose application programming. They include the
general-purpose registers (GPRs), segment registers, flags
register, and instruction-pointer register. The number and
width of available registers depends on the operating mode.
The registers and register ranges shaded light gray in Figure 3-1
are available only in 64-bit mode. Those shaded dark gray are
available only in legacy mode and compatibility mode. Thus, in
64-bit mode, the 32-bit general-purpose, flags, and instructionpointer registers available in legacy mode and compatibility
mode are extended to 64-bit widths, eight new GPRs are
available, and the DS, ES, and SS segment registers are ignored.
When naming registers, if reference is made to multiple
register widths, a lower-case r notation is used. For example, the
notation rAX refers to the 16-bit AX, 32-bit EAX, or 64-bit RAX
register, depending on an instruction’s effective operand size.
Chapter 3: General-Purpose Programming27
Chapter 3: General-Purpose Programming27
Page 62
AMD64 Technology24592—Rev. 3.10—March 2005
General-Purpose Registers (GPRs)
rAX
rBX
rCX
rDX
rBP
rSI
rDI
rSP
R8
R9
R10
Segment
Registers
CS
DS
ES
FS
6331032
Flags and Instruction Pointer Registers
GS
SS
150
Available to sofware in all modes
Available to sofware only in 64-bit mode
Ignored by hardware in 64-bit mode
6331032
Figure 3-1.General-Purpose Programming Registers
R11
R12
R13
R14
R15
rFLAGS
rIP
513-131.eps
3.1.1 Legacy RegistersIn legacy and compatibility modes, all of the legacy x86
registers are available. Figure 3-2 shows a detailed view of the
GPR, flag, and instruction-pointer registers.
28Chapter 3: General-Purpose Programming
Page 63
24592—Rev. 3.10—March 2005AMD64 Technology
register
encoding
0
3
1
2
6
7
5
4
low
high
8-bit
8-bit32-bit
AH (4)
BH (7)
CH (5)
DH (6)
3115016
310
AL
BL
CL
DL
SI
DI
BP
SP
FLAGS
IP
16-bit
AX
BX
CX
DX
SI
DI
BP
SP
FLAGSIPEFLAGS
EAX
EBX
ECX
EDX
ESI
EDI
EBP
ESP
EIP
513-311.eps
Figure 3-2.General Registers in Legacy and Compatibility Modes
Eight 32-bit registers (EAX, EBX, ECX, EDX, EDI, ESI, EBP,
ESP).
The size of register used by an instruction depends on the
effective operand size or, for certain instructions, the opcode,
address size, or stack size. The 16-bit and 32-bit registers are
encoded as 0 through 7 in Figure 3-2. For opcodes that specify a
byte operand, registers encoded as 0 through 3 refer to the lowbyte registers (AL, BL, CL, DL) and registers encoded as 4
through 7 refer to the high-byte registers (AH, BH, CH, DH).
The 16-bit FLAGS register, which is also the low 16 bits of the
32-bit EFLAGS register, shown in Figure 3-2, contains control
and status bits accessible to application software, as described
in Section 3.1.4, “Flags Register,” on page 38. The 16-bit IP or
Chapter 3: General-Purpose Programming29
Chapter 3: General-Purpose Programming29
Page 64
AMD64 Technology24592—Rev. 3.10—March 2005
32-bit EIP instruction-pointer register contains the address of
the next instruction to be executed, as described in Section 2.5,
“Instruction Pointer,” on page 24.
3.1.2 64-Bit-Mode
Registers
In 64-bit mode, eight new GPRs are added to the eight legacy
GPRs, all 16 GPRs are 64 bits wide, and the low bytes of all
registers are accessible. Figure 3-3 on page 31 shows the GPRs,
flags register, and instruction-pointer register available in 64bit mode. The GPRs include:
The size of register used by an instruction depends on the
effective operand size or, for certain instructions, the opcode,
address size, or stack size. For most instructions, access to the
extended GPRs requires a REX prefix (Section 3.5.2, “REX
Prefixes,” on page 91). The four high-byte registers (AH, BH,
CH, DH) available in legacy mode are not addressable when a
REX prefix is used.
In general, byte and word operands are stored in the low 8 or 16
bits of GPRs without modifying their high 56 or 48 bits,
respectively. Doubleword operands, however, are normally
stored in the low 32 bits of GPRs and zero-extended to 64 bits.
The 64-bit RFLAGS register, shown in Figure 3-3 on page 31,
contains the legacy EFLAGS in its low 32-bit range. The high 32
bits are reserved. They can be written with anything but they
always read as zero (RAZ). The 64-bit RIP instruction-pointer
register contains the address of the next instruction to be
executed, as described in Section 3.1.5, “Instruction Pointer
Register,” on page 42.
30Chapter 3: General-Purpose Programming
Page 65
24592—Rev. 3.10—March 2005AMD64 Technology
not modified for 8-bit operands
not modified for 16-bit operands
register
encoding
zero-extended
for 32-bit operands
low
16-bit32-bit64-bit
8-bit
0
3
1
2
AH*
BH*
CH*
DH*
6
7
5
4
8
9
10
11
12
13
14
AL
BL
CL
DL
SIL**
DIL**
BPL**
SPL**
R8B
R9B
R10B
R11B
R12B
R13B
R14B
AX
BX
CX
DX
SI
DI
BP
SP
R8W
R9W
R10W
R11W
R12W
R13W
R14W
EAX
EBX
ECX
EDX
ESI
EDI
EBP
ESP
R8D
R9D
R10D
R11D
R12D
R13D
R14D
RAX
RBX
RCX
RDX
RSI
RDI
RBP
RSP
R8
R9
R10
R11
R12
R13
R14
15
6331157081632
0
R15B
R15W
RFLAGS
R15D
R15
513-309.eps
RIP
6331032
* Not addressable when
a REX prefix is used.
** Only addressable when
a REX prefix is used.
Figure 3-3.General Registers in 64-Bit Mode
Figure 3-4 on page 32 illustrates another way of viewing the 64bit-mode GPRs, showing how the legacy GPRs overlap the
extended GPRs. Gray-shaded bits are not modified in 64-bit
mode.
Chapter 3: General-Purpose Programming31
Chapter 3: General-Purpose Programming31
Page 66
AMD64 Technology24592—Rev. 3.10—March 2005
63323116 158 70
Gray areas are not modified in 64-bit mode.AH*AL
0
3
1
2
6
7
Register Encoding
5
4
0EAX
RAX
0EBX
RBX
CH*CL
0ECX
RCX
DH*DL
0EDX
RDX
0ESI
RSI
0EDI
RDI
0EBP
RBP
0ESP
RSP
AX
BH*BL
BX
CX
DX
SIL**
SI
DIL**
DI
BPL**
BP
SPL**
SP
R8B
R8W
R15B
R15W
15
8
0R8D
…
0R15D
* Not addressable when a REX prefix is used.** Only addressable when a REX prefix is used.
R8
R15
Figure 3-4.GPRs in 64-Bit Mode
32Chapter 3: General-Purpose Programming
Page 67
24592—Rev. 3.10—March 2005AMD64 Technology
Default Operand Size. For most instructions, the default operand
size in 64-bit mode is 32 bits. To access 16-bit operand sizes, an
instruction must contain an operand-size prefix (66h), as
described in Section 3.2.2, “Operand Sizes and Overrides,” on
page 45. To access the full 64-bit operand size, most instructions
must contain a REX prefix.
For details on operand size, see Section 3.2.2, “Operand Sizes
and Overrides,” on page 45.
Byte Registers. 64-bit mode provides a uniform set of low-byte,
low-word, low-doubleword, and quadword registers that is wellsuited for register allocation by compilers. Access to the four
new low-byte registers in the legacy-GPR range (SIL, DIL, BPL,
SPL), or any of the low-byte registers in the extended registers
(R8B–R15B), requires a REX instruction prefix. However, the
legacy high-byte registers (AH, BH, CH, DH) are not accessible
when a REX prefix is used.
Zero-Extension of 32-Bit Results. As Figure 3-3 and Figure 3-4 show,
when performing 32-bit operations with a GPR destination in
64-bit mode, the processor zero-extends the 32-bit result into
the full 64-bit destination. 8-bit and 16-bit operations on GPRs
preserve all unwritten upper bits of the destination GPR. This
is consistent with legacy 16-bit and 32-bit semantics for partialwidth results.
Software should explicitly sign-extend the results of 8-bit, 16bit, and 32-bit operations to the full 64-bit width before using
the results in 64-bit address calculations.
The following four code examples show how 64-bit, 32-bit, 16bit, and 8-bit ADDs work. In these examples, “48” is a REX
prefix specifying 64-bit operand size, and “01C3” and “00C3”
are the opcode and ModRM bytes of each instruction (see
“Opcode Syntax” in Volume 3 for details on the opcode and
ModRM encoding).
Example 1: 64-bit Add:
Before:RAX =0002_0001_8000_2201
RBX =0002_0002_0123_3301
48 01C3 ADD RBX,RAX ;48 is a REX prefix for size.
Result:RBX = 0004_0003_8123_5502
Chapter 3: General-Purpose Programming33
Chapter 3: General-Purpose Programming33
Page 68
AMD64 Technology24592—Rev. 3.10—March 2005
Example 2: 32-bit Add:
Before:RAX = 0002_0001_8000_2201
RBX = 0002_0002_0123_3301
01C3 ADD EBX,EAX ;32-bit add
Result:RBX = 0000_0000_8123_5502
(32-bit result is zero extended)
Example 3: 16-bit Add:
Before:RAX = 0002_0001_8000_2201
RBX = 0002_0002_0123_3301
66 01C3 ADD BX,AX ;66 is 16-bit size override
Result:RBX = 0002_0002_0123_5502
(bits 63:16 are preserved)
Example 4: 8-bit Add:
3.1.3 Implicit Uses of
GPRs
Before:RAX = 0002_0001_8000_2201
RBX = 0002_0002_0123_3301
00C3 ADD BL,AL ;8-bit add
Result:RBX = 0002_0002_0123_3302
(bits 63:08 are preserved)
GPR High 32 Bits Across Mode Switches. The processor does not
preserve the upper 32 bits of the 64-bit GPRs across switches
from 64-bit mode to compatibility or legacy modes. When using
32-bit operands in compatibility or legacy mode, the high 32
bits of GPRs are undefined. Software must not rely on these
undefined bits, because they can change from one
implementation to the next or even on a cycle-to-cycle basis
within a given implementation. The undefined bits are not a
function of the data left by any previously running process.
Most instructions can use any of the GPRs for operands.
However, as Table 3-1 shows, some instructions use some GPRs
implicitly. Details about implicit use of GPRs are described in
“General-Purpose Instruction Reference” in Volume 3.
Table 3-1 on page 35 shows implicit register uses only for
application instructions. Certain system instructions also make
implicit use of registers. These system instructions are
described in “System Instruction Reference” in Volume 3.
34Chapter 3: General-Purpose Programming
Page 69
24592—Rev. 3.10—March 2005AMD64 Technology
Table 3-1.Implicit Uses of GPRs
Registers
1
Low 8-Bit16-Bit32-Bit64-Bit
ALAXEAX
BLBXEBX
CLCXECX
RAX
RBX
RCX
NameImplicit Uses
• Operand for decimal arithmetic,
multiply, divide, string, compareand-exchange, table-translation,
and I/O instructions.
2
Accumulator
• Special accumulator encoding for
ADD, XOR, and MOV instructions.
• Used with EDX to hold doubleprecision operands.
• CPUID processor-feature
information.
• Address generation in 16-bit
code.
2
Base
• Memory address for XLAT
instruction.
• CPUID processor-feature
information.
• Bit index for shift and rotate
instructions.
• Iteration count for loop and
2
Count
repeated string instructions.
• Jump conditional if zero.
• CPUID processor-feature
information.
• Operand for multiply and divide
instructions.
• Port number for I/O instructions.
DLDXEDX
RDX
2
I/O Address
• Used with EAX to hold doubleprecision operands.
• CPUID processor-feature
information.
Note:
1. Gray-shaded registers have no implicit uses.
2. Accessible only in 64-bit mode.
Chapter 3: General-Purpose Programming35
Chapter 3: General-Purpose Programming35
Page 70
AMD64 Technology24592—Rev. 3.10—March 2005
Table 3-1.Implicit Uses of GPRs (continued)
Registers
1
Low 8-Bit16-Bit32-Bit64-Bit
2
SIL
2
DIL
2
BPL
2
SPL
R8B–R10B
2
R11B
2
SIESI
DIEDI
BPEBP
SPESP
R8W–R10W
R11W
2
2
R8D–R10D
2
R11D
RSI
RDI
RBP
RSP
2
R8–R10
R11
NameImplicit Uses
• Memory address of source
2
Source Index
operand for string instructions.
• Memory index for 16-bit
addresses.
• Memory address of destination
2
Destination
Index
operand for string instructions.
• Memory index for 16-bit
addresses.
2
2
2
Base Pointer
Stack Pointer
2
NoneNo implicit uses
None
• Memory address of stack-frame
base pointer.
• Memory address of last stack
entry (top of stack).
• Holds the value of RFLAGS on
SYSCALL/SYSRET.
R12B–R15B2R12W–R15W2R12D–R15D
Note:
1. Gray-shaded registers have no implicit uses.
2. Accessible only in 64-bit mode.
Arithmetic Operations. Several forms of the add, subtract, multiply,
and divide instructions use AL or rAX implicitly. The multiply
and divide instructions also use the concatenation of rDX:rAX
for double-sized results (multiplies) or quotient and remainder
(divides).
Sign-Extensions. The instructions that double the size of operands
by sign extension (for example, CBW, CWDE, CDQE, CWD,
CDQ, CQO) use rAX register implicitly for the operand. The
CWD, CDQ, and CQO instructions also uses the rDX register.
Special MOVs. The MOV instruction has several opcodes that
implicitly use the AL or rAX register for one operand.
String Operations. Many types of string instructions use the
accumulators implicitly. Load string, store string, and scan
2
R12–R15
2
NoneNo implicit uses
36Chapter 3: General-Purpose Programming
Page 71
24592—Rev. 3.10—March 2005AMD64 Technology
string instructions use AL or rAX for data and rDI or rSI for the
offset of a memory address.
I/O-Address-Space Operations. The I/O and string I/O instructions
use rAX to hold data that is received from or sent to a device
located in the I/O-address space. DX holds the device I/Oaddress (the port number).
Table Translations. The table translate instruction (XLATB) uses
AL for an memory index and rBX for memory base address.
Compares and Exchanges. Compare and exchange instructions
(CMPXCHG) use the AL or rAX register for one operand.
Decimal Arithmetic. The decimal arithmetic instructions (AAA,
AAD, AAM, AAS, DAA, DAS) that adjust binary-coded decimal
(BCD) operands implicitly use the AL and AH register for their
operations.
Shifts and Rotates. Shift and rotate instructions can use the CL
register to specify the number of bits an operand is to be shifted
or rotated.
Conditional Jumps. Special conditional-jump instructions use the
rCX register instead of flags. The JCXZ and JrCXZ instructions
check the value of the rCX register and pass control to the
target instruction when the value of rCX register reaches 0.
Repeated String Operations. With the exception of I/O string
instructions, all string operations use rSI as the source-operand
pointer and rDI as the destination-operand pointer. I/O string
instructions use rDX to specify the input-port or output-port
number. For repeated string operations (those preceded with a
repeat-instruction prefix), the rSI and rDI registers are
incremented or decremented as the string elements are moved
from the source location to the destination. Repeat-string
operations also use rCX to hold the string length, and
decrement it as data is moved from one location to the other.
Stack Operations. Stack operations make implicit use of the rSP
register, and in some cases, the rBP register. The rSP register is
used to hold the top-of-stack pointer (or simply, stack pointer).
rSP is decremented when items are pushed onto the stack, and
incremented when they are popped off the stack. The ENTER
and LEAVE instructions use rBP as a stack-frame base pointer.
Chapter 3: General-Purpose Programming37
Chapter 3: General-Purpose Programming37
Page 72
AMD64 Technology24592—Rev. 3.10—March 2005
Here, rBP points to the last entry in a data structure that is
passed from one block-structured procedure to another.
The use of rSP or rBP as a base register in an address
calculation implies the use of SS (stack segment) as the default
segment. Using any other GPR as a base register without a
segment-override prefix implies the use of the DS data segment
as the default segment.
The push all and pop all instructions (PUSHA, PUSHAD, POPA,
POPAD) implicitly use all of the GPRs.
CPUID Information. The CPUID instruction makes implicit use of
the EAX, EBX, ECX, and EDX registers. Software loads a
function code into EAX, executes the CPUID instruction, and
then reads the associated processor-feature information in
EAX, EBX, ECX, and EDX.
3.1.4 Flags RegisterFigure 3-5 on page 39 shows the 64-bit RFLAGS register and the
flag bits visible to application software. Bits 15–0 are the
FLAGS register (accessed in legacy real and virtual-8086
modes), bits 31–0 are the EFLAGS register (accessed in legacy
protected mode and compatibility mode), and bits 63–0 are the
RFLAGS register (accessed in 64-bit mode). The name rFLAGS
refers to any of the three register widths, depending on the
current software context.
38Chapter 3: General-Purpose Programming
Page 73
24592—Rev. 3.10—March 2005AMD64 Technology
3263
Reserved, Read as Zero (RAZ)
1516
See Volume 2 for System Flags
Reserved or System Flag
SymbolDescriptionBit
OFOverflow Flag11
DFDirection Flag10
SFSign Flag 7
ZFZero Flag 6
AFAuxiliary Carry Flag 4
PFParity Flag 2
CFCarry Flag 0
F
Figure 3-5.rFLAGS Register—Flags Visible to Application Software
The low 16 bits (FLAGS portion) of rFLAGS are accessible by
application software and hold the following flags:
One control flag (the direction flag DF).
987654321010111231
A
DFO
ZFS
F
P
F
C
F
F
Six status flags (carry flag CF, parity flag PF, auxiliary carry
flag AF, zero flag ZF, sign flag SF, and overflow flag OF).
The direction flag (DF) flag controls the direction of string
operations. The status flags provide result information from
logical and arithmetic operations and control information for
conditional move and jump instructions.
Bits 31–16 of the rFLAGS register contain flags that are
accessible only to system software. These flags are described in
“System Registers” in Volume 2. The highest 32 bits of
RFLAGS are reserved. In 64-bit mode, writes to these bits are
ignored. They are read as zeros (RAZ). The rFLAGS register is
initialized to 02h on reset, so that all of the programmable bits
are cleared to zero.
Chapter 3: General-Purpose Programming39
Chapter 3: General-Purpose Programming39
Page 74
AMD64 Technology24592—Rev. 3.10—March 2005
The effects that rFLAGS bit-values have on instructions are
summarized in the following places:
Conditional Moves (CMOVcc)—Table 3-4 on page 52.
Conditional Jumps (Jcc)—Table 3-5 on page 67.
Conditional Sets (SETcc)—Table 3-6 on page 72.
The effects that instructions have on rFLAGS bit-values are
summarized in “Instruction Effects on RFLAGS” in Volume 3.
The sections below describe each application-visible flag. All of
these flags are readable and writable. For example, the POPF,
POPFD, POPFQ, IRET, IRETD, and IRETQ instructions write
all flags. The carry and direction flags are writable by dedicated
application instructions. Other application-visible flags are
written indirectly by specific instructions. Reserved bits and
bits whose writability is prevented by the current values of
system flags, current privilege level (CPL), or the current
operating mode, are unaffected by the POPFx instructions.
Carry Flag (CF). Bit 0. Hardware sets the carry flag to 1 if the last
integer addition or subtraction operation resulted in a carry
(for addition) or a borrow (for subtraction) out of the mostsignificant bit position of the result. Otherwise, hardware clears
the flag to 0.
The increment and decrement instructions—unlike the
addition and subtraction instructions—do not affect the carry
flag. The bit shift and bit rotate instructions shift bits of
operands into the carry flag. Logical instructions like AND, OR,
XOR clear the carry flag. Bit-test instructions (BTx) set the
value of the carry flag depending on the value of the tested bit
of the operand.
Software can set or clear the carry flag with the STC and CLC
instructions, respectively. Software can complement the flag
with the CMC instruction.
Parity Flag (PF). Bit 2. Hardware sets the parity flag to 1 if there is
an even number of 1 bits in the least-significant byte of the last
result of certain operations. Otherwise (i.e., for an odd number
of 1 bits), hardware clears the flag to 0. Software can read the
flag to implement parity checking.
Auxiliary Carry Flag (AF). Bit 4. Hardware sets the auxiliary carry
flag to 1 if the last binary-coded decimal (BCD) operation
40Chapter 3: General-Purpose Programming
Page 75
24592—Rev. 3.10—March 2005AMD64 Technology
resulted in a carry (for addition) or a borrow (for subtraction)
out of bit 3. Otherwise, hardware clears the flag to 0.
The main application of this flag is to support decimal
arithmetic operations. Most commonly, this flag is used
internally by correction commands for decimal addition (AAA)
and subtraction (AAS).
Zero Flag (ZF). Bit 6. Hardware sets the zero flag to 1 if the last
arithmetic operation resulted in a value of zero. Otherwise (for
a non-zero result), hardware clears the flag to 0. The compare
and test instructions also affect the zero flag.
The zero flag is typically used to test whether the result of an
arithmetic or logical operation is zero, or to test whether two
operands are equal.
Sign Flag (SF). Bit 7. Hardware sets the sign flag to 1 if the last
arithmetic operation resulted in a negative value. Otherwise
(for a positive-valued result), hardware clears the flag to 0.
Thus, in such operations, the value of the sign flag is set equal
to the value of the most-significant bit of the result. Depending
on the size of operands, the most-significant bit is bit 7 (for
bytes), bit 15 (for words), bit 31 (for doublewords), or bit 63 (for
quadwords).
Direction Flag (DF). Bit 10. The direction flag determines the order
in which strings are processed. Software can set the direction
flag to 1 to specify decrementing the data pointer for the next
string instruction (LODSx, STOSx, MOVSx, SCASx, CMPSx,
OUTSx, or INSx). Clearing the direction flag to 0 specifies
incrementing the data pointer. The pointers are stored in the
rSI or rDI register. Software can set or clear the flag with the
STD and CLD instructions, respectively.
Overflow Flag (OF). Bit 11. Hardware sets the overflow flag to 1 to
indicate that the most-significant (sign) bit of the result of the
last signed integer operation differed from the signs of both
source operands. Otherwise, hardware clears the flag to 0. A set
overflow flag means that the magnitude of the positive or
negative result is too big (overflow) or too small (underflow) to
fit its defined data type.
The OF flag is undefined after the DIV instruction and after a
shift of more than one bit. Logical instructions clear the
overflow flag.
Chapter 3: General-Purpose Programming41
Chapter 3: General-Purpose Programming41
Page 76
AMD64 Technology24592—Rev. 3.10—March 2005
3.1.5 Instruction
Pointer Register
The instruction pointer register—IP, EIP, or RIP, or simply rIP
for any of the three depending on the context—is used in
conjunction with the code-segment (CS) register to locate the
next instruction in memory. See Section 2.5, “Instruction
Pointer,” on page 24 for details.
3.2Operands
Operands are either referenced by an instruction's encoding or
included as an immediate value in the instruction encoding.
Depending on the instruction, referenced operands can be
located in registers, memory locations, or I/O ports.
3.2.1 Data TypesFigure 3-6 on page 43 shows the register images of the general-
purpose data types. In the general-purpose programming
environment, these data types can be interpreted by instruction
syntax or the software context as the following types of
numbers and strings:
Signed (two's-complement) integers.
Unsigned integers.
BCD digits.
Packed BCD digits.
Strings, including bit strings.
The double quadword data type is supported in the RDX:RAX
registers by the MUL, IMUL, DIV, IDIV, and CQO instructions.
Software can interpret the data types in ways other than those
shown in Figure 3-6 on page 43 but the AMD64 instruction set
does not directly support such interpretations and software
must handle them entirely on its own.
Table 3-2 on page 44 shows the range of representable values
for the general-purpose data types.
42Chapter 3: General-Purpose Programming
Page 77
24592—Rev. 3.10—March 2005AMD64 Technology
127
s
127
Signed Integer
16 bytes (64-bit mode only)
s
63
Unsigned Integer
16 bytes (64-bit mode only)
63
8 bytes (64-bit mode only)
s
31
8 bytes (64-bit mode only)
31
4 bytes
s
15
4 bytes
15
2 bytes
s
70
2 bytes
0
Double
Quadword
Quadword
Doubleword
Word
Byte
0
Double
Quadword
Quadword
Doubleword
Word
Byte
Packed BCD
513-326.eps
Figure 3-6.General-Purpose Data Types
Signed and Unsigned Integers. The architecture supports signed and
unsigned 1 byte, 2 bytes, 4 byte and 8 byte integers. The sign bit
is stored in the most significant bit.
73
BCD Digit
Bit
0
Chapter 3: General-Purpose Programming43
Chapter 3: General-Purpose Programming43
Page 78
AMD64 Technology24592—Rev. 3.10—March 2005
Table 3-2.Representable Values of General-Purpose Data Types
Data TypeByteWordDoublewordQuadword
1
Signed Integers
Unsigned Integers
Packed BCD Digits
BCD Digit
Note:
1. The sign bit is the most-significant bit (e.g., bit 7 for a byte, bit 15 for a word, etc.).
2. The double quadword data type is supported in the RDX:RAX registers by the MUL, IMUL, DIV, IDIV, and CQO instructions.
-27 to +(27 -1)-215 to +(215 -1)-231 to +(231 -1)-263 to +(263 -1)-2
8
0 to +2
(0 to 255)
-1
00 to 99multiple packed BCD-digit bytes
0 to 9multiple BCD-digit bytes
0 to +216-1
(0 to 65,535)
0 to +232-1
(0 to 4.29 x 10
9
(0 to 1.84 x 10
)
0 to +2
64
-1
19
)
Double
Quadword
127
to +(2
0 to +2
(0 to 3.40 x 10
Binary-Coded-Decimal (BCD) Digits. BCD digits have values ranging
from 0 to 9. These values can be represented in binary encoding
with four bits. For example, 0000b represents the decimal
number 0 and 1001b represents the decimal number 9. Values
ranging from 1010b to 1111b are invalid for this data type.
Because a byte contains eight bits, two BCD digits can be stored
in a single byte. This is referred to as packed-BCD. If a single
BCD digit is stored per byte, it is referred to as unpacked-BCD.
In the x87 floating-point programming environment (described
in Section 6, “x87 Floating-Point Programming,” on page 293)
an 80-bit packed BCD data type is also supported, along with
conversions between floating-point and BCD data types, so that
data expressed in the BCD format can be operated on as
floating-point values.
128
127
-1
2
-1)
38
)
Integer add, subtract, multiply, and divide instructions can be
used to operate on single (unpacked) BCD digits. The result
must be adjusted to produce a correct BCD representation. For
unpacked BCD numbers, the ASCII-adjust instructions are
provided to simplify that correction. In the case of division, the
adjustment must be made prior to executing the integer-divide
instruction.
Similarly, integer add and subtract instructions can be used to
operate on packed-BCD digits. The result must be adjusted to
produce a correct packed-BCD representation. Decimal-adjust
44Chapter 3: General-Purpose Programming
Page 79
24592—Rev. 3.10—March 2005AMD64 Technology
instructions are provided to simplify packed-BCD result
corrections.
Strings. Strings are a continuous sequence of a single data type.
The string instructions can be used to operate on byte, word,
doubleword, or quadword data types. The maximum length of a
string of any data type is 232–1 bytes, in legacy or compatibility
modes, or 264–1 bytes in 64-bit mode. One of the more common
types of strings used by applications are byte data-type strings
known as ASCII strings, which can be used to represent
character data.
Bit strings are also supported by instructions that operate
specifically on bit strings. In general, bit strings can start and
end at any bit location within any byte, although the BTx bitstring instructions assume that strings start on a byte boundary.
The length of a bit string can range in size from a single bit up
to 232–1 bits, in legacy or compatibility modes, or 264-–1 bits in
64-bit mode.
3.2.2 Operand Sizes
and Overrides
Default Operand Size. In legacy and compatibility modes, the
default operand size is either 16 bits or 32 bits, as determined
by the default-size (D) bit in the current code-segment
descriptor (for details, see “Segmented Virtual Memory” in
Volume 2). In 64-bit mode, the default operand size for most
instructions is 32 bits.
Application software can override the default operand size by
using an operand-size instruction prefix. Table 3-3 on page 46
shows the instruction prefixes for operand-size overrides in all
operating modes. In 64-bit mode, the default operand size for
most instructions is 32 bits. A REX prefix (see Section 3.5.2,
“REX Prefixes,” on page 91) specifies a 64-bit operand size, and
a 66h prefix specifies a 16-bit operand size. The REX prefix
takes precedence over the 66h prefix.
Chapter 3: General-Purpose Programming45
Chapter 3: General-Purpose Programming45
Page 80
AMD64 Technology24592—Rev. 3.10—March 2005
Table 3-3.Operand-Size Overrides
Default
Operating Mode
64-Bit
Mode
Long
Mode
Compatibility
Mode
Legacy Mode
(Protected, Virtual-8086,
or Real Mode)
Note:
1. A “no” indicates that the default operand size is used. An “x” means “don’t care.”
2. Near branches, instructions that implicitly reference the stack pointer, and certain other
instructions default to 64-bit operand size. See “General-Purpose Instructions in 64-Bit Mode”
in Volume 3
Operand
Size (Bits)
2
32
32
16
32
16
Effective
Operand
Size
(Bits)
64xyes
32nono
16yesno
32no
16yes
32yes
16n o
32no
16yes
32yes
16n o
Instruction Prefix
1
66h
Applicable
REX
Not
There are several exceptions to the 32-bit operand-size default
in 64-bit mode, including near branches and instructions that
implicitly reference the RSP stack pointer. For example, the
near CALL, near JMP, Jcc, LOOPcc, POP, and PUSH
instructions all default to a 64-bit operand size in 64-bit mode.
Such instructions do not need a REX prefix for the 64-bit
operand size. For details, see “General-Purpose Instructions in
64-Bit Mode” in Volume 3.
Effective Operand Size. The term effective operand size describes the
operand size for the current instruction, after accounting for
the instruction’s default operand size and any operand-size
override or REX prefix that is used with the instruction.
46Chapter 3: General-Purpose Programming
Page 81
24592—Rev. 3.10—March 2005AMD64 Technology
Immediate Operand Size. In legacy mode and compatibility modes,
the size of immediate operands can be 8, 16, or 32 bits,
depending on the instruction. In 64-bit mode, the maximum size
of an immediate operand is also 32 bits, except that 64-bit
immediates can be copied into a 64-bit GPR using the MOV
instruction.
When the operand size of a MOV instruction is 64 bits, the
processor sign-extends immediates to 64 bits before using them.
Support for true 64-bit immediates is accomplished by
expanding the semantics of the MOV reg, imm16/32 instructions.
In legacy and compatibility modes, these instructions—opcodes
B8h through BFh—copy a 16-bit or 32-bit immediate
(depending on the effective operand size) into a GPR. In 64-bit
mode, if the operand size is 64 bits (requires a REX prefix),
these instructions can be used to copy a true 64-bit immediate
into a GPR.
3.2.3 Operand
Addressing
Operands for general-purpose instructions are referenced by
the instruction's syntax or they are incorporated in the
instruction as an immediate value. Referenced operands can be
in registers, memory, or I/O ports.
Register Operands. Most general-purpose instructions that take
register operands reference the general-purpose registers
(GPRs). A few general-purpose instructions reference operands
in the RFLAGS register, XMM registers, or MMX™ registers.
The type of register addressed is specified in the instruction
syntax. When addressing GPRs or XMM registers, the REX
instruction prefix can be used to access the extended GPRs or
XMM registers, as described in Section 3.5, “Instruction
Prefixes,” on page 87.
Memory Operands. Many general-purpose instructions can access
operands in memory. Section 2.2, “Memory Addressing,” on
page 16 describes the general methods and conditions for
addressing memory operands.
I/O Ports. Operands in I/O ports are referenced according to the
conventions described in Section 3.8, “Input/Output,” on
page 111.
Immediate Operands. In certain instructions, a source operand—
called an immediate operand, or simply immediate—is included
Chapter 3: General-Purpose Programming47
Chapter 3: General-Purpose Programming47
Page 82
AMD64 Technology24592—Rev. 3.10—March 2005
as part of the instruction rather than being accessed from a
register or memory location. For details on the size of
immediate operands, see “Immediate Operand Size” on
page 47.
3.2.4 Data AlignmentA data access is aligned if its address is a multiple of its operand
size, in bytes. The following examples illustrate this definition:
Byte accesses are always aligned. Bytes are the smallest
addressable parts of memory.
Word (two-byte) accesses are aligned if their address is a
multiple of 2.
Doubleword (four-byte) accesses are aligned if their address
is a multiple of 4.
Quadword (eight-byte) accesses are aligned if their address
is a multiple of 8.
The AMD64 architecture does not impose data-alignment
requirements for accessing data in memory. However,
depending on the location of the misaligned operand with
respect to the width of the data bus and other aspects of the
hardware implementation (such as store-to-load forwarding
mechanisms), a misaligned memory access can require more
bus cycles than an aligned access. For maximum performance,
avoid misaligned memory accesses.
Performance on many hardware implementations will benefit
from observing the following operand-alignment and operandsize conventions:
Avoid misaligned data accesses.
Maintain consistent use of operand size across all loads and
stores. Larger operand sizes (doubleword and quadword)
tend to make more efficient use of the data bus and any
data-forwarding features that are implemented by the
hardware.
When using word or byte stores, avoid loading data from the
same doubleword of memory, other than the identical start
addresses of the stores.
48Chapter 3: General-Purpose Programming
Page 83
24592—Rev. 3.10—March 2005AMD64 Technology
3.3Instruction Summary
This section summarizes the functions of the general-purpose
instructions. The instructions are organized by functional
group—such as, data-transfer instructions, arithmetic
instructions, and so on. Details on individual instructions are
given in the alphabetically organized “General-Purpose
Instruction Reference” in Volume 3.
3.3.1 SyntaxEach instruction has a mnemonic syntax used by assemblers to
specify the operation and the operands to be used for source
and destination (result) data. Figure 3-7 shows an example of
the mnemonic syntax for a compare (CMP) instruction. In this
example, the CMP mnemonic is followed by two operands, a 32bit register or memory operand and an 8-bit immediate
operand.
CMP reg/mem32, imm8
Mnemonic
First Source Operand
and Destination Operand
Second Source Operand
513-139.eps
Figure 3-7.Mnemonic Syntax Example
In most instructions that take two operands, the first (left-most)
operand is both a source operand and the destination operand.
The second (right-most) operand serves only as a source.
Instructions can have one or more prefixes that modify default
instruction functions or operand properties. These prefixes are
summarized in Section 3.5, “Instruction Prefixes,” on page 87.
Instructions that access 64-bit operands in a general-purpose
register (GPR) or any of the extended GPR or XMM registers
require a REX instruction prefix.
Unless otherwise stated in this section, the word register means
a general-purpose register (GPR). Several instructions affect
the flag bits in the RFLAGS register. “Instruction Effects on
Chapter 3: General-Purpose Programming49
Chapter 3: General-Purpose Programming49
Page 84
AMD64 Technology24592—Rev. 3.10—March 2005
RFLAGS” in Volume 3 summarizes the effects that instructions
have on rFLAGS bits.
3.3.2 Data TransferThe data-transfer instructions copy data between registers and
memory.
Move.
MOV—Move
MOVSX—Move with Sign-Extend
MOVZX—Move with Zero-Extend
MOVD—Move Doubleword or Quadword
MOVNTI—Move Non-Temporal Doubleword or Quadword
MOVx copies a byte, word, doubleword, or quadword from a
register or memory location to a register or memory location.
The source and destination cannot both be memory locations.
An immediate constant can be used as a source operand with
the MOV instruction. For MOV, the destination must be of the
same size as the source, but the MOVSX and MOVZX
instructions copy values of smaller size to a larger size by using
sign-extension or zero-extension. The MOVD instruction copies
a doubleword or quadword between a general-purpose register
or memory and an XMM or MMX register.
The MOV instruction is in many aspects similar to the
assignment operator in high-level languages. The simplest
example of their use is to initialize variables. To initialize a
register to 0, rather than using a MOV instruction it may be
more efficient to use the XOR instruction with identical
destination and source operands.
The MOVNTI instruction stores a doubleword or quadword from
a register into memory as “non-temporal” data, which assumes
a single access (as opposed to frequent subsequent accesses of
“temporal data”). The operation therefore minimizes cache
pollution. The exact method by which cache pollution is
minimized depends on the hardware implementation of the
instruction. For further information, see Section 3.9, “Memory
Optimization,” on page 115.
Conditional Move.
CMOVcc—Conditional Move If condition
50Chapter 3: General-Purpose Programming
Page 85
24592—Rev. 3.10—March 2005AMD64 Technology
The CMOVcc instructions conditionally copy a word,
doubleword, or quadword from a register or memory location to
a register location. The source and destination must be of the
same size.
The CMOVcc instructions perform the same task as MOV but
work conditionally, depending on the state of status flags in the
RFLAGS register. If the condition is not satisfied, the
instruction has no effect and control is passed to the next
instruction. The mnemonics of CMOVcc instructions indicate
the condition that must be satisfied. Several mnemonics are
often used for one opcode to make the mnemonics easier to
remember. For example, CMOVE (conditional move if equal)
and CMOVZ (conditional move if zero) are aliases and compile
to the same opcode. Table 3-4 on page 52 shows the RFLAGS
values required for each CMOVcc instruction.
In assembly languages, the conditional move instructions
correspond to small conditional statements like:
IF a = b THEN x = y
CMOVcc instructions can replace two instructions—a
conditional jump and a move. For example, to perform a highlevel statement like:
IF ECX = 5 THEN EAX = EBX
without a CMOVcc instruction, the code would look like:
cmp ecx, 5; test if ecx equals 5
jnz Continue; test condition and skip if not met
mov eax, ebx; move
Continue:; continuation
but with a CMOVcc instruction, the code would look like:
cmp ecx, 5; test if ecx equals to 5
cmovz eax, ebx; test condition and move
Replacing conditional jumps with conditional moves also has
the advantage that it can avoid branch-prediction penalties that
may be caused by conditional jumps.
Support for CMOVcc instructions depends on the processor
implementation. To find out if a processor is able to perform
CMOVcc instructions, use the CPUID instruction.
Chapter 3: General-Purpose Programming51
Chapter 3: General-Purpose Programming51
Page 86
AMD64 Technology24592—Rev. 3.10—March 2005
Table 3-4.rFLAGS for CMOVcc Instructions
Mnemonic
CMOVOOF = 1Conditional move if overflow
CMOVNOOF = 0Conditional move if not overflow
CMOVB
CMOVC
CMOVNAE
CMOVAE
CMOVNB
CMOVNC
CMOVE
CMOVZ
CMOVNE
CMOVNZ
CMOVBE
CMOVNA
CMOVA
CMOVNBE
Required Flag
State
CF = 1
CF = 0
ZF = 1
ZF = 0
CF = 1 or
ZF = 1
CF = 0 and
ZF = 0
Description
Conditional move if below
Conditional move if carry
Conditional move if not above or equal
Conditional move if above or equal
Conditional move if not below
Conditional move if not carry
Conditional move if equal
Conditional move if zero
Conditional move if not equal
Conditional move if not zero
Conditional move if below or equal
Conditional move if not above
Conditional move if not below or equal
Conditional move if not below or equal
CMOVSSF = 1Conditional move if sign
CMOVNSSF = 0Conditional move if not sign
CMOVP
CMOVPE
CMOVNP
CMOVPO
CMOVL
CMOVNGE
CMOVGE
CMOVNL
CMOVLE
CMOVNG
CMOVG
CMOVNLE
PF = 1
PF = 0
SF <> OF
SF = OF
ZF = 1 or
SF <> OF
ZF = 0 and
SF = OF
Conditional move if parity
Conditional move if parity even
Conditional move if not parity
Conditional move if parity odd
Conditional move if less
Conditional move if not greater or equal
Conditional move if greater or equal
Conditional move if not less
Conditional move if less or equal
Conditional move if not greater
Conditional move if greater
Conditional move if not less or equal
52Chapter 3: General-Purpose Programming
Page 87
24592—Rev. 3.10—March 2005AMD64 Technology
Stack Operations.
POP—Pop Stack
POPA—Pop All to GPR Words
POPAD—Pop All to GPR Doublewords
PUSH—Push onto Stack
PUSHA—Push All GPR Words onto Stack
PUSHAD—Push All GPR Doublewords onto Stack
ENTER—Create Procedure Stack Frame
LEAVE—Delete Procedure Stack Frame
PUSH copies the specified register, memory location, or
immediate value to the top of stack. This instruction
decrements the stack pointer by 2, 4, or 8, depending on the
operand size, and then copies the operand into the memory
location pointed to by SS:rSP.
POP copies a word, doubleword, or quadword from the memory
location pointed to by the SS:rSP registers (the top of stack) to a
specified register or memory location. Then, the rSP register is
incremented by 2, 4, or 8. After the POP operation, rSP points
to the new top of stack.
PUSHA or PUSHAD stores eight word-sized or doublewordsized registers onto the stack: eAX, eCX, eDX, eBX, eSP, eBP,
eSI and eDI, in that order. The stored value of eSP is sampled at
the moment when the PUSHA instruction started. The resulting
stack-pointer value is decremented by 16 or 32.
POPA or POPAD extracts eight word-sized or doubleword-sized
registers from the stack: eDI, eSI, eBP, eSP, eBX, eDX, eCX and
eAX, in that order (which is the reverse of the order used in the
PUSHA instruction). The stored eSP value is ignored by the
POPA instruction. The resulting stack pointer value is
incremented by 16 or 32.
It is a common practice to use PUSH instructions to pass
parameters (via the stack) to functions and subroutines. The
typical instruction sequence used at the beginning of a
subroutine looks like:
pushebp; save current EBP
movebp, esp; set stack frame pointer value
subesp, N; allocate space for local variables
Chapter 3: General-Purpose Programming53
Chapter 3: General-Purpose Programming53
Page 88
AMD64 Technology24592—Rev. 3.10—March 2005
The rBP register is used as a stack frame pointer—a base address
of the stack area used for parameters passed to subroutines and
local variables. Positive offsets of the stack frame pointed to by
rBP provide access to parameters passed while negative offsets
give access to local variables. This technique allows creating reentrant subroutines.
The ENTER and LEAVE instructions provide support for
procedure calls, and are mainly used in high-level languages.
The ENTER instruction is typically the first instruction of the
procedure, and the LEAVE instruction is the last before the
RET instruction.
The ENTER instruction creates a stack frame for a procedure.
The first operand, size, specifies the number of bytes allocated
in the stack. The second operand, depth, specifies the number of
stack-frame pointers copied from the calling procedure’s stack
(i.e., the nesting level). The depth should be an integer in the
range 0–31.
Typically, when a procedure is called, the stack contains the
following four components:
Parameters passed to the called procedure (created by the
calling procedure).
Return address (created by the CALL instruction).
Array of stack-frame pointers (pointers to stack frames of
procedures with smaller nesting-level depth) which are used
to access the local variables of such procedures.
Local variables used by the called procedure.
All these data are called the stack frame. The ENTER
instruction simplifies management of the last two components
of a stack frame. First, the current value of the rBP register is
pushed onto the stack. The value of the rSP register at that
moment is a frame pointer for the current procedure: positive
offsets from this pointer give access to the parameters passed to
the procedure, and negative offsets give access to the local
variables which will be allocated later. During procedure
execution, the value of the frame pointer is stored in the rBP
register, which at that moment contains a frame pointer of the
calling procedure. This frame pointer is saved in a temporary
register. If the depth operand is greater than one, the array of
depth-1 frame pointers of procedures with smaller nesting level
is pushed onto the stack. This array is copied from the stack
54Chapter 3: General-Purpose Programming
Page 89
24592—Rev. 3.10—March 2005AMD64 Technology
frame of the calling procedure, and it is addressed by the rBP
register from the calling procedure. If the depth operand is
greater than zero, the saved frame pointer of the current
procedure is pushed onto the stack (forming an array of depth
frame pointers). Finally, the saved value of the frame pointer is
copied to the rBP register, and the rSP register is decremented
by the value of the first operand, allocating space for local
variables used in the procedure. See “Stack Operations” on
page 53 for a parameter-passing instruction sequence using
PUSH that is equivalent to ENTER.
The LEAVE instruction removes local variables and the array of
frame pointers, allocated by the previous ENTER instruction,
from the stack frame. This is accomplished by the following two
steps: first, the value of the frame pointer is copied from the
rBP register to the rSP register. This releases the space
allocated by local variables and an array of frame pointers of
procedures with smaller nesting levels. Second, the rBP register
is popped from the stack, restoring the previous value of the
frame pointer (or simply the value of the rBP register, if the
depth operand is zero). Thus, the LEAVE instruction is
equivalent to the following code:
mov rSP, rBP
pop rBP
3.3.3 Data ConversionThe data-conversion instructions perform various
transformations of data, such as operand-size doubling by sign
extension, conversion of little-endian to big-endian format,
extraction of sign masks, searching a table, and support for
operations with decimal numbers.
Sign Extension.
CBW—Convert Byte to Word
CWDE—Convert Word to Doubleword
CDQE—Convert Doubleword to Quadword
CWD—Convert Word to Doubleword
CDQ—Convert Doubleword to Quadword
CQO—Convert Quadword to Octword
The CBW, CWDE, and CDQE instructions sign-extend the AL,
AX, or EAX register to the upper half of the AX, EAX, or RAX
register, respectively. By doing so, these instructions create a
double-sized destination operand in rAX that has the same
Chapter 3: General-Purpose Programming55
Chapter 3: General-Purpose Programming55
Page 90
AMD64 Technology24592—Rev. 3.10—March 2005
numerical value as the source operand. The CBW, CWDE, and
CDQE instructions have the same opcode, and the action taken
depends on the effective operand size.
The CWD, CDQ and CQO instructions sign-extend the AX, EAX,
or RAX register to all bit positions of the DX, EDX, or RDX
register, respectively. By doing so, these instructions create a
double-sized destination operand in rDX:rAX that has the same
numerical value as the source operand. The CWD, CDQ, and
CQO instructions have the same opcode, and the action taken
depends on the effective operand size.
Flags are not affected by these instructions. The instructions
can be used to prepare an operand for signed division
(performed by the IDIV instruction) by doubling its storage
size.
The MOVMSKPS instruction moves the sign bits of four packed
single-precision floating-point values in an XMM register to the
four low-order bits of a general-purpose register, with zeroextension. MOVMSKPD does a similar operation for two
packed double-precision floating-point values: it moves the two
sign bits to the two low-order bits of a general-purpose register,
with zero-extension. The result of either instruction is a sign-bit
mask.
Translate.
XLAT—Translate Table Index
The XLAT instruction replaces the value stored in the AL
register with a table element. The initial value in AL serves as
an unsigned index into the table, and the start (base) of table is
specified by the DS:rBX registers (depending on the effective
address size).
This instruction is not recommended. The following instruction
serves to replace it:
MOV AL,[rBX + AL]
56Chapter 3: General-Purpose Programming
Page 91
24592—Rev. 3.10—March 2005AMD64 Technology
ASCII Adjust.
AAA—ASCII Adjust After Addition
AAD—ASCII Adjust Before Division
AAM—ASCII Adjust After Multiply
AAS—ASCII Adjust After Subtraction
The AAA, AAD, AAM, and AAS instructions perform
corrections of arithmetic operations with non-packed BCD
values (i.e., when the decimal digit is stored in a byte register).
There are no instructions which directly operate on decimal
numbers (either packed or non-packed BCD). However, the
ASCII-adjust instructions correct decimal-arithmetic results.
These instructions assume that an arithmetic instruction, such
as ADD, was performed on two BCD operands, and that the
result was stored in the AL or AX register. This result can be
incorrect or it can be a non-BCD value (for example, when a
decimal carry occurs). After executing the proper ASCII-adjust
instruction, the AX register contains a correct BCD
representation of the result. (The AAD instruction is an
exception to this, because it should be applied before a DIV
instruction, as explained below). All of the ASCII-adjust
instructions are able to operate with multiple-precision decimal
values.
AAA should be applied after addition of two non-packed
decimal digits. AAS should be applied after subtraction of two
non-packed decimal digits. AAM should be applied after
multiplication of two non-packed decimal digits. AAD should be
applied before the division of two non-packed decimal numbers.
Although the base of the numeration for ASCII-adjust
instructions is assumed to be 10, the AAM and AAD
instructions can be used to correct multiplication and division
with other bases.
BCD Adjust.
DAA—Decimal Adjust after Addition
DAS—Decimal Adjust after Subtraction
The DAA and DAS instructions perform corrections of addition
and subtraction operations on packed BCD values. (Packed BCD
values have two decimal digits stored in a byte register, with the
higher digit in the higher four bits, and the lower one in the
Chapter 3: General-Purpose Programming57
Chapter 3: General-Purpose Programming57
Page 92
AMD64 Technology24592—Rev. 3.10—March 2005
lower four bits.) There are no instructions for correction of
multiplication and division with packed BCD values.
DAA should be applied after addition of two packed-BCD
numbers. DAS should be applied after subtraction of two
packed-BCD numbers.
DAA and DAS can be used in a loop to perform addition or
subtraction of two multiple-precision decimal numbers stored
in packed-BCD format. Each loop cycle would operate on
corresponding bytes (containing two decimal digits) of
operands.
Endian Conversion.
BSWAP—Byte Swap
The BSWAP instruction changes the byte order of a doubleword
or quadword operand in a register, as shown in Figure 3-8. In a
doubleword, bits 7–0 are exchanged with bits 31–24, and bits
15–8 are exchanged with bits 23–16. In a quadword, bits 7–0 are
exchanged with bits 63–56, bits 15–8 with bits 55–48, bits 23–16
with bits 47–40, and bits 31–24 with bits 39–32. See the
following illustration.
Figure 3-8.BSWAP Doubleword Exchange
A second application of the BSWAP instruction to the same
operand restores its original value. The result of applying the
BSWAP instruction to a 16-bit register is undefined. To swap
bytes of a 16-bit register, use the XCHG instruction.
The BSWAP instruction is used to convert data between littleendian and big-endian byte order.
0781516233124
0781516233124
58Chapter 3: General-Purpose Programming
Page 93
24592—Rev. 3.10—March 2005AMD64 Technology
3.3.4 Load Segment
Registers
These instructions load segment registers.
LDS, LES, LFS, LGS, LSS—Load Far Pointer
MOV segReg—Move Segment Register
POP segReg—Pop Stack Into Segment Register
The LDS, LES, LFD, LGS, and LSS instructions atomically load
the two parts of a far pointer into a segment register and a
general-purpose register. A far pointer is a 16-bit segment
selector and a 16-bit or 32-bit offset. The load copies the
segment-selector portion of the pointer from memory into the
segment register and the offset portion of the pointer from
memory into a general-purpose register.
The effective operand size determines the size of the offset
loaded by the LDS, LES, LFD, LGS, and LSS instructions. The
instructions load not only the software-visible segment selector
into the segment register, but they also cause the hardware to
load the associated segment-descriptor information into the
software-invisible (hidden) portion of that segment register.
The MOV segReg and POP segReg instructions load a segment
selector from a general-purpose register or memory (for MOV
segReg) or from the top of the stack (for POP segReg) to a
segment register. These instructions not only load the softwarevisible segment selector into the segment register but also
cause the hardware to load the associated segment-descriptor
information into the software-invisible (hidden) portion of that
segment register.
In 64-bit mode, the POP DS, POP ES, and POP SS instructions
are invalid.
3.3.5 Load Effective
Address
LEA—Load Effective Address
The LEA instruction calculates and loads the effective address
(offset within a given segment) of a source operand and places
it in a general-purpose register.
LEA is related to MOV, which copies data from a memory
location to a register, but LEA takes the address of the source
operand, whereas MOV takes the contents of the memory
location specified by the source operand. In the simplest cases,
LEA can be replaced with MOV. For example:
lea eax, [ebx]
Chapter 3: General-Purpose Programming59
Chapter 3: General-Purpose Programming59
Page 94
AMD64 Technology24592—Rev. 3.10—March 2005
has the same effect as:
mov eax, ebx
However, LEA allows software to use any valid addressing mode
for the source operand. For example:
lea eax, [ebx+edi]
loads the sum of EBX and EDI registers into the EAX register.
This could not be accomplished by a single MOV instruction.
LEA has a limited capability to perform multiplication of
operands in general-purpose registers using scaled-index
addressing. For example:
lea eax, [ebx+ebx*8]
loads the value of the EBX register, multiplied by 9, into the
EAX register.
operations, such as addition, subtraction, multiplication, and
division on integer operands.
Add and Subtract.
ADC—Add with Carry
ADD—Signed or Unsigned Add
SBB—Subtract with Borrow
SUB—Subtract
NEG—Two’s Complement Negation
The ADD instruction performs addition of two integer
operands. There are opcodes that add an immediate value to a
byte, word, doubleword, or quadword register or a memory
location. In these opcodes, if the size of the immediate is
smaller than that of the destination, the immediate is first signextended to the size of the destination operand. The arithmetic
flags (OF, SF, ZF, AF, CF, PF) are set according to the resulting
value of the destination operand.
The ADC instruction performs addition of two integer
operands, plus 1 if the carry flag (CF) is set.
The SUB instruction performs subtraction of two integer
operands.
60Chapter 3: General-Purpose Programming
Page 95
24592—Rev. 3.10—March 2005AMD64 Technology
The SBB instruction performs subtraction of two integer
operands, and it also subtracts an additional 1 if the carry flag is
set.
The ADC and SBB instructions simplify addition and
subtraction of multiple-precision integer operands, because
they correctly handle carries (and borrows) between parts of a
multiple-precision operand.
The NEG instruction performs negation of an integer operand.
The value of the operand is replaced with the result of
subtracting the operand from zero.
Multiply and Divide.
MUL—Multiply Unsigned
IMUL—Signed Multiply
DIV—Unsigned Divide
IDIV—Signed Divide
The MUL instruction performs multiplication of unsigned
integer operands. The size of operands can be byte, word,
doubleword, or quadword. The product is stored in a destination
which is double the size of the source operands (multiplicand
and factor).
The MUL instruction's mnemonic has only one operand, which
is a factor. The multiplicand operand is always assumed to be an
accumulator register. For byte-sized multiplies, AL contains the
multiplicand, and the result is stored in AX. For word-sized,
doubleword-sized, and quadword-sized multiplies, rAX contains
the multiplicand, and the result is stored in rDX and rAX.
The IMUL instruction performs multiplication of signed integer
operands. There are forms of the IMUL instruction with one,
two, and three operands, and it is thus more powerful than the
MUL instruction. The one-operand form of the IMUL
instruction behaves similarly to the MUL instruction, except
that the operands and product are signed integer values. In the
two-operand form of IMUL, the multiplicand and product use
the same register (the first operand), and the factor is specified
in the second operand. In the three-operand form of IMUL, the
product is stored in the first operand, the multiplicand is
specified in the second operand, and the factor is specified in
the third operand.
Chapter 3: General-Purpose Programming61
Chapter 3: General-Purpose Programming61
Page 96
AMD64 Technology24592—Rev. 3.10—March 2005
The DIV instruction performs division of unsigned integers.
The instruction divides a double-sized dividend in AH:AL or
rDX:rAX by the divisor specified in the operand of the
instruction. It stores the quotient in AL or rAX and the
remainder in AH or rDX.
The IDIV instruction performs division of signed integers. It
behaves similarly to DIV, with the exception that the operands
are treated as signed integer values.
Division is the slowest of all integer arithmetic operations and
should be avoided wherever possible. One possibility for
improving performance is to replace division with
multiplication, such as by replacing i/j/k with i/(j*k). This
replacement is possible if no overflow occurs during the
computation of the product. This can be determined by
considering the possible ranges of the divisors.
Increment and Decrement.
DEC—Decrement by 1
INC—Increment by 1
The INC and DEC instructions are used to increment and
decrement, respectively, an integer operand by one. For both
instructions, an operand can be a byte, word, doubleword, or
quadword register or memory location.
These instructions behave in all respects like the corresponding
ADD and SUB instructions, with the second operand as an
immediate value equal to 1. The only exception is that the carry
flag (CF) is not affected by the INC and DEC instructions.
Apart from their obvious arithmetic uses, the INC and DEC
instructions are often used to modify addresses of operands. In
this case it can be desirable to preserve the value of the carry
flag (to use it later), so these instructions do not modify the
carry flag.
3.3.7 Rotate and ShiftThe rotate and shift instructions perform cyclic rotation or non-
cyclic shift, by a given number of bits (called the count), in a
given byte-sized, word-sized, doubleword-sized or quadwordsized operand.
When the count is greater than 1, the result of the rotate and
shift instructions can be considered as an iteration of the same
62Chapter 3: General-Purpose Programming
Page 97
24592—Rev. 3.10—March 2005AMD64 Technology
1-bit operation by count number of times. Because of this, the
descriptions below describe the result of 1-bit operations.
The count can be 1, the value of the CL register, or an
immediate 8-bit value. To avoid redundancy and make rotation
and shifting quicker, the count is masked to the 5 or 6 leastsignificant bits, depending on the effective operand size, so that
its value does not exceed 31 or 63 before the rotation or shift
takes place.
Rotate.
RCL—Rotate Through Carry Left
RCR—Rotate Through Carry Right
ROL—Rotate Left
ROR—Rotate Right
The RCx instructions rotate the bits of the first operand to the
left or right by the number of bits specified by the source
(count) operand. The bits rotated out of the destination
operand are rotated into the carry flag (CF) and the carry flag is
rotated into the opposite end of the first operand.
The ROx instructions rotate the bits of the first operand to the
left or right by the number of bits specified by the source
operand. Bits rotated out are rotated back in at the opposite
end. The value of the CF flag is determined by the value of the
last bit rotated out. In single-bit left-rotates, the overflow flag
(OF) is set to the XOR of the CF flag after rotation and the
most-significant bit of the result. In single-bit right-rotates, the
OF flag is set to the XOR of the two most-significant bits. Thus,
in both cases, the OF flag is set to 1 if the single-bit rotation
changed the value of the most-significant bit (sign bit) of the
operand. The value of the OF flag is undefined for multi-bit
rotates.
Bit-rotation instructions provide many ways to reorder bits in an
operand. This can be useful, for example, in character
conversion, including cryptography techniques.
Shift.
SAL—Shift Arithmetic Left
SAR—Shift Arithmetic Right
SHL—Shift Left
Chapter 3: General-Purpose Programming63
Chapter 3: General-Purpose Programming63
Page 98
AMD64 Technology24592—Rev. 3.10—March 2005
SHR—Shift Right
SHLD—Shift Left Double
SHRD—Shift Right Double
The SHx instructions (including SHxD) perform shift
operations on unsigned operands. The SAx instructions operate
with signed operands.
SHL and SAL instructions effectively perform multiplication of
an operand by a power of 2, in which case they work as moreefficient alternatives to the MUL instruction. Similarly, SHR
and SAR instructions can be used to divide an operand (signed
or unsigned, depending on the instruction used) by a power of
2.
Although the SAR instruction divides the operand by a power
of 2, the behavior is different from the IDIV instruction. For
example, shifting –11 (FFFFFFF5h) by two bits to the right (i.e.
divide –11 by 4), gives a result of FFFFFFFDh, or –3, whereas
the IDIV instruction for dividing –11 by 4 gives a result of –2.
This is because the IDIV instruction rounds off the quotient to
zero, whereas the SAR instruction rounds off the remainder to
zero for positive dividends, and to negative infinity for negative
dividends. This means that, for positive operands, SAR behaves
like the corresponding IDIV instruction, and for negative
operands, it gives the same result if and only if all the shiftedout bits are zeroes, and otherwise the result is smaller by 1.
The SAR instruction treats the most-significant bit (msb) of an
operand in a special way: the msb (the sign bit) is not changed,
but is copied to the next bit, preserving the sign of the result.
The least-significant bit (lsb) is shifted out to the CF flag. In the
SAL instruction, the msb is shifted out to CF flag, and the lsb is
cleared to 0.
The SHx instructions perform logical shift, i.e. without special
treatment of the sign bit. SHL is the same as SAL (in fact, their
opcodes are the same). SHR copies 0 into the most-significant
bit, and shifts the least-significant bit to the CF flag.
The SHxD instructions perform a double shift. These
instructions perform left and right shift of the destination
operand, taking the bits to copy into the most-significant bit
(for the SHRD instruction) or into the least-significant bit (for
the SHLD instruction) from the source operand. These
instructions behave like SHx, but use bits from the source
64Chapter 3: General-Purpose Programming
Page 99
24592—Rev. 3.10—March 2005AMD64 Technology
operand instead of zero bits to shift into the destination
operand. The source operand is not changed.
3.3.8 Compare and
Test
The compare and test instructions perform arithmetic and
logical comparison of operands and set corresponding flags,
depending on the result of comparison. These instruction are
used in conjunction with conditional instructions such as Jcc or
SETcc to organize branching and conditionally executing blocks
in programs. Assembler equivalents of conditional operators in
high-level languages (do…while, if…then…else, and similar)
also include compare and test instructions.
Compare.
CMP—Compare
The CMP instruction performs subtraction of the second
operand (source) from the first operand (destination), like the
SUB instruction, but it does not store the resulting value in the
destination operand. It leaves both operands intact. The only
effect of the CMP instruction is to set or clear the arithmetic
flags (OF, SF, ZF, AF, CF, PF) according to the result of
subtraction.
The CMP instruction is often used together with the conditional
jump instructions (Jcc), conditional SET instructions (SETcc)
and other instructions such as conditional loops (LOOPcc)
whose behavior depends on flag state.
Test.
TEST—Test Bits
The TEST instruction is in many ways similar to the AND
instruction: it performs logical conjunction of the
corresponding bits of both operands, but unlike the AND
instruction it leaves the operands unchanged. The purpose of
this instruction is to update flags for further testing.
The TEST instruction is often used to test whether one or more
bits in an operand are zero. In this case, one of the instruction
operands would contain a mask in which all bits are cleared to
zero except the bits being tested. For more advanced bit testing
and bit modification, use the BTx instructions.
Chapter 3: General-Purpose Programming65
Chapter 3: General-Purpose Programming65
Page 100
AMD64 Technology24592—Rev. 3.10—March 2005
Bit Scan.
BSF—Bit Scan Forward
BSR—Bit Scan Reverse
The BSF and BSR instructions search a source operand for the
least-significant (BSF) or most-significant (BSR) bit that is set
to 1. If a set bit is found, its bit index is loaded into the
destination operand, and the zero flag (ZF) is set. If no set bit is
found, the zero flag is cleared and the contents of the
destination are undefined.
Bit Test.
BT—Bit Test
BTC—Bit Test and Complement
BTR—Bit Test and Reset
BTS—Bit Test and Set
The BTx instructions copy a specified bit in the first operand to
the carry flag (CF) and leave the source bit unchanged (BT), or
complement the source bit (BTC), or clear the source bit to 0
(BTR), or set the source bit to 1 (BTS).
These instructions are useful for implementing semaphore
arrays. Unlike the XCHG instruction, the BTx instructions set
the carry flag, so no additional test or compare instruction is
needed. Also, because these instructions operate directly on
bits rather than larger data types, the semaphore arrays can be
smaller than is possible when using XCHG. In such semaphore
applications, bit-test instructions should be preceded by the
LOCK prefix.
Set Byte on Condition.
SETcc—Set Byte if condition
The SETcc instructions store a 1 or 0 value to their byte operand
depending on whether their condition (represented by certain
rFLAGS bits) is true or false, respectively. Table 3-5 on page 67
shows the rFLAGS values required for each SETcc instruction.
66Chapter 3: General-Purpose Programming
Loading...
+ hidden pages
You need points to download manuals.
1 point = 1 manual.
You can buy points or you can get point for every manual you upload.