Download as pdf or txt
Download as pdf or txt
You are on page 1of 54

+

William Stallings
Computer Organization
and Architecture
9th Edition
Chapter 14
Processor Structure and Function
LEARNING OBJECTIVES
After studying this chapter, you should be
able to:

 Distinguish between user-visible and


control/status registers, and discuss the
purposes of registers in each category.
 Summarize the instruction cycle.
 Discuss the principle behind instruction
pipelining and how it works in practice.
+ Compare and contrast the various forms of

pipeline hazards.
 Present an overview of the x86 processor
structure.
 Present an overview of the ARM processor
structure.
+
Processor Organization
Processor Requirements:
 Fetch instruction
 The processor reads an instruction from memory (register, cache, main memory)

 Interpret instruction
 The instruction is decoded to determine what action is required

 Fetch data
 The execution of an instruction may require reading data from memory or an I/O
module

 Process data
 The execution of an instruction may require performing some arithmetic or logical
operation on data

 Write data
 The results of an execution may require writing data to memory or an I/O module

 In order to do these things the processor needs to store some data


temporarily and therefore needs a small internal memory
CPU With the System Bus

Registers

ALU

Control
Unit

Control Data Address


Bus Bus Bus

System
Bus

Figure 14.1 The CPU with the System Bus


CPU Internal Structure
Arithmetic and Logic Unit

Status Flags
Registers

Shifter

I nternal CPU Bus


Complementer

Arithmetic
and
Boolean
Logic

Control
Unit

Control
Paths

Figure 14.2 I nternal Structur e of the CPU


+
Register Organization
 Within the processor there is a set of registers that function as a
level of memory above main memory and cache in the
hierarchy

 The registers in the processor perform two roles:

User-Visible Registers Control and Status Registers

 Enable the machine or  Used by the control unit to


assembly language control the operation of the
programmer to minimize main processor and by privileged
memory references by operating system programs to
optimizing use of registers control the execution of
programs
User-Visible Registers

Categories:
Referenced by means of • General purpose
the machine language • Can be assigned to a variety of functions by
the programmer
that the processor • Data
executes • May be used only to hold data and cannot
be employed in the calculation of an
operand address
• Address
• May be somewhat general purpose or may
be devoted to a particular addressing mode
• Examples: segment pointers, index
registers, stack pointer
• Condition codes
• Also referred to as flags
• Bits set by the processor hardware as the
result of operations
Table 14.1
Condition Codes
Advantages Disadvantages
1. Because condition codes are set by normal 1. Condition codes add complexity, both to
arithmetic and data movement instructions, the hardware and software. Condition code
they should reduce the number of bits are often modified in different ways
COMPARE and TEST instructions needed. by different instructions, making life more
2. Conditional instructions, such as BRANCH difficult for both the microprogrammer
are simplified relative to composite and compiler writer.
instructions, such as TEST AND 2. Condition codes are irregular; they are
BRANCH. typically not part of the main data path, so
3. Condition codes facilitate multiway they require extra hardware connections.
branches. For example, a TEST instruction 3. Often condition code machines must add
can be followed by two branches, one on special non-condition-code instructions for
less than or equal to zero and one on special situations anyway, such as bit
greater than zero. checking, loop control, and atomic
semaphore operations.
4. Condition codes can be saved on the stack 4. In a pipelined implementation, condition
during subroutine calls along with other codes require special synchronization to
register information. avoid conflicts.
+
Control and Status Registers
Four registers are essential to instruction execution:

 Program counter (PC)


 Contains the address of an instruction to be fetched

 Instruction register (IR)


 Contains the instruction most recently fetched

 Memory address register (MAR)


 Contains the address of a location in memory

 Memory buffer register (MBR)


 Contains a word of data to be written to memory or the word most
recently read
+
Program Status Word (PSW)

Register or set of registers that


contain status information

Common fields or flags include:


• Sign
• Zero
• Carry
• Equal
• Overflow
• Interrupt Enable/Disable
• Supervisor
Data registers General registers General Registers
D0 AX Accumulator EAX AX
D1 BX Base EBX BX
D2 CX Count ECX CX
D3 DX Data EDX DX
D4
D5 Pointers & index ESP SP
D6 SP Stack ptr EBP BP
D7 BP Base ptr ESI SI
SI Source index EDI DI
Address registers DI Dest index
A0 Program Status
A1 Segment FLAGS Register
A2 CS Code I nstruction Pointer
A3 DS Data
A4 SS Stack (c) 80386 - Pentium 4
A5 ES Extrat
A6
A7´ Program status
Flags Example
Microprocessor
I nstr ptr
Program status
Program counter (b) 8086
Status register
Register
(a) M C68000
Organizations
Figure 14.3 Example M icroprocessor Register Organizations
Includes the following Instruction
stages:
Cycle

Fetch Execute Interrupt

If interrupts are
Read the next enabled and an
Interpret the opcode
instruction from interrupt has occurred,
and perform the
memory into the save the current
indicated operation
processor process state and
service the interrupt
Instruction Cycle

Fetch

I nterrupt I ndirect

Execute

Figure 14.4 The I nstruction Cycle


Instruction Cycle State Diagram
I ndirection I ndirection

I nstruction Operand Operand


fetch fetch store

Multiple Multiple
operands results

I nstruction I nstruction Operand Operand


Data I nterrupt
address operation address address I nterrupt
Operation check
calculation decoding calculation calculation

No
Instruction complete, Return for string interrupt
fetcth next instruction or vector data

Figure 14.5 I nstruction Cycle State Diagram


Data Flow, Fetch Cycle
CPU

PC M AR
M emory

Control
Unit

IR M BR

Address Data Control


Bus Bus Bus
MBR = Memory buffer register
MAR = Memory address register
IR = Instruction register
PC = Program counter

Figure 14.6 Data Flow, Fetch Cycle


Data Flow, Indirect Cycle
CPU

M AR
M emory

Control
Unit

M BR

Address Data Control


Bus Bus Bus

Figure 14.7 Data Flow, I ndirect Cycle


Data Flow, Interrupt Cycle
CPU

PC M AR
M emory

Control
Unit

M BR

Address Data Control


Bus Bus Bus

Figure 14.8 Data Flow, I nterrupt Cycle


Pipelining: Its Natural!
Example (Hennessy & Patterson, Computer
Architecture. Third Edition)

Laundry Example
 Ann, Brian, Cathy, Dave each have one
A B C D
load of clothes to wash, dry, and fold
 Washer takes 30 minutes

 Dryer takes 40 minutes

 “Folder” takes 20 minutes


Sequential Laundry
6 PM 7 8 9 10 11 Midnight
Time
30 40 20 30 40 20 30 40 20 30 40 20
T
a A
s
k
B
O
r C
d
e D
r
• Sequential laundry takes 6 hours for 4 loads

• If they learned pipelining, how long would laundry take?


Pipelined Laundry

6 PM 7 8 9 10 11 Midnight
Time
30 40 40 40 40 20
T
A
a
s
k B

O C
r
d D
e
r • Pipelined laundry takes 3.5 hours for 4 loads
Pipelining Lessons

 Pipelining doesn’t help


6 PM 7 8 9 latency of single task, it helps
throughput of entire
workload
Time
T  Pipeline rate limited by
30 40 40 40 40 20 slowest pipeline stage
a
 Multiple tasks operating
s A simultaneously
k  Potential speedup = Number
pipe stages
B
O  Unbalanced lengths of pipe
r stages reduces speedup
C
d  Time to “fill” pipeline and
time to “drain” it reduces
e D speedup
r
Pipelining Strategy

To apply this concept


to instruction
execution we must
Similar to the use of recognize that an
an assembly line in a instruction has a
manufacturing plant number of stages

New inputs are


accepted at one end
before previously
accepted inputs
appear as outputs at
the other end
Two-Stage Instruction Pipeline

I nstruction I nstruction Result


Fetch Execute

(a) Simplified view

Wait New address Wait

I nstruction I nstruction Result


Fetch Execute

Discard
(b) Expanded view

Figure 14.9 Two-Stage I nstruction Pipeline


+
Additional Stages
 Fetch instruction (FI)
 Fetch operands (FO)
 Read the next expected
instruction into a buffer  Fetch each operand from
memory
 Decode instruction (DI)  Operands in registers need
 Determine the opcode and not be fetched
the operand specifiers
 Execute instruction (EI)
 Calculate operands (CO)  Perform the indicated
 Calculate the effective operation and store the
address of each source result, if any, in the specified
operand destination operand location
 This may involve
 Write operand (WO)
displacement, register
indirect, indirect, or other  Store the result in memory
forms of address calculation
Timing Diagram for Instruction
Pipeline Operation
Time

1 2 3 4 5 6 7 8 9 10 11 12 13 14
I nstruction 1 FI DI CO FO EI WO

I nstruction 2 FI DI CO FO EI WO

I nstruction 3 FI DI CO FO EI WO

I nstruction 4 FI DI CO FO EI WO

I nstruction 5 FI DI CO FO EI WO

I nstruction 6 FI DI CO FO EI WO

I nstruction 7 FI DI CO FO EI WO

I nstruction 8 FI DI CO FO EI WO

I nstruction 9 FI DI CO FO EI WO

Figure 14.10 Timing Diagram for I nstruction Pipeline Operation


The Effect of a Conditional Branch
on Instruction Pipeline Operation
Time Branch Penalty

1 2 3 4 5 6 7 8 9 10 11 12 13 14
I nstruction 1 FI DI CO FO EI WO

I nstruction 2 FI DI CO FO EI WO

I nstruction 3 FI DI CO FO EI WO

I nstruction 4 FI DI CO FO

I nstruction 5 FI DI CO

I nstruction 6 FI DI

I nstruction 7 FI

I nstruction 15 FI DI CO FO EI WO

I nstruction 16 FI DI CO FO EI WO

Figure 14.11 The Effect of a Conditional Branch on I nstruction Pipeline Operation


+
Fetch
I nstruction
FI

Decode
DI I nstruction

Calculate
CO Operands

Yes Uncon-
ditional
Branch?

Six Stage No

Instruction Pipeline FO
Fetch
Operands

Execute
EI I nstruction

Update Write
PC
WO Operands

Empty
Pipe Yes Branch No
or
I nter
-rupt?

Figure 14.12 Six-Stage I nstruction Pipeline


+ FI DI CO FO EI WO FI DI CO FO EI WO

1 I1 1 I1

2 I2 I1 2 I2 I1

3 I3 I2 I1 3 I3 I2 I1

4 I4 I3 I2 I1 4 I4 I3 I2 I1

5 I5 I4 I3 I2 I1 5 I5 I4 I3 I2 I1

6 I6 I5 I4 I3 I2 I1 6 I6 I5 I4 I3 I2 I1

Time
7 I7 I6 I5 I4 I3 I2 7 I7 I6 I5 I4 I3 I2
Alternative Pipeline 8 I8 I7 I6 I5 I4 I3 8 I 15 I3

Depiction 9 I9 I8 I7 I6 I5 I4 9 I 16 I 15

10 I9 I8 I7 I6 I5 10 I 16 I 15

11 I9 I8 I7 I6 11 I 16 I 15

12 I9 I8 I7 12 I 16 I 15

13 I9 I8 13 I 16 I 15

14 I9 14 I 16

(a) No branches (b) With conditional branch

Figure 14.13 An Alternative Pipeline Depiction


Pipeline Performance
Retardo Retardo Retardo
Retardo Retardo Retardo Retardo Retardo
Retardo
Retardo Retardo Retardo Retardo
tmax d ti d ti d tmax d
d ti d ti d

D D D EI D WO D
Instruction i D DI D CO FO (Write
FI (Calculate (Fetch
(Execute

(Fetch (Decode Instruction ) Operand )


Operand Out
Operand)
Instruction) CK
Instruction)
CK Direction) CK CK CK CK Instructions
Clock CK


El Tiempo de Ciclo es el tiempo necesario para que un conjunto de
instrucciones avance una etapa a través del cauce

  max[i]+ d  m + d 1≤ i ≤ k
Donde:
i : retardo de tiempo de circuitería de la i-ésima etapa del cauce
m : máximo retardo de etapa
k: número de etapas del cauce de instrucciones
d: retardo de tiempo de un registro (equivalente a un pulso de reloj ), m  d
Pipeline Performance
Para ejecutar n instrucciones sin saltos en un cauce de k etapas
se requiere un tiempo total Tkn

Tkn  [k+(n-1)]

Sk: Factor de aceleración del cauce


Sk  T1,n / Tk,n de instrucciones segmentado
comparado con la ejecución sin
segmentación

Sk  nk / [k+(n-1)]  nk / [k+(n-1)]


+

Speedup Factors
with Instruction Aquí, en el limite (n ) la velocidad se
Pipelining incrementa en un factor k.

Aquí, el factor de aceleración se aproxima


al número de instrucciones que se pueden
introducir en el cauce sin saltos.
Pipeline Hazards
Occur when the
pipeline, or some
portion of the There are three
pipeline, must stall types of hazards:
because conditions • Resource
do not permit • Data
continued execution • Control

Also referred to as a
pipeline bubble
+
Clock cycle
1 2 3 4 5 6 7 8 9
I1 FI DI FO EI WO

I nstrutcion
I2 FI DI FO EI WO
Resource I3 FI DI FO EI WO
Hazards I4 FI DI FO EI WO

(a) Five-stage pipeline, ideal case

A resource hazard occurs when two Clock cycle


or more instructions that are already 1 2 3 4 5 6 7 8 9
in the pipeline need the same
resource I1 FI DI FO EI WO
I nstrutcion
I2 FI DI FO EI WO
The result is that the instructions must
be executed in serial rather than I3 I dle FI DI FO EI WO
parallel for a portion of the pipeline
I4 FI DI FO EI WO
A resource hazard is sometimes
referred to as a structural hazard (b) I 1 source operand in memory

Figure 14.15 Example of Resource Hazard


Clock cycle
1 2 3 4 5 6 7 8 9 10
ADD EAX, EBX FI DI FO EI WO RAW
SUB ECX, EAX FI DI I dle FO EI WO

I3 FI DI FO EI WO

I4 FI DI FO EI WO

Hazard
Figure 14.16 Example of Data Hazard

+
Data Hazards
A data hazard occurs when there is a conflict in the
access of an operand location
+
Types of Data Hazard
 Read after write (RAW), or true dependency
 An instruction modifies a register or memory location
 Succeeding instruction reads data in memory or register location
 Hazard occurs if the read takes place before write operation is
complete

 Write after read (WAR), or antidependency


 An instruction reads a register or memory location
 Succeeding instruction writes to the location
 Hazard occurs if the write operation completes before the read
operation takes place

 Write after write (WAW), or output dependency


 Two instructions both write to the same location
 Hazard occurs if the write operations take place in the reverse order
of the intended sequence
+
Control Hazard

 Also known as a branch hazard

 Occurs when the pipeline makes the wrong decision on a


branch prediction

 Brings instructions into the pipeline that must subsequently


be discarded

 Dealing with Branches:


 Multiple streams
 Prefetch branch target
 Loop buffer
 Branch prediction
 Delayed branch
Multiple Streams

A simple pipeline suffers a penalty for a


branch instruction because it must choose
one of two instructions to fetch next and may
make the wrong choice

A brute-force approach is to replicate the


initial portions of the pipeline and allow the
pipeline to fetch both instructions, making
use of two streams

Drawbacks:
• With multiple pipelines there are contention delays
for access to the registers and to memory
• Additional branch instructions may enter the pipeline
before the original branch decision is resolved
Prefetch Branch Target

 When a conditional branch is recognized, the


target of the branch is prefetched, in addition
to the instruction following the branch

 Target is then saved until the branch


instruction is executed

 If the branch is taken, the target has already


been prefetched

 IBM 360/91 uses this approach


+
+
Loop Buffer
 Small, very-high speed memory maintained by the instruction fetch
stage of the pipeline and containing the n most recently fetched
instructions, in sequence

 Benefits:
 Instructions fetched in sequence will be available without the usual
memory access time
 If a branch occurs to a target just a few locations ahead of the address of
the branch instruction, the target will already be in the buffer
 This strategy is particularly well suited to dealing with loops
 Very good for small loops or jumps (secuencias IF-THEN e IF-THEN-ELSE)
 Used by CRAY-1, Motorola 68010

 Similar in principle to a cache dedicated to instructions


 Differences:
 The loop buffer only retains instructions in sequence
 Is much smaller in size and hence lower in cost
Loop Buffer

Branch address • Contiene, secuencialmente,


las n instrucciones captadas
recientemente
I nstruction to be
8
Loop Buffer decoded in case of hit • Si se va a dar un salto, el
(256 bytes) hardware comprueba primero
si el destino de salto está en el
buffer
M ost significant address bits
compared to determine a hit
• Si está allí, la siguiente
instrucción se capta del buffer

Figure 14.17 Loop Buffer


+
Branch Prediction

 Various techniques can be used to predict whether a branch


will be taken:

1. Predict never taken  These approaches are static


2. Predict always taken  They do not depend on the
execution history up to the time of
3. Predict by opcode the conditional branch instruction

1. Taken/not taken switch


 These approaches are dynamic
2. Branch history table
 They depend on the execution history
+
Read next Read next
conditional conditional
branch instr branch instr

Predict taken Predict not taken

Branch Prediction Yes Branch


taken?
No Branch
taken?
Flow Chart No Yes

Read next Read next


conditional conditional
branch instr branch instr

Predict taken Predict not taken

Yes Branch No No Branch


taken? taken?

Yes

Figure 14.18 Branch Prediction Flow Chart


Branch Prediction State Diagram
Not Taken
Taken Predict Predict
Taken Taken
Taken

Not Taken
Taken

Not Taken
Predict Predict Not Taken
Not Taken Not Taken
Taken

Figure 14.19 Branch Prediction State Diagram


Next sequential

+
address

Select
M emory

E Branch M iss
Handling

(a) Predict never taken strategy

Next sequential
address
I PFAR Branch

Dealing With instruction


address
Target
address State

Select
Lookup

Branches
M emory

Add new I PFAR = instruction


entry prefix address register

Update
state

Branch M iss Redirect


Handling
E

(b) Branch history table strategy

Figure 14.20 Dealing with Branches


+
Intel 80486 Pipelining
 Fetch
 Objective is to fill the prefetch buffers with new data as soon as the old data
have been consumed by the instruction decoder
 Operates independently of the other stages to keep the prefetch buffers full

 Decode stage 1
 All opcode and addressing-mode information is decoded in the D1 stage
 3 bytes of instruction are passed to the D1 stage from the prefetch buffers
 D1 decoder can then direct the D2 stage to capture the rest of the instruction

 Decode stage 2
 Expands each opcode into control signals for the ALU
 Also controls the computation of the more complex addressing modes

 Execute
 Stage includes ALU operations, cache access, and register update

 Write back
 Updates registers and status flags modified during the preceding execute
stage
+ Fetch D1
Fetch
D2
D1
Fetch
EX
D2
D1
WB
EX
D2
WB
EX WB
M OV Reg1, M em1
M OV Reg1, Reg2
M OV M em2, Reg1

(a) No Data Load Delay in the Pipeline

80486
Instruction Fetch D1 D2 EX WB M OV Reg1, M em1

Pipeline
Fetch D1 D2 EX M OV Reg2, (Reg1)

Examples (b) Pointer Load Delay

Fetch D1 D2 EX WB CM P Reg1, I mm
Fetch D1 D2 EX Jcc Target
Fetch D1 D2 EX Target

(c) Branch I nstruction Timing

Figure 14.21 80486 I nstruction Pipeline Examples


Table 14.2
x86 Processor Registers
Type Number Length (bits) Purpose
General 8 32 General-purpose user registers
Segment 6 16 Contain segment selectors
EFLAGS 1 32 Status and control bits
Instruction Pointer 1 32 Instruction pointer

(a) I nteger Unit in 32-bit Mode

Type Number Length (bits) Purpose


General 16 32 General-purpose user registers
Segment 6 16 Contain segment selectors
RFLAGS 1 64 Status and control bits
Instruction Pointer 1 64 Instruction pointer

(b) Integer Unit in 64-bit Mode


Table 14.2
x86 Processor Registers

Type Number Length (bits) Purpose


Numeric 8 80 Hold floating-point numbers
Control 1 16 Control bits
Status 1 16 Status bits
Tag Word 1 16 Specifies contents of numeric
registers
Instruction Pointer 1 48 Points to instruction interrupted
by exception
Data Pointer 1 48 Points to operand interrupted by
exception

(c) Floating-Point Unit


x86 EFLAGS Register

31 21 16 15 0
I V V A V R N IO O D I T S Z A P C
I I
D P F C M F T PL F F F F F F F F F

ID = Identification flag DF = Direction flag


VIP = Virtual interrupt pending IF = Interrupt enable flag
VIF = Virtual interrupt flag TF = Trap flag
AC = Alignment check SF = Sign flag
VM = Virtual 8086 mode ZF = Zero flag
RF = Resume flag AF = Auxiliary carry flag
NT = Nested task flag PF = Parity flag
IOPL = I/O privilege level CF = Carry flag
OF = Overflow flag

Figure 14.22 x86 EFLAGS Register


Control
Registers
Mapping of MMX Registers to
Floating-Point Registers
Floating-Point
Tag Floating-Point Registers
79 63 0
00
00
00
00
63 0
00 MM7
00 MM6
00 MM5
00 MM4
MM3
MM2
MM1
MM0

M M X Registers

Figure 14.24 M apping of M M X Registers to Floating-Point Registers


+
Interrupt Processing
Interrupts and Exceptions
 Interrupts
 Generated by a signal from hardware and it may occur at random
times during the execution of a program
 Maskable
 Nonmaskable

 Exceptions
 Generated from software and is provoked by the execution of an
instruction
 Processor detected
 Programmed

 Interrupt vector table


 Every type of interrupt is assigned a number
 Number is used to index into the interrupt vector table
Table 14.3
x86 Exception and Interrupt Vector Table
Vector Number Description
0 Divide error; division overflow or division by zero
1 Debug exception; includes various faults and traps related to debugging
2 NMI pin interrupt; signal on NMI pin
3 Breakpoint; caused by INT 3 instruction, which is a 1-byte instruction useful for debugging
4 INTO-detected overflow; occurs when the processor executes INTO with the OF flag set
5 BOUND range exceeded; the BOUND instruction compares a register with boundaries stored in
memory and generates an interrupt if the contents of the register is out of bounds.
6 Undefined opcode
7 Device not available; attempt to use ESC or WAIT instruction fails due to lack of external device
8 Double fault; two interrupts occur during the same instruction and cannot be handled serially
9 Reserved
10 Invalid task state segment; segment describing a requested task is not initialized or not valid
11 Segment not present; required segment not present
12 Stack fault; limit of stack segment exceeded or stack segment not present
13 General protection; protection violation that does not cause another exception (e.g., writing to a
read-only segment)
14 Page fault
15 Reserved
16 Floating-point error; generated by a floating-point arithmetic instruction
17 Alignment check; access to a word stored at an odd byte address or a doubleword stored at an
address not a multiple of 4
18 Machine check; model specific
19-31 Reserved
32-255 User interrupt vectors; provided when INTR signal is activated
Unshaded: exceptions Shaded: interrupts
+ Summary Processor Structure
and Function
Chapter 14
 Instruction pipelining
 Processor organization
 Pipelining strategy
 Register organization  Pipeline performance
 User-visible registers  Pipeline hazards
 Control and status registers  Dealing with branches
 Intel 80486 pipelining
 Instruction cycle
 The indirect cycle
 Data flow

 The x86 processor family


 Register organization
 Interrupt processing

You might also like