Xilinx
Block Preview

Introduction

The block computes the reciprocal of IEEE-754 floating-point inputs. On every rising edge of CLK, if CE = 1, the reciprocal unit computes

$$ \mathrm{F}(n) = \frac{1}{\mathrm{A}(n)}, $$

where input A and output F follow IEEE-754 single or double precision format.

The reciprocal is implemented with the Xilinx floating_point IP core configured for Reciprocal operation using Newton-Raphson iteration with blocking flow control, fixed 30-cycle latency, and configurable DSP primitive usage.

Pin Description

A Input Variable bit BIT VECTOR
Floating-point input value (IEEE-754). Width: 32 bits (Single) or 64 bits (Double). Must not be zero unless division-by-zero behavior is acceptable. Accepted when CE = 1 and READY_OUT = 1.
Default: Must be connected
CE Input 1 bit BIT
Clock Enable (tvalid), active high. When CE = 1, input A is accepted and reciprocal computation begins. Can be tied to ‘1’ for continuous operation.
Default: 1
READY_IN Input 1 bit BIT
Downstream ready signal (tready input), active high. Indicates if downstream logic can accept new data. Can be tied to ‘1’ if backpressure is not needed.
CLK Input 1 bit BIT
Global clock. Every rising edge triggers pipeline advancement. Connected to system acquisition clock.
Default: Default Board Clock
F Output 32 bit BIT VECTOR
Floating-point reciprocal output (IEEE-754). Width: 32 bits (Single) or 64 bits (Double). Valid when DV = 1. Result = 1/A.
DV Output 1 bit BIT
Data Valid output (tvalid), active high. Indicates when output F contains a valid reciprocal. Asserts 30 clock cycles after corresponding CE = 1.
READY_OUT Output 1 bit BIT
Upstream ready signal (tready output), active high. Indicates this block can accept new input data. Used for flow control in streaming pipelines.

Properties

Property window

Float Format FloatFormat

Select between single precision 32 bit and double precision 64 bit

Floating-point precision for both input and output:

  • Single → 32-bit (8-bit exponent, 24-bit mantissa including implicit bit)
  • Double → 64-bit (11-bit exponent, 53-bit mantissa including implicit bit)

The operation preserves the precision format end-to-end.

Default: Single

Options: Single Double

DSP Usage DSPUsage

DSP Usage. Single precision: No [0], Full[8]. Double precision: No[0], Full[14]

DSP primitive allocation for Newton-Raphson iterations:

  • No_Usage → LUT-only implementation (0 DSPs, lower speed)
  • Full_Usage → DSP-optimized (8/14 DSPs for Single/Double, recommended, higher speed)

Full usage provides significantly better timing performance.

Default: Full_Usage

Options: No_Usage Full_Usage

Functional description

The component computes the multiplicative inverse:

$$ F = \frac{1}{A} $$

The implementation uses Newton-Raphson iteration to refine an initial approximation:

$$ x_{n+1} = x_n (2 - A \cdot x_n) $$

converging to $1/A$. This is often faster than full division when the numerator is constant (1).

Special cases

IEEE-754 special value handling:

  • 1/1 = 1
  • 1/0 = ±Inf (division by zero, sign preserved)
  • 1/(±Inf) = ±0
  • 1/NaN = NaN (NaN propagation)

DSP Usage

The DSP Usage property controls Newton-Raphson iteration resources:

Single precision:

  • No_Usage → Pure LUT implementation (0 DSPs, lower speed)
  • Full_Usage → DSP-optimized (8 DSPs, higher speed)

Double precision:

  • No_Usage → Pure LUT implementation (0 DSPs, lower speed)
  • Full_Usage → DSP-optimized (14 DSPs, higher speed)

Full DSP usage significantly improves timing at the cost of DSP48 primitives.

Timing

The IP has a fixed 30-cycle pipeline latency:

Clock cycle Event
0 Input A presented with CE = 1
30 Output F valid with DV = 1

The READY_IN/READY_OUT handshake signals enable backpressure control for streaming applications.

Typical use cases

  • Normalization (x / max = x × (1/max))
  • Inverse scaling operations
  • Precomputing division constants
  • Iterative algorithms requiring repeated division by same value

Performance note

For general A/B division, use the dedicated division block. Use reciprocal when:

  • Dividing many values by the same constant (compute 1/constant once)
  • Implementing x/y as x × (1/y) in multiply-heavy pipelines