Xilinx
HLS
Block Preview

Introduction

The Block Mean Square block chops the input stream into consecutive blocks of N samples and, at the end of each block, publishes the mean of the squares of that block:

$$ S_2 = \sum_{i=0}^{N-1} x_i^2 , \qquad \overline{x^2} = \frac{S_2}{N} $$

This is the average power of the block. It is the second moment about zero, and it is the RMS before the square root:

$$ \mathrm{rms} = \sqrt{\overline{x^2}} $$

N is a runtime input, not a property. You drive the exponent on the EXP pin and the block size is $N = 2^{\mathrm{EXP}}$:

EXP N EXP N
4 16 12 4096
6 64 16 65536
8 256 20 1048576

Because N is a power of two, the division by N is an exact arithmetic shift. There is no divider, no reciprocal ROM and no rounding beyond the single final requantisation into your Q format - which is also why the block size can be changed while the design is running, for free.

Why power, and not RMS

The square root is the only expensive thing in this family: it is a serial digit recurrence that takes tens of clocks and lengthens the block latency. It is also, very often, unnecessary.

Whenever the consumer of the number is going to compare it against a threshold - a level meter, a squelch, a pile-up veto, an AGC decision, a “is this channel alive” test - you can compare powers instead of amplitudes and skip the root entirely. The comparison is monotonic:

$$ \mathrm{rms} > T \iff \overline{x^2} > T^2 $$

so you square the threshold once, at design time or in a register, and the hardware never takes a root at all. That is the case this block is for.

Reach for Block RMS instead when you genuinely need the number in the units of the input - when a human reads it, when it scales something, or when it is divided by another amplitude. Block RMS is this same accumulator with the serial root bolted on: if you need both the mean square and the RMS, one Block RMS costs less than this block plus a Block RMS, because the two would otherwise duplicate the multiplier and the accumulator.

Note also that the mean square includes the DC level. If what you want is the spread about the mean rather than the power about zero, that is the variance - see Block Variance and Block Std Dev.

Cost

Exactly one multiplier. $x \cdot x$ at one sample per clock is a real IN_SW $\times$ IN_SW multiply (IN_SW = the input width, +1 bit if the input is unsigned) and it is unavoidable at full rate. That single DSP is the entire cost difference against Block Mean. Everything else is shift-add: one accumulator, one barrel shifter for the divide-by-N, one requantiser. No divider, no square root, no serial arithmetic.

When to use this instead of Block Statistics

The all-in-one Block Statistics block is not deprecated and can emit this same MEAN_SQ among twenty other statistics. The rule is simple:

  • you want several statistics of the SAME block - the mean square and the mean and the min/max of the same N samples - use Block Statistics. They share one accumulator and one serial tail, so the second and third statistic are nearly free.
  • you want exactly one number - use this block. Then you synthesise only that number: the pin list, the logic and the tail are all that the mean square needs, and nothing else reaches the synthesiser.

Pin Description

IN Input IN_BitsInt + IN_BitsFract bit BIT VECTOR
Input samples, fixed point in the IN Q format. Squared and accumulated only on the clocks where IN_DV is high.
Default: Must be connected
IN_DV Input 1 bit BIT
Per-sample qualifier, active high, and the ONLY qualifier this block has. A sample is squared, accumulated, and counts towards N, exactly on the clocks where this is high; the tail keeps running regardless. Unconnected defaults to '1'. (There is deliberately no CE pin - to stall the block, gate this.)
EXP Input 6 bit BIT VECTOR
Block size exponent, runtime programmable: the block is $N = 2^{\text{EXP}}$ samples long, and the division by N is the arithmetic shift this exponent names. 6 bits unsigned, accepted range 0 .. Max Block Exponent; larger values are clamped to Max Block Exponent. Sampled on the first accepted sample of a block and held for that whole block, so a change takes effect on the NEXT block. Use EXP >= 1 (see “Timing”). Unconnected defaults to 10 (N = 1024).
MEAN_SQ Output MEAN_SQ_BitsInt + MEAN_SQ_BitsFract bit BIT VECTOR
$S_2 / N$, the mean of the SQUARES - the average power of the block - in the MEAN_SQ Q format. Updated on the OUT_DV clock and on no other; it holds the previous block’s result until then. It is never negative, so an UNSIGNED format costs nothing and buys a bit.
OUT_DV Output 1 bit BIT
One-clock pulse marking a valid result. It fires L = 2 clocks after the clock on which the N-th sample of the block was accepted, not when that sample arrives. MEAN_SQ is updated on this clock and on no other. BUSY is still high here and falls on the next clock.
CLK 1 bit
Clock.
RESET 1 bit
Synchronous reset: clears the accumulator, the block counter, the sample count and the tail.
BUSY 1 bit
High from the start of a block - its first accumulated sample - until its result is out: it covers the tail as well. Its last high clock is the OUT_DV pulse, and it falls on the clock after. On a continuous stream it simply stays high. Present on the symbol only when Enable BUSY = YES.
INTEGRATING 1 bit
High only while the block is accumulating: it rises on the clock after the first sample of a block is accepted and falls on the clock after the N-th. On a continuous stream it dips for exactly one clock per block boundary, which makes it a free block marker. Present on the symbol only when Enable INTEGRATING = YES.
SAMPLE_COUNT 32 bit

How many samples have been accumulated so far in the current block: 1 after the first, N after the N-th. It is NOT cleared at the end of a block

  • it HOLDS the final count through the tail and past OUT_DV, until the first sample of the next block takes it back to 1, so on the OUT_DV clock it reads the length of the block being presented. Only RESET clears it to 0. Fixed 32 bits. Present on the symbol only when Enable SAMPLE_COUNT = YES.

Properties

Property window

IN Integer Bits IN_BitsInt

Number of INTEGER bits of the input sample (the sign, when present, uses one of them).

Integer bits of the input sample (the sign, when present, uses one of them). 1..64. Default 16.

Default: 16

Options: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64

IN Fractional Bits IN_BitsFract

Number of FRACTIONAL bits of the input sample, i.e. the bits to the right of the binary point. Total width = integer + fractional bits, and must not exceed 64.

Fractional bits of the input sample. 0..64. Total input width must be 2..64 bits. Default 0.

Default: 0

Options: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64

IN Sign IN_Sign

Select whether the input sample is signed (two’s complement) or unsigned.

SIGNED (two’s complement) or UNSIGNED input. Default SIGNED. An UNSIGNED input costs one extra bit internally, because a sample has to be promoted to signed before it is squared - and that promoted width is also the width of the multiplier.

Default: SIGNED

Options: UNSIGNED SIGNED

Max Block Exponent MaxBlockExponent

Largest block-size exponent the accumulators are sized for: the block can be up to 2^MaxBlockExponent samples long. The EXP input is clamped to this value at run time. Raising it widens the internal accumulators, and ON THE BLOCKS WHOSE SERIAL ENGINES ARE SIZED FROM THOSE ACCUMULATORS (Coefficient of Variation, SNR, Skewness, Kurtosis, Correlation, Autocorrelation, Linear Regression) it also LENGTHENS THE SERIAL TAIL – even when the runtime EXP is small. Keep it at the largest block you actually use. The default of 20 covers blocks of up to 1048576 samples.

Largest block-size exponent the accumulator is sized for: the block can be up to $2^{\text{MaxBlockExponent}}$ samples long, and the EXP input is clamped to this value at run time. Raising it widens the internal sum-of-squares register by one bit per unit; it does NOT lengthen the latency of this block, which is a constant 2 clocks, and it does not widen the multiplier. Keep it at the largest block you actually use. 1..31, default 20, i.e. blocks of up to 1048576 samples out of the box.

Default: 20

Options: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31

MEAN_SQ Integer Bits MEAN_SQ_BitsInt

Number of INTEGER bits of the MEAN SQUARE output (the sign, when present, uses one of them).

Integer bits of the MEAN SQUARE output. 1..64, default 32. It is a SQUARE: allow about twice the input integer bits. 32 is exactly the square of a 16 bit sample. It does NOT have to grow with the block exponent - averaging cannot exceed the largest single squared sample.

Default: 32

Options: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64

MEAN_SQ Fractional Bits MEAN_SQ_BitsFract

Number of FRACTIONAL bits of the MEAN SQUARE output, i.e. the bits to the right of the binary point. Total width = integer + fractional bits, and must not exceed 64.

Fractional bits of the MEAN SQUARE output. 0..64, total width 2..64 bits, default 0. A square has twice the fractional bits of the input, so allow about $2F$ if you want to keep them.

Default: 0

Options: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64

MEAN_SQ Sign MEAN_SQ_Sign

Select whether the MEAN SQUARE output is signed (two’s complement) or unsigned.

SIGNED or UNSIGNED MEAN SQUARE output. Default UNSIGNED: a mean of squares is never negative, so UNSIGNED buys one bit and a signed format would simply waste it.

Default: UNSIGNED

Options: UNSIGNED SIGNED

Enable BUSY EnableBusy

YES: the BUSY (high from the first sample of a block until its result is out – it COVERS THE SERIAL TAIL, and its last high clock IS the OUT_DV pulse) pin is present. NO: the pin AND all of its logic are removed BEFORE synthesis, so nothing is paid for it.

YES: the BUSY pin exists. It is high from the first sample of a block until its result is out, tail included, and its last high clock is the OUT_DV pulse. NO: the pin and its register are removed before synthesis. Default NO.

Default: NO

Options: NO YES

Enable INTEGRATING EnableIntegrating

YES: the INTEGRATING (high only while the block is ACCUMULATING; it drops as soon as the N-th sample has been taken and the tail starts, so BUSY-and-not-INTEGRATING means ‘computing’) pin is present. NO: the pin AND all of its logic are removed BEFORE synthesis, so nothing is paid for it.

YES: the INTEGRATING pin exists. It is high only while the block is accumulating, so BUSY high with INTEGRATING low means “the samples are all in, I am computing”. NO: the pin and its register are removed. Default NO.

Default: NO

Options: NO YES

Enable SAMPLE_COUNT EnableSampleCount

YES: the SAMPLE_COUNT (32 bit, how many samples have been accumulated so far in the current block: 1 after the first, N after the N-th. It is NOT cleared at the block end – it holds N until the NEXT block’s first accepted sample takes it back to 1. On a CONTINUOUS stream that happens DURING the serial tail, so at OUT_DV it reads how far into the next block the input has already got, NOT N. To capture the length of the block being presented, latch SAMPLE_COUNT on the clock INTEGRATING falls – that one always reads N) pin is present. NO: the pin AND all of its logic are removed BEFORE synthesis, so nothing is paid for it.

YES: the SAMPLE_COUNT pin exists - a fixed 32 bit count of the samples accumulated so far in the current block, holding the final count through the tail and past OUT_DV. NO: the pin and its counter are removed. Default NO.

Default: NO

Options: NO YES

Rounding Rounding

ROUND: round to nearest when a result has to be requantised into a coarser output format. TRUNCATE: drop the bits (cheaper, adds a negative bias).

ROUND: round to nearest when the result has to be requantised into a coarser output format. TRUNCATE: drop the bits (cheaper, adds a negative bias). Default ROUND.

Default: ROUND

Options: TRUNCATE ROUND

Saturation EnableSaturation

YES: clip to the largest representable value of each output format (symmetric for signed formats). NO: wrap around.

YES: clip to the largest representable value of the MEAN_SQ format (symmetric bounds for signed formats). NO: wrap around. Only matters when the output format is too narrow for the value - which, for a square given fewer than twice the input integer bits, is easy to arrange. Default YES.

Default: YES

Options: NO YES

Accuracy

The accumulator $S_2$ is an exact integer - each $x^2$ is an exact product of two integers, accumulated at full width - and the division by N is a shift, so the only error in this block is the single final requantisation into the Q format you chose for the MEAN_SQ pin. There is no accumulated rounding, no truncated intermediate and no approximation anywhere in the datapath.

That is not an aspiration. The host regression (tb/block-ops/run_tb.ps1) demands tolerance ZERO against a Python golden (tb/block-ops/ gen_golden.py) that evaluates the definition above in exact rational arithmetic - not “within 1 LSB”, not “within a few counts”. Any deviation at all fails the build. (Contrast Block RMS, whose serial square root is specified to 1 LSB: taking the root is what introduces the slack, and this block does not take one.)

Accumulation and IN_DV

IN_DV is the only qualifier. It says “this clock carries a sample”: a sample is squared, accumulated, and counts towards N, exactly on the clocks where IN_DV is high. Clocks with IN_DV low are ignored completely - whatever sits on IN during them cannot corrupt the block - while the tail keeps running, which is what you want: the tail has nothing to do with the input stream.

Unconnected, IN_DV ties to '1' and EXP ties to 10 (N = 1024), so the block free-runs with nothing wired except IN.

There is deliberately no CE pin. On the all-in-one Block Statistics block an earlier revision had one, and it did not survive synthesis: with nothing but internal state gated by it, Vitis could reason the frozen path away and delete the port from the generated entity while SciCompiler’s wrapper still wired it, which failed a real Vivado build with [VRFC 10-718] formal port <ce> does not exist in entity. The whole per-operator family was built without one. To stall this block, gate its IN_DV - a block that only accumulates on IN_DV has no need to be frozen.

When EXP changes

EXP is clamped to Max Block Exponent and then latched on the first accepted sample of a block, and held for that whole block. A change therefore takes effect on the NEXT block: a block in progress always finishes against the N it was started with, and a block is never emitted against a different N than the one it was accumulated with. The divide-by-N shift at the end uses the exponent that was latched, not whatever happens to be on the pin when the result comes out.

Timing: the latency contract

OUT_DV pulses for one clock, L clocks after the clock on which the N-th sample of the block was accepted - not when that sample arrives. MEAN_SQ is updated on that same clock and on no other. For this block

$$ L = 2 $$

and it is a constant: there is no serial arithmetic here at all - the multiply happens in the accumulation, at one sample per clock, and the divide is a shift - so L does not depend on the input width, on the output width or on EXP. The two clocks are one to enter the final state and one to present the registered result.

The rule that governs the whole family is that the tail of one block must finish before the next block completes, i.e.

$$ 2^{\mathrm{EXP}} \ge L $$

If a block completes while the previous tail is still running, that block’s result is DROPPED: no OUT_DV for it, the accumulator is unaffected and later blocks come out correctly, but a result is silently skipped. There is no error pin for it.

With $L = 2$ that condition is $2^{\mathrm{EXP}} \ge 2$, i.e. EXP $\ge$ 1, so it cannot bite here: the only value that violates it is EXP = 0, a block of a single sample. The blocks where this rule really matters are the ones with a serial tail - Block RMS, Block Variance, Block Std Dev and Block Crest Factor, whose L runs to tens of clocks and whose minimum usable exponent the compiler prints in the compilation log. It is one more reason to prefer the mean square over the RMS when a threshold comparison is all you need: short blocks stay usable.

Knowing where the block is: BUSY, INTEGRATING and SAMPLE_COUNT

Three optional status outputs, all defaulting to NO. They answer different questions:

INTEGRATING BUSY
accumulating the block 1 1
tail computing 0 1
idle 0 0

Every output of this block is a register, so each status bit is observed on the clock after the event that sets it:

  • INTEGRATING rises on the clock after the FIRST sample of a block is accepted and falls on the clock after the N-th - it is high exactly while the block is ACCUMULATING.
  • BUSY covers the accumulation and the tail. It rises with INTEGRATING, stays high across the tail, and its LAST HIGH CLOCK IS THE OUT_DV PULSE; it falls on the clock after.
  • On a continuous stream the next block starts before the previous tail ends, so BUSY never drops and INTEGRATING dips for exactly one clock per block boundary - which makes it a free block marker.
  • SAMPLE_COUNT is a fixed 32 bits and reads 1 after the first accepted sample, N after the N-th. It is NOT cleared at the block end: it HOLDS N through the tail and past OUT_DV, until the first sample of the next block takes it back to 1. So on the OUT_DV clock it reads the length of the block being presented - which is the useful thing to latch alongside the result. Only RESET clears it to 0.

Q formats

Both ports carry their own fixed point format (integer bits, fractional bits, sign), the same convention as the Fixed P. family. The result is requantised into the MEAN_SQ format with the selected rounding (nearest / truncate) and overflow policy (saturate / wrap); saturation is symmetric for signed formats, as everywhere else in the toolchain.

Sizing, for an input of $W$ bits with $F$ fractional bits:

  • MEAN_SQ is a SQUARE. It reaches the square of the largest sample - allow about $2W$ integer bits, or $2F$ fractional bits, if you do not want it to saturate. The default is 32 bits UNSIGNED, which is exactly the square of a 16 bit sample.
  • It is never negative, so leave it UNSIGNED and buy a bit of range.
  • Averaging cannot make it larger than the biggest single squared sample, so unlike Block Sum the output width does not have to grow with the block exponent.

If you are going to compare it against a squared threshold, remember to square the threshold in the same Q format as this pin.

Verification

The core is regression tested by a host-side csim harness (tb/block-ops/run_tb.ps1) that runs one simulated clock at a time and follows OUT_DV. The expected values come from tb/block-ops/gen_golden.py, which evaluates $S_2/N$ in exact rational arithmetic and shares no algorithm with the core; the tolerance is 0. Mean square coverage includes pseudo-random input, a sine into a fractional output format, unsigned input (the case where a sample has to be promoted before it is squared), and an EXP that changes half way through a block. The status outputs are checked clock by clock against the contract above. A cross-check compiles this core and the all-in-one block_stats.cpp into the same binary, drives them with identical stimulus, and compares the two clock by clock.