DSP - BLOCK MEAN ABS
The mean absolute deviation of a block of N consecutive samples about a RUNTIME level driven on the LEVEL pin: mean_abs = (sum of |x - LEVEL|) / N. N is a power of two chosen at RUN TIME on the EXP input pin (EXP = 10 means N = 1024), so the division by N is an exact arithmetic shift and the block contains no divider at all. This is the cheap, ROBUST alternative to Block Std Dev: no multiplier, no square root, no serial tail, and far less sensitive to a single outlier - for a Gaussian signal it is a fixed 0.7979 = sqrt(2/pi) times the standard deviation, so a threshold calibrated on one converts to the other by a constant. The result is BIT EXACT and is presented exactly 2 clocks after the block’s last sample. LEVEL unconnected ties to 0, which makes it the mean absolute VALUE. IN_DV is the only qualifier and there is deliberately no CE pin. Optional BUSY / INTEGRATING / SAMPLE_COUNT status outputs. Blocks of up to 2^20 samples out of the box, 2^31 if you ask for it.
Introduction
The Block Mean Abs block chops the input stream into consecutive blocks of N
samples and, at the end of each block, publishes the mean absolute deviation
of that block about a level $\ell$ driven on the LEVEL pin. With
$d_i = x_i - \ell$:
$$ \mathrm{mean_abs} = \frac{1}{N} \sum_{i=0}^{N-1} |d_i| = \frac{1}{N} \sum_{i=0}^{N-1} |x_i - \ell| $$
One subtractor, two accumulators, one subtraction and one shift. No multiplier, no square root, no divider.
N is a runtime input, not a property. You drive the exponent on the
EXP pin and the block size is $N = 2^{\mathrm{EXP}}$:
| EXP | N | EXP | N |
|---|---|---|---|
| 4 | 16 | 12 | 4096 |
| 6 | 64 | 16 | 65536 |
| 8 | 256 | 20 | 1048576 |
Because N is a power of two, the division by N is an exact arithmetic shift. There is no divider, no reciprocal ROM and no rounding beyond the single final requantisation into your Q format - which is also why the block size can be changed while the design is running, for free.
What it is FOR
The mean absolute deviation is a robust spread / noise estimate that needs no square root at all. It is the cheap alternative to the standard deviation, and on most channels it is the better one:
-
for a Gaussian signal it is a fixed multiple of the standard deviation,
$$ \mathrm{mean_abs} = \sqrt{\tfrac{2}{\pi}};\sigma \approx 0.7979,\sigma $$
so a threshold calibrated on one converts to the other by a constant - if you were reaching for Block Std Dev only to get “how noisy is this”, multiply by 1.2533 (or do not bother, and calibrate in these units);
-
a single large outlier moves it linearly instead of quadratically, so a stray spike does not blow up the noise estimate the way it does with $\sigma$;
-
it costs one adder and a shift instead of a serial multiply plus a serial square root, and its latency is a constant 2 clocks instead of tens.
With LEVEL at 0 - which is what an unconnected pin gives you - it is the mean
absolute value of the block, a straightforward activity measure. Feed the
block’s own mean back into LEVEL and you get the textbook mean absolute
deviation about the mean, delayed by one block, which is the usual way to
use it on a stationary signal.
Cost
One subtractor for $d = x - \ell$, two area accumulators (each (input working width + 1 + Max Block Exponent) bits), one subtraction and one requantising shift in the final state. No multiplier, no divider, no square root, no serial arithmetic.
When to use this instead of Block Statistics
The all-in-one Block Statistics block is not deprecated and computes this same mean absolute deviation among twenty other statistics. The rule is simple:
- you want several statistics of the SAME block - mean abs and mean and the areas of the same N samples - use Block Statistics. They share one accumulator and one serial tail, so the second and third statistic are nearly free.
- you want exactly one number - use this block. Then you synthesise only that number: the pin list, the logic and the tail are all that the mean absolute deviation needs, and nothing else reaches the synthesiser.
Two Block Statistics blocks side by side would duplicate the accumulators; two per-operator blocks side by side duplicate them too. One Block Statistics block never does.
Pin Description
IN_DV is high.
'1'. (There is deliberately no CE pin - to stall the block,
gate this.)
EXP, so a mid-block change takes
effect on the NEXT block. Unconnected defaults to 0, which makes the result
the mean absolute VALUE; drive it with the block’s own mean to get the mean
absolute deviation about the mean, delayed by one block.
OUT_DV clock and on no
other; it holds the previous block’s result until then.
MEAN_ABS is updated on this clock and on no other.
BUSY is still high here and falls on the next clock.
OUT_DV pulse, and it falls on the clock after. On a continuous stream it
simply stays high. Present on the symbol only when Enable BUSY = YES.
How many samples have been accumulated so far in the current block: 1 after the first, N after the N-th. It is NOT cleared at the end of a block
- it HOLDS the final count through the tail and past
OUT_DV, until the first sample of the next block takes it back to 1, so on theOUT_DVclock it reads the length of the block being presented. OnlyRESETclears it to 0. Fixed 32 bits. Present on the symbol only when Enable SAMPLE_COUNT = YES.
Properties
Number of INTEGER bits of the input sample (the sign, when present, uses one of them).
Integer bits of the input sample (the sign, when present, uses one of them). 1..64. Default 16. The LEVEL pin uses this same format.Default: 16
Options: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64
Number of FRACTIONAL bits of the input sample, i.e. the bits to the right of the binary point. Total width = integer + fractional bits, and must not exceed 64.
Fractional bits of the input sample. 0..64. Total input width must be 2..64 bits. Default 0. The LEVEL pin uses this same format.Default: 0
Options: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64
Select whether the input sample is signed (two’s complement) or unsigned.
SIGNED (two’s complement) or UNSIGNED input. Default SIGNED. An UNSIGNED input costs one extra bit internally, because a sample has to be promoted to signed before the level is subtracted from it.Default: SIGNED
Options: UNSIGNED SIGNED
Largest block-size exponent the accumulators are sized for: the block can be up to 2^MaxBlockExponent samples long. The EXP input is clamped to this value at run time. Raising it widens the internal accumulators, and ON THE BLOCKS WHOSE SERIAL ENGINES ARE SIZED FROM THOSE ACCUMULATORS (Coefficient of Variation, SNR, Skewness, Kurtosis, Correlation, Autocorrelation, Linear Regression) it also LENGTHENS THE SERIAL TAIL – even when the runtime EXP is small. Keep it at the largest block you actually use. The default of 20 covers blocks of up to 1048576 samples.
Largest block-size exponent the accumulators are sized for: the block can be up to $2^{\text{MaxBlockExponent}}$ samples long, and theEXP input is
clamped to this value at run time. Raising it widens the two internal area
registers by one bit per unit; it does NOT lengthen the latency of this
block, which is a constant 2 clocks. Keep it at the largest block you
actually use. 1..31, default 20, i.e. blocks of up to 1048576 samples
out of the box.
Default: 20
Options: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31
Number of INTEGER bits of the MEAN_ABS output (the sign, when present, uses one of them).
Integer bits of the MEAN_ABS output. 1..64, default 16. The mean absolute deviation is in the units of the input and never exceeds the largest representable sample magnitude, so the input integer bits are always enough.Default: 16
Options: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64
Number of FRACTIONAL bits of the MEAN_ABS output, i.e. the bits to the right of the binary point. Total width = integer + fractional bits, and must not exceed 64.
Fractional bits of the MEAN_ABS output. 0..64, total width 2..64 bits, default 0. Ask for more than the input has if you want sub-LSB resolution on the noise estimate - the accumulators carry those bits exactly.Default: 0
Options: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64
Select whether the MEAN_ABS output is signed (two’s complement) or unsigned.
SIGNED or UNSIGNED MEAN_ABS output. Default UNSIGNED: the mean of $|d|$ can never be negative, so UNSIGNED buys one bit. Pick SIGNED only if the number has to feed a signed datapath downstream.Default: UNSIGNED
Options: UNSIGNED SIGNED
YES: the BUSY (high from the first sample of a block until its result is out – it COVERS THE SERIAL TAIL, and its last high clock IS the OUT_DV pulse) pin is present. NO: the pin AND all of its logic are removed BEFORE synthesis, so nothing is paid for it.
YES: theBUSY pin exists. It is high from the first sample of a block
until its result is out, tail included, and its last high clock is the
OUT_DV pulse. NO: the pin and its register are removed before synthesis.
Default NO.
Default: NO
Options: NO YES
YES: the INTEGRATING (high only while the block is ACCUMULATING; it drops as soon as the N-th sample has been taken and the tail starts, so BUSY-and-not-INTEGRATING means ‘computing’) pin is present. NO: the pin AND all of its logic are removed BEFORE synthesis, so nothing is paid for it.
YES: theINTEGRATING pin exists. It is high only while the block is
accumulating, so BUSY high with INTEGRATING low means “the samples are
all in, I am computing”. NO: the pin and its register are removed.
Default NO.
Default: NO
Options: NO YES
YES: the SAMPLE_COUNT (32 bit, how many samples have been accumulated so far in the current block: 1 after the first, N after the N-th. It is NOT cleared at the block end – it holds N until the NEXT block’s first accepted sample takes it back to 1. On a CONTINUOUS stream that happens DURING the serial tail, so at OUT_DV it reads how far into the next block the input has already got, NOT N. To capture the length of the block being presented, latch SAMPLE_COUNT on the clock INTEGRATING falls – that one always reads N) pin is present. NO: the pin AND all of its logic are removed BEFORE synthesis, so nothing is paid for it.
YES: theSAMPLE_COUNT pin exists - a fixed 32 bit count of the samples
accumulated so far in the current block, holding the final count through
the tail and past OUT_DV. NO: the pin and its counter are removed.
Default NO.
Default: NO
Options: NO YES
ROUND: round to nearest when a result has to be requantised into a coarser output format. TRUNCATE: drop the bits (cheaper, adds a negative bias).
ROUND: round to nearest when the result has to be requantised into a coarser output format. TRUNCATE: drop the bits (cheaper, adds a negative bias). Default ROUND.Default: ROUND
Options: TRUNCATE ROUND
YES: clip to the largest representable value of each output format (symmetric for signed formats). NO: wrap around.
YES: clip to the largest representable value of the output format (symmetric bounds for signed formats). NO: wrap around. Only matters when the output format is too narrow for the value - which, for a mean absolute deviation sized like the input, cannot happen. Default YES.Default: YES
Options: NO YES
Accuracy
The two accumulators are exact integers, the subtraction that combines them
is exact, and the division by N is a shift, so the only error in this
block is the single final requantisation into the Q format you chose for the
MEAN_ABS pin. There is no accumulated rounding, no truncated intermediate and
no approximation anywhere in the datapath.
That is not an aspiration. The host regression (tb/block-ops/run_tb.ps1)
demands tolerance ZERO against a Python golden (tb/block-ops/ gen_golden.py) that evaluates the definition above in exact rational
arithmetic - not “within 1 LSB”, not “within a few counts”. Any deviation at
all fails the build.
How the sum of magnitudes is built
$\sum |d_i|$ is not a third accumulator. The block keeps the positive and the negative areas of the block,
$$ A^{+} = !!\sum_{d_i > 0}!! d_i ;;(\ge 0), \qquad A^{-} = !!\sum_{d_i < 0}!! d_i ;;(\le 0) $$
and subtracts them in the final state:
$$ \sum_i |d_i| = A^{+} - A^{-} $$
One subtraction instead of an extra accumulator and an absolute-value stage in the sample path. The identity is EXACT, because a sample sitting exactly on the level ($d = 0$) contributes 0 to both areas - so it is neither counted twice nor lost.
The divide by N and the requantisation into your output format are then folded into one shift: the block never materialises the sum and then divides it, it shifts once by $\mathrm{IN_fract} + \mathrm{EXP} - \mathrm{MEAN_ABS_{fract}}$ and rounds once.
The LEVEL pin
LEVEL is an ordinary input pin in the input Q format, and it is always on
the symbol. It defaults to 0 when left unconnected, which makes the result
the mean absolute value rather than a deviation about anything.
LEVEL is latched on the first accepted sample of a block, exactly like
EXP, and held for that whole block. A mid-block change therefore takes effect
on the NEXT block: a block is never split between two levels, and every
result is referred to one single, well defined level.
The classic wiring is a feedback loop: take MEAN from a Block Mean block
(or from Block Statistics) driven by the same stream and the same EXP, and
drive it into this block’s LEVEL. Because the level is latched per block,
what you get is the mean absolute deviation about the previous block’s
mean - which on a stationary signal is exactly the textbook quantity, one block
late.
Accumulation and IN_DV
IN_DV is the only qualifier. It says “this clock carries a sample”: a
sample is accumulated, and counts towards N, exactly on the clocks where
IN_DV is high. Clocks with IN_DV low are ignored completely - whatever
sits on IN during them cannot corrupt the block - while the tail keeps
running, which is what you want: the tail has nothing to do with the input
stream.
Unconnected, IN_DV ties to '1', EXP ties to 10 (N = 1024) and LEVEL
ties to all zeros, so the block free-runs with nothing wired except IN.
There is deliberately no CE pin. On the all-in-one Block Statistics block an earlier revision had one, and it did not survive synthesis: with nothing but internal state gated by it, Vitis could reason the frozen path away and delete the port from the generated entity while SciCompiler’s wrapper still wired it, which failed a real Vivado build with [VRFC 10-718] formal port <ce> does not exist in entity. The whole per-operator family was built without one. To stall this block, gate its
IN_DV- a block that only accumulates onIN_DVhas no need to be frozen.
When EXP changes
EXP is clamped to Max Block Exponent and then latched on the first
accepted sample of a block, and held for that whole block. A change
therefore takes effect on the NEXT block: a block in progress always
finishes against the N it was started with, and a block is never emitted
against a different N than the one it was accumulated with. The divide-by-N
shift at the end uses the exponent that was latched, not whatever happens to be
on the pin when the result comes out.
Timing: the latency contract
OUT_DV pulses for one clock, L clocks after the clock on which the N-th
sample of the block was accepted - not when that sample arrives. MEAN_ABS
is updated on that same clock and on no other. For this block
$$ L = 2 $$
and it is a constant: there is no serial arithmetic here at all - the two
accumulators run at one sample per clock, and the final state is one
subtraction plus one requantising shift - so L does not depend on the input
width, on the output width or on EXP. The two clocks are one to enter the
final state and one to present the registered result.
The rule that governs the whole family is that the tail of one block must finish before the next block completes, i.e.
$$ 2^{\mathrm{EXP}} \ge L $$
If a block completes while the previous tail is still running, that block’s
result is DROPPED: no OUT_DV for it, the accumulators are unaffected and
later blocks come out correctly, but a result is silently skipped. There is
no error pin for it.
With $L = 2$ that condition is $2^{\mathrm{EXP}} \ge 2$, i.e. EXP $\ge$ 1, so it cannot bite here: the only value that violates it is EXP = 0, a block of a single sample. The blocks where this rule really matters are the ones with a serial tail - Block RMS, Block Variance, Block Std Dev and Block Crest Factor, whose L runs to tens of clocks and whose minimum usable exponent the compiler prints in the compilation log. It is one more reason to prefer the mean absolute deviation over the standard deviation when a noise number is all you need: short blocks stay usable.
Knowing where the block is: BUSY, INTEGRATING and SAMPLE_COUNT
Three optional status outputs, all defaulting to NO. They answer different questions:
INTEGRATING |
BUSY |
|
|---|---|---|
| accumulating the block | 1 | 1 |
| tail computing | 0 | 1 |
| idle | 0 | 0 |
Every output of this block is a register, so each status bit is observed on the clock after the event that sets it:
INTEGRATINGrises on the clock after the FIRST sample of a block is accepted and falls on the clock after the N-th - it is high exactly while the block is ACCUMULATING.BUSYcovers the accumulation and the tail. It rises withINTEGRATING, stays high across the tail, and its LAST HIGH CLOCK IS THEOUT_DVPULSE; it falls on the clock after.- On a continuous stream the next block starts before the previous tail
ends, so
BUSYnever drops andINTEGRATINGdips for exactly one clock per block boundary - which makes it a free block marker. SAMPLE_COUNTis a fixed 32 bits and reads 1 after the first accepted sample, N after the N-th. It is NOT cleared at the block end: it HOLDS N through the tail and pastOUT_DV, until the first sample of the next block takes it back to 1. So on theOUT_DVclock it reads the length of the block being presented - which is the useful thing to latch alongside the result. OnlyRESETclears it to 0.
Q formats
IN, LEVEL and MEAN_ABS carry fixed point formats in the usual convention
of the Fixed P. family - LEVEL shares the input format, since it is
compared with samples. The result is requantised into the MEAN_ABS format
with the selected rounding (nearest / truncate) and overflow policy (saturate /
wrap); saturation is symmetric for signed formats, as everywhere else in the
toolchain.
Sizing is easy here: the mean absolute deviation is in the units of the input and can never exceed the largest representable magnitude of a sample, so the input format is always safe for the output. It is never negative, so the default format is UNSIGNED, which buys one bit; choose SIGNED only if the number has to feed a signed datapath downstream. Ask for more fractional bits than the input has if you want sub-LSB resolution on the noise estimate - the accumulators carry those bits exactly, they are simply discarded by the final shift if you do not ask for them.
Verification
The core is regression tested by a host-side csim harness
(tb/block-ops/run_tb.ps1) that runs one simulated clock at a time and
follows OUT_DV. The expected values come from tb/block-ops/gen_golden.py,
which evaluates $\sum|x-\ell|/N$ in exact rational arithmetic and shares no
algorithm with the core; the tolerance is 0. Mean abs coverage includes
pseudo-random input at level 0, a non-zero level, a signal with samples
sitting exactly on the level (the case that pins the “contributes 0 to
both areas” convention and makes the pos-minus-neg identity exact), a LEVEL
pin that changes half way through every block in both directions - which is
what proves the per-block latch - and a fractional output format. The status
outputs are checked clock by clock against the contract above. A
cross-check compiles this core and the all-in-one block_stats.cpp into the
same binary, drives them with identical stimulus, and compares the two clock
by clock.