DSP - BLOCK SUM
The raw sum of a block of N consecutive samples - the mean WITHOUT the final divide-by-N shift, and the cheapest block of the family: one adder and one requantiser. N is a power of two chosen at RUN TIME on the EXP input pin (EXP = 10 means N = 1024); the sum never divides by N, but N is still what decides where a block ENDS, so the block length can be changed at run time for free. SATURATION MATTERS MORE HERE THAN ANYWHERE ELSE IN THE FAMILY - the internal accumulator is sized for the worst case, but the SUM output format is yours to choose and a long block of same-signed samples overflows a narrow one easily. IN_DV is the only qualifier and there is deliberately no CE pin. Optional BUSY / INTEGRATING / SAMPLE_COUNT status outputs. Blocks of up to 2^20 samples out of the box, 2^31 if you ask for it.
Introduction
The Block Sum block chops the input stream into consecutive blocks of N samples and, at the end of each block, publishes the sum of that block:
$$ S_1 = \sum_{i=0}^{N-1} x_i $$
That is the raw first accumulator, presented as it stands. It is Block Mean without the final shift, and therefore the cheapest block of the whole family: one adder and one requantiser, no multiplier, no divider and no serial arithmetic.
N is a runtime input, not a property. You drive the exponent on the
EXP pin and the block size is $N = 2^{\mathrm{EXP}}$:
| EXP | N | EXP | N |
|---|---|---|---|
| 4 | 16 | 12 | 4096 |
| 6 | 64 | 16 | 65536 |
| 8 | 256 | 20 | 1048576 |
The sum itself never divides by N - but N is what decides where a block ends, so it is still a runtime input, and changing the block length while the design is running costs nothing. In the rest of the family the same property has a second consequence: because N is a power of two every division by N is an exact arithmetic shift, which is why none of these blocks contains a divider.
What it is FOR
The sum is the integral of the block. In pulse processing, the sum of the samples over a gate IS the integrated charge; in a slow control loop it is the quantity you accumulate and let the CPU scale itself; in a measurement chain it is the numerator you want when the denominator is going to be something other than N (a live-time, a number of triggers, a second block’s sum).
Take it, rather than the mean, when:
- you want the raw accumulator and intend to do your own arithmetic downstream - combining several blocks, normalising by something other than N, or accumulating further;
- you want the full precision of the accumulation with no shift at all;
- you want the absolute cheapest block-rate measurement there is.
Take Block Mean instead when the number you actually want is a baseline or a level: it is this block plus one barrel shifter, and its output stays in the range of a sample, which makes it far easier to size.
Cost
One accumulator (input working width + Max Block Exponent bits) and one requantiser. No multiplier, no divider, not even a shifter in the datapath. Nothing in this family is cheaper.
When to use this instead of Block Statistics
The all-in-one Block Statistics block is not deprecated and can emit
this same SUM among twenty other statistics. The rule is simple:
- you want several statistics of the SAME block - the sum and the mean and the RMS of the same N samples - use Block Statistics. They share one accumulator and one serial tail, so the second and third statistic are nearly free.
- you want exactly one number - use this block. Then you synthesise only that number: the pin list, the logic and the tail are all that the sum needs, and nothing else reaches the synthesiser.
Pin Description
IN_DV is high.
'1'. (There is deliberately no CE pin - to stall the block,
gate this.)
OUT_DV clock and on no other; it holds the previous block’s result until
then. This is the pin that can clip: the internal accumulator cannot
overflow, but this format is yours to choose - size it as (input width +
block exponent) bits if you want it never to.
SUM is updated on this clock and on no other. BUSY is
still high here and falls on the next clock.
OUT_DV pulse, and it falls on the clock after. On a continuous stream it
simply stays high. Present on the symbol only when Enable BUSY = YES.
How many samples have been accumulated so far in the current block: 1 after the first, N after the N-th. It is NOT cleared at the end of a block
- it HOLDS the final count through the tail and past
OUT_DV, until the first sample of the next block takes it back to 1, so on theOUT_DVclock it reads the length of the block being presented - the divisor that turns this sum into a mean. OnlyRESETclears it to 0. Fixed 32 bits. Present on the symbol only when Enable SAMPLE_COUNT = YES.
Properties
Number of INTEGER bits of the input sample (the sign, when present, uses one of them).
Integer bits of the input sample (the sign, when present, uses one of them). 1..64. Default 16.Default: 16
Options: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64
Number of FRACTIONAL bits of the input sample, i.e. the bits to the right of the binary point. Total width = integer + fractional bits, and must not exceed 64.
Fractional bits of the input sample. 0..64. Total input width must be 2..64 bits. Default 0.Default: 0
Options: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64
Select whether the input sample is signed (two’s complement) or unsigned.
SIGNED (two’s complement) or UNSIGNED input. Default SIGNED. An UNSIGNED input costs one extra bit internally, because a sample has to be promoted to signed before it can be accumulated - and note that an unsigned stream is the worst case for the SUM format, since every sample pushes the total the same way.Default: SIGNED
Options: UNSIGNED SIGNED
Largest block-size exponent the accumulators are sized for: the block can be up to 2^MaxBlockExponent samples long. The EXP input is clamped to this value at run time. Raising it widens the internal accumulators, and ON THE BLOCKS WHOSE SERIAL ENGINES ARE SIZED FROM THOSE ACCUMULATORS (Coefficient of Variation, SNR, Skewness, Kurtosis, Correlation, Autocorrelation, Linear Regression) it also LENGTHENS THE SERIAL TAIL – even when the runtime EXP is small. Keep it at the largest block you actually use. The default of 20 covers blocks of up to 1048576 samples.
Largest block-size exponent the accumulator is sized for: the block can be up to $2^{\text{MaxBlockExponent}}$ samples long, and theEXP input is
clamped to this value at run time. Raising it widens the internal sum
register by one bit per unit; it does NOT lengthen the latency of this
block, which is a constant 2 clocks. Keep it at the largest block you
actually use. 1..31, default 20, i.e. blocks of up to 1048576 samples
out of the box.
Default: 20
Options: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31
Number of INTEGER bits of the SUM output (the sign, when present, uses one of them).
Integer bits of the SUM output. 1..64, default 24. It holds N samples: allow input integer bits + block exponent if you want it never to clip. 24 covers a 16 bit input over blocks up to $2^{8}$.Default: 24
Options: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64
Number of FRACTIONAL bits of the SUM output, i.e. the bits to the right of the binary point. Total width = integer + fractional bits, and must not exceed 64.
Fractional bits of the SUM output. 0..64, total width 2..64 bits, default 0. A sum needs no more fractional bits than the input has - the accumulation adds range, not resolution.Default: 0
Options: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64
Select whether the SUM output is signed (two’s complement) or unsigned.
SIGNED or UNSIGNED SUM output. Default SIGNED. A sum of signed samples needs SIGNED; an unsigned input can use UNSIGNED and buy one bit of range, which is worth having on this particular pin.Default: SIGNED
Options: UNSIGNED SIGNED
YES: the BUSY (high from the first sample of a block until its result is out – it COVERS THE SERIAL TAIL, and its last high clock IS the OUT_DV pulse) pin is present. NO: the pin AND all of its logic are removed BEFORE synthesis, so nothing is paid for it.
YES: theBUSY pin exists. It is high from the first sample of a block
until its result is out, tail included, and its last high clock is the
OUT_DV pulse. NO: the pin and its register are removed before synthesis.
Default NO.
Default: NO
Options: NO YES
YES: the INTEGRATING (high only while the block is ACCUMULATING; it drops as soon as the N-th sample has been taken and the tail starts, so BUSY-and-not-INTEGRATING means ‘computing’) pin is present. NO: the pin AND all of its logic are removed BEFORE synthesis, so nothing is paid for it.
YES: theINTEGRATING pin exists. It is high only while the block is
accumulating, so BUSY high with INTEGRATING low means “the samples are
all in, I am computing”. NO: the pin and its register are removed.
Default NO.
Default: NO
Options: NO YES
YES: the SAMPLE_COUNT (32 bit, how many samples have been accumulated so far in the current block: 1 after the first, N after the N-th. It is NOT cleared at the block end – it holds N until the NEXT block’s first accepted sample takes it back to 1. On a CONTINUOUS stream that happens DURING the serial tail, so at OUT_DV it reads how far into the next block the input has already got, NOT N. To capture the length of the block being presented, latch SAMPLE_COUNT on the clock INTEGRATING falls – that one always reads N) pin is present. NO: the pin AND all of its logic are removed BEFORE synthesis, so nothing is paid for it.
YES: theSAMPLE_COUNT pin exists - a fixed 32 bit count of the samples
accumulated so far in the current block, holding the final count through
the tail and past OUT_DV. NO: the pin and its counter are removed.
Default NO.
Default: NO
Options: NO YES
ROUND: round to nearest when a result has to be requantised into a coarser output format. TRUNCATE: drop the bits (cheaper, adds a negative bias).
ROUND: round to nearest when the result has to be requantised into a coarser output format. TRUNCATE: drop the bits (cheaper, adds a negative bias). Default ROUND. It only has an effect if the SUM format has fewer fractional bits than the input.Default: ROUND
Options: TRUNCATE ROUND
YES: clip to the largest representable value of each output format (symmetric for signed formats). NO: wrap around.
YES: clip to the largest representable value of the SUM format (symmetric bounds for signed formats). NO: wrap around. This is the block where the setting earns its keep - a long block of same-signed samples overflows a narrow SUM format easily, and neither policy raises a flag. Default YES.Default: YES
Options: NO YES
Accuracy
The accumulator $S_1$ is an exact integer, so the only error in this
block is the single final requantisation into the Q format you chose for the
SUM pin. There is no accumulated rounding and no approximation anywhere in
the datapath.
That is not an aspiration. The host regression (tb/block-ops/run_tb.ps1)
demands tolerance ZERO against a Python golden (tb/block-ops/ gen_golden.py) that evaluates the definition above in exact rational
arithmetic - not “within 1 LSB”, not “within a few counts”. Any deviation at
all fails the build.
Sizing the SUM output: saturation matters here
This is the one block of the family where the output format needs real
thought. The internal accumulator is sized for the worst case - the input
working width plus Max Block Exponent bits, so it cannot overflow no matter
what you feed it. But the SUM pin carries the Q format you choose,
and the requantisation into it is the last step. A long block of same-signed
samples reaches numbers that a narrow format simply cannot hold:
- a 16 bit signed input summed over a $2^{10}$ block reaches $\pm 2^{15}\cdot 2^{10} = \pm 2^{25}$ - a 26 bit signed quantity;
- the same input over a $2^{20}$ block reaches $\pm 2^{35}$.
Size SUM as (input width + block exponent) bits if you want it never to
clip. If you know your signal never sits at full scale for a whole block you
can be tighter, but there is no cheap safety net: with Saturation = YES the
pin clamps to the format maximum (symmetric bounds for a signed format) and
with NO it wraps, and neither raises a flag. Note that the exponent in
that rule is the exponent you actually run at, not Max Block Exponent - but
EXP is a runtime input, so if you intend to sweep it, size for the largest
value you will drive.
The default is 24 bits signed, which is exactly enough for a 16 bit signed input over blocks up to $2^{8}$.
Accumulation and IN_DV
IN_DV is the only qualifier. It says “this clock carries a sample”: a
sample is accumulated, and counts towards N, exactly on the clocks where
IN_DV is high. Clocks with IN_DV low are ignored completely - whatever
sits on IN during them cannot corrupt the block - while the tail keeps
running, which is what you want: the tail has nothing to do with the input
stream.
Unconnected, IN_DV ties to '1' and EXP ties to 10 (N = 1024), so the
block free-runs with nothing wired except IN.
There is deliberately no CE pin. On the all-in-one Block Statistics block an earlier revision had one, and it did not survive synthesis: with nothing but internal state gated by it, Vitis could reason the frozen path away and delete the port from the generated entity while SciCompiler’s wrapper still wired it, which failed a real Vivado build with [VRFC 10-718] formal port <ce> does not exist in entity. The whole per-operator family was built without one. To stall this block, gate its
IN_DV- a block that only accumulates onIN_DVhas no need to be frozen.
When EXP changes
EXP is clamped to Max Block Exponent and then latched on the first
accepted sample of a block, and held for that whole block. A change
therefore takes effect on the NEXT block: a block in progress always
finishes against the N it was started with, and a block is never emitted
against a different N than the one it was accumulated with. For a sum that
is the difference between a meaningful number and a partial one, so it is
worth stating plainly: no result you ever see is the sum of a number of
samples other than the N its SAMPLE_COUNT reports.
Timing: the latency contract
OUT_DV pulses for one clock, L clocks after the clock on which the N-th
sample of the block was accepted - not when that sample arrives. SUM is
updated on that same clock and on no other. For this block
$$ L = 2 $$
and it is a constant: there is no serial arithmetic here at all, so L does
not depend on the input width, on the output width or on EXP. The two clocks
are one to enter the final state and one to present the registered result.
The rule that governs the whole family is that the tail of one block must finish before the next block completes, i.e.
$$ 2^{\mathrm{EXP}} \ge L $$
If a block completes while the previous tail is still running, that block’s
result is DROPPED: no OUT_DV for it, the accumulator is unaffected and
later blocks come out correctly, but a result is silently skipped. There is
no error pin for it.
With $L = 2$ that condition is $2^{\mathrm{EXP}} \ge 2$, i.e. EXP $\ge$ 1, so it cannot bite here: the only value that violates it is EXP = 0, a block of a single sample. The blocks where this rule really matters are the ones with a serial tail - Block RMS, Block Variance, Block Std Dev and Block Crest Factor, whose L runs to tens of clocks and whose minimum usable exponent the compiler prints in the compilation log.
Knowing where the block is: BUSY, INTEGRATING and SAMPLE_COUNT
Three optional status outputs, all defaulting to NO. They answer different questions:
INTEGRATING |
BUSY |
|
|---|---|---|
| accumulating the block | 1 | 1 |
| tail computing | 0 | 1 |
| idle | 0 | 0 |
Every output of this block is a register, so each status bit is observed on the clock after the event that sets it:
INTEGRATINGrises on the clock after the FIRST sample of a block is accepted and falls on the clock after the N-th - it is high exactly while the block is ACCUMULATING.BUSYcovers the accumulation and the tail. It rises withINTEGRATING, stays high across the tail, and its LAST HIGH CLOCK IS THEOUT_DVPULSE; it falls on the clock after.- On a continuous stream the next block starts before the previous tail
ends, so
BUSYnever drops andINTEGRATINGdips for exactly one clock per block boundary - which makes it a free block marker. SAMPLE_COUNTis a fixed 32 bits and reads 1 after the first accepted sample, N after the N-th. It is NOT cleared at the block end: it HOLDS N through the tail and pastOUT_DV, until the first sample of the next block takes it back to 1. So on theOUT_DVclock it reads the length of the block being presented - which is exactly the divisor a consumer needs if it wants to turn this sum back into a mean. OnlyRESETclears it to 0.
Q formats
Both ports carry their own fixed point format (integer bits, fractional bits,
sign), the same convention as the Fixed P. family. The result is requantised
into the SUM format with the selected rounding (nearest / truncate) and
overflow policy (saturate / wrap); saturation is symmetric for signed formats,
as everywhere else in the toolchain.
A sum of signed samples is signed, and it reaches N times the amplitude of a sample - see “Sizing the SUM output” above, which is the part of this guide worth re-reading before you place the block.
Verification
The core is regression tested by a host-side csim harness
(tb/block-ops/run_tb.ps1) that runs one simulated clock at a time and
follows OUT_DV. The expected values come from tb/block-ops/gen_golden.py,
which evaluates $\sum x$ in exact rational arithmetic and shares no algorithm
with the core; the tolerance is 0. Sum coverage includes pseudo-random and
ramp inputs, a deliberately narrow saturating output format, the same
configuration with saturation off so it wraps, and an EXP that changes
half way through a block. The status outputs are checked clock by clock
against the contract above. A cross-check compiles this core and the all-in-one
block_stats.cpp into the same binary, drives them with identical stimulus,
and compares the two clock by clock.