DSP - BLOCK CORRELATION TM
TM (time multiplexed) twin of Block Correlation: TM Factor (2..16) independent channel PAIRS packed on TWO wide buses, each pair computing the Pearson correlation coefficient r = cov(a,b)/(sigma_a*sigma_b) - dimensionless, gain independent, bounded to [-1,+1] - over its own block of N = 2^EXP sample pairs (EXP on a runtime pin, shared by every lane). Lane k of IN_A pairs with lane k of IN_B, and ONE IN_DV qualifies both buses. One shared frame - one EXP, one IN_DV, one OUT_DV per block - with per-lane accumulators at II=1 and ONE serial tail serving the lanes one after the other, so the tail resources stay those of the scalar block whatever the lane count: this is THE LONGEST TAIL OF THE TM FAMILY, so budget a long block. Lane 0 sits in the LOW bits of every packed bus. Same per-lane numbers as the scalar twin, bit for bit. Optional shared BUSY / INTEGRATING / SAMPLE_COUNT status outputs.
Introduction
Per lane pair, over each block of N sample pairs:
$$ r_i = \frac{N\sum_k a_i b_i - \sum_k a_i \sum_k b_i}{\sqrt{\bigl(N\sum_k a_i^2 - (\sum_k a_i)^2\bigr)\bigl(N\sum_k b_i^2 - (\sum_k b_i)^2\bigr)}} $$
This is the time multiplexed (TM) twin of the scalar Block Correlation block: TM Factor independent channel pairs packed on two wide buses, each computing the Pearson correlation coefficient of its own block of N consecutive sample pairs. Nothing is shared between the channel pairs except the frame.
Per lane pair, r answers “do these two channels move together”, not “by how much” - which is exactly what Block Covariance TM cannot give you - and it is gain independent, so lane pairs at different gains compare directly.
The TM contract
IN_AandIN_Bare BOTH packed TM buses: TM Factor lanes of input width bits each, lane 0 in the LOW bits of BOTH buses (lane 0 is the oldest sample of the clock) - the same packing as every other TM block in the toolchain. Lane k ofIN_Apairs with lane k ofIN_B: each lane pair is an independent two-channel correlation with its own five accumulators (Sa, Sb, Saa, Sbb, Sab) and its own result.- ONE shared
EXPpin, ONEIN_DV, one frame: the singleIN_DVqualifies EVERY lane of BOTH buses at once, so the a/b pairing is positional per lane and the two streams cannot slip inside a block - in any lane.BUSY,INTEGRATINGandSAMPLE_COUNTtherefore stay scalar -SAMPLE_COUNTcounts per-lane sample pairs, which are identical in every lane by construction. - Accumulation runs at II=1 on the packed buses, one accumulator set per lane. TM Factor times the accumulator registers is the unavoidable cost of TM state - plus THREE real multipliers per lane (aa, bb and a*b) in the accumulation path.
- The serial post-processing is ONE shared engine set serving the lanes one after the other: four serial products, one digit-recurrence square root and one restoring division PER LANE, all out of ONE shift-add stage and ONE compare-subtract stage - not TM Factor of each. That is what keeps the tail resources of the SCALAR block, at the cost of tail latency x TM Factor.
- ONE
OUT_DVper block, after the LAST lane’s tail completes. All output lanes are staged as each lane finishes and committed together on theOUT_DVclock, so every packed output moves on that clock and no other.
N is a runtime input: the block size is $N = 2^{\mathrm{EXP}}$, EXP clamped to Max Block Exponent and latched on the first accepted sample of a block, so a change takes effect on the NEXT block - for every lane at once. Because N is a power of two, every N scaling in the identity is an exact shift - and both the input Q scaling and N cancel exactly in r.
When to use this instead of TM Factor scalar blocks
One TM block and TM Factor scalar blocks compute the same numbers. The TM block pays the per-lane accumulators (unavoidable either way) but shares ONE frame, ONE control FSM and ONE serial tail across all lanes - and this block has the most expensive scalar tail of the family, so it is where the TM sharing saves the most. The price is tail latency (TM Factor times the scalar tail - hundreds of clocks here) and the coupling of the frame: all lane pairs must share the same block length and the same sample cadence. Channel pairs that need different block sizes need scalar blocks.
Pin Description
IN_B; both
are sampled together on the clocks where IN_DV is high.
IN_A.
Lane k pairs with lane k of IN_A.
Properties
Number of INTEGER bits of the input sample (per lane, both streams) (the sign, when present, uses one of them).
Integer bits of ONE LANE of both inputs (the sign, when present, uses one of them). 1..64. Default 16. IN_A and IN_B share one lane format; r is gain independent, so the format only sets range, not scale.Default: 16
Options: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64
Number of FRACTIONAL bits of the input sample (per lane, both streams), i.e. the bits to the right of the binary point. Total width = integer + fractional bits, and must not exceed 64.
Fractional bits of one lane of both inputs. 0..64, total lane width 2..64 bits. Default 0. They cancel exactly in r.Default: 0
Options: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64
Select whether the input sample (per lane, both streams) is signed (two’s complement) or unsigned.
SIGNED (two’s complement) or UNSIGNED lanes. Default SIGNED. Applies to every lane of both buses.Default: SIGNED
Options: UNSIGNED SIGNED
Number of INDEPENDENT time-multiplexed channels packed on the IN bus and on every result bus. Lane 0 occupies the LOW bits (lane 0 = the oldest sample of the clock), the same packing as every other TM block. All lanes share one EXP / IN_DV / frame; each lane gets its own accumulators, but the serial post-processing is ONE engine serving the lanes one after the other, so the tail latency (and the minimum usable EXP) grows with this factor.
Number of independent channel pairs packed on the buses, 2..16. Default 4. Multiplies BOTH input widths, the packed output width AND the serial tail length (the shared tail serves the lanes one after the other), so it also raises the minimum usable EXP - on this block the tail is the family’s longest, so raise it with care.Default: 4
Options: 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
Largest block-size exponent the accumulators are sized for: the block can be up to 2^MaxBlockExponent samples long. The EXP input is clamped to this value at run time. Raising it widens the internal accumulators, and ON THE BLOCKS WHOSE SERIAL ENGINES ARE SIZED FROM THOSE ACCUMULATORS (Coefficient of Variation, SNR, Skewness, Kurtosis, Correlation, Autocorrelation, Linear Regression) it also LENGTHENS THE SERIAL TAIL – even when the runtime EXP is small. Keep it at the largest block you actually use. The default of 20 covers blocks of up to 1048576 samples.
Largest block-size exponent the per-lane accumulators are sized for; the EXP input is clamped to it at run time. Raising it widens every lane’s accumulators AND lengthens the shared products, root and divider - the per-lane tail grows by roughly nine clocks per unit. 1..31, default 20.Default: 20
Options: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31
Number of INTEGER bits of the CORRELATION output (per lane) (the sign, when present, uses one of them).
Integer bits of ONE LANE of the CORRELATION output. 1..64, default 2 - r never exceeds 1 in magnitude, so 2 integer bits (sign included) are enough.Default: 2
Options: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64
Number of FRACTIONAL bits of the CORRELATION output (per lane), i.e. the bits to the right of the binary point. Total width = integer + fractional bits, and must not exceed 64.
Fractional bits of one lane. 0..64, default 14 - r is a pure fraction, so this is where it lives. Each one also lengthens the shared divider by one clock per lane.Default: 14
Options: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64
Select whether the CORRELATION output (per lane) is signed (two’s complement) or unsigned.
SIGNED or UNSIGNED lanes. Default SIGNED - anticorrelation is negative and is the interesting case.Default: SIGNED
Options: UNSIGNED SIGNED
YES: the BUSY (high from the first sample of a block until its result is out – it COVERS THE SERIAL TAIL, and its last high clock IS the OUT_DV pulse) pin is present. NO: the pin AND all of its logic are removed BEFORE synthesis, so nothing is paid for it.
YES: the shared BUSY pin exists. Default NO.Default: NO
Options: NO YES
YES: the INTEGRATING (high only while the block is ACCUMULATING; it drops as soon as the N-th sample has been taken and the tail starts, so BUSY-and-not-INTEGRATING means ‘computing’) pin is present. NO: the pin AND all of its logic are removed BEFORE synthesis, so nothing is paid for it.
YES: the shared INTEGRATING pin exists. Default NO.Default: NO
Options: NO YES
YES: the SAMPLE_COUNT (32 bit, how many samples have been accumulated so far in the current block: 1 after the first, N after the N-th. It is NOT cleared at the block end – it holds N until the NEXT block’s first accepted sample takes it back to 1. On a CONTINUOUS stream that happens DURING the serial tail, so at OUT_DV it reads how far into the next block the input has already got, NOT N. To capture the length of the block being presented, latch SAMPLE_COUNT on the clock INTEGRATING falls – that one always reads N) pin is present. NO: the pin AND all of its logic are removed BEFORE synthesis, so nothing is paid for it.
YES: the shared 32 bit SAMPLE_COUNT pin exists. Default NO.Default: NO
Options: NO YES
ROUND: round to nearest when a result has to be requantised into a coarser output format. TRUNCATE: drop the bits (cheaper, adds a negative bias).
ROUND / TRUNCATE at each lane’s final requantisation. Effectively a NO-OP on this block: the numerator is pre-shifted so each lane’s quotient already carries exactly the output resolution and nothing is left to shift. Saturation still applies. Default ROUND.Default: ROUND
Options: TRUNCATE ROUND
YES: clip to the largest representable value of each output format (symmetric for signed formats). NO: wrap around.
YES: clip each lane to its output format (symmetric for signed). NO: wrap. |r| <= 1 by arithmetic, so a wide-enough lane format never engages it. Default YES.Default: YES
Options: NO YES
Accuracy, per lane
Each lane carries the scalar twin’s bound: everything up to the division is exact integer arithmetic, the denominator root carries 4 guard bits which the numerator’s pre-shift cancels exactly, and the quotient is floored, so each lane’s r is within 1 LSB - and each lane is bit-identical to the scalar twin run on the same lane pair. As in the scalar twin, per lane: |r| <= 1 is guaranteed by the arithmetic (Cauchy-Schwarz on exact integers), so the default Q2.14 lane format never saturates on a legitimate value and +1.0 lands exactly on 16384; a constant channel in a lane pair makes that lane report 0, reached by the arithmetic itself, not by a special case.
Because each lane’s quotient already comes out at the output resolution, the final requantisation has nothing left to shift: the Rounding property is effectively a NO-OP on this block (saturation still applies).
Timing: the TM latency contract
OUT_DV pulses ONCE per block, L clocks after the clock on which the
N-th sample was accepted, where
$$ L = 1 + \mathrm{TM},\bigl(3(\mathrm{IN_SW} + \mathrm{EXP}) + \mathrm{VNW} + \mathrm{RTW} + \mathrm{CNUMW} + 4\bigr) $$
(BCOT_TAIL in the core). Per lane: the three variance/covariance
numerator products, the radicand product, the root and the division, all
out of the shared engines, lane by lane. The tail DEPENDS ON THE RUNTIME
EXPONENT, like the scalar twin. This is the family’s lane-multiplexed rule
$L_{tm}(e) = \mathrm{TM},(L_{scalar}(e)-1)+1$ - applied to the longest
scalar tail in the family. At the defaults (16 bit signed inputs, Q2.14
CORRELATION, Max Block Exponent 20, TM 4, EXP 10) VNW = 72, RTW = 76 and
CNUMW = 90, so L = 1 + 4*320 = 1281 clocks - which does NOT fit a 2^10
block: at those defaults the minimum usable EXP is 11. Budget a long
block, or trim the input width / Max Block Exponent.
The family drop rule applies with the TM tail: the tail of one block must
finish before the NEXT block completes, $2^{{\mathrm{{EXP}}}} \ge L$, or the
completing block’s result is silently DROPPED (no OUT_DV, accumulators
unaffected, no error pin). The TM tail makes the minimum usable EXP larger
than the scalar twin’s - the compiler prints both the worst-case tail and
the minimum EXP in the compilation log, and the property window refuses a
configuration whose minimum exceeds Max Block Exponent.
Knowing where the block is: BUSY, INTEGRATING and SAMPLE_COUNT
Identical to the scalar family, and SHARED by all lanes: INTEGRATING is
high exactly while the block is accumulating (it dips one clock per block
boundary on a continuous stream), BUSY also covers the (TM-long) tail and
its last high clock IS the OUT_DV pulse, SAMPLE_COUNT reads 1 after the
first accepted sample pair and N after the N-th.
SAMPLE_COUNTis NOT cleared at the block end - it holds N until the NEXT block’s first accepted sample takes it back to 1. On a CONTINUOUS stream that happens DURING the serial tail (which on a TM block is TM Factor times longer), so atOUT_DVit reads how far into the next block the input has got, NOT N. The clock that always reads N is the oneINTEGRATINGfalls on - latch it there.
Verification
The core is regression tested by a host-side csim harness
(tb/block-ops-tm/run_tb_tm.ps1) with per-lane goldens computed by
gen_golden_tm.py in exact rational arithmetic ON EACH LANE PAIR’S STREAMS
ALONE (lanes deliberately carry different signals - the generator refuses
identical lanes), plus the strongest available lane-independence check:
after every run, the SCALAR twin is replayed on each lane pair’s streams by
themselves and lane k of every TM result must match it BIT FOR BIT. The
status waveform is checked clock by clock, every packed output is checked
to move only on OUT_DV, and the TM-specific mutant classes (shared
accumulator, lane swaps on either bus, wrong-lane tail reads, early commit,
drop rule) are killed. What no host harness can prove - that Vitis accepts
and schedules the core at II=1, with four serial products, a root and a
division closed through one stage each - is stated in AGENT/block_ops.log,
not silently implied. Note also that wide configurations (the packed bus
above 127 bits, or IN_SW + Max Block Exponent large enough to push the
radicand past the host stub’s ceiling) are exercised only by csynth, never
by the host harness.