Generated reference › Scaled Dot Product Attention — Machine Learning/Neural Networks
kind: generated#block#machine-learning-neural-networks

Scaled Dot Product Attention — Machine Learning/Neural Networks

Machine_Learning/Neural_Networks/Scaled_Dot_Product_Attention · 1 input / 1 output port(s) at insert · exports to Python, MATLAB, Java, Rust, C, C++, VHDL, Verilog, SystemVerilog, PLC Structured Text

Description#

The block's own DESCRIPTION_HTML, rendered verbatim — the same text the config dialog's info panel and the library navigator show. Fix a wrong sentence in the block's .cpp (R-D9), never here.

Scaled Dot-Product Attention

Machine Learning / Neural Networks

H attention heads over a window of T positions:

Q = X·Wq, K = X·Wk, V = X·Wv, A = softmax(Q·K' ÷ √dk) row-wise, Y = A·V – each head over its own slice of the three projections, and the heads' outputs concatenated across columns.

Each output row is a weighted average of the value vectors, with the weights decided by how well that position's query matches every position's key. It is the mechanism a transformer is built from, and the reason a model can relate two positions directly rather than passing information along a chain of recurrences.

The window arrives as a matrix rather than being buffered inside the block, which is what lets Positional Encoding feed straight into it – attention alone is order-blind, so that block exists to mark position, and the two are meant to be used together.

Ports

  • X – the window, [T, dmodel]: one row per position, one column per feature. T and dmodel are read from the signal, not configured.
  • Y – the attended window, [T, H·dv] – the heads side by side, head n writing columns [n·dv, (n+1)·dv). That is all of Wv's columns, so the output width does not change with the head count: splitting a projection into more heads changes how the window is mixed, not its shape.

Parameters

  • Query Weights – Wq, [dmodel, H·dk]: the heads' query projections side by side, head n taking the contiguous column range [n·dk, (n+1)·dk). That is how PyTorch's nn.MultiheadAttention packs in_proj_weight, so a trained set slices in unchanged. Its width must divide by Heads.
  • Key Weights – Wk, the same shape as Wq and sliced the same way: queries and keys are compared by dot product, so they must live in the same space. A mismatch is reported rather than broadcast.
  • Value Weights – Wv, [dmodel, H·dv], sliced the same way. Its full width is the output width and is free of the other two; it too must divide by Heads.
  • Heads – H, how many attention heads share the three projections; a whole number ≥ 1, and 1 by default. ⚠ dk is per head, so the scale 1 ÷ √dk moves when you change this even though no matrix changed shape: two heads over a [dmodel, 4] Wq have dk = 2, not 4. Each head attends independently – its own scores, its own softmax – which is the whole point: one head can only express one weighting of the window per position, and several let different heads track different relations.
  • Masking – whether a position may attend to later ones.
    • None – every position attends to the whole window. The encoder-style reading, and the right one when the window is a complete snapshot.
    • Causal – position i attends only to j ≤ i, so no output row depends on a position later than itself. The decoder-style reading, and the one to use when the window is a sliding history and a later row must not leak into an earlier one.
  • Sampling Time (s) – zero or less inherits the solver's rate; a positive value runs the block at that period.

Code export

All ten targets: Python, MATLAB, Java, Rust, C, C++, VHDL, Verilog, SystemVerilog and PLC Structured Text. Every projection, score and weight is unrolled: T, dmodel, H, dk and dv are known at export time, so no backend contains a loop, an index type or a mask array – and no head index either: each head is emitted as its own straight-line block of statements. The scale 1 ÷ √dk is folded into the score's coefficients once, and under Causal the masked terms are simply not emitted – there is no −∞ and no comparison anywhere.

⚠ The emitted body grows as H·T²·dk terms. That is inherent to unrolling attention, and it is the cost of having no loops: a window of 4 is a few hundred lines per backend, a window of 32 is tens of thousands. Splitting a fixed projection into more heads is close to free (H·dk is unchanged); raising T is not. Keep T to what the model actually needs.

The three HDL targets are simulation-only real arithmetic – the softmax needs exp and a division, neither of which belongs in a Q16.16 datapath – quantizing only at the port boundary, exactly as Dense Layer does for its Tanh and Sigmoid activations.

Simulink bridge

None. Simulink's attention layers live inside its Deep Learning blocks, which take a trained network object rather than three weight matrices, and no config value can carry an object across the bridge. The bridge reports the block rather than dropping it silently, and it has no parity testbench, which is the documented consequence of Support::None rather than a gap.

Notes

  • Algebraic and stateless: the whole window comes from the current input, so nothing is carried between steps. Buffering the window is the upstream block's job.
  • The row maximum is subtracted before exponentiating, inherited from Softmax for the same reason: without it a logit a user will really produce overflows in single precision and in the HDL paths. It cancels exactly in the ratio, so it changes the arithmetic's range and not its value.
  • Multi-head attention is the Heads parameter, not a second block and not a wiring pattern: the heads share one input window and one set of matrices, so carrying them here costs one integer and keeps the three projections in the shape a trained model exports them in. What this block does not carry is the output projection Wo that a transformer applies to the concatenated heads – that is a plain [H·dv, dmodel] matrix multiply, which Dense Layer already is.

Code facts#

FactValue
registered typeMachine_Learning/Neural_Networks/Scaled_Dot_Product_Attention
familyMachine_Learning/Neural_Networks
solver environment classICoreBlock_0_Machine_Learning_1_Neural_Networks_2_Scaled_Dot_Product_Attention
sourcesrc/ICoreBlocks/ICoreBlockLibrary/Blocks/Machine_Learning/Neural_Networks/Scaled_Dot_Product_Attention/ICoreBlock_0_Machine_Learning_1_Neural_Networks_2_Scaled_Dot_Product_Attention.cpp
headersrc/ICoreBlocks/ICoreBlockLibrary/Blocks/Machine_Learning/Neural_Networks/Scaled_Dot_Product_Attention/ICoreBlock_0_Machine_Learning_1_Neural_Networks_2_Scaled_Dot_Product_Attention.h
default size on canvas150 × 90 px
ports at insert1 in, 1 out
code generators implementedPython, MATLAB, Java, Rust, C, C++, VHDL, Verilog, SystemVerilog, PLC Structured Text

Ports#

#DirectionSignal typeDescription label
1inICoreDoubleX
2outICoreDoubleY

Ports the constructor creates. A block whose port list changes with its configuration adds or removes ports at load time; the count above is the one a freshly inserted block has.

Configuration variables#

Config variableDefaultSimulink parameter
Query Weights[0.6 -0.2; 0.1 0.5; -0.4 0.3]—
Key Weights[0.3 0.45; -0.5 0.2; 0.25 -0.35]—
Value Weights[0.7 0.15; 0.2 -0.6; -0.1 0.4]—
MaskingNone%~%Causal~~None—
Heads1—

Every block also carries Sampling Time (s) from ICoreBlockSolverEnvironment: zero or less inherits the solver's rate, a positive value runs the block at that period.

supportSupport::None
Simulink path—
port-count rulePortsParam::None
SampleTime parameteryes

Caveat (shown to the user): no Simulink equivalent as a signal block: its attention layers exist only inside the Deep Learning blocks, which take a trained network OBJECT rather than the three weight matrices this block carries as config, and no ParamRule can move an object across the bridge

Catalog contract: src/ICoreBlocks/ICoreCoder/ICoreCommandSystem/SimulinkBridge/ICoreSimulinkBlockCatalog.h

Description vs code#

The checker has a blind spot here — it could not resolve something (a grouped port bullet, a computed config name), which is reported and never counted as a pass. A reader has to settle it:

  • B0 every stimulus in the sample errored — cross-checks skipped

The verdict above is tools/docs/check_block_descriptions.py (P7.1), which compares LISTS. It cannot read a sentence: "stateless" on a block with a state, an initial-value semantic the recursion does not implement, a "not synthesizable" caveat the HDL banner contradicts. That is the agent audit (P7.3) on BLOCK_DESCRIPTION_AUDIT.md, and this tool's green is not a substitute for one.

File banner (developer view)#

The top comment of the block's .cpp — the maths, the realization and the export strategy, addressed to whoever changes it. It must not contradict the description above (P7.5).

Scaled Dot-Product Attention — H heads over a [T, d_model] window per head n, over the COLUMN RANGE [n*d_k, (n+1)*d_k) of Wq/Wk and [n*d_v, ...) of Wv: Q = X*Wq K = X*Wk V = X*Wv S = Q*K' / sqrt(d_k) d_k is PER HEAD, so the scale moves with Heads A = softmax(S) row-wise (the row max is subtracted -- Softmax's rule) Y = [ A_0*V_0 | ... | A_{H-1}*V_{H-1} ] [T, H*d_v], the heads concatenated

Heads closes Multi_Head_Attention's board row and changes no config's shape: the head projections sit side by side in the matrices that were always there, PyTorch's in_proj_weight packing. At H = 1 every emitted body is byte-for-byte what it was.

Stateless: the window arrives as a matrix rather than being buffered here, so Positional_Encoding -- which exists precisely so this block is not order-blind -- feeds straight in. Everything is unrolled at export time, and the causal mask is expressed by NOT EMITTING the masked terms rather than by a runtime comparison.

The three HDL targets are simulation-only real: this needs exp() and a division.

Sample results#

No stimulus produced a sampled output in this rig — Invalid "Query Weights" at: ICore Blocks/Home/Scaled Dot Product Attention. That is a fact about the single-block rig, not a verdict on the block: an offline batch fit, a block whose output only appears at onSolverFinish, or one that needs a driven environment cannot be exercised alone.

Category unsampled · sample time 0.1 · 60 steps · commit c01902987 · produced by docsSample --out <folder> --steps 60

Sample data: docs/generated/samples/Machine_Learning__Neural_Networks__Scaled_Dot_Product_Attention.json