Gradient Boosted Trees — Machine Learning/Classical Models
Machine_Learning/Classical_Models/Gradient_Boosted_Trees · 1 input / 1 output port(s) at insert · exports to Python, MATLAB, Java, Rust, C, C++, VHDL, Verilog, SystemVerilog, PLC Structured Text
Description#
The block's own DESCRIPTION_HTML, rendered verbatim — the same text the config dialog's info panel and the library navigator show. Fix a wrong sentence in the block's .cpp (R-D9), never here.
Gradient Boosted Trees
Machine Learning / Classical Models
Evaluates an additive ensemble of regression trees:
score = base + rate · Σt wt · leaft(u), then y = score, y = 1/(1 + e−score) or y = ±1 on its sign.
This is the inference half of XGBoost, LightGBM and sklearn's GradientBoosting – all three dump a model that reduces to exactly this. Boosting adds where a random forest averages, and that one difference is the only thing separating this block from Random Forest: they share the stacked node table, the leaf marker and the walker.
It is also AdaBoost. Give each tree its own weight in Tree Weights and take the verdict off Sign, and the sum above is AdaBoost's weighted vote H(u) = sign(Σt αt ht(u)) written out – the same stacked node table, the same walker, one more column of numbers.
Ports
- u – the feature column [d,1]. d is read off the largest Feature Index the tree table actually uses, so it needs no setting of its own.
- y – a scalar [1,1]: the prediction, after whatever Output selects.
Parameters
- Feature Index, Threshold, Left Child, Right
Child, Leaf Value – the node table, one row per node,
stacked across every tree, exactly as Decision Tree and Random
Forest take it. A sample goes left when
u[Feature Index] ≤ Threshold, which is scikit-learn's sense. A negative Left Child marks a leaf (TREE_LEAF), and at a leaf the feature and threshold are ignored, sotree_pastes in unchanged. - Tree Roots – one row per tree, giving that tree's first node. Its length is the number of trees.
- Base Score – the ensemble's starting prediction, before any
tree. XGBoost calls it
base_score; sklearn's initial estimator contributes it. - Learning Rate – the shrinkage applied to the tree sum. ⚠ Which
exporter you came from decides this, and nothing in the numbers can tell you:
- XGBoost and LightGBM bake the shrinkage into the leaf values when they dump a model, so leave this at 1.
- sklearn's GradientBoosting keeps raw leaves and applies
learning_rateat predict time – paste that value here.
- Tree Weights – the weight each tree votes with, wt
above. Either a single row, which gives every tree that same weight, or a
column of exactly one row per tree, in Tree Roots' order. The
default is 1: plain boosting, every tree counting the same. Paste
AdaBoost's
estimator_weights_here to get its weighted vote. Anything else – a column that is neither 1 nor T rows – is refused rather than padded. - Output – the link applied to the score:
- Raw (regression) – the score itself. The default.
- Logistic (binary probability) – 1/(1 + e−s), which is what a binary classifier's margin means.
- Sign (weighted vote) – +1 when the score is
strictly positive and −1 otherwise, which is AdaBoost's
verdict. A score of exactly zero gives −1, the same side
scikit-learn's
classes_[(decision > 0)]sends it.
- Sampling Time (s) – zero or less inherits the solver's rate; a positive value runs the block at that period.
Code export
All ten targets: Python, MATLAB, Java, Rust, C, C++, VHDL, Verilog, SystemVerilog and PLC Structured Text. The whole ensemble is baked into the body: the traversal is emitted as a nested if cascade rather than a loop over the table, so each node appears exactly once, no index arithmetic exists anywhere, and the emitted code is O(nodes) rather than O(nodes×depth).
The Output settings differ in hardware, and it is worth knowing which you
are exporting. In Raw the block is comparisons, adds and one multiply
– exact in every target and genuine synthesizable Q16.16 on the
three HDLs. Sign adds one comparison and so stays there too.
Logistic adds an exponential, so those three targets fall back to
simulation-only real arithmetic for that setting alone,
quantizing only at the port boundary.
Tree Weights costs nothing anywhere. Both the weight and the leaf value are known when the code is written, so their product is folded into the one constant the cascade already carried: no target multiplies per tree, the three HDLs quantize once instead of twice, and at the default weight of 1 the emitted body is exactly what it was without the setting.
Simulink bridge
None. Boosted-tree inference belongs to the Statistics and Machine
Learning Toolbox, which is not installed here, and its interface takes a fitted
model object rather than parameters – which no parameter rule could carry
even if it were present. The bridge reports the block rather than dropping it
silently, and it has no parity testbench, which is the documented
consequence of Support::None rather than a gap.
Notes
- Algebraic and stateless: the prediction depends only on the current sample.
- Nonlinear and discontinuous, and deliberately carries no state space: the output steps at every split, so no A/B/C/D describes it and model reduction correctly refuses the block.
- Children must have a higher row number than their parent, which every fitted exporter satisfies. It is checked, because a backward-pointing table would send the code emitter into unbounded recursion – a stack overflow at export time, far from the configuration that caused it.
- The rate multiplies the SUM, not each leaf – the same number in exact arithmetic, and one multiply instead of one per tree. The tree weight cannot be treated that way and is not: it is per tree, so it is folded into that tree's leaves instead.
- Sign is the one setting a fixed-point export can disagree with the live run about. The score is a sum of baked constants, so it takes only as many values as the ensemble has leaf combinations; place them clear of zero and every target agrees exactly, put one within Q16.16's 1.5×10−5 of zero and the three HDL targets may call it the other way.
Code facts#
| Fact | Value |
|---|---|
| registered type | Machine_Learning/Classical_Models/Gradient_Boosted_Trees |
| family | Machine_Learning/Classical_Models |
| solver environment class | ICoreBlock_0_Machine_Learning_1_Classical_Models_2_Gradient_Boosted_Trees |
| source | src/ICoreBlocks/ICoreBlockLibrary/Blocks/Machine_Learning/Classical_Models/Gradient_Boosted_Trees/ICoreBlock_0_Machine_Learning_1_Classical_Models_2_Gradient_Boosted_Trees.cpp |
| header | src/ICoreBlocks/ICoreBlockLibrary/Blocks/Machine_Learning/Classical_Models/Gradient_Boosted_Trees/ICoreBlock_0_Machine_Learning_1_Classical_Models_2_Gradient_Boosted_Trees.h |
| default size on canvas | 150 × 80 px |
| ports at insert | 1 in, 1 out |
| code generators implemented | Python, MATLAB, Java, Rust, C, C++, VHDL, Verilog, SystemVerilog, PLC Structured Text |
Ports#
| # | Direction | Signal type | Description label |
|---|---|---|---|
| 1 | in | ICoreDouble | u |
| 2 | out | ICoreDouble | y |
Ports the constructor creates. A block whose port list changes with its configuration adds or removes ports at load time; the count above is the one a freshly inserted block has.
Configuration variables#
| Config variable | Default | Simulink parameter |
|---|---|---|
Feature Index | [0; -2; -2; 1; -2; -2] | — |
Threshold | [0; 0; 0; 0.2; 0; 0] | — |
Left Child | [1; -1; -1; 4; -1; -1] | — |
Right Child | [2; -1; -1; 5; -1; -1] | — |
Leaf Value | [0; 0.5; -0.4; 0; -0.3; 0.6] | — |
Tree Roots | [0; 3] | — |
Base Score | 0 | — |
Learning Rate | 1 | — |
Tree Weights | 1 | — |
Output | Raw (regression)%~%Logistic (binary probability)%~%Sign (… | — |
Every block also carries Sampling Time (s) from ICoreBlockSolverEnvironment: zero or less inherits the solver's rate, a positive value runs the block at that period.
Simulink bridge#
| support | Support::None |
| Simulink path | — |
| port-count rule | PortsParam::None |
SampleTime parameter | yes |
Caveat (shown to the user): no Simulink equivalent available here: boosted-tree inference belongs to the Statistics and Machine Learning Toolbox, which is not installed on this machine, and its interface takes a fitted MODEL OBJECT rather than parameters -- which no ParamRule could carry even if it were present. There is therefore no parameter set to map onto and no reference to run a parity testbench against
Catalog contract: src/ICoreBlocks/ICoreCoder/ICoreCommandSystem/SimulinkBridge/ICoreSimulinkBlockCatalog.h
Description vs code#
The checker has a blind spot here — it could not resolve something (a grouped port bullet, a computed config name), which is reported and never counted as a pass. A reader has to settle it:
B0every stimulus in the sample errored — cross-checks skipped
The verdict above is
tools/docs/check_block_descriptions.py(P7.1), which compares LISTS. It cannot read a sentence: "stateless" on a block with a state, an initial-value semantic the recursion does not implement, a "not synthesizable" caveat the HDL banner contradicts. That is the agent audit (P7.3) on BLOCK_DESCRIPTION_AUDIT.md, and this tool's green is not a substitute for one.
File banner (developer view)#
The top comment of the block's .cpp — the maths, the realization and the export strategy, addressed to whoever changes it. It must not contradict the description above (P7.5).
Gradient Boosted Trees — an ADDITIVE ensemble over Decision_Tree's stacked table score = base + rate * SUM_t w_t * leaf_t(u) y = score, or 1/(1 + exp(-score)), or +1/-1 on its sign
Random_Forest's block with one line changed: the accumulator adds instead of averaging, so there is no 1/T to fold. Everything else -- the stacked node columns, the leaf marker, the nested-if cascade, the left/right sense -- is Decision_Tree's, deliberately, so the three blocks cannot disagree about which way a sample falls at a split.
⚠
Tree Weightsand theSignlink are what make this AdaBoost as well: a per-tree alpha_t and the verdict taken off the weighted vote's sign. Neither costs anything at run time -- w_t * leaf is FOLDED at export time into the constant the cascade already carried, and the sign is one comparison, so Sign stays genuine Q16.16 on the three HDL targets where Logistic does not.
Sample results#
No stimulus produced a sampled output in this rig — Invalid input size at: ICore Blocks/Home/Gradient Boosted Trees. That is a fact about the single-block rig, not a verdict on the block: an offline batch fit, a block whose output only appears at onSolverFinish, or one that needs a driven environment cannot be exercised alone.
Category unsampled · sample time 0.1 · 60 steps · commit c01902987 · produced by docsSample --out <folder> --steps 60
Sample data: docs/generated/samples/Machine_Learning__Classical_Models__Gradient_Boosted_Trees.json