Label Encoder — Machine Learning/Preprocessing
Machine_Learning/Preprocessing/Label_Encoder · 1 input / 1 output port(s) at insert · exports to Python, MATLAB, Java, Rust, C, C++
Description#
The block's own DESCRIPTION_HTML, rendered verbatim — the same text the config dialog's info panel and the library navigator show. Fix a wrong sentence in the block's .cpp (R-D9), never here.
Label Encoder
Machine Learning / Preprocessing
Turns a text label into the class code a model expects: y = the position of u in the label list, counting from Index Base. Text that is in the list nowhere comes out as Unknown Code. This is the block that lets a name – a mode, a state, a category read off a bus – drive arithmetic.
Ports
- Input – the label u (
str,ICoreString), one value. Only text: a number arriving here has no label to be, and the one way to change a signal's type is a Data Type Conversion in the wire. - Output – the code y (
i32,ICoreInt32), one value. A code is a whole number and it is frequently negative –-1is the usual Unknown Code – so it is a signed integer.
Parameters
- Labels – the fitted label list, separated by semicolons:
Idle;Run;Fault. Order is what defines the codes, so it must be the order the model was fitted with. Surrounding spaces are kept, because a label may legitimately have one; a label may not itself contain a semicolon. An empty list is allowed and means every input is unknown. - Unknown Code – the code for text that matches no label. Whole number, and it is not shifted by Index Base: it names something that is not a class.
- Index Base – what the first label is called:
- Zero-based (PyTorch, numpy) – the first label is code 0.
- One-based (MATLAB) – the first label is code 1.
- Sampling Time (s) – zero or less inherits the solver's rate; a positive value runs the block at that period.
Matching is exact, byte for byte
A label matches only text that is identical to it – same letters, same case, same spaces. There is deliberately no case-insensitive setting: the six languages this block exports to disagree about what lowercasing means the moment the text stops being plain English (one of them makes the text longer), and a rule that held in five targets and not the sixth would be worse than no rule. To accept two spellings, list them both; both codes can then be treated as one downstream.
If the same label is listed twice, the first position wins – the lowest code. That is the tie rule Argmax Decision and One-Hot Encoder already use, so a chain of the three never has to be read twice.
Code export
C, C++, Python, MATLAB, Java and Rust all carry this block, and all six answer the same code as the simulation for every input. The label list is baked into the generated code at export time rather than being offered as a tunable parameter: it is what the comparison IS, not a number to retune afterwards.
VHDL, Verilog, SystemVerilog and PLC Structured Text do not carry a text signal, so an export to one of them stops and names this block and the reason rather than emitting something that does not run.
Simulink bridge
None – Support::None. Simulink's String library compares,
searches and converts text, and none of it maps a fitted label set onto a class
code; its Statistics blocks take a fitted model object rather than a parameter
set, which no bridge parameter can carry. The block is reported as unsupported
rather than silently dropped.
Notes
- Algebraic, with no state: the output depends only on the current input.
- An unconnected input is empty text. That is unknown unless the empty string
is itself one of the labels, which is allowed –
Idle;;Faultmakes the empty string code 1. - Text is carried and compared to 256 bytes, the limit every text signal in this library is cut at, so a label longer than that can never be matched.
Code facts#
| Fact | Value |
|---|---|
| registered type | Machine_Learning/Preprocessing/Label_Encoder |
| family | Machine_Learning/Preprocessing |
| solver environment class | ICoreBlock_0_Machine_Learning_1_Preprocessing_2_Label_Encoder |
| source | src/ICoreBlocks/ICoreBlockLibrary/Blocks/Machine_Learning/Preprocessing/Label_Encoder/ICoreBlock_0_Machine_Learning_1_Preprocessing_2_Label_Encoder.cpp |
| header | src/ICoreBlocks/ICoreBlockLibrary/Blocks/Machine_Learning/Preprocessing/Label_Encoder/ICoreBlock_0_Machine_Learning_1_Preprocessing_2_Label_Encoder.h |
| default size on canvas | 100 × 60 px |
| ports at insert | 1 in, 1 out |
| code generators implemented | Python, MATLAB, Java, Rust, C, C++ |
Ports#
| # | Direction | Signal type | Description label |
|---|---|---|---|
| 1 | in | ICoreString | — |
| 2 | out | ICoreInt32 | — |
Ports the constructor creates. A block whose port list changes with its configuration adds or removes ports at load time; the count above is the one a freshly inserted block has.
Configuration variables#
| Config variable | Default | Simulink parameter |
|---|---|---|
Labels | Idle;Run;Fault | — |
Unknown Code | -1 | — |
Index Base | Zero-based (PyTorch, numpy)%~%One-based (MATLAB)~~Zero-ba… | — |
Every block also carries Sampling Time (s) from ICoreBlockSolverEnvironment: zero or less inherits the solver's rate, a positive value runs the block at that period.
Simulink bridge#
| support | Support::None |
| Simulink path | — |
| port-count rule | PortsParam::None |
SampleTime parameter | yes |
Caveat (shown to the user): no Simulink equivalent: its String library compares, searches and converts text but maps no fitted label set onto a class code, and the Statistics blocks take a fitted model OBJECT rather than a parameter set. There is no parameter set this block could map onto
Catalog contract: src/ICoreBlocks/ICoreCoder/ICoreCommandSystem/SimulinkBridge/ICoreSimulinkBlockCatalog.h
Description vs code#
The lists agree. check_block_descriptions.py finds no disagreement between the description's Ports, Parameters, Code export and Simulink bridge lists and the code's.
The verdict above is
tools/docs/check_block_descriptions.py(P7.1), which compares LISTS. It cannot read a sentence: "stateless" on a block with a state, an initial-value semantic the recursion does not implement, a "not synthesizable" caveat the HDL banner contradicts. That is the agent audit (P7.3) on BLOCK_DESCRIPTION_AUDIT.md, and this tool's green is not a substitute for one.
File banner (developer view)#
The top comment of the block's .cpp — the maths, the realization and the export strategy, addressed to whoever changes it. It must not contradict the description above (P7.5).
Label Encoder -- text in, a class code out y = the position of u in a configured label list, or the configured Unknown Code when u is in no position at all. sklearn's LabelEncoder.transform, as a block.
This is the first block in Machine_Learning whose INPUT is text, and it is the reason its board row stopped being a fold. While every port in the tree was an ICoreDouble a label was a number wearing a name, and Lookup_Tables/Direct_Lookup_Table_nD really did express the whole of it. A table is indexed BY a number; it cannot be indexed by "Fault".
Read the header first: the byte-equality rule, the first-match tie rule, why the unknown code is not shifted by Index Base, and why four of the ten targets are absent, are all decided there.
Sample results#
This block carries a ICoreString signal, whose value is text rather than a number and is not something a plot has an axis for. The samples are in the table below, exactly as the run recorded them.
| t | in ICoreString-Out-0 | out ICoreInt32-Out-0 |
|---|---|---|
| 0 | u0 | -1 |
| 0.04 | u0 | -1 |
| 0.08 | u0 | -1 |
| 0.12 | u0 | -1 |
| 0.16 | u0 | -1 |
| 0.2 | u0 | -1 |
| 0.24 | u0 | -1 |
| 0.28 | u0 | -1 |
| 0.32 | u0 | -1 |
| 0.36 | u0 | -1 |
| 0.4 | u0 | -1 |
| 0.44 | u0 | -1 |
| 0.48 | u0 | -1 |
| 0.52 | u0 | -1 |
Every 4th of 60 samples, from the table stimulus.
The same rig also ran:
| Stimulus | What it is | Output range |
|---|---|---|
impulse | Impulse: one sample of 1 at k = 5, 0 elsewhere (Repeating Sequence Stair) | -1 … -1 |
ramp | Ramp: slope 1 from t = 0 | -1 … -1 |
sine | Sine Wave: amplitude 1, 2 rad/s, no phase, no bias | -1 … -1 |
step | Step: 0 -> 1 at t = 1 s | -1 … -1 |
Plotted: table — Repeating Sequence Stair: [-2 -1 -0.5 0 0.5 1 2 3], one entry per sample
Category static · sample time 0.01 · 60 steps · commit 66ab17e58 · produced by docsSample --out <folder> --blocks Label_Encoder --steps 60 · data docs/generated/samples/Machine_Learning__Preprocessing__Label_Encoder.json · the SVG is generated from those numbers by tools/docs/plot_svg.py, so it is a run and not a drawing (R-D10).