Raw Byte Code Elimination
Raw Byte Code Elimination Plan
Status: IN PROGRESS. All thirteen gated images are byte-identical to their physical
dumps, but that is a build-match guarantee, not a disassembly guarantee: a .byte run that spells real
code reassembles to the same bytes as an .incbin of those bytes, and passes the gate
identically either way. Measured code-as-.byte remains in the Main CPU and HDAE5000 ROMs;
the Sub CPU Payload reached zero.
Goal: Convert all executable code currently represented as raw .byte sequences to native TLCS-900 assembly mnemonics.
Context
All thirteen gated images rebuild byte-identically from source. That is necessary for
correct disassembly but not sufficient: a .byte run that spells real code, or a data
region disassembled into plausible-but-wrong instruction mnemonics, both reassemble to the
same bytes and pass the gate cleanly. Whether a .byte run is undecoded code can only be
settled by a per-byte disassembly attempt (Step 1’s method, below), not by counting
.incbin directives — the table under Current Status states what that measurement
currently finds, ROM by ROM.
The sub-CPU payload is now at zero code-as-.byte. Its last two blocked families are
both closed: DSP_Bytecode_Op01/02/03 (569 B) needed a decoder fix before the encoder’s own
output could round-trip (see LLVM Semantic Instructions), and the TaskEvent/FIFO/TaskSched family’s apparent
~407 B gap turned out to be a measurement bug in the round-trip prober, not a decoder
limitation — of the 672 B it flagged, 50 B is genuine data (loaded as an address, never
executed) and the remaining 622 B decodes and reassembles byte-exact once the prober’s own
undercount is corrected. Conversions in this image are proved by round trip: disassemble the
run, re-assemble that exact text, require the original bytes back.
Counting .incbin distorts in both directions. It hides real debt written as .byte,
and it equally rewards pushing legitimate data into .byte — respelling a viewable
PNG-backed image as hex reduces the metric while destroying the better representation.
Neither direction is progress.
Scope: .byte sequences that encode native TLCS-900 CPU instructions across all
thirteen gated images. Data tables, strings, bitmaps, firmware bytecode for software interpreters, and padding are out of scope (correct as-is).
Current Status
By ROM
| ROM | Code .byte remaining |
Instruction statements | Status |
|---|---|---|---|
| Main CPU (v9 / v10, each) | confirmed-region backlog 0 B + misframed islands (partly converted; not a fixed pool — see below) | 358,993 (v9) / 336,011 (v10) | Not complete |
| Main CPU (v7) | 236,713 B in 789 confirmed regions, plus 29,032 B in 2,456 misframed islands — the largest code-as-.byte debt in the project |
268,186 | Not complete |
| Sub CPU Payload | 0 | 45,478 | Complete |
| Sub CPU Boot | 0 | 1,358 | Complete |
| Table Data | 0 | 5,758 | Complete |
| Custom Data | 0 (data only) | 0 | Complete |
| HDAE5000 | 13,168 B | 35,285 | Not complete |
SX-WSA1R prom_a / prom_b |
0 — byte runs audited and typed | 121,691 / 67,195 | Complete |
SX-WSA1R prom_c / prom_d |
0 — survived a falsification attack | 75,701 / 0 | Complete |
The instruction statements column is, in the disassembly repository,
python3 notes/syntax-convergence-probes/mnemonic_census.py --root <dir>
with <dir> one of v10/maincpu, v9/maincpu, v7/maincpu, v142/subcpu, subcpu/boot,
hdae5000, table_data, custom_data, wsa1/prom_a … wsa1/prom_d; the TOTAL line is
the figure. The census counts every instruction-shaped statement under the root — native
mnemonics, tree-local synthetic names and tree macros alike — so it is not a count of
native instructions. Custom data’s 0 is a fact about a pure data ROM, not a gap. The SX-WSA1R
figures cover the files under each prom_* root: the kernel and DSP-driver sources both
processors share (wsa1/kernel/, 938 statements; wsa1/dsp/, 97) are assembled into prom_a
and prom_c but live outside both roots.
The misframed islands in the Main CPU rows are left as .byte on purpose: fixing one means
re-framing an instruction already present in a neighbouring converted region, not filling a
gap, and this work requires a round-trip proof per region rather than a bulk relabel.
Converting a confirmed region creates new islands at its boundary, so the island count is
not a fixed pool — it moves with the tree state and should be re-measured, not quoted from
this page, before being used for planning.
scripts/analysis/v9_v10_undisassembled_census.py is the v9/v10 and v7 census; it needs a
scratch tree and is run as --prepare, then --judge v7 and --islands v7 --max-island 63.
hdae5000/tools/measure_debt.py is the HDAE5000 one and prints to stdout with no arguments.
The HD-AE5000 figure rises as work lands, and that is correct. Twelve regions of mis-disassembled data were retyped back into
.byte, which moved the measurement up rather than down. A number that only ever falls would mean the instrument cannot see data-framed-as-code, which is the third kind of debt and the one the byte gate is blind to in both directions.
Why the column says “instruction statements”
The figures that once occupied that column — 239,683 for the main CPU, 35,721, 1,357, 1,678 and 502 for the others, labelled native instruction counts — could not be reproduced from anything in the repository: they predate the LLVM toolchain, and they were not merely stale. The HD-AE5000’s was off by a factor of about seventy (502 against the census’s 35,285). A figure with no producer does not get restored; the column is filled from the census and labelled by what the census counts, which is every instruction-shaped statement, native or synthetic.
Two instruments exist and measure different things, so they cannot share a column:
notes/syntax-convergence-probes/mnemonic_census.py --root <dir>counts instruction statements per source root — the column above. Its bucket breakdown for v10 is on ROM Reconstruction.llvm-mc -gemits one DWARF line row per instruction statement and none for data, which gives a per-image count that includes every source the image pulls in. Values on record in the disassembly repository’s notes: HD-AE5000 36,391, custom data 0,wsa1/prom_c76,647,prom_d0 — close to the census, not equal to it.
Neither is a count of native instructions; how far the tree is from that is the subject of LLVM Semantic Instructions. This page states debt, which is measured, rather than progress.
LLVM Backend Encodings Added
All previously missing instruction encodings have been implemented in the LLVM TLCS-900 backend:
| Category | Prefix | Count Converted | LLVM Status |
|---|---|---|---|
| JR/JRL/CALR branch instructions | 0x1E etc. |
1,214 | Fixed (label-based) |
| Compact register loads (d8 prefix) | 0xD8-0xEF |
2,680 reg-reg + 831 ALU/LD/BIT | Implemented |
| PrevBank (D7 prefix) | 0xD7 |
147 | Implemented |
| Memory R+d8 addressing | various | 3,616 | Implemented |
| Compact dst (CALL/JP/CPW/LD) | various | 1,038 | Implemented |
| Short LD (compact load) | 0x20-0x3F |
523 | Implemented |
| Compact imm32 loads | various | 684 | Implemented |
| ld A, (R+d16) source loads | 0xC3 |
~970 | Implemented (Mar 14) |
| ld (R+d16), A stores | 0xF3 |
~400 | Implemented |
| Shifts/Rotates/MUL/DIV | various | 246 | Implemented |
HDAE5000: not complete
hdae5000/tools/measure_debt.py counts 13,168 B of undocumented .byte/.word operand
bytes — overwhelmingly scattered single-byte numeric fields rather than one contiguous block
— plus 3,815 B more that carries a decoding comment but has not been converted to real
instructions. Against a 512 KB ROM that is 2.5 % debt and 96.8 % real source. Most remaining
.byte in the tree is genuine data (custom-filesystem templates — the HD-AE5000 filesystem
is not FAT16 — string constants, UI bitmaps, etc.), but these figures are not yet zero.
There are 0 B of raw .incbin with no rebuild rule.
Execution Plan
Step 1: Precise Automated Audit
Write a Python script (scripts/audit_byte_code.py) that:
- Parses all
.sfiles across all ROMs - Identifies
.bytesequences between native instructions (code context) - Attempts
llvm-mc --triple=tlcs900 --disassembleon each sequence - Classifies results: (a) already decodable by LLVM → immediate conversion, (b) needs LLVM backend addition, (c) confirmed data
- Groups code
.byteby first byte (opcode prefix) to identify LLVM encoding families - Outputs a report with: file, line, bytes, category, status
Step 2: Convert Already-Decodable Instructions
Some .byte sequences may already have LLVM support but were written as .byte historically. Convert them directly to native mnemonics using the disassembler output.
Verification: make clean && make all + compare_roms.py after each batch.
Step 3: LLVM Backend — Compact Register Loads
Add encoding support for compact register load instructions:
ld wa, 0(D8 A8),ld xde, 0(EA A8),ld xwa, 1(E8 A9), etc.- These are 2-byte compact forms vs the 3-4 byte extended forms
Step 4: LLVM Backend — Compact Stack Pointer Arithmetic
Add encoding support for:
dec N, xsp(EF 6A/6E) — decrement stack pointer by Ninc N, xsp(EF 62/66) — increment stack pointer by N
Step 5: LLVM Backend — calr Fix
Fix calr with numeric address targets. Currently broken — emits absolute bytes instead of relative offset. Either fix the encoder to compute the relative offset, or add a new mnemonic variant.
Step 6: LLVM Backend — F2 Immediate-to-Memory Stores
Add encoding support for the F2-prefix ld (mem), imm instructions that store immediate values to memory addresses. ~358 occurrences.
Step 7: LLVM Backend — C3 R+d16 Source Loads
Add encoding support for ld A, (R+d16) source addressing (C3 prefix). ~216 occurrences. Note: the D3/E3/F3 destination variants already work; this is the source (load) direction.
Step 8: LLVM Backend — Remaining D7 Prevbank
Add cps qiz, 0 and any other prevbank instructions not yet supported (~4 occurrences for cps, ~182 total D7-prefix).
Step 9: Batch Convert .byte → Native Mnemonics
After each LLVM backend addition (Steps 3-8), convert the corresponding .byte sequences in the disassembly to native instructions. Use Python scripts with binary I/O (Latin-1 safety policy). Verify byte match after each batch.
Step 10: Disassemble FDC Raw Byte Blocks
The FDC routines in maincpu/storage/fdc_routines.s contain ~434 lines of raw .byte that are actual instruction sequences. These need:
- Disassembly using
llvm-mc --disassembleorllvm-objdump - Analysis of each routine’s purpose
- Semantic labeling (no
LABEL_XXXXXXallowed) - Documentation header comments
Step 11: Disassemble Flash/Floppy Handler Blocks
maincpu/storage/flash_floppy_handlers.s and maincpu/storage/single_load.s contain ~2,272 lines of raw byte blocks that need full disassembly, semantic labeling, and documentation.
Step 12: Iterative Jump/Call Table Discovery & Disassembly
Newly disassembled code may reveal previously unidentified jump tables or call tables. These must be found and their targets disassembled, repeating until exhaustion:
- Scan for undiscovered tables in all newly disassembled code blocks (sequences of
.longvalues in ROM range,lda+jp (xwa)patterns, indexed dispatch) - Verify table entries are code targets (not already-disassembled or false positives)
- Disassemble newly discovered code targets with semantic labeling and documentation
- Recurse — newly disassembled code may contain more tables
- Terminate when no new tables or code targets are found
Scope: All ROMs (maincpu, subcpu, hdae5000, table_data).
Step 13: Final Verification & Website Sync
make gate-all— thirteen images byte-identical, or the change does not land- Run
scripts/sync_docs_labels.py --applyto update any new labels on the website - Update
rom-reconstruction.md
Ordering & Priorities
Do first: Step 1 (audit) — gives precise scope for everything else. Then: Steps 2 (free wins), 3-4 (easy LLVM additions, high impact). Then: Steps 5-8 (harder LLVM work, each unblocks batch conversions). Then: Steps 9-11 (conversion work, depends on LLVM additions). Then: Step 12 (iterative table discovery — feeds back into Steps 9-11). Finally: Step 13 (verification & sync).
Steps 3-8 are independent of each other and can be parallelized. Step 12 is iterative and may cycle back through Steps 3-11.
Verification
After each step:
cd kn5000-roms-disasm && make gate-all— thirteen images, byte-exact or non-zero exit- LLVM tests:
cd llvm-project && build/bin/llvm-lit llvm/test/CodeGen/TLCS900/
Do not accept a percentage in place of the gate.
compare_roms.pyprintsSimilarity: 100.00%rounded to two decimals, which in a 2 MB ROM covers up to 104 differing bytes, and it silently skips any section whose built file is missing — so a run that never assembled the six ASL mirror sections prints nine sections, all reading100.00%, and looks identical to a passing full run. See Disassembly Workflow.
Policy Compliance
All newly disassembled code MUST:
- Have semantic label names (no
LABEL_XXXXXX) - Have documentation header comments explaining what each routine does
- Be verified with byte-match builds before committing
- Have website docs updated if labels appear on documentation pages