sd:isa
Differences
This shows you the differences between two versions of the page.
| Both sides previous revisionPrevious revisionNext revision | Previous revision | ||
| sd:isa [2026/06/16 12:31] – appledog | sd:isa [2026/09/08 04:07] (current) – external edit 127.0.0.1 | ||
|---|---|---|---|
| Line 178: | Line 178: | ||
| === Tier 4: Acceleration (LLVM / Forth / Hardware) | === Tier 4: Acceleration (LLVM / Forth / Hardware) | ||
| - | Instructions added to remove work from a specific hot path rather than to add | + | Instructions added to remove work from a specific hot path rather than to add expressiveness. Each one has a primary consumer noted below. |
| - | expressiveness. Each one has a primary consumer noted below. | + | |
| + | Canidates for inclusion: LEA, fused CMP-Bcc and CMP-Jcc, conditional move, LD_IDX16, [[# | ||
| | # | hex | Mnemonic | Example | | # | hex | Mnemonic | Example | ||
| | 28 | $1C | [[# | | 28 | $1C | [[# | ||
| | 29 | $1D | [[# | | 29 | $1D | [[# | ||
| + | | 60 | $3C | [[# | ||
| + | | 61 | $3D | [[# | ||
| | 210 | $D2 | [[# | | 210 | $D2 | [[# | ||
| | 211 | $D3 | [[# | | 211 | $D3 | [[# | ||
| Line 388: | Line 391: | ||
| <wrap #cmpc /> | <wrap #cmpc /> | ||
| **''# | **''# | ||
| - | Non-zero byte compare, useful for strings. Compares up to C characters; C returns | + | Non-zero byte compare, useful for strings. Compares up to C characters; C returns either the index of the first mismatch or the matched length. Sets ZERO on a full match; otherwise CARRY distinguishes the -1 / +1 ordering. |
| - | either the index of the first mismatch or the matched length. Sets ZERO on a full | + | |
| - | match; otherwise CARRY distinguishes the -1 / +1 ordering. | + | |
| - | CMPC allows early termination when '' | + | CMPC allows early termination when '' |
| - | semantics. If both strings reach a terminator at the same position with all prior | + | |
| - | bytes equal, the loop exits matched (Z=1, C=1). Only '' | + | |
| - | that point '' | + | |
| - | '' | + | |
| - | null-terminated strcmp in one instruction. | + | |
| <wrap #skpc /> | <wrap #skpc /> | ||
| Line 493: | Line 489: | ||
| <wrap #casetab /> | <wrap #casetab /> | ||
| **''# | **''# | ||
| - | Jump table held at a base address rather than inline. Indexes an address from the | + | Jump table held at a base address rather than inline. Indexes an address from the table at '' |
| - | table at '' | + | |
| - | //The precise indexing and out-of-bounds policy should be confirmed against the | + | |
| - | current emulator handler: the older CASE3 (computed table at base) and CASEB | + | |
| - | (scanned '' | + | |
| - | table-at-base form.// | + | |
| <wrap #cvtan /> | <wrap #cvtan /> | ||
| **''# | **''# | ||
| - | Convert ASCII to number. Maps ' | + | Convert ASCII to number. Maps ' |
| - | a decimal digit (0-9) and Carry if it is not a hex digit (0-15). Also a fast digit | + | |
| - | test: '' | + | |
| - | Designed for zoned decimal, and works for zoned hex. | + | |
| <wrap #cvtna /> | <wrap #cvtna /> | ||
| Line 513: | Line 501: | ||
| <wrap #idx /> | <wrap #idx /> | ||
| - | **''# | + | **''# |
| - | Indexed load/store: a 24-bit pointer register plus a displacement. The '' | + | |
| - | take a signed byte immediate (-128..+127); | + | |
| - | These back the compiler' | + | |
| - | pointer base must be a 24-bit register; the effective address reads the base at its | + | |
| true width before applying the offset. | true width before applying the offset. | ||
| Line 535: | Line 519: | ||
| <codify armasm> | <codify armasm> | ||
| - | LDFLX $F000 ; loop start | + | LDFLX $F000 ; |
| - | LDA #1 | + | LDA #1 ; start at 1 |
| - | STA [FLX] | + | STA [FLX] ; write starting counter at first four bytes |
| - | LDA #1000 | + | LDA #1000 ; go until 1000 |
| - | STA [FLX+4] | + | STA [FLX+4] |
| loop: | loop: | ||
| - | LSTEPM | + | LSTEPM |
| JNZ @loop | JNZ @loop | ||
| </ | </ | ||
| - | Unlike a DEC loop it counts upward and supports an arbitrary start. | + | Unlike a DEC loop it counts upward and supports an arbitrary start. |
| - | for Forth and has been generalized by [[# | + | |
| - | LSTEPM | + | Only used by Forth. It is recommended to use LSTEP instead; this instruction |
| <wrap #ttos /> | <wrap #ttos /> | ||
| Line 568: | Line 552: | ||
| LSTEP C, @loop | LSTEP C, @loop | ||
| </ | </ | ||
| + | |||
| + | |||
| + | <wrap #mac /> | ||
| + | **''# | ||
| + | Multiply-accumulate: | ||
| + | written. The result takes the accumulator' | ||
| + | read at their own -- so '' | ||
| + | |||
| + | **Why it exists.** Every tile lookup in a tile game computes | ||
| + | '' | ||
| + | rendered cell and every walkability test; it was the single most executed | ||
| + | arithmetic sequence in the program. '' | ||
| + | and the scratch register the product needed disappears with it. | ||
| + | |||
| + | <codify armasm> | ||
| + | ; ELM = tile_base + (y * width + x) * 2 | ||
| + | LDB [@level_dim] | ||
| + | MOV Z, X ; Z = column | ||
| + | MAC Z, Y, BL ; + row * width | ||
| + | ADD Z, Z ; two bytes per tile | ||
| + | LDELM [@level_tiles] | ||
| + | ADD ELM, Z | ||
| + | </ | ||
| + | |||
| + | **Where else it helps.** Anything shaped '' | ||
| + | two MACs and no temporary at all: | ||
| + | |||
| + | <codify armasm> | ||
| + | LDA #0 | ||
| + | MAC A, I, I ; A = dx*dx | ||
| + | MAC A, J, J ; A += dy*dy | ||
| + | </ | ||
| + | |||
| + | Also dot products, fixed-point scaling, Horner' | ||
| + | inner loop of any filter, matrix or checksum routine. | ||
| + | |||
| + | **Recognising the pattern.** Look for a '' | ||
| + | to something and then never used again, or a '' | ||
| + | scratch register. If that register exists only to carry the product from the | ||
| + | multiply to the add, MAC removes the register and an instruction together. | ||
| + | |||
| + | **Measured.** Adding MAC to rogueima' | ||
| + | 5.0% off the cost of a turn. | ||
| + | |||
| + | |||
| + | <wrap #absd /> | ||
| + | **''# | ||
| + | Absolute difference: '' | ||
| + | always clear and Z is set exactly when the two operands are equal. '' | ||
| + | read at its own width. | ||
| + | |||
| + | **Why it exists.** Distance work needs the size of a difference, not its sign, | ||
| + | and getting one without this instruction costs four to six instructions of | ||
| + | compare, branch, and subtract-the-other-way round -- **per axis**. Rogueima' | ||
| + | line-of-sight does it twice for every square it examines, a few thousand times | ||
| + | a turn. | ||
| + | |||
| + | <codify armasm> | ||
| + | ; dx = |x - px|, without a branch in sight | ||
| + | LDA #0 | ||
| + | LDAL [@PX] | ||
| + | LDI #0 | ||
| + | MOV IL, XL | ||
| + | ABSD I, A ; I = |x - px| | ||
| + | </ | ||
| + | |||
| + | **Where else it helps.** Manhattan and Chebyshev distance, "are these two within | ||
| + | N of each other", | ||
| + | audio or images. It doubles as a branch-free equality test: Z is set precisely | ||
| + | when the operands match, so you get '' | ||
| + | |||
| + | **Recognising the pattern.** Look for a compare followed by two subtractions in | ||
| + | opposite orders down the two arms of a branch -- any | ||
| + | '' | ||
| + | |||
| + | **Measured.** 7.2% off the cost of a turn in rogueima -- more than MAC, because | ||
| + | the branches it removes were unpredictable ones inside the hottest loop. | ||
| + | |||
| + | |||
| + | <wrap #ptrace /> | ||
| + | **'' | ||
| + | //No opcode assigned. This is a specification for review, not a description of | ||
| + | the machine.// | ||
| + | |||
| + | Walk a line through a tile array, stopping at the first tile that blocks, and | ||
| + | optionally marking every tile visited on the way. A PPU instruction in the same | ||
| + | family as '' | ||
| + | |||
| + | ^ ^ ^ | ||
| + | | **in** | '' | ||
| + | | ::: | '' | ||
| + | | ::: | '' | ||
| + | | ::: | '' | ||
| + | | ::: | '' | ||
| + | | **out** | '' | ||
| + | | ::: | '' | ||
| + | | ::: | '' | ||
| + | |||
| + | At every tile along the line it ORs '' | ||
| + | '' | ||
| + | the edge of the array. A '' | ||
| + | projectile or a "can this monster see me" test wants. | ||
| + | |||
| + | <codify armasm> | ||
| + | ; one line-of-sight ray, marking SEEN|VISIBLE and stopping at anything opaque | ||
| + | LDELM [@level_tiles] | ||
| + | LDB [@level_dim] | ||
| + | MOV K, B | ||
| + | LDBL #2 ; two bytes per tile | ||
| + | LDBH #1 ; the flags are the second of them | ||
| + | LDCL @TF_OPAQUE | ||
| + | LDCH $0C ; TF_SEEN | TF_VISIBLE | ||
| + | LDXL [@PX] | ||
| + | LDYL [@PY] | ||
| + | LDI [@target_x] | ||
| + | LDJ [@target_y] | ||
| + | PTRACE | ||
| + | </ | ||
| + | |||
| + | **Why.** Rogueima' | ||
| + | 45,000 of the 62,577 instructions a turn costs -- 48 ms at 1.3 MIPS. Each ray | ||
| + | is roughly 540 instructions of Bresenham stepping and per-tile testing. As one | ||
| + | instruction, | ||
| + | instead of 540, and a turn drops to roughly 15 ms. | ||
| + | |||
| + | **Where else.** The masks are what make it general. Projectile and thrown-object | ||
| + | paths, wand and breath-weapon beams, monster targeting, swept collision in a | ||
| + | tile platformer, and post-pathfinding line smoothing ("can I walk straight from | ||
| + | here to there" | ||
| + | '' | ||
| + | |||
| + | The application that makes it a platform capability rather than one game's | ||
| + | helper is a **Wolfenstein-style raycaster**: | ||
| + | tile map, 320 rays a frame at 60 Hz. In software that is several million | ||
| + | instructions a second, more than this machine has, so the genre is off the | ||
| + | table. As an instruction it is about 19,000 a second plus the column drawing. | ||
| + | |||
| + | **A design fork worth settling first.** Two different things wear similar | ||
| + | clothes here: | ||
| + | |||
| + | * **Bresenham** visits the tiles a line passes through. Symmetric, integer, | ||
| + | and what line of sight, projectiles, | ||
| + | * **DDA** steps to each grid // | ||
| + | distance. That is what a perspective raycaster needs -- Wolfenstein used it | ||
| + | because Bresenham' | ||
| + | a wrong distance is a wrong wall height and a fisheye. | ||
| + | |||
| + | '' | ||
| + | wants a sibling -- '' | ||
| + | of the tile was struck, for texture mapping. Trying to make one instruction do | ||
| + | both would make both worse. | ||
| + | |||
| + | The name deliberately says the job rather than the algorithm, so that the pair | ||
| + | reads '' | ||
| + | |||
| + | **Open questions.** | ||
| + | - Six registers in is a lot. Is a descriptor block at '' | ||
| + | '' | ||
| + | level-scoped rather than per-call? | ||
| + | - Should it also report the tile it stopped ON versus the last clear tile | ||
| + | before it? Line of sight wants the blocker marked; a projectile wants the | ||
| + | square in front of it. | ||
| + | - Should a zero-length line (start == target) mark the start square? | ||
| + | |||
| + | **Precedent.** No CPU precedent that I know of, but a strong coprocessor one: | ||
| + | this is what blitters did. The Amiga blitter (1985) had a hardware Bresenham | ||
| + | line mode, and the TMS34010 (1986) was a graphics processor with '' | ||
| + | instruction. It belongs to that tradition, which is where the rest of the PPU | ||
| + | already lives. | ||
| + | |||
sd/isa.1781613071.txt.gz · Last modified: by appledog
