Smallsome MicroRISC: uRISC-T1 Architecture, ISA, Compiler and Formats
As of: 10 October 2026 Theoretical specification: Version 1.0, document revision 4
Branding note: The public brand is Smallsome MicroRISC, with the technical
spelling microrisc. The normative ISA/model remains uRISC-T1, and the
compiler target remains urisc-t1. The public entry is
https://microrisc.smallsome.com/, forwarding to the canonical project home
https://smallsome.com/microrisc/. The source language is Smallsome 3.0, based
on Cymple 2.0. Its source modes are Smallsome Code (ASCII syntax) and
Smallsome Symbols (Unicode-symbol syntax). The format family
is Smallsome Formats; individual format names and source filenames remain
unchanged.
Document revision 2 applied English text and confirmed Smallsome branding. Revision 3 aligned the language references with the shared 3.0 specification. Revision 4 corrects the num/word boundary, ownership, cancellation and language manifest contracts. ISA version, encodings and hardware arithmetic are unchanged. All compiler, runtime and lab contracts remain theoretical; no implementation or physical validation is commissioned by this document.
1. Purpose and Status
This document consolidates the current design of the uRISC CPU family. It distinguishes binding target decisions, justified target proposals, and the historical state of the existing uRISC-Lab-v4.1 emulator.
The binding theory now also covers all opcodes, their encoding, compiler/ABI rules, integration of all eight in-house format families, and the contract for a future 8x8 lab app with a Scrapbook. No app, RTL, or hardware prototype is being built at this stage. Fully specified does not mean implemented or physically proven.
The central idea is not a particularly large single processor. Many small, deterministic cores are scheduled into a static task and data flow by the compiler itself. Specialized hardware is limited to time-critical transport, particularly display scanout, audio modulation, DMA, RAM access, and USB at the bit level. Audio, video, graphics, and application algorithms remain programmable.
Current target scope (decided): exactly one CPUlet with 64 cores and 256 MiB of shared RAM. All decisions in this document refer to this target system. Systems with multiple CPUlets remain open and are collected as future work in section 23.
Terms used in this document:
- decided: target decision for the hardware architecture.
- normative: binding behavior of the theoretical T1 model.
- target proposal: architecturally preferred, but not yet encoded as RTL.
- emulator state: behavior of the existing v4.1 model; not automatically a decision for future hardware.
Reading order:
- Sections 2..8: computer structure and cores.
- Sections 9..11: complete binary ISA.
- Sections 12..20 and 27: memory, contact, events, and devices.
- Section 28: compiler, ABI, and runtime.
- Section 29: all in-house formats with their actual dependencies.
- Sections 30..31: lab/Scrapbook, evidence, and source baseline.
- Sections 21..25: historical hardware guidance and future work, not prerequisites for the theoretical lab.
The historical reference is located at:
Tools/Diverse Tools/murisc/murisc_v41_mod_webm_opcode_audit.zip
2. Binding Guidelines
- A core is small, deterministic, and fully in-order.
- Normal instructions must not cause wait cycles. Fixed, documented execution times such as the branch penalty are not wait cycles.
- The result of an instruction that writes a register must be usable by the immediately following instruction.
- External latencies are not pulled into the instruction pipeline.
- Only the explicit WAIT family may put a running task to sleep. DONE ends it; FAULT and dispatcher halt are separate control states.
- Local memory is an explicit scratchpad, not a cache.
- Shared RAM, display, USB, disk I/O, and other peripherals use the same contact contract.
- A uniform contact means a uniform protocol, not a single electrically shared bus.
- There is no cache coherence. Data ownership and synchronization are managed by the compiler, dispatcher, messages, and explicit transfers.
- The ISA remains independent of the number of available cores.
- No codec-, audio-, or graphics-specific opcodes are introduced.
- Future larger systems are built from identical 64-core units (future work, section 23). Current decisions must not prevent this scaling.
- Each local 1-KiB slot has exactly one owner per cycle.
- Response data for a transaction initiated by a core lands exclusively in the local SRAM reserved for it. Autonomous bus masters such as USB, display, and audio DMA may use shared RAM directly. Every completion, event, and error reaches the responsible core as a message. There is no second notification path.
- No additional protocol or graphics bridges are intended. DRAM, power supply, clock source, connectors, and the protection, level-shifting, and load components required for each connection are unavoidable.
3. Structure
3.1 Core
The core is the smallest programmable computing unit.
3.2 CPUlet
A CPUlet consists of exactly 64 cores and a local contact router. The target system consists of exactly one CPUlet.
1 CPUlet = 8 groups x 8 cores = 64 cores
The term CPUlet initially denotes a logical and physical functional unit on a chip. It does not imply a future division into separately manufactured chiplets.
Internal structure (target proposal):
8 cores --+
... +-- Group arbiter (8:1 / 1:8, round robin within the class) --+
8 cores --+ |
|
8 group ports + RAM/peripherals + CPUlet-SRAM ---------------------+
= 10x10 crossbar, 64 bits
- A full 64-port crossbar is ruled out because its wiring grows quadratically.
- A group port provides 4 GB/s of raw bandwidth and at most 3.2 GB/s of full-data payload bandwidth for eight cores: when shared fairly, at most 400 MB/s or 1.6 bytes per RUN cycle and core. The shared RAM port and group/router contention jointly determine the bottleneck; the topology alone does not prove freedom from bottlenecks.
- Messages between cores in the same group do not use the router.
- The group is also the unit for clock gating and shutdown.
- Four mesh ports for multi-CPUlet systems belong to future work (section 23) and are not present in the target system.
4. Target System
The decided configuration is:
Cores: 64 (8 groups x 8)
CPUlets: 1
System cores: 4 permanently reserved
Application cores: 60
Clock levels: 0 / 50 / 250 MHz
Local memory: 12 KiB per core
CPUlet-SRAM: 0 bytes in the T1 standard profile, address window reserved
Shared RAM: 256 MiB
Graphics output: up to 1280 x 720 at 60 Hz via DVI-D
Audio output: up to 7.1 via PDM, alternatively I2S/TDM
Input devices: 3 x USB 1.1 Full-Speed with integrated LS/FS PHYs
Mass storage: 1 x USB 2.0 High-Speed
The four system cores handle input, output, USB, the dispatcher, and other system tasks. The remaining 60 cores are available to applications and media. The exact internal distribution of the four system tasks remains a matter for the compiler and operating system.
Minimum board population: uRISC chip, one LPDDR4 component (section 15.2), power supply, clock source, connectors, and the required passive protection and termination components. USB port power and an analog audio output capable of driving its load may require additional load switches, buffers, or amplifiers. DRAM in the same package (SiP) is a possible future variant.
5. Clock and Energy Model
The previously discussed 750-MHz level is removed. The target hardware has three states:
| State | Clock | Meaning |
|---|---|---|
OFF |
0 MHz | shut down or only necessary retention |
IDLE |
50 MHz | light real-time and standby work |
RUN |
250 MHz | full computing performance |
Omitting 750 MHz is a deliberate architectural decision:
- A cycle at 250 MHz lasts 4 ns rather than 1.33 ns.
- Single-cycle multiplication, local memory, and complete forwarding are substantially more realistic.
- Supply voltage, clock tree, area, and power dissipation decrease.
- The compiler distributes bulk data work across multiple cores instead of forcing peak serial performance.
The contact fabric, RAM controllers, and display unit may independently
operate at 500 MHz. Frequency changes and shutdown should be controllable at
least at CPUlet or group level (8 cores); an individual
core must be capable of separate clock gating and being set to OFF.
6. Computing Requirements and Evidence Status
The previously discussed core counts are planning assumptions. None has yet been confirmed by a cycle-accurate T1 guest program. Native Mac measurements in the format matrix demonstrate the listed decoder runs, not the instruction costs of our ISA.
| Task | Hypothesis for T1 validation |
|---|---|
| RAU | one IDLE core for a specified mono/stereo stream |
| RFXL | one IDLE core for a specified image size and loading deadline |
| PMF0 | four RUN cores for a specified video/bitrate profile |
| 3D in the class of Tomb Raider 1 | about six to eight RUN cores for a bounded scene case |
| System and I/O | four reserved cores; specific system load must be budgeted |
| Office/media load | application-specific; 16/64 cores indicate scale |
RAU requires more than local SRAM for its unchanged native main stereo buffers. RFXL streams may have serial dependencies. PMF0 requires reference RAM. Section 29 describes these limits and the specific form of validation for each format.
64 cores at 250 MHz yield a theoretical 16 GInstr./s at full local throughput; after reserving the four system cores, 15 GInstr./s remain for applications. Branches, serial portions, slot planning, contact, and RAM limit actual performance. Estimating an instruction count from an ARM runtime without a trace is not proof of T1 performance.
7. Compiler and Execution Model
The compiler is a central part of the architecture. During compilation it should already:
- split loops and independent data regions,
- distribute tasks across cores or core groups,
- schedule local 1-KiB slots and define their ownership for each phase,
- allocate RAM regions and resources,
- schedule communication and completion events,
- check real-time budgets in cycles, including branch penalties,
- define data ownership so that no cache coherence is required.
The hardware dispatcher starts the prepared tasks. It should not be replaced by complex dynamic out-of-order or migration logic.
Serial portions remain possible. USB state machines, filesystem logic, linked structures, and heavily branching control do not have to be artificially parallelized. 250 MHz is considered sufficient for these control tasks; DMA and specialized transport handle bulk data. Light system tasks such as mouse and keyboard run on the system cores and start suitable routines on other cores as needed.
8. Target Core Structure
8.1 Registers and State
The v4.1 emulator provides the current basis:
- 12 general-purpose 32-bit registers
R0throughR11, - 4 resource registers
S0throughS3, - 32-bit program counter,
- 64-bit accumulator for
MAC, - persistence and active masks for local memory banks,
- timestamps for
WAIT, - instruction counter,
- message FIFO with 16 entries.
In the emulator, a message contains a 16-bit sender identifier and a 32-bit value. 16 bits are also sufficient for future multi-CPUlet systems with up to 4096 cores; the final subdivision of the identifier by group and core is not separately encoded by CPUlet/group in the historical emulator. T1 uses coreID=8*y+x; the future extension is not part of T1.
T1 adds kind8/status8; the message is exactly 64 bits. Core and resource identifiers and buffers are defined in section 27.
The register file contains twelve 32-bit values, three read ports for SEL, and up to three write capabilities for RECV. Physical banking, bypass, and port area remain to be determined; they are not a measured negligible quantity.
8.2 Local Memory
Each core has a fixed:
12 slots x 1 KiB = 12 KiB local SRAM
Properties:
- guaranteed local and deterministic,
- target: one-cycle access time,
- explicitly managed by program and compiler,
- individually activatable banks,
- persistence mask for
IDLEandWAIT, - no transparent cache,
- no automatic cache fill or write-back.
Slot ownership (decided):
- Each 1-KiB slot is a separate SRAM bank.
- Per cycle, a slot has exactly one owner: instruction fetch, load/store unit, or contact/DMA.
- Code starts at slot 0 by default but may reside in any slot. The compiler determines placement during optimization.
- The compiler guarantees that a slot in a given phase is either code, data, or a DMA destination.
- A violation causes a
FAULTrather than a wait cycle (section 19).
Four MiB of shared RAM per core is only a capacity rule for external RAM. It is not installed as physically local SRAM per core and is not permanently assigned to a core.
8.3 Pipeline
A four-stage in-order pipeline is decided:
| Stage | Task |
|---|---|
| IF | instruction fetch from the local code slot |
| ID | decoding, register reads, and forwarding preparation |
| EX | ALU, multiplication, MAC, address calculation, local SRAM access, and branch decision |
| WB | write-back |
Observable contract:
- at most, and normally, one instruction per active cycle,
- no out-of-order execution,
- no speculative program execution that changes state,
- no normal instruction causes a wait cycle,
- the result is available to the immediately following instruction, including after
LD, - complete forwarding for register results,
- fixed and documented execution time for all normal instructions.
Branch penalties (decided):
| Instruction | Additional lost issue cycles |
|---|---|
BEQ/BNE not taken |
0 |
BEQ/BNE taken |
2 |
JMP |
2 |
The branch decision occurs in EX. Instructions fetched up to that point may
be decoded, but must not execute when the branch is taken; they are
discarded before any state change. An LD immediately before a branch
supplies its result through WB-to-EX forwarding. This causes neither a
wait cycle nor a combined SRAM-compare-PC path within the same
cycle.
Fetch/decode errors of a younger instruction are only recorded as pending until its valid EX acceptance. A discarded wrong path causes no FAULT. Branches and faults check their own state before any side effect.
Branches with magnitude comparisons are not introduced; CMP
and BNE serve that purpose. The compiler knows the fixed penalty and can
express short paths without branches using SEL.
More area for parallel computing paths, fast multiplication, and forwarding is accepted. Variable or iterative instruction execution times are not accepted.
9. Instruction Format and ISA Version (Normative)
The theoretical ISA is named uRISC-T1, version 1.0. It has exactly
31 primary opcodes. T1 denotes the fully specified theoretical
contract, not an implemented chip or one validated by RTL.
Old v4.1 machine code is not binary compatible. Source programs are
reassembled; there is no automatic legacy execution mode.
An instruction word is 32 bits and is stored little-endian in local SRAM:
31 26 25 22 21 18 17 14 13 0
+-------------+----------+----------+----------+----------------+
| Opcode 6 Bit| A 4 Bit | B 4 Bit | C 4 Bit | Imm 14 Bit |
+-------------+----------+----------+----------+----------------+
I14 is the signed immediate in bits 13:0; U14 is the same
bit value without a sign. I22/U22 use bits 21:0, U18 bits 17:0.
Register identifiers 0 through 11 denote R0 through R11. R0 is a normal,
writable register. 12 through 15 are invalid for GPR operands.
Resource fields separately address S0 through S3; they are not GPR aliases.
Every undocumented mode, reserved opcode, and set
reserved bit causes FAULT_ILLEGAL. Unused fields must be zero.
Invalid register fields are not silently masked.
Addressing is byte-based. The PC is the address of the current instruction;
nextPC = PC + 4. A relative branch calculates
target = nextPC + 4 * signExtend(displacement) without 32-bit overflow.
PC and branch target must be divisible by four and reside in an enabled,
active code slot. An invalid target causes FAULT_PC before
the link register or PC is changed.
10. Historical v4.1 Opcode Set (Informative)
The historical isa.pbi numbers opcodes starting at 1:
NOP MOVI MOV ADD ADDI SUB AND OR XOR SHL SHR MUL MAC
LD32 ST32 LD16 ST16 BEQ BNE JMP WAIT WAITUNTIL
SEND RECV SIGNAL RWRITE RREAD DONE
These are 28 primary opcodes. v4.1 is a functional starting point: 14-bit MOVI, separate load/store opcodes, partially blocking resource instructions, and a simulation quantum instead of a pipeline. This meaning applies exclusively to the archived version. Section 11 is the sole normative T1 ISA.
11. Complete Opcode Set (Normative)
11.1 General Arithmetic Rules
u32(x) denotes the lower 32 bits; s32(x) is the same bit value in
two's complement. Normal addition/subtraction and immediate addition
operate modulo 2^32. They produce neither flags nor exceptions on overflow.
There is no flags register. A comparison explicitly produces a mask.
Lane 0 occupies the least significant bits. Width w and lane count n
satisfy w*n=32. Carries, saturation, and shifts operate per lane;
they never extend into an adjacent lane.
The 64-bit accumulator ACC is a two's-complement bit vector.
Accumulation is modulo 2^64, without implicit saturation.
All operands are read before the result is written. This also applies
when destination and source are the same register.
11.2 Opcode Table
Numbers 16 and 17 remain reserved so that historical LD16/ST16 words do not accidentally denote a new valid instruction. 0 and 34 through 63 are also reserved: 31 assigned, 33 reserved.
| Decimal / Hex | Instruction | Canonical operands | Effect |
|---|---|---|---|
| 1 / 01 | NOP | none | nextPC only |
| 2 / 02 | MOVI | Ra, I22 | Ra = signExtend(I22) |
| 3 / 03 | MOV | Ra, Rb | Ra = Rb |
| 4 / 04 | ADD | Ra, Rb, Rc, mode | per-lane addition |
| 5 / 05 | ADDI | Ra, Rb, I14 | scalar modulo addition |
| 6 / 06 | SUB | Ra, Rb, Rc, mode | per-lane subtraction |
| 7 / 07 | AND | Ra, Rb, Rc | bitwise AND |
| 8 / 08 | OR | Ra, Rb, Rc | bitwise OR |
| 9 / 09 | XOR | Ra, Rb, Rc | bitwise exclusive OR |
| 10 / 0A | SHL | Ra, Rb, count, mode | per-lane left shift |
| 11 / 0B | SHR | Ra, Rb, count, mode | per-lane right shift |
| 12 / 0C | MUL | Ra, Rb, Rc, mode | half of a scalar 64-bit product |
| 13 / 0D | MAC | Ra, Rb, Rc, mode | write/accumulate/read ACC |
| 14 / 0E | LD | Ra, [Rb + I14], mode | local 8-/16-/32-bit load |
| 15 / 0F | ST | Ra, [Rb + I14], mode | local 8-/16-/32-bit store |
| 18 / 12 | BEQ | Ra, Rb, I14 | branch on bit equality |
| 19 / 13 | BNE | Ra, Rb, I14 | branch on bit inequality |
| 20 / 14 | JMP | form, operand | relative/indirect jump, optional link |
| 21 / 15 | WAIT | form, operand | explicit wait for message/time |
| 22 / 16 | WAITUNTIL | Ra, Rb | explicit wait until absolute time |
| 23 / 17 | SEND | Ra, Rb, Rc | status to Ra, destination from Rb, value from Rc |
| 24 / 18 | RECV | Ra, Rb, Rc | status, value, and metadata from Rx FIFO |
| 25 / 19 | SIGNAL | Ra, Rb | SEND with fixed value 1 |
| 26 / 1A | RWRITE | Ra, Sb, Rc | status, resource, descriptor address |
| 27 / 1B | RREAD | Ra, Sb, Rc | status, resource, descriptor address |
| 28 / 1C | DONE | none | end task and stop core |
| 29 / 1D | CMP | Ra, Rb, Rc, mode | per-lane zero/full mask |
| 30 / 1E | SEL | Ra, Rb, Rc | bitwise selection using old Ra as mask |
| 31 / 1F | PERM | Ra, Rb, Rc, pattern | select four bytes from eight source bytes |
| 32 / 20 | MOVHI | Ra, U18 | replace upper 18 bits |
| 33 / 21 | CLZ | Ra, Rb | count leading zero bits |
Ra/Rb/Rc are fields A/B/C, except where the following tables define
a different field assignment. A field not required for the particular
instruction is zero. Modes are bit values, not additional instruction words.
11.3 Constants, Moves, and Boolean Operations
- NOP: A=B=C=U14=0.
- MOVI: A is the destination, bits 21:0 are I22; range -2097152 to 2097151.
- MOV: A destination, B source; C=U14=0.
- MOVHI: A destination, B=0, bits 17:0 are U18.
Ra = (U18 << 14) | (oldRa & 0x3FFF). - AND/OR/XOR: A destination, B/C sources; U14=0.
- ADDI: A destination, B source, C=0; I14 ranges from -8192 to 8191.
An arbitrary 32-bit constant K is canonically constructed using:
MOVI Ra, K & 0x3FFF
MOVHI Ra, K >> 14
For K in the I22 range, MOVI with the signed value is sufficient. The assembler may optimize the two-instruction sequence but must not generate additional undocumented opcodes.
11.4 ADD, SUB, and CMP
ADD/SUB: bits U14[1:0] encode lane width: 0=32, 1=16, 2=8, 3=invalid. Bits [3:2] encode 0=modulo, 1=unsigned saturating, 2=signed saturating, 3=invalid. Bits [13:4] are zero.
Unsigned saturation clamps to [0, 2^w-1], signed saturation to [-2^(w-1), 2^(w-1)-1]. SUB with unsigned saturation returns zero if the difference would be negative.
CMP: bits [1:0] are the same width identifier. Bits [4:2] select:
| Value | Comparison per lane |
|---|---|
| 0 | equal |
| 1 | unequal |
| 2 | signed less than |
| 3 | signed less than/equal |
| 4 | unsigned less than |
| 5 | unsigned less than/equal |
6/7 are invalid; bits [13:5] are zero. A true condition returns all w bits set, a false condition all zero. Greater than and greater than/equal are expressed by swapping sources.
SEL has U14=0 and computes
Ra = (oldRa & Rb) | (~oldRa & Rc).
A scalar CMP mask selects a whole word; lane masks select lanes,
arbitrary masks select individual bits. SEL is not a truth-value test.
11.5 Shifts and PERM
SHL/SHR use the following bits:
| Bits in U14 | Meaning |
|---|---|
| [1:0] | lane width as in ADD |
| [2] | SHR: 0 logical, 1 arithmetic; SHL: must be 0 |
| [3] | 0 immediate count, 1 register count |
| [8:4] | immediate count, only when [3]=0 |
| [13:9] | zero |
For an immediate, C=0. For a register, C is the count source and [8:4]=0.
All lanes use the same count; its lower log2(w) bits apply.
Immediate values >=w are invalid. Register counts are used modulo w.
Vacated positions are filled with zero for logical shifts and with the old
lane sign for arithmetic SHR. SAR is an assembler representation
for SHR with bit [2]=1, not a separate opcode.
PERM: bits [2:0], [5:3], [8:6], [11:9] select result bytes 0 through 3. Index 0..3 selects Rb byte 0..3, 4..7 Rc byte 0..3. Bit [12]=1 reads Rc as four zero bytes; C must then be 0 and Rc is not read. Bit [13]=0. Example byte swap of Rb: indices 3,2,1,0, pattern 0x0053 with the zero source optionally disabled.
11.6 MUL and MAC
MUL has exclusively scalar 32-bit sources. Lane MUL is not part of T1. U14[0]=0 unsigned, 1 signed; [1]=0 lower, 1 upper product half. [13:2]=0. Signed MUL multiplies two s32 values, unsigned MUL two u32 values. The lower half is identical for both; the upper half depends on signedness. There is no automatic rounding or saturation.
MAC: [0] selects unsigned/signed product as in MUL, [2:1] selects:
| Value | Action |
|---|---|
| 0 | ACC = u64(ACC + product); Ra = low32(ACC_new) |
| 1 | ACC = product; Ra = low32(ACC_new) |
| 2 | Ra = low32(ACC); ACC is retained |
| 3 | Ra = high32(ACC); ACC is retained |
[13:3]=0. For action 2/3, B=C=[0]=0. Action 1 with a zero-valued register as source sets ACC to zero. This is the canonical initialization; no hidden reset opcode is introduced. Immediately successive MAC instructions see the new ACC through forwarding.
CLZ: C=U14=0, Ra is the number of leading zero bits in Rb; CLZ(0)=32. DIV, MOD, CTZ, LOOP, floating point, and specialized cryptographic opcodes are not part of T1. Software libraries implement them where permitted by the target profile. A LOOP proposal is therefore rejected for T1.
11.7 Local Memory
For LD/ST, C is a mode field, not a register:
| C | LD | ST |
|---|---|---|
| 0 | u8 to u32 | lower 8 bits |
| 1 | u16 to u32 | lower 16 bits |
| 2 | u32 | 32 bits |
| 3 | s8 to s32 | invalid |
| 4 | s16 to s32 | invalid |
The effective address is formed as the mathematical sum u32(Rb)+I14. Negative values, overflow, an end beyond 12288, and missing permissions fault. 16-/32-bit accesses are aligned to 2/4 bytes. An access must not cross a slot boundary. Byte order is little-endian. An ST writes only after all checks; on FAULT, SRAM remains unchanged.
The IF code slot and EX data slot in a cycle must be different. A DMA-owned slot must not be read or written by IF or EX. Violation: FAULT_SLOT. Code may execute only after transfer completion and explicit release by the dispatcher (section 27).
11.8 Branches and Software Functions
BEQ/BNE: A/B sources, C=0, I14 relative to nextPC. No flags are read. CMP plus comparison with a known zero-valued register implements magnitude comparisons.
JMP uses A as the form:
| A | B/C/Imm | Effect |
|---|---|---|
| 0 | bits 21:0 = I22 | relative jump |
| 1 | B=target register, C=Imm=0 | absolute indirect jump |
| 2 | bits 21:0 = I22 | relative; R11 = nextPC |
| 3 | B=target register, C=Imm=0 | indirect; R11 = nextPC |
Other A values are illegal. For form 3, an old R11 used as the jump source
is read before writing the link. All taken forms cost exactly two
additional issue cycles. call label and ret are unambiguous
assembler pseudoinstructions for form 2 and form 1 with B=11.
There is neither a hardware stack nor separate CALL/RET opcodes.
11.9 Waiting and Task Completion
WAIT and WAITUNTIL form the only explicit wait instruction family. They complete once; nextPC is then the resume point. An already present message prevents entry into WAITING. An event arriving simultaneously must not be lost.
WAIT uses A as the form:
| A | Operand | Effect |
|---|---|---|
| 0 | B=C=Imm=0 | wait until Rx is not empty |
| 1 | bits 21:0 = U22 | wait until now+U22 system ticks or a message |
| 2 | B/C GPR, Imm=0 | duration = (u32(Rc)<<32) OR u32(Rb); time or message |
WAITUNTIL: A is low32, B high32 of an absolute system-tick deadline, C=Imm=0. Here too, a message may wake the core early. Comparisons are modulo 2^64; a time interval must be smaller than 2^63. Zero duration or an already reached deadline does not wait. In T1, a system tick is 2 ns (500 MHz), independent of the core clock. An absolute time is read through the system resource; no invented host time influences the guest.
A timer wake is latched as a wake reason, creates no artificial message, and consumes no Rx entry. WAIT observes state; this is not a second external data or notification channel. Software that wants to wait exclusively until the deadline checks the time again after a message wake.
DONE: A=B=C=Imm=0. All older instructions are completed; younger ones are discarded. With open transfers, occupied Tx, or reserved completions, DONE causes FAULT_INFLIGHT. Otherwise the core enters STOPPED. The dispatcher takes over resources and state. STOPPED is not sleeping until the next message. Completion generates TASK_DONE through the same contact. A reserved completion latch retains it until ACK; a new task image may start only afterwards. This latch is not a normal task Tx entry.
11.10 Message Instructions
SEND: A status destination, B destination core, C message value, Imm=0. Resource/system messages cannot be forged through SEND. SIGNAL: A status destination, B destination core, C=Imm=0; value is 1. Success means insertion into the local Tx buffer, not yet acceptance by the destination.
RECV: A status destination, B value destination, C metadata destination, Imm=0. All three destination registers must be different. On success, exactly one FIFO entry is removed: Rb = value32; Rc = source16 | (kind8<<16) | (eventStatus8<<24). Ra = OK. If the FIFO is empty, Ra=EMPTY; Rb/Rc remain unchanged. SEND/SIGNAL may overlap destination and source registers because sources are read first. No message instruction blocks.
Message formats, statuses, and reserved completion slots are defined once in section 27. There is no implicit IRQ routine.
11.11 Resource Instructions
RREAD/RWRITE: A status destination, B resource identifier 0..3, C GPR containing a local descriptor address, Imm=0. The descriptor is 16 bytes, divisible by 4, entirely in a DATA slot, and readable by EX. It is read synchronously; only then may it be changed.
| Offset | Type | Field |
|---|---|---|
| 0 | u32 | offset within the resource |
| 4 | u16 | local SRAM address |
| 6 | u16 | length, 1..1024 bytes |
| 8 | u32 | cookie for completion message |
| 12 | u32 | reserved, must be 0 |
The data resides entirely within a single enabled DATA slot. Descriptor and data must occupy different slots. An RREAD destination is exclusively owned by the contact until completion, as is an RWRITE source. IF/EX may continue using other slots. Hardware atomically validates resource bounds, rights, destination, length, a free tag, and a free completion slot before acceptance. On rejection, the data slot remains free and no completion message is generated.
An accepted request returns Ra=OK and exactly one later completion. Local range/rights errors return a status; invalid instruction fields or a descriptor slot already owned by DMA cause FAULT. RREAD never writes to a register later.
11.12 Timing and Assembler Contract
All normal instructions have an EX issue window of one core cycle. The four-stage pipeline has four cycles of latency from IF to WB, but a throughput of at most one instruction per cycle. A taken branch discards the two younger windows. WAIT and DONE drain the pipeline in order. FAULT is precise.
A descriptor access internally requires a 128-bit read operation from its bank; ordinary LD/ST provide at most 32 bits. This is an explicit area/bank requirement of the theoretical model. A physical design requiring multiple cycles here does not satisfy T1 unchanged.
The assembler supports labels, decimal/hex literals, registers,
symbolic mode names, .code, .data, .align, .byte, .word,
and defined pseudoinstructions:
load32=constant sequence, call/ret=JMP forms, SAR=SHR mode.
It checks all values before encoding and reports address, source line, and
rejection reason. It does not silently expand oversized relative branches;
the compiler explicitly generates load32 plus indirect JMP for them.
Encoding examples for registers R0/R1/R2:
| Assembler | Hex word | Bytes in SRAM |
|---|---|---|
| NOP | 04000000 | 00 00 00 04 |
| MOVI R0, 1 | 08000001 | 01 00 00 08 |
| ADD R0, R1, R2, scalar.wrap | 10048000 | 00 80 04 10 |
| JMP.rel -1 | 50000000 OR 003FFFFF = 503FFFFF | FF FF 3F 50 |
These examples are theoretical encoding requirements, not recorded assembler test runs.
12. Wait-Free Instructions and External Events
Normal arithmetic instructions, local memory instructions, and message operations do not block the pipeline. External resources physically cannot guarantee an immediate response. Therefore:
RREADandRWRITEspecify a local SRAM address and a length.- An
RREADis sent only if a free transaction tag exists. The response slot is therefore already reserved when the request is made. - Response data lands directly in the specified local slot, never asynchronously in a register. Late responses cause no register hazards.
- Completion generates a message in the Rx FIFO.
SEND,RECV,RREAD, andRWRITEimmediately return a status in a register. A full or empty FIFO and a missing tag do not block the instruction.- Software may then deliberately wait, process another task, or retry later.
- Only the explicit WAIT family puts a task to sleep; messages and the documented timer conditions may wake it.
This is a change from the v4.1 emulator. There, RECV,
RREAD, and RWRITE reset the program counter when data is missing and
put the core to sleep.
13. Messages and Synchronization
13.1 Buffers per Core (Decided)
| Structure | Size | Meaning |
|---|---|---|
| Rx message FIFO | 16 | shared acceptance order |
| normal Rx quota | 8 | USER/system input |
| reserved completion quota | 8 | open transfers and unread completions |
| Tx buffer | 8 | outgoing messages until ACK |
| transaction tags | 8 | RREAD/RWRITE including completion consumption |
The exact quota, reservation, and status semantics are defined in sections 27.1 and 27.5. Full structures produce a local status; they do not stop any normal instruction.
13.2 Delivery
SENDreports success when there is space in the Tx buffer.- Each occupied Tx entry carries a unique message identifier until completion. ACK and NACK contain this identifier.
- The entry remains occupied until the destination acknowledges with ACK. A retry uses the same identifier; the destination must not insert an already accepted message into its Rx FIFO a second time.
- If the destination FIFO is full, the destination responds with NACK. Hardware automatically retries delivery with fair arbitration.
- Messages from the same sender to the same destination are accepted in sending order. There is no global order between different senders.
- Messages are rejected at the destination rather than queued in the network. They never permanently block links. This rules out protocol deadlock due to a full destination FIFO. Logical software deadlock, where tasks wait for each other's events, remains possible and must be prevented by the compiler or runtime system.
13.3 Use
Messages are used for:
- task completion,
- data ownership changes,
- events,
- timers,
- DMA and transaction completion,
- resource readiness,
- error reports (
FAULT), - waking a sleeping core.
RAM regions writable by multiple participants should not be synchronized through implicit cache or lock logic. The preferred sequence is:
- The compiler or dispatcher assigns a region to a writing task.
- The task writes its data.
- A completion event transfers ownership.
- Subsequent tasks read or take over the region.
Atomic RAM operations are not yet decided. They should be introduced only if messages and static data ownership are demonstrably insufficient.
14. Uniform Contact
14.1 Basic Parameters (Decided)
Data width: 64 bits
Contact clock: 500 MHz
Raw bandwidth: 4 GB/s per link and direction
Maximum transaction: 1024 bytes
Header per packet: 16 bytes
Payload per packet: at most 64 bytes
Full transfer data direction: 16 x (2+8) = 160 beats = 320 ns
Full-data payload bandwidth: at most 3.2 GB/s before arbitration
Address width: 40 bits
Smaller transfers with byte precision are allowed. Final beats do not become visible as additional payload bytes. A full 1-KiB transfer corresponds to one slot, but shorter transfers also reserve the entire affected bank.
14.2 Channels
Request and response are separated logically and in terms of credits. Source, destination, address, tag, length, class, and status reside in the uniform header. Its complete bit allocation is defined once in section 27.7. Messages use the request channel, ACK/NACK the response channel. Pure control responses have a reserved opportunity to make progress.
14.3 Packets and Virtual Channels (Decided)
Transactions are split into at most sixteen packets, each with up to 64 payload bytes. With a 16-byte header, the header accounts for 20 percent of transmitted bytes in full data packets.
| Virtual channel | Contents |
|---|---|
| request, real-time | classes 0 and 1 |
| request, normal | classes 2 and 3 |
| response, real-time | classes 0 and 1 |
| response, normal | classes 2 and 3 |
Real-time traffic may interrupt a normal packet at beat boundaries; packet assembly and credits remain separate for each VC. Four flit slots of eight bytes per VC yield 128 bytes of link buffering per direction, in addition to header/assembler/tag state. This is not a total area calculation for a router port.
In the 64-core target system, the crossbar does not use XY mesh routing. XY routing belongs exclusively to future systems (section 23). Protocol deadlock freedom requires separate credits, consumable control responses, and the admission rules in section 27; four VC names alone do not prove it.
14.4 Real-Time Classes
Intended order:
| Class | Use |
|---|---|
| 0 | audio, hard real-time, error reports |
| 1 | display scanout |
| 2 | RAM DMA and normal task data |
| 3 | disk, USB, network, and background traffic |
Audio and display receive guaranteed time slots. Free time slots are used by lower classes.
15. Shared RAM
15.1 Capacity (Decided)
The target system has 256 MiB of shared RAM, or 4 MiB per core, supplied by exactly one DRAM component (section 15.2).
A useful measure of size is how quickly RAM can be filled in practice: a USB-2 flash drive supplies about 35 MB/s and fills 256 MiB in about 7 seconds.
The 4 MiB per core is a capacity figure. It does not create fixed private partitions for each core.
15.2 DRAM Component (Decided)
LPDDR4 or LPDDR4X is decided as a single-channel component with
2 Gbit x16 in a single die (SDP). The reference component is a
Winbond W66BP6NB in the x16 VFBGA100 version; an ISSI
IS43LQ16128A or its LPDDR4X variant is a second source to be
qualified.
Organization: 1 channel x 16 DQ x 8 banks, 2 Gbit = 256 MiB
Package: JEDEC-100-Ball-BGA
Target speed grade: 3200 MT/s, i.e. 6.4 GB/s theoretical peak
Operation: 3200 MT/s; for 3.2 GB/s full-data contact payload,
at least 50 percent DRAM efficiency is required
Reference: Winbond W66BP6NB, x16, VFBGA100, 3200 MT/s
Second source: ISSI IS43LQ16128A/AL, x16, BGA100, 3200 MT/s
Rationale:
- 2 Gbit gives exactly the target size of 256 MiB. Smaller densities exist from individual manufacturers but do not meet the decided capacity.
- One channel with 16 bits requires only about 34 to 38 signal pins, fewer than DDR3 or DDR4 x16.
- The theoretical DRAM peak exceeds the contact maximum. Sufficient sustained bandwidth for scanout and computation is a model/controller condition, not evidence already established.
- Winbond and ISSI offer suitable LPDDR4 families with long-term or industrial variants. The specific speed grade and temperature range must be checked against the orderable part numbers available at the time before layout approval.
- DDR3 was not selected as the target interface because of its higher pin count, higher I/O power, and long-term procurement risk.
- PSRAM (HyperRAM/OctalRAM) was rejected because configurations available with suitable capacity and pin count do not achieve the intended payload bandwidth.
LPDDR4 requires 1.8 V and 1.1 V supplies (LPDDR4X: VDDQ 0.6 V) and a memory-controller training phase at startup. The DRAM PHY is licensed IP (section 21).
Price, package option, temperature class, and availability must be checked again before manufacturing. The architecture depends only on 2 Gbit, x16, and at least 3200 MT/s, not on a single orderable part number.
15.3 Full Access
All cores must be able to address the entire shared RAM. Physical bank or channel organization must not partition the visible address space.
The compiler and dispatcher allocate dynamic RAM resources with base, length, and access rights. Within an assigned resource, a core uses a 32-bit offset. The contact can nevertheless transport physical addresses of at least 40 bits.
15.4 CPUlet-SRAM (Future Option)
T1 reserves an address window but has no additional CPUlet-SRAM in the standard profile: size 0 bytes. Tables, reference images, and shared structures reside in DRAM. The cores' local SRAM remains unchanged.
256/512 KiB may be investigated in a later profile revision. Even then, the same contact, explicit transfers, and ownership rules apply; no cache is created. A fixed end-to-end latency must not be inferred from an SRAM access time: the router and competing requests remain relevant. Area and benefit are demonstrated only in the corresponding profile.
16. Graphics Output
16.1 Decided Maximum
Maximum output: 1280 x 720 at 60 Hz
Maximum internal color depth: 32 bits per pixel
1080p is intended for future multi-CPUlet systems (section 23). 1440p and 4K are not targets.
16.2 Framebuffer
The framebuffer must be a normal region of shared RAM. There is no separate VRAM and no parallel graphics memory path.
| Buffer at 720p | Size |
|---|---|
| 32 bits, one image | 3.5 MiB |
| 32 bits, double buffer | 7.0 MiB |
| 32 bits, triple buffer | 10.5 MiB |
| RGB565, double buffer | 3.5 MiB |
| Metric at 720p, 32 bits, and 60 Hz | Value |
|---|---|
| Scanout | 221 MB/s |
| 1-KiB transfers per image | 3600 |
| Link share for scanout alone, against 3.2 GB/s payload maximum | 6.9 % |
| Link share for scanout plus rewriting, same bottleneck direction | 13.8 % |
The sum is a link share only if the bottleneck direction is shared: read/write data may physically use different directions. At DRAM, both are data traffic; headers and pauses are counted separately.
A 720p triple buffer occupies about 4 percent of the 256 MiB.
16.3 Display Unit
The display unit is not a GPU replacement. It handles only deterministic transport:
- framebuffer start address,
- line length,
- resolution,
- pixel format,
- switching image buffers during vertical blanking,
- request FIFO and line buffers,
- TMDS encoding and serialization for DVI-D.
The display unit is designed for single-link DVI up to a 165 MHz pixel clock. 720p is the deliberately decided product and quality limit of the target system. The PHY could technically transport 1080p60 as well; however, the target system still lacks measured end-to-end budgets for renderer, RAM, contact, display, and simultaneous media load.
Tiles, sprites, text, composition, scaling, and effects remain core tasks. T1 has no hardware scaler. An application scales explicitly through guest code into the completed framebuffer; costs and filter selection belong to its budget.
16.4 Digital Output: DVI-D
Single-link DVI-D (TMDS) is technically decided. Approval for a production product remains subject to a final legal and compliance review:
3 differential data pairs
1 differential clock pair
= 8 signal pins
plus DDC (2 pins, I2C for EDID) and hotplug (1 pin)
| Mode | Pixel clock | Bitrate per data pair |
|---|---|---|
| 720p60 | 74.25 MHz | 742.5 Mbit/s |
| 1080p60 (future) | 148.5 MHz | 1.485 Gbit/s |
Rationale:
- No external protocol converter is intended. TMDS encoding is digital logic; however, serializer, line driver, ESD protection, and controlled output impedance form a characterized high-speed I/O macro and are part of the PHY and signal-integrity design.
- Interoperability with HDMI inputs via a passive DVI-HDMI adapter is a development goal, not a guarantee for every display device. EDID is read; unsupported modes are not output.
- Publication of the DVI-1.0 specification itself grants no IP license, but refers to a reciprocal royalty-free Adopter Agreement. Patent status, applicability of this agreement, and trademark use must be legally clarified before production approval.
- The device uses a DVI connector or a DVI-HDMI cable, not an HDMI connector. The HDMI specification, HDMI trademark, and HDMI logo remain excluded. DisplayPort remains excluded because of membership, document, and compliance dependencies.
- Audio is not transmitted over DVI; it has separate pins (section 17).
VGA was rejected: about 20 pins for 18-bit color through resistor ladders, and modern monitors require active adapters. LVDS/OpenLDI was rejected because it requires a dedicated receiver at the monitor.
1440p60 (about 241 MHz) would require dual-link DVI and does not work through passive HDMI adapters. It is at most an option for dual-link DVI monitors.
16.5 Rendering Model for 3D (Target Proposal)
The target is 3D in the class of Tomb Raider 1 at 720p, not modern 3D environments.
Estimate for 720p60: 1280 x 720 pixels with overdraw 2 yield 110.592 million pixels per second. An assumed 10 to 15 instructions per textured pixel yield 1.106 to 1.659 GInstr./s, mathematically about 5 to 7 fully utilized RUN cores. Six to eight with reserve remain a hypothesis. Geometry, clipping, tile lists, texture misses, and game logic must be budgeted additionally.
The bottleneck is texture access, not instructions. An RREAD per pixel
is ruled out. The renderer therefore works in tiles. Example
slot allocation for a core:
| Slots | Contents |
|---|---|
| 4 KiB | image tile 32 x 32 at 32 bits |
| 2 KiB | tile Z buffer or polygon sorting |
| 2 KiB | current texture window 32 x 64 with an 8-bit palette |
| 3 KiB | code |
| 1 KiB | stack, descriptors, palette, and small state |
- Completed tiles are written to the framebuffer using
RWRITE. - The texture working set resides in DRAM; windows are loaded locally. The example allocation has no free double-buffer slot. Transfer and computation must therefore occur in phases, or the tile must be reduced.
- Lane modes (4 x 8 bits) serve Gouraud shading and blending.
- Perspective correction uses
CLZ, a reciprocal table, and a Newton step.
17. Audio
Audio operates as a hard real-time stream. Up to 8 channels (7.1) are intended.
Digital chip output without an external audio DAC (decided):
- PDM: one sigma-delta bitstream per channel. Interpolation (CIC), sigma-delta modulator, and the temporally characterized output cell reside on the chip.
- 7.1 requires 8 pins.
- In software, the modulator at about 3 MHz per channel would occupy nearly an entire core.
- The cores supply completed PCM samples through DMA.
- A passive RC low-pass filter provides only an analog experimental signal suitable for a high-impedance load. It is not a guaranteed line output and can drive neither headphones nor speakers. An analog output capable of driving its load requires at least a low-pass filter, buffer or amplifier, and protection circuitry outside the chip.
- Achievable quality is determined by modulator order, oversampling, output cell, I/O supply, clock jitter, filter, and load. The audio pins receive a separate filtered I/O supply; dynamic range, noise, and distortion are measured on the prototype rather than promised in advance as a bit count.
Alternative mode of the same audio unit on the same pins: I2S/TDM (TDM8 on one data line) for boards with an external hi-fi DAC.
S/PDIF was rejected because it carries multichannel audio only in compressed, license-requiring formats.
Audio bandwidth is small compared with graphics. Eight channels at 24 bits and 192 kHz require about 4.6 MB/s; with 32-bit transport, about 6.1 MB/s. Audio nevertheless receives the highest traffic class because latency and interruptions matter more than absolute data volume.
One IDLE core for RAU or RFXL is a previous planning hypothesis. Rate, image size, deadline, slot/RAM requirements, and guest instruction costs must be demonstrated according to section 29.
18. USB and Mass Storage
The target platform uses USB for input devices, MIDI, and mass storage. SATA, PCIe, USB 3, USB4, and Thunderbolt are omitted.
There is exactly one USB host logic with multiple ports and two PHY classes.
18.1 USB Low-/Full-Speed with Integrated PHY (Decided)
- Low- and Full-Speed (1.5 and 12 Mbit/s) use two data pins per port and an integrated, characterized LS/FS transceiver cell. It ensures, among other things, levels, receiver thresholds, output impedance, edge shape, and SE0 detection. Passive components and ESD protection required by the USB design are added on the board.
- Hardware: bit level, i.e. NRZI, bit stuffing, CRC5/CRC16, 1-ms SOF timing, and packet buffers.
- Software on the system cores: transaction and device level (HID, MIDI).
- One port per device. No hub chip is required.
- Target system: 3 ports for mouse, keyboard, and MIDI.
- The 5-V VBUS supply, current limiting, and overcurrent detection reside in external load switches or protection components and are controlled and queried by the chip.
18.2 USB 2.0 High-Speed (Decided)
- One High-Speed port (480 Mbit/s) for mass storage, typically a USB flash drive.
- High-Speed requires an analog PHY. It is integrated as licensed IP to avoid an additional chip. A USB-2 PHY is mature and substantially simpler than a USB-3 PHY.
- The same port also supports Full-Speed.
- Actual flash-drive throughput: about 35 MB/s.
- A USB-2 flash drive can operate on an LS/FS port in Full-Speed mode. Its usable throughput there is typically well below 12 Mbit/s; this suffices for many compressed media streams, but not for quickly loading large programs.
USB must block neither display scanout nor audio. Bulk transfers for disk and network run in the lowest contact class.
18.3 Real-Time Decoding of Compressed Data
Compressed data is decoded directly during reading:
Drive --USB--> RAM ring buffer --Contact cl. 2--> Decoder core slot --> Audio / Framebuffer / RAM
(cl. 3) (double buffer: slot A fills, slot B decodes)
- A configured RAM ring buffer of, for example, 1 to 2 seconds bridges bounded storage pauses. It provides no guarantee against unbounded USB stalls. The profile specifies bitrate, buffer bytes, and maximum permitted pause.
- Parallelization follows the existing bitstream boundaries and dependencies of each format family (section 29). RAU frame copies, RFXL predictors, and PMF0 history are not treated as independent. Decoder and USB can both be bottlenecks.
- Example: 35 MB/s compressed at a factor of 3 yields over 100 MB/s of payload.
19. Error Handling, Boot, and Debug (Decided)
There are no interrupts or exceptions in the classical sense.
Fault causes:
- invalid opcode,
- access outside an assigned resource,
- slot ownership conflict.
Sequence:
- The core enters the
FAULTstate; its state is preserved. - It sends a class 0 fault message to a designated system core. The message value contains cause and flags.
- The system core reads PC, registers, and status through status registers addressable through the contact.
The same mechanism serves boot and debug: a system core can halt, inspect, load, and start any core through the contact. There is no parallel debug path.
20. Package and Pins
Preliminary minimum estimate for 64 cores, one LPDDR4 component, DVI-D, PDM audio, and USB:
| Area | Pins or balls |
|---|---|
| LPDDR4, one channel x16 | about 34 to 38 |
| USB 2.0 High-Speed including port control | about 4 to 5 |
| USB Low-/Full-Speed, 3 ports including VBUS control | about 9 to 12 |
| DVI-D including DDC and hotplug | 11 |
| Audio PDM or I2S/TDM | 8 |
| Clock and reset | about 3 |
| Boot and minimal debug | about 2 to 4 |
| Calibration and references | about 2 to 4 |
| Functional signals | about 73 to 85 |
| Power and ground | about 40 to 55 |
| Total | about 113 to 140 |
A 144-ball BGA remains the preferred lower bound, but this functional calculation leaves only about 4 to 31 spare balls. Whether power, return-current paths, LPDDR4 escape routing, and separation of sensitive I/O domains can be accommodated reliably must be shown by a specific package and pinout study. If the reserve is insufficient, 169 BGA is the next target size; this does not change the architecture itself. The core signals do not leave the chip.
21. Rough Area and Energy Guidance
Without a target process, SRAM macros, synthesis, and physical layout, no reliable area or power figures are possible.
An earlier rough 28-nm estimate for a core still envisaged as capable of 750 MHz yielded:
- about 35 to 60 kGE of logic,
- about 0.06 to 0.12 mm2 including 12 KiB SRAM,
- about 3 to 8 mW at 250 MHz as a rough activity scale.
The final design limited to 250 MHz should become smaller and more energy-efficient because the fast clock tree, 1.33-ns critical paths, and 750-MHz voltage level are eliminated. A CPUlet-SRAM with 256 KiB would add roughly 0.4 mm2. These values must not be used as specifications until actual PPA synthesis has been performed.
Regarding the manufacturing process: open PDKs (SKY130, GF180MCU, IHP SG13G2) fit the open design ambition, but barely reach 250 MHz with single-cycle multiplication, and for 64 cores yield roughly 80 to 170 mm2 and do not offer the required LPDDR4, USB-2-HS, USB-LS/FS, and TMDS I/O macros as ready-made, characterized standard solutions. A commercial process around 28 nm with licensed or specifically qualified PHY IP is realistic. ISA, RTL, and protocols can remain open; manufacturing, PHY macros, and the use of external standards must be assessed separately.
22. Future Physical Prototypes (Outside the T1 Lab)
An FPGA is a possible future investigation before an ASIC. The current project scope remains complete theory followed by a Mac lab app; an FPGA is neither purchased nor required for it. The following historical guidance figures are not a promise.
| FPGA | Toolchain | Cores (estimate) | Clock (estimate) |
|---|---|---|---|
| Lattice ECP5-85 | fully open | about 16 to 24 | about 80 to 120 MHz |
| AMD Artix-7 200T | Vivado, free of charge | about 32 to 48 | about 100 to 150 MHz |
| Kintex class | Vivado | 64 | - |
The estimate is based on 2 to 4 k LUTs per core and must be confirmed by the first synthesis. A partial configuration with 16 to 32 cores is sufficient because the ISA is independent of core count.
Rules:
- The contract applies in cycles: one instruction per cycle and fixed branch times. With proportional scaling, RUN:Fabric=1:2 and IDLE:Fabric=1:10 remain, for example 20/100 MHz for IDLE/RUN and 200 MHz fabric. Instruction budgets then remain comparable; actual RAM/PHY latencies, image/sample rates, and absolute real-time deadlines must be budgeted anew.
- FPGA boards usually have DDR3, and ECP5 and Artix-7 do not directly support LPDDR4. The prototype may therefore use DDR3. The memory controller sits behind the contact; the contact contract is identical.
- DVI, PDM audio, and USB LS/FS can be functionally prototyped with suitable FPGA I/O cells. This does not yet demonstrate USB compliance or TMDS signal integrity; if the board I/Os do not meet the electrical requirements, an external transceiver solution is used for the prototype. USB 2.0 High-Speed requires an external PHY chip with a ULPI connection on the FPGA.
- The emulator is first brought into line with the target behavior in this document and remains the sole reference model. RTL is checked against it in co-simulation.
- The FPGA is used to measure whether CPUlet-SRAM is needed and, if so, at what size.
23. Future: Multi-CPUlet Systems (Open)
This section is not part of the target system. It collects previous considerations so that current decisions do not prevent future larger systems. None of this is decided.
Guidelines for the target system:
- The ISA remains independent of core count.
- The sender identifier of messages (16 bits) leaves room for a CPUlet identifier.
- The contact transports addresses of at least 40 bits.
- The contact contract also applies unchanged between CPUlets.
23.1 Macroblock and Maximum Configuration
A macroblock arranges 16 CPUlets as a 4-by-4 mesh (1024 cores). Each CPUlet receives four additional mesh ports (north, south, east, west); the crossbar grows from 10x10 to 14x14. Simple deterministic XY routing is preferred. Four macroblocks (2 x 2) yield the previous upper limit of 4096 cores.
| Cores | CPUlets | Macroblocks | RAM at 4 MiB per core | Total local SRAM | Display output |
|---|---|---|---|---|---|
| 64 (target system) | 1 | - | 256 MiB | 768 KiB | 720p60 |
| 256 | 4 | - | 1 GiB | 3 MiB | 1080p60 |
| 1024 | 16 | 1 | 4 GiB | 12 MiB | 1080p60 |
| 4096 | 64 | 4 | 16 GiB | 48 MiB | 1080p60 |
Theoretical peak performance at 250 MHz: 64 GInstr./s with 256 cores, 256 GInstr./s with 1024, and 1024 GInstr./s with 4096 cores.
23.2 RAM and Banking
A single RAM port is insufficient for large systems. Previous consideration for a 1024-core macroblock:
- 16 logical RAM banks,
- multiple physical memory controllers,
- interleaving at 1-KiB slot boundaries,
- every address reachable from every CPUlet.
As an initial working assumption, four physical memory controllers with four logical banks each were considered. This means 256 rather than 64 cores share a controller; bandwidth per core drops to one quarter of that in the target system. In addition to capacity, a bandwidth rule per core would therefore need to be defined.
23.3 Input and Output
- From 4 CPUlets onward, 1080p60 is intended through the same single-link DVI display unit. A 1080p framebuffer at 32 bits: 7.9 MiB per image, scanout 498 MB/s, 8100 1-KiB transfers per image, 15.6 percent of the 3.2-GB/s full-data payload maximum without additional load.
- Larger variants require larger packages because of more memory channels and higher current.
24. Differences Between the v4.1 Emulator and Hardware Target
The emulator is a valuable functional starting point, but not a cycle-accurate hardware model.
| Topic | v4.1 emulator | Target hardware |
|---|---|---|
| Core count | models 4096 | 64, one CPUlet; multi-CPUlet systems are future work |
| Cluster | 64 cores | 64 cores equal one CPUlet, 8 groups of 8 |
| Clocks | symbolic ULTRA/MIDI/TURBO |
0/50/250 MHz, no 750 MHz |
| Execution | one call per simulation quantum | one instruction per active cycle |
| Pipeline | not cycle-accurate | four stages; untaken branch 0, taken branch 2 additional issue cycles |
| Local memory | 12 KiB, lazily allocated | 12 x 1 KiB banks with slot ownership |
| External RAM | not a core component | shared RAM, 256 MiB |
| Load/store | separate 16/32-bit opcodes | unified 8/16/32 bits |
| Constants | MOVI with 14 bits |
MOVI with 22 bits plus MOVHI |
| Resources | block and sleep | nonblocking, response in local slot, completion as a message |
| Messages | FIFO without delivery acknowledgment | Tx buffer, ACK/NACK, automatic retry |
| Errors | not modeled | FAULT plus fault message |
| Graphics | lab/audit environment | DVI-D display unit with RAM framebuffer, 720p60 |
| Audio | model resource | PDM or I2S/TDM, up to 7.1 |
| Peripherals | model resources | uniform contact, integrated USB-LS/FS and USB-2-HS PHYs |
New hardware decisions should therefore be reflected first in this specification and then specifically in the emulator. Old behavior must not persist as a second permanent architectural path.
25. Decided Theory and Open Physical Validation
ISA, compiler concept, ABI, resource/message contract, format profiles, and lab contract are decided for T1 in sections 9..14 and 27..31. The theoretical specification no longer leaves an optional LOOP, unknown signed MUL, or unspecified transfer format open.
The following remain open for a future physical implementation:
- Process, cells, SRAM/128-bit descriptor banks, and PHY/I/O macros.
- Timing of MUL/MAC, SRAM, forwarding, and resource acceptance.
- Clock domains, reset/wake transitions, and power supply.
- DRAM controller, training, refresh, and sustained bandwidth.
- Electrical boot source and package/pinout study.
- USB/TMDS compliance and PDM output quality under load.
- DVI legal/Adopter review before production approval.
- Measured area, power, and temperature budgets.
Compiler, guest codecs, and reference model still need to be implemented or aligned according to this theory. Their functional, timing, and worst-case validation follows the contract in section 31; this work has not been completed in this round.
26. Compact Target Summary
The uRISC target system is a 64-core system consisting of one CPUlet with eight groups of eight cores. Four cores are reserved for system and I/O tasks, 60 for applications. Each core has twelve 32-bit registers, four resource connections, a 64-bit MAC accumulator, a message FIFO, and twelve local 1-KiB slots with exactly one owner per cycle. Cores operate at 0, 50, or at most 250 MHz.
The four-stage pipeline executes one instruction per cycle. Normal instructions
are deterministic and do not wait for each other; taken branches cost a fixed
two additional issue cycles. External work is organized through nonblocking
requests with reserved response space. Responses land in local SRAM;
completions, events, and errors arrive as messages, for which only an
explicit WAIT waits. The compiler splits programs into tasks, slots,
and data regions during compilation itself.
Shared RAM comprises 256 MiB, or 4 MiB per core, from a single LPDDR4 component with 2 Gbit x16. All cores can address the entire RAM. The framebuffer resides exclusively in this RAM. CPUlet-SRAM is reserved as an address window but in the T1 standard profile is 0 bytes.
A uniform 64-bit contact at 500 MHz, with transactions up to 1024 bytes, 64-byte packets, and four virtual channels, connects core groups, RAM, DMA, display, audio, and USB. Audio and display receive guaranteed real-time windows.
The target system outputs at most 720p60 through DVI-D. Interoperability with HDMI inputs through a passive adapter is intended and verified using EDID and prototypes; production approval requires legal and electrical DVI review. Audio runs at up to 7.1 as PDM without an external DAC or alternatively as I2S/TDM; an analog output capable of driving its load requires external filter and output stages. Mouse, keyboard, and MIDI connect to three ports with integrated USB-LS/FS PHYs, mass storage to one port with an integrated USB-2.0-High-Speed PHY. Proprietary formats execute according to their existing stream/packet/root boundaries and dependencies. HDMI, DisplayPort, SATA, PCIe, USB 3, USB4, and Thunderbolt are not part of the target architecture.
Systems with multiple CPUlets up to 4096 cores remain open and future work (section 23). They should use the same ISA, cores, and contact contract.
27. System, Memory, and Event Contract (Normative)
27.1 Configuration, Numbering, and Limits
The T1 standard profile has 64 cores and 256 MiB RAM. The visible
8x8 arrangement numbers row y and column x as coreID=8*y+x.
Each fabric group is a row of eight cores.
The screen position is a representation, not an eight-way meshed
physical connection. Cores 0..3 are system cores.
| Quantity | Decided value |
|---|---|
| local bytes per core | 12288 |
| GPRs | 12 x 32 bits |
| ACC | 64 bits |
| resource bindings | 4 |
| local slot count | 12 |
| Rx FIFO | 16 x 64 bits |
| of which normal input slots | 8 |
| of which completion slots | 8 |
| Tx buffer | 8 entries |
| simultaneously open RREAD/RWRITE | 8 |
| transaction identifier on contact | 16 bits |
| shared RAM size | 268435456 bytes |
| CPUlet-SRAM in standard profile | 0 bytes |
| system tick | 500000000 ticks/s |
| RUN / IDLE clock | 250 / 50 MHz |
An Rx FIFO contains the combined accepted sequence of both classes. The quotas are admission limits, not two separately readable FIFOs. A resource request reserves one of the eight completion slots. The slot remains occupied until RECV consumes the completion, even if the transaction has already completed. There are therefore never more than eight unconsumed resource completions.
27.2 Visible States and Reset
The core has: OFF, READY, RUNNING, WAITING, STOPPED, and FAULT. OFF is power-/clock-gated and cannot be awakened by normal messages. READY is prepared but not yet started. RUNNING operates at the configured frequency; IDLE is a frequency, not a state. WAITING retains registers, ACC, PC, bound resources, and all required slots. A message or timer can wake it. STOPPED and FAULT require an explicit dispatcher action.
Reset sets GPRs, ACC, PC, FIFO/Tx/tag state, and timers to zero, revokes all resource rights, and sets cores 1..63 to OFF. Core 0 starts only after a validated boot task image is loaded. SRAM and DRAM may be physically uninitialized; the loader initializes every readable region before enabling it. The model treats access to uninitialized regions as FAULT_UNINIT. It supplies no host-dependent random bytes.
The theoretical boot path is an externally supplied boot image: loader -> local code of core 0 -> RAM initialization -> system tasks 1..3 -> application tasks. The eventual electrical boot source has no influence on the theoretical specification; the app will load the same validated image across the host boundary.
27.3 Slot Rights, Owners, and Code Changes
Each slot has the role CODE or DATA and an initialization state. A role is set only by the dispatcher while the core is halted. DATA may additionally be DMA-owned during a transfer. The owner per cycle is IF, EX, or DMA, never more than one.
Persistence means retaining contents, not granting access to others. A WAIT with open transfers must lose neither descriptor/data state nor Rx/tag state. A transition to OFF requires that no transfers or message acknowledgments remain open. Control logic for wake and reception remains reachable in WAITING.
Self-modifying code is not allowed in a running task. Code overlay: task halts -> dispatcher waits for open transfers -> affected bank becomes DATA -> RREAD and completion -> code validation -> CODE role -> start at an enabled PC. No asynchronous code change during IF.
The sum of code, data, stack, and DMA banks must fit in twelve slots. A 12-KiB working set with no room for code/stack is not a valid task image. The CPUlet's total SRAM capacity is 768 KiB.
27.4 Resources and Access Rights
S0..S3 are opaque bindings to validated resource entries. A task cannot change their bits with MOV. The dispatcher binds base address, length, READ/WRITE rights, generation, owner, and traffic class. The physical base has up to 40 bits, the descriptor a 32-bit offset. A T1 resource window is at most 256 MiB.
An offset/length pair is checked with offset <= size and
length <= size-offset, never through an overflowing addition.
A relinquished or reassigned resource receives a new generation.
Generation exhaustion retires the ID instead of wrapping and revalidating an
old token. The same rule applies to task, endpoint and destination generations.
Old bindings return CAPABILITY and must not trigger RAM access.
Resources are RAM windows or memory windows of a device service. Control commands for display/audio/USB are written as validated descriptors into such windows. A core has no second unrestricted MMIO access and cannot arbitrarily address other cores or device registers.
System resources for debug, clock, and task control are bound exclusively to system cores or the halted host loader. Application code receives only the required data/service windows. A classical MMU is not required: LD/ST remain local, external accesses are exclusively capability-checked.
27.5 Status and Messages
Instruction status is a u32:
| Value | Name | Meaning |
|---|---|---|
| 0 | OK | accepted or successfully read |
| 1 | BUSY | Tx, tag, completion quota, or required data slot occupied |
| 2 | EMPTY | RECV without a message |
| 3 | TARGET | destination identifier invalid or destination permanently unreachable |
| 4 | DESCRIPTOR | length, alignment, or local descriptor region invalid |
| 5 | CAPABILITY | resource unbound, stale, or lacking rights |
| 6 | RANGE | resource range invalid |
| 7 | IO | accepted transfer fails at the device |
| 8 | CANCELLED | accepted transfer terminated by system cancellation |
Normal rejections change only the status register and nextPC. An accepted transfer reports a later IO/CANCELLED completion instead of a second local return value. After a partial error, the entire RREAD destination must be treated as invalid; no partial success is assumed. RWRITE is not atomic in external RAM.
The 64-bit message consists of source16, kind8, status8, value32. USER=0, READ_DONE=1, WRITE_DONE=2, FAULT_EVENT=3, SYSTEM=4, TASK_DONE=5. USER has status zero; value is the SEND value. READ_DONE/WRITE_DONE carry the descriptor cookie as value and the completion status. FAULT_EVENT has source=coreID, status=0, and value=fault code; the privileged debug service reads fault PC and instruction. TASK_DONE has source=coreID, status=0, and value=task ID. Core sources 0..63 are USER/FAULT_EVENT/TASK_DONE; resource sources FF00..FF03 denote the bound S0..S3, FFFE the dispatcher. The remaining identifiers are reserved for future systems.
Cookies are unique among a task's open requests. The tag is reused only when transfer and completion consumption have ended. A data slot is released before its completion message becomes visible. The core may read or transfer it again only after RECV of the matching successful completion.
ACK/NACK are fabric protocol, not USER messages. An ACK confirms insertion, not RECV. Tx sequence identifiers are not reused while outstanding. A destination contains a finite dedup table for the outstanding windows of all sources; the specific memory organization is an implementation matter. Source-destination order is maintained even with NACK and retry. T1 models internal links as lossless: an accepted packet does not disappear, ACKs are not lost. A diagnosed link/model error ends the experiment with DEVICE/INTERNAL. Retry is triggered only after NACK, not by an invented ACK timeout. Sequences run modulo 65536 per source-destination pair; at most eight are outstanding, and the destination retains the last eight accepted sequences as its dedup window. Old in-flight packets must not survive a destination generation change. A later model with packet loss requires its own protocol revision and must not silently replace this assumption.
A faulted or stopped receiver is not blindly sent data forever: system cancellation invalidates the destination generation and terminates affected Tx entries with a diagnosable error. Such an error causes FAULT_DELIVERY at the sending core because SEND has no reserved register slot for a later return value. Fault state is latched separately and cannot be lost because of a full Rx FIFO.
27.6 Ordering, Visibility, and Completion
There is no general implicit RAM fence. The contract is:
- A successful RWRITE completion means all bytes are visible to subsequent accesses by other masters.
- Only then does a message transfer data ownership.
- The receiver consumes the message and may start RREAD.
- A successful RREAD completion means all destination bytes have been written locally and are visible.
Requests from a core may complete out of order if there is no dependency; cookies identify them. Two simultaneously writing tasks with overlapping regions are an invalid ownership plan. An RWRITE that is only buffered in the controller must not yet trigger a visibility ACK.
Acceptance, linear byte access, and global transaction atomicity are different. A 1024-byte transfer is not atomic; data ownership prevents observation of partly written regions. Atomic RAM ALU instructions are not part of T1.
27.7 Contact Packet and Actual Payload Bandwidth
The raw bandwidth of 64 bits x 500 MHz is 4 GB/s per direction. For T1, a packet header of 16 bytes is decided:
| Bit range of the 128-bit header | Field |
|---|---|
| 39:0 | physical address or service offset |
| 55:40 | source |
| 71:56 | destination |
| 87:72 | tag/sequence |
| 94:88 | payload length 0..64 bytes |
| 98:95 | kind |
| 100:99 | traffic class 0..3 |
| 105:101 | status |
| 109:106 | packet index 0..15 |
| 113:110 | packet count minus 1 |
| 121:114 | byte mask |
| 127:122 | flags, in T1 only bit 122=LAST allowed |
Kind: 0 READ_REQ, 1 READ_DATA, 2 WRITE_DATA, 3 WRITE_ACK, 4 MESSAGE, 5 ACK, 6 NACK, 7 ERROR; 8..15 reserved. The byte mask applies only to a short single beat; for more than eight payload bytes it is FF, and payload length limits the final beat. Zero payload length applies only to pure control packets. For READ_REQ, length denotes the requested packet payload; no payload beats follow there. All other data types provide exactly ceil(length/8) beats; unused final bytes are zero.
A full data block has 2 header beats + 8 payload beats = 10 beats = 20 ns. A 1-KiB transfer has 16 data packets = 160 beats = 320 ns without arbitration/response. Its full unidirectional bandwidth is therefore 3.2 GB/s rather than 4 GB/s; headers are 20 percent of total bytes. READ requests and WRITE completions additionally occupy the opposite direction. Small transfers, control traffic, and arbitration further reduce payload bandwidth. Raw and payload values are never equated.
Real-time traffic may interrupt normal data only at beat boundaries. An interrupted packet retains parser state, VC, and tag; the other class has separate packet assembly. Packets of the same VC are not interleaved with each other in the middle of a payload. Request/response have separate credits. An endpoint must be able to consume pure ACK/NACK/ERROR even under data pressure.
27.8 Arbitration and Real-Time
Group arbiters use round robin per class, routers also per output. In 100 fabric ticks, the standard model reserves: 4 ticks for class 0, 12 for class 1, 64 for class 2, 20 for class 3. Unused ticks are available to other classes; minimum quotas remain under competing load. Audio/display data uses packets of at most 64 bytes. Real-time-class control traffic is rate-limited so that faulty software cannot claim the media quota without bounds.
At a continuously usable output, the reservation theoretically provides up to 128 MB/s for class 0 and 384 MB/s for class 1 with full data packets. Other traffic in the same class and controller pauses must be subtracted. Audio requires only a small fraction; 720p60 display requires 221.184 MB/s of active pixel data. A link quota is not yet a DRAM worst-case proof.
RAM models refresh, bank conflicts, and access latencies as explicit events. Hard real-time is proven only for a profile that bounds worst-case pauses and permitted traffic loads. Mass storage with unbounded delays has no hard real-time guarantee; buffers reduce failure risk but do not eliminate it.
27.9 Errors and Cancellation
Precise fault codes:
| Value | Code | Cause |
|---|---|---|
| 1 | ILLEGAL | invalid opcode, field, or mode |
| 2 | PC | invalid fetch/branch target |
| 3 | ALIGN | misaligned local access |
| 4 | LOCAL_RANGE | local access outside SRAM |
| 5 | SLOT | role/ownership conflict |
| 6 | UNINIT | uninitialized readable region |
| 7 | INFLIGHT | DONE/OFF with open operations |
| 8 | DELIVERY | accepted message undeliverable |
| 9 | DEVICE | fatal device error not treatable as transfer status |
| 10 | INTERNAL | violation of a model invariant |
Fault PC and instruction word are latched. Older instructions remain completed; the faulting instruction and all younger ones have no local side effects. External errors after acceptance are reported as transfer status or later DELIVERY, not retroactively projected onto an old PC.
A running DMA retains its slot on fault and completes the transfer or produces CANCELLED. Debug can read the slot only afterwards. External writes already performed are not rolled back. Fault notification to the system service is retried from a reserved error latch until accepted; status inspection remains possible.
Cooperative cancellation stops new requests, drains completions/Tx, returns ownership, and ends the task. Forced cancellation is saved in the Scrapbook as cancelled, not as a successful experiment.
27.10 Device Services and Pixel/Audio Contract
Display presents only a fully written buffer. PRESENT supplies resource binding, offset, width, height, stride, format, and sequence number. It switches at VBlank and reports the actual presentation as a SYSTEM message. Only one pending present entry is allowed; further ones return BUSY. The old front buffer becomes writable again only after the switch.
RGBA8 denotes the byte sequence R,G,B,A. An LD32 therefore reads 0xAABBGGRR; internal notation must not suggest ARGB. Scanout ignores A after prior composition. RGB565 is little-endian with R[15:11], G[10:5], B[4:0]. Stride is in bytes, at least width*pixelbytes, and 8-byte-aligned. Framebuffer addresses are aligned to 64 bytes.
Audio takes interleaved PCM from a RAM ring: 1..8 channels, signed 16 or signed 24 in 32 bits, explicit sample rate, channel order FL,FR,FC,LFE,SL,SR,BL,BR. Mono/stereo use the front entries. 24 bits are right-aligned with correct sign extension in the 32-bit word; PDM/I2S pack from this. Sample rates 8000,16000,24000,32000,44100,48000,96000,192000 are service profiles, not a guaranteed RAU codec limit.
The central audio mixer produces a final PCM stream, not format-specific direct connections to the DAC. Underflow supplies silence, counts an event, and causes a real-time check to fail. Display underflow repeats the last valid image or supplies a defined error state, never unchecked RAM.
USB supplies bytes and time events into RAM ring buffers. Parsers receive bounded memory windows, not host file pointers. The theoretical USB service models throughput and pauses; it claims no electrically validated USB implementation.
27.11 Uniform Service Commands
The following structures are T1 runtime contracts, not new codec file formats. A service window is bound only to an authorized task. RWRITE at offset 0 supplies a command and payload as a contiguous local memory block.
The 32-byte command header is little-endian:
| Offset | Type | Meaning |
|---|---|---|
| 0 | u16 | protocol version, 1 |
| 2 | u16 | service operation |
| 4 | u32 | command cookie |
| 8 | u32 | payload bytes, 0..992 |
| 12 | u32 | flags, T1=0 |
| 16 | u32 | task ID of the authorized caller |
| 20 | u32 | resource/task generation |
| 24 | u32 | reserved, 0 |
| 28 | u32 | reserved, 0 |
| 32 | bytes | payload of exactly the declared length |
Header and payload remain within a 1-KiB slot. Service identity follows from the binding, not from an arbitrarily selectable global device pointer. Task ID/generation are checked against the bound caller identity. An error after transfer acceptance produces an erroneous WRITE_DONE; it must not execute a partly validated command.
WRITE_DONE=OK confirms insertion into the bounded service queue. A longer-running service later generates SYSTEM with the command cookie and status. Any result is visible beforehand in the bound result window. The normally bounded Rx quota uses ACK/NACK like other service messages; the service retains pending events. Admission limits service queues to eight commands per calling task. Full queues respond to the write with BUSY. A task must distinguish both stages: WRITE_DONE and service result.
Payloads contain logical resource IDs, not physical pointers. The service resolves ID plus generation according to the caller's rights. IDs are allocated only by the dispatcher. The four S bindings suffice for data traffic; a task must not expand its rights through a command.
| Service / Operation | Payload and result |
|---|---|
| System / 1 GET_TIME | no payload; result window: time u64, coreID u32, wakeReason u32 |
| System / 2 GET_TX_STATE | no payload; result: occupied Tx u32, open tags u32, unread completions u32, reserved u32 |
| System / 3 DIAGNOSTIC | up to 992 validated UTF-8 bytes; SYSTEM=OK after insertion into the bounded diagnostic ring |
| Dispatcher / 1 START_TASK | task image ID, generation, coreID, entry PC as four u32 |
| Dispatcher / 2 STOP_TASK | task ID, generation as two u32; orderly halt/cancellation |
| Display / 1 PRESENT | resource ID, generation, offset, width, height, stride, format, sequence as eight u32 |
| Audio / 1 CONFIGURE | resource ID, generation, offset, ring bytes, rate, channels, sample format, sequence as eight u32 |
| Audio / 2 SUBMIT | sequence, write position, valid frames, reserved as four u32 |
| USB / 1 READ_BLOCKS | resource ID, generation, offset, byte count, logical block index u64 as 24 bytes |
GET_TIME freezes the time at service acceptance. The result window is located in the respective system service starting at offset 1024, has 16 bytes, and is disjoint for each caller. It may be read using RREAD only after SYSTEM=OK. Until that read completes, no second GET_TIME/ GET_TX_STATE is allowed on the same binding. Thus two LD instructions cannot tear a 64-bit time, and no later command overwrites unread data. wakeReason: 0=no wake since task start, 1=message, 2=timer, 3=both in the same tick; GET_TIME does not consume the reason.
Display format: 0=RGBA8, 1=RGB565. Audio format: 0=PCM16, 1=PCM24-in-32. CONFIGURE is possible only without an old running stream. SUBMIT is monotonic modulo ring size and reports only fully visible audio frames. The service must not release a ring region still being read before the read cursor has passed it. USB block index and sizes are checked against the modeled medium.
START_TASK/STOP_TASK are privileged; an application task reaches them only through the designated system dispatcher. The header is a validation boundary, not a security substitute for resource rights.
28. Compiler, ABI, and Runtime (Normative as a Concept)
28.1 One Frontend, One uRISC Target
Smallsome 3.0, in Code or Symbols source mode, is the language. The existing
frontend is reused;
a future urisc-t1 target is introduced alongside the existing Ozon target
without copying the grammar or parser.
Lowering, resource analysis, and code generation are target-specific.
The ISA may also be used directly through the assembler.
Both source modes normalize into one AST/typed IR. The current frontend implements only its established Code subset; Symbols and the expanded 3.0 grammar require future work in that same frontend. The complete language's binary64 num is not redefined by T1: the table below is the explicitly restricted integer profile, not full-language conformance. Endpoint directions, ownership and bounded admission follow the shared language specification.
A compiler project contains source text, target profile, input resources, input/output contracts, and declared budgets. These produce machine code, slot images, task graph, resource plan, and proof report. An executable image without this metadata is not an approved real-time program.
28.2 Supported Language
| Language feature | T1 contract |
|---|---|
| word | native signed 32 bits; addition/subtraction/multiplication modulo 2^32 |
| num | binary64 language semantics; only proven equivalent bounded integer operations, otherwise rejected |
| bool | canonical 0 or 1; CMP masks kept separate internally |
| fixed arrays, structs, bytes | static size, local or bound RAM storage |
| text | UTF-8 literals and existing LeXA resources |
| fn, if, match, for, while | normal control-flow lowering |
| on start / on frame | one-time initialization task or task bound to VBlank |
| bounded UTF-8 diagnostic block through system service, counted guest work | |
| spawn, channel, send, receive | bounded tasks, fixed channel/message capacities |
| borrow/move | statically checked ownership across local and RAM regions |
| word / and % | shared software division, truncated toward zero; no DIV opcode |
| list/map, dynamic text | general dynamic values unsupported; bounded quantum outcome storage is statically reserved |
| extern | no guest-host function calls; resources instead of native libraries |
| race/collect/timeout | bounded task set, cancellation/join plan |
| recursion | only with statically bounded depth; otherwise rejected |
| floating point | unsupported in T1; explicit fixed-point libraries |
| detach, host file/network, arbitrary timer allocation | rejected by the static real-time profile |
Changing targets never changes num into word. num 1/2 cannot lower to integer zero; a num addition whose exact binary64 result exceeds the declared lowering range is rejected rather than wrapped. word arithmetic explicitly requests the machine contract in either source mode. Fixed-point formats are documented word-based library contracts (width, sign, fraction bits), not ungrammatical new source types. Explicit to_word/to_num conversions preserve the language's checked boundary. Nullable fields require initialized presence tags or proof that null is unreachable; fixed arrays may not silently replace null by zero. Range induction stops before exceeding its mathematical bound, not by wrapping INT_MAX back to INT_MIN. A bounded-target report accounts for that check.
A language feature not yet supported by the frontend remains a specific extension task for the same parser. This text does not claim that the existing Ozon backend already has these T1 capabilities.
Runtime errors such as division by zero are reported as defined task errors to the dispatcher. Hardware FAULT is a machine error and is not automatically treated as a catchable Smallsome Code exception. The exact Smallsome 3.0 language semantics are retained when aligning the backend; if language and target arithmetic differ, lowering must explicitly compensate for the difference or reject the program.
Collection indices remain 1-based according to Smallsome Code. Lowering checks the source index before converting (index-1)*elementsize into a byte address. struct field offsets and padding are recorded in the build report, not in the bitstream of an embedded format. A new on-frame event does not start a second overlapping writer of the same frame state: if the task is still running, the profile reports a deadline violation and skips the restart. on-start completes before normal application tasks begin.
guru/throw/rethrow are implemented as explicit control-flow/task-error edges according to Smallsome Code semantics. A handler requires a statically bounded context and a known target; it is not a hardware IRQ. An error in a handler that can no longer catch it ends the task with ERROR. print serializes only permitted values into a bounded diagnostic buffer; unbounded string allocation is not allowed for this.
28.3 Compilation Steps
- The existing frontend produces the syntax tree and source positions.
- Type checking resolves sizes, resources, and value ranges.
- A typed intermediate representation models basic blocks, phi values, effects, and explicit ownership transitions.
- Bounds, alias, and dependency analysis determines read and write intervals and necessary task edges.
- Parallelization splits only proven independent work.
- Slot planning assigns code, data, stack, and ping-pong transfers.
- Register allocation and ISA lowering produce T1 instructions.
- Scheduling defines cores, clock level, transfers, and budgets.
- The linker checks branch targets, overlays, resources, and the overall image.
- A report states assumptions, limits, worst-case paths, and open runtime conditions. Only then is the image considered admissible.
The intermediate representation is a compiler product, not a second executable machine. Only T1 machine code is executed.
28.4 Parallelization and Granularity
A loop is split only if: iteration bounds are bounded, write intervals are disjoint, read data is stable, and no unhandled dependency exists. Pointer aliasing, variable indices, and successive LZSS/predictor states may prevent splitting. Execution then remains serial; an assumed speedup does not count as proof.
Preferred task sizes are whole image tiles, PMF0 roots, bounded audio sections, or existing resource chunks. A task per pixel or sample is ruled out because of message and transfer overhead.
Reductions receive local partial results and an ordered join. word modulo addition may be reordered after algebraic proof; signed saturation is generally not associative. The compiler must not freely reorder a saturating reduction. Identical results must not depend on the incidental number of host threads.
28.5 Task Graph and Dispatcher
A task descriptor contains: task ID, code image and entry PC, parameter/stack layout, slot roles, four resource bindings, permitted cores, clock level, dependencies, deadline, WCET budget, and cancellation path.
A task becomes READY only when all inputs are visible and its destination core/resource plan is free. The dispatcher starts READY tasks. Normal application tasks are run-to-completion; WAIT permits asynchronous continuation of the same task, not automatic migration.
A communication graph must have at least one free reception/transfer window per cyclic pipeline or a proven initial token. The compiler checks capacities, blocking WAITs, and join orders per channel. General deadlock freedom of arbitrary programs is undecidable; unknown cases receive no proof status.
Static assignment is the real-time standard. Dynamic selection of a free application core is allowed for tasks without hard deadlines, but their waiting time must then not be considered statically proven. The four system cores remain reserved for dispatcher/I/O.
28.6 Local ABI
| Register | Role |
|---|---|
| R0..R3 | parameters, R0 return value; caller-saved |
| R4..R5 | temporary; caller-saved |
| R6..R9 | callee-saved |
| R10 | local stack pointer |
| R11 | link address; caller must save it before a nested call |
| ACC | caller-saved |
| S0..S3 | task bindings, unchanged by normal fn calls |
The stack grows downward from a DATA-slot upper bound defined by the linker and is 4-byte-aligned. The linker determines the maximum from the call graph, spill requirements, and declared recursion depth. ST/LD check stack accesses just like all local data.
Four parameters are passed in R0..R3, further ones in a static argument block. Values above 32 bits use explicit low/high pairs or a local result block; no hidden 128-bit register state is assumed. External data is resources plus offset/length, not local pointers.
A code image may occupy several CODE slots. Functions within the same image use JMP link forms. An overlay change is a dispatcher/task change, not a normal call through a DMA-owned slot. The ABI is identical for assembler, compiler, and future app.
28.7 Libraries and Numeric Processing
Software division uses one shared bounded 32-step algorithm for unsigned quotient/remainder. Signed division handles signs separately; rounding is toward zero. Division by zero is a task error. INT_MIN/-1 returns INT_MIN under the word modulo contract; this behavior must be explicit in the backend report.
CTZ can be expressed through bit isolation and CLZ: if x=0 ->32, otherwise 31-CLZ(x AND (0-x)). ABS/MIN/MAX/CLAMP use CMP, SEL, and SUB; signed INT_MIN remains INT_MIN under modulo ABS, while saturating ABS is an explicit library choice.
Fixed-point library formats carry width, sign, and fraction bits in their declared library contract; their source values are word or bounded word arrays. Multiplication uses MUL-low/high or ACC; rounding, shift, and saturation are explicit operations. Reciprocal/root tables and Newton steps are approximations with documented input ranges and error budgets, not exact division.
Decoder bit readers, byte order, and bounds checks form shared compiler/guest library components. Creating a uRISC version follows the same format semantics and golden vectors of the active modules; no new incompatible bitstream is introduced.
28.8 Slot Planning and Ping-Pong
Example for a small streaming-capable task:
| Slots | Use |
|---|---|
| 0..2 | CODE, at most 3 KiB for the phase |
| 3 | stack, descriptors, small state |
| 4 | input A |
| 5 | input B |
| 6 | output A |
| 7 | output B |
| 8..11 | table/history/working data per phase |
Descriptor and data bank are separate. During RREAD into input B, EX works on input A; after completion, the roles switch. RWRITE from output A permits computation into output B. This allocation is an example, not a guarantee for every decoder.
The slot planner tracks live intervals and DMA leases per bank. Spilling to an active DMA slot is prohibited. If the working set does not fit, it reduces the phase, places state in a RAM resource, or rejects the target profile with required/free size. It cannot silently turn a local LD into a RAM access.
28.9 Budget Calculation and Admissibility
A task budget covers: pipeline startup/drain, issued instructions, two windows per taken branch, software libraries, message attempts, transfer waiting time, overlays, and dispatcher work.
For a finite trace without WAIT, the control value
cycles = issued + 3 + 2*takenBranches must be used, provided
pipeline startup/drain and branch windows are not double-counted.
For multiple sections, boundaries are determined from the event model;
a blanket formula does not replace the trace.
RUN: 4 ns per core cycle, IDLE: 20 ns. A 60-Hz frame lasts 16.666666... ms, mathematically about 4.166 million RUN cycles per core. Four cores supply about 16.666 million issue windows per frame before control/wait overhead. 60 application cores supply a theoretical maximum of 15 GInstr./s, all 64 together 16 GInstr./s.
A PCM block of 1024 samples per channel at 48 kHz lasts 21.333333... ms: about 1.066 million IDLE cycles. At 192 kHz, it is only 5.333333... ms or 266666 IDLE cycles. An RAU IDLE validation must name rate, channels, and profile.
A hard-deadline report requires bounded loops, admissible worst-case I/O pauses, bounded retry/backpressure, and buffered media starts. If a bound is missing, the result is UNPROVEN with its cause, not PASS.
28.10 Diagnostics and Build Product
Each rejection states source location, task, affected resource, calculated requirements, and target limit. Examples:
task decode_stereo: local working set 28672 bytes, available 12288
loop reconstruct: dependency on previous sample, not parallelizable
task video_root: deadline not proven; reference RAM pause unbounded
spawn workers: at most 72 simultaneously active tasks, target permits 60
A deterministic build contains: ISA ID, target profile ID, source/compiler hash, input resource hashes, CODE/DATA images, relocations, entry PCs, task graph, slot/ownership plan, resource rights, cycle model version, debug source mapping, and report. Identical sources, versions, and options produce identical guest bytes. Timestamps and host paths belong in separate provenance metadata.
28.11 Ownership, Channels, and Task Errors
A read-only borrow is task-local and may not escape to another task. Physically shared read-only RAM is a separate immutable region capability: each reader has a validated binding and no writer may exist during its lease. A mutable borrow exclusively owns its entire declared interval. Send reservation immediately excludes the sender from use, mutation or a second move. A successful visible transfer commits revocation/generation reassignment and gives the receiver the new binding. Failure restores the sender's ownership before language unwind. A DMA completion alone is not an endpoint move. Local addresses are core-specific and cannot be moved as pointers into another core context.
A channel has static element size and capacity. 32-bit scalars fit directly in USER. Larger elements reside in a bounded RAM ring; USER transmits slot/sequence as a token. The compiler runtime validates token, generation, and element size. Queue capacity is the declared channel capacity, not automatically the Rx FIFO size. The source channel consists of distinct move-only send/receive endpoints; USER carries a token, never an unvalidated shared source owner. Direction, element type, closed endpoint state and channel closed/drained state are checked separately. Capacity zero is rendezvous; the runtime must reserve bounded pending-send/receive state and must not misrepresent hardware FIFO buffering as a language message queue. No null messages are transported.
Smallsome Code send may wait according to the language contract; its lowering consists of nonblocking SEND/status checking plus explicit WAIT. Smallsome Code receive is correspondingly a RECV check plus a WAIT loop. Rx may also contain completion/service messages: the runtime demultiplexes by kind and source into bounded task-state regions. It must neither discard another recipient's message nor arbitrarily often reinsert it into the same hardware FIFO.
A task error invalidates its output regions only after orderly cancellation. Successors receive ERROR instead of successful ownership. race/stop terminates the loser cooperatively; already visible external writes are not undone. The ownership plan must therefore separate loser outputs from regions in active use.
Race selects the earliest successful completion tick; equal ticks use the lowest one-based SOURCE INPUT index, not coreID or Rx arrival order. The event model's coreID order does not override this language tie-breaker. A completion at or before the deadline precedes timeout, even if the dispatcher observes it later. Later completions cannot win over an expired deadline. Source timestamps, input indices and delivery status must fit the bounded dispatcher descriptor.
Cancellation is signaled immediately but a language scope does not finish until children, timers, Tx and DMA leases reach terminal states. External effects already visible are never rolled back. Every admissible hard-deadline program includes cancellation/drain WCET and device stall bounds; otherwise the result is UNPROVEN. A waiting task does not automatically release its core/local slots. Releasing execution capacity requires an explicit dispatcher safe point and resource plan; no implicit scheduler migration is introduced.
Returning word/scalar, tuples or owned resources protects the result region before RAII; caller R0 or a bounded result block receives it only after cleanup. Guru cannot access released locals. Structured Error/Result records use a bounded ABI representation including state, numeric code and source location. Language error codes (e.g. 1102 ArithmeticError) and hardware fault IDs (1..10) remain distinct fields. Truncated diagnostic text must be flagged, not silently represented as a complete host Error. Hardware FAULT requires dispatcher action and cannot resume inside a source Guru handler.
28.12 Assembly Syntax and Concrete Examples
Canonical modes are constructed from the tables in section 11. The dotted names are fixed assembler words:
| Instruction | Mode words |
|---|---|
| ADD/SUB | scalar, lane16, lane8; wrap, usat, ssat |
| CMP | scalar, lane16, lane8; eq, ne, slt, sle, ult, ule |
| SHL | scalar, lane16, lane8; imm or reg |
| SHR | scalar, lane16, lane8; logical or arithmetic; imm or reg |
| MUL | unsigned or signed; low or high |
| MAC | unsigned or signed; add or set; readlow or readhigh |
| LD | u8, u16, u32, s8, s16 |
| ST | u8, u16, u32 |
| JMP | rel, indirect, linkrel, linkindirect |
| WAIT | message, ticksimm, ticksreg |
Mode names with multiple parts are joined by dots in table order,
for example lane8.usat or
scalar.arithmetic.reg. Inapplicable parts are rejected.
For LD/ST, the mode is attached to the mnemonic, for example LD.u16.
Register counts in shift instructions appear as Rc instead of an
immediate value. JMP.rel/JMP.linkrel use a label or I22,
JMP.indirect/JMP.linkindirect exactly one register.
WAIT.message has no operand; WAIT.ticksimm U22,
WAIT.ticksreg lowRegister,highRegister.
WAITUNTIL lowRegister,highRegister remains a canonical mnemonic.
MAC.readlow/readhigh has only Ra; the assembler sets B/C and
the signed bit to zero. All other MAC forms have three registers.
Lexical contract:
ASCII mnemonics, case-insensitive mnemonics,
case-sensitive labels, R0..R11/S0..S3, comma as operand separator,
# as comment, hex values with 0x, optional minus before numbers.
No implicit octal interpretation. One statement per line.
Label addresses are byte addresses; the assembler calculates
relative word distances from them. .word writes u32 little-endian,
.byte exactly the checked u8 values. .align n is allowed for
powers of two up to 1024 and fills with zero.
CODE padding uses NOP words instead of opcode 0.
Example A, scalar sum 1..4, slot 0 CODE and slot 3 DATA. R3 is a deliberately constructed zero-valued register; R0 remains writable:
.code 0
start:
MOVI R0, 0
MOVI R1, 1
MOVI R2, 5
MOVI R3, 0
again:
ADD R0, R0, R1, scalar.wrap
ADDI R1, R1, 1
BNE R1, R2, again
MOVI R4, 3072
ST.u32 R0, [R4 + 0]
DONE
Expectation: R0=10, DATA[3072..3075]=0A 00 00 00, three taken BNE and one untaken BNE, STOPPED without open transfers. 19 issued instructions plus 3 pipeline windows and 6 branch windows yield 28 RUN cycles = 112 ns in the basic model. DONE drain is included in the shared pipeline drain here. A later trace must reproduce this specific reference.
Example B, branchless maximum of a signed value and zero:
MOVI R0, 0
CMP R3, R1, R0, scalar.slt
SEL R3, R0, R1
R3 is 0 for negative R1, otherwise the unchanged value of R1. The result also handles INT_MIN correctly, without ABS overflow.
Example C, data-bank RREAD and cookie: R4 points to a validated descriptor in slot 3, whose data destination is slot 4 and whose cookie is 17. S0 denotes the readable data window. R9 contains zero; R6 is only the status destination.
request:
RREAD R6, S0, R4
BNE R6, R9, request_failed
poll:
RECV R6, R7, R8
BEQ R6, R9, dispatch_event
WAIT.message
JMP.rel poll
dispatch_event checks kind/source in R8 and cookie in R7; READ_DONE with status OK makes slot 4 usable. Other events go to the designated bounded demultiplexer. request_failed handles BUSY/errors; immediate infinite retry is not an admissible real-time plan. The two target labels are application continuations, not hidden hardware instructions.
28.13 Build Manifest and Theoretical Programs
The build product consists of a canonical UTF-8 JSON manifest plus referenced CODE/DATA/resource bytes. It is a program description and changes none of the Generation-26 bitstreams. WORM0 can transport the required individual resources as before; no second universal media container is defined.
Binding manifest fields:
| Field | Meaning |
|---|---|
| schema | urisc-t1-program-2; version 1 manifests lack the audited language/runtime contract |
| isa | urisc-t1-1.0 |
| model | specific event model version |
| profile | 64-core standard profile and explicit latency/error parameters |
| language | Smallsome 3.0 revision, modes by canonical source ID, num/word capabilities, pinned Unicode version |
| inputs | IDs, lengths, SHA-256, and format generation |
| tasks | IDs, core mask, entry, code/data images, ABI parameters |
| slots | twelve roles per task, initialization bytes, and leases |
| resources | logical ID, generation, size, rights, traffic class |
| edges | producer, consumer, region, ownership/event condition |
| budgets | deadline, instruction/transfer limits, and proof status |
| runtime | contract urisc-t1-runtime-2, task/queue/timer caps, bounded Error/Result layouts, cancellation/drain limits |
| debug | source mapping and named breakpoints |
| expected | output hashes, final states, and expected errors |
Integers are stored in JSON as decimal numbers; 64-bit time and 40-bit addresses as decimal strings so that a JSON reader with IEEE-754 numbers does not round them. Binary images are little-endian. Manifest lists have a fixed order by ID, fields fixed names; hash objects are not identified by a host path.
The loader checks schema, ISA, latency profile, lengths and hashes, disjoint rights, initialized code/data regions, entry alignment, and admissible system/application cores before changing state. An erroneous manifest is rejected as a whole. Importing v4.1 requires reassembly for T1, not automatic interpretation of old opcode bytes.
29. Smallsome Formats on uRISC (Normative Integration Contract)
29.1 Source, Generation, and Host/Guest Boundary
All eight format families use Generation 26. Their binary grammar in the active modules and release contracts under Formats remains authoritative. This chapter specifies their execution on T1 and does not copy a second divergent file-format standard.
A theoretical decoder is a T1 guest program made of normal opcodes. The native Retro decoder supplies golden results and metadata for comparison. Calling it is not a simulated core task and must account for no uRISC cycles or core counts. A decoder not yet ported is marked HOST_REFERENCE/UNPORTED. The future app may use it for preview; no guest performance report is then provided.
Codec engine logic knows only memory. Host files, preview, audio output, and Scrapbook persistence use the existing project boundaries. Audio goes through the central mixer, images through central image composition. Containers delegate embedded codecs; they do not implement RAU/RFXL a second time.
29.2 Format Matrix and Task Boundaries
| Family | Active contract | T1 work unit | Dependency |
|---|---|---|---|
| RAU / RC26 | mono/stereo PCM | sample section within a frame or RC packet | bit reader, predictor/LMS, possibly previous frame |
| RFXL | RGB24/RGBA32 image | existing stream, then line/tile | LZSS, MTF, and spatial predictor |
| RFXA / RFXZ | animation, optional RAU sprites | frame/delta; decompressed group | persistent canvas and palette |
| PMF0 / PMFZ | video, optionally continuous RAU | existing root per frame; group for PMFZ | history frames, copy-current within a root |
| RPMC | song, patterns, sample/instrument state | validated loading sections and time events | pattern/voice/mixer state |
| LeXA | UTF-8 and bitmap-font resources | existing text/font block | compression stream and glyph boundaries |
| FormA / FMA1 | Geometry-Core and Scene | geometry/scene loading phase, later render tile | delta values, indices, materials |
| WORM0 | single typed resource | existing resource chunk | chunk decoder, then embedded codec |
A 1-KiB slot is a transport/working bank, not a new bitstream block format. 16..64-KiB blocks are not retrospectively invented as a universal guarantee of independence. Byte and bitstream boundaries remain exactly preserved.
29.3 RAU-26 and RC26
Sources: RAU module, release audit. RA26 has a 17-byte header with flags, explicit rate, and sample count; frames contain up to 1024 samples per channel. The current decoder reconstructs fixed predictor/Rice, optional LMS, constant frames, and copies of the previous frame. A current RA26 file is therefore not generally frame-parallel.
RC26 contains packets with validated, matching metadata. Each embedded RA26 packet starts its own decoder state; packet parallelism is possible after header/length validation. Within a packet, decoder dependencies remain. The final PCM join preserves sample order.
Native reference output is signed PCM16 interleaved, with sample counts per channel. The format supports mono/stereo; 7.1 results from mixing multiple sources or other PCM resources, not from eight-channel RA26.
The existing stereo decoder already uses: history 210242*4=16384 bytes, residuals 8192 bytes, and transform buffers 8192 bytes, totaling 32768 bytes before code/additional state. Unchanged, it does not fit in 12 KiB locally. The T1 plan uses RAM for history and bounded local sample sections; an optimized streaming liveness plan must demonstrate correct predictor/LMS order. These are porting requirements, not a claim that a 12-KiB implementation already exists.
Data flow: RAM input ring -> header/bit reader -> residual section -> predictor/LMS -> inverse stereo transform -> PCM16 -> shared audio input ring -> central mixer -> final output ring. DMA completions protect slot changes.
Suitable opcodes: LD.u8/u16/u32, SHR/SHL, AND, CLZ for bit-reader helpers, ADD/SUB, signed MUL/MAC, CMP/SEL, and ST. Rice escape is bounded by the format; nevertheless, bit-reader boundaries are validated before every refill. 64-bit intermediate values are represented through register pairs or ACC.
The validation case must cover lossless and all lossy profiles, mono/stereo, final short frames, previous-frame mode, RC packet boundaries, seek, and erroneous/truncated streams. The hypothesis of one core at 50 MHz is confirmed only after T1 instruction and worst-case transfer measurements for a specified rate.
29.4 RFXL-26
Source: RFXL module. Output is RGB24 or RGBA32; alpha is reconstructed according to the existing format. Palette, Planar, Raster, and Sparse-Screen remain existing modes. Crappy/Low/Mid/High/Lossless are encoder profiles, not new uRISC decoder opcodes.
Codec paths include raw, RLE, BytePack, Rice, and LZSS. LZSS has a 12-bit window; back-references may overlap and must be reconstructed in a defined copying order. A Memcpy for nonoverlapping regions does not replace this case. Predictors require previous pixels/lines; MTF has a continuously mutated palette-index order. Such streams are serial until a proven reset/substream boundary is reached.
A 720p RGBA image has 3686400 bytes. One RGBA line has 5120 bytes; two lines already occupy 10240 bytes. Together with 4096 bytes of LZSS history, palette, code, and stack, this does not fit locally. Decompressed streams, history, or lines therefore reside in RAM; local sections and DMA windows are planned individually. A planner must check T1 throughput including these transfers.
Data flow: stream validation -> serial decompression/unpredict where required -> pixel/palette reconstruction -> RGBA RAM -> parallel composition on disjoint tiles -> framebuffer. Parallelizing encoder candidate search does not parallelize the same decoder bitstream.
PERM sorts bytes, CMP/SEL produce masks, lane ADD/SUB support provably suitable pixel arithmetic. Signed intermediate values for MED and other predictor rules must not be distorted by unsigned lane saturation.
A single image has no intrinsic 60-Hz deadline. The application specifies image size, loading/presentation deadline, and permitted decoding phases. A universal statement that RFXL runs at IDLE in real time without an image size/deadline is invalid.
29.5 RFXA-26 and RFXZ-26
Source: animation module. Full frames use RFXL; delta frames modify a persistent canvas. Existing global palettes, rectangles, and 16x16 delta regions remain the format boundaries. Canvas decoder order is binding.
After validation, a delta may be distributed across multiple cores only if its write regions are disjoint or the original ordering effect is preserved. A later frame starts only after completion of the previous canvas state. Optional RAU frame chunks are audio sprites; they must be distinguished from PMF0's continuous audio track.
RFXZ must decompress the required existing group. Group size and peak RAM are checked before starting. Neither wrapper nor root index eliminates decompression costs. The canvas is marked presentable only after a complete frame.
29.6 PMF0-26 and PMFZ-26
Sources: video module, release audit. PMF0 uses existing frame/GOP/root indices, quadtree leaves, Copy-current, Copy-reference, Reference-Residual, Motion, Flat, BiColor, and QuadColor. RGB/YCoCg and prepared chroma profiles are inversely transformed exactly according to active Generation 26.
Frame history may use lags 1,2,4,8. A reference is immutable until its final reader finishes. At 720p, eight RGB24 history images occupy 81280720*3=22118400 bytes (21.094 MiB), plus the current RGB image at 2.637 MiB and RGBA framebuffer at 3.516 MiB. These values fit in 256 MiB, but not in local slots. Further wrapper, bitstream, audio, and rendering buffers are additional.
Before a frame: validate indices, root geometry, stream boundaries, and reference lags; then distribute independent roots to free application cores. The active encoder limits Copy-current to the current root. The decoder, by contrast, checks preceding image positions in particular; an externally supplied stream must therefore not be treated as root-independent without validation. The guest validator checks Copy-current source regions and creates dependency edges for cross-root accesses or decodes the affected section serially. Within a root, quadtree/bit-reader order and Copy-current dependencies remain; a compiler must not split these pixel by pixel. It must check the Copy-current source region against the already reconstructed region under the existing decoder contract.
A root with 64x64 RGB requires 12288 bytes for pixels alone; with code and state, it does not fit entirely locally. The guest works in subsections and RAM history without turning these into new independent bitstream roots. Even a 32x32 root requires 3072 bytes of RGB plus state/reference windows.
Frame barrier: all roots successful -> color/alpha/composition phase -> all writes visible -> PRESENT. Frame n+1 may use references from n only after this reconstruction barrier. Audio is forwarded independently according to sample time.
PMF0 audio is continuous RAU-26/RC26 with fragments between video frames, not a sequence of audio sprites. Fragment boundaries are not automatically codec reset points. The demuxer feeds bytes into the existing RAU streaming state.
PMFZ groups independent PMF0 units and optionally uses BriefLZ for an entire group. Decompression requires the whole group block beforehand; startup latency and peak RAM must be budgeted explicitly. Documented encoder practice may choose plain PMF0 for audio. No lab may account for PMFZ memory or decompression time as zero.
Four RUN cores for 720p60 theoretically correspond to about 16.666 million issue windows per frame, about 18.08 per pixel before branches, bit reader, reference transfers, and synchronization. This is a hypothesis to be checked, not decoding evidence. Known ARM/PureBasic timings are not converted to uRISC cycles.
29.7 RPMC-26
Source: music module. RPMC is loaded as a validated song with orders, patterns, events, samples, and instrument/surround state. It is not a music bitstream to be decoded anew for every audio frame.
The runtime separates time events from the central sample mixer: pattern/tick service -> voice parameters -> existing sample data -> mixer sections -> final PCM ring. Sample/instrument slots remain semantically sparse according to RPMC, not artificially densely duplicated. Effect state, order changes, and tempo changes advance serially according to music semantics.
Voice computation may be distributed into disjoint local partial results. The join mixes in a fixed order with defined width/rounding. 7.1 assignment is mixer/channel state, not a new RPMC or RAU opcode.
29.8 LeXA-26
Source: text/font module. Text resources supply UTF-8 bytes and existing indices; bitmap-font resources supply validated geometry and glyph blocks. Unicode decoding, glyph selection, and screen rasterization are application/font tasks. TTF/OTF import belongs to preparation, not to an invented LeXA runtime parser.
Guest sequence: header/lengths -> existing compression path -> text/glyph RAM -> layout -> disjoint text tiles -> central image composition. Compression streams are parallelized only at actual resource boundaries. 1bpp glyphs use LD.u8, masks, and SEL; row/column bit arrangement remains that of the existing font contract.
29.9 FormA-26 / FMA1
Source: geometry module. Int32-mm coordinates, triangle indices, delta-varint/byteplanes/PackBits, scene attributes, and embedded RFXL textures remain unchanged. Geometry-Core and the Scene appendix do not duplicate geometry.
Data flow: scene validation -> geometry/attribute decoding -> geometry RAM -> transform/clip -> tile lists -> tile renderer -> framebuffer. Delta/varint order remains binding in the loading path. Rendering is a subsequent use, not a property of the codec.
Fixed-point transformation, perspective, and depth specify value ranges, rounding, and error budget. A 16-bit Z buffer is a scene decision, not a general guarantee for Int32-mm geometry. Texture RFXL is loaded through the same codec path. Texture tiles are made available locally through DMA before use.
29.10 WORM0-26
Source: resource container. WORM0 contains exactly one typed resource and the existing chunk boundaries/methods. The guest validates the overall payload, decompresses existing chunks, and delegates to RAU, RFXL, LeXA, FormA, or RAW according to resource type.
Independent existing chunks may be distributed across multiple cores where their decoder state starts anew for each chunk. A decoded payload must reach exactly the declared size. A decompressed but semantically invalid embedded resource remains an error, not a decoder-side RAW fallback.
Raw-data fallback is an existing encoder decision; it must not conceal corrupt codec data. Compression and embedded codec are timed separately, then added as a complete loading path.
29.11 Porting, Equality, and Measurement Protocol
The same validation sequence applies to every family:
- Record the current native golden corpus and module hash.
- Capture binary boundaries and numeric intermediate values from the active decoder; do not port a historical format version.
- Develop the same algorithms as T1 guest code or compiler lowering and assemble them for the specified slot plan.
- Check output against the native reference: lossless byte-exact, lossy also byte-exact against the decoder of the same bitstream.
- Check EOF, truncation, erroneous lengths/offsets, boundary dimensions, and resource failure with defined errors.
- Log instructions, cycles, transfers, peak RAM, peak slots, messages, underflows, and deadline violations.
A codec profile is UNPORTED, FUNCTIONAL, TIMED, or BOUNDED. FUNCTIONAL requires output equality; TIMED a modeled trace; BOUNDED additionally requires the documented worst-case assumptions. One passing clip does not make the entire format family BOUNDED.
30. Theoretical Lab and Scrapbook Contract (Specification)
30.1 Scope of the Future App
The app will be built according to this specification. It executes T1 code, visualizes an 8x8 CPUlet, and saves reproducible experiments. RTL, FPGA, physical manufacturing, and automatic tool installation are not prerequisites for this theoretical project.
The current PureBasic emulator is consolidated into a single T1 reference core. GUI, assembler, compiler, execution display, and headless verification use the same core. Format modules remain native references, and the existing memory contracts remain in place.
30.2 Event Time and Determinism
A T1 event model uses integer system ticks. RUN issues every two fabric ticks, IDLE every ten. The model order at a point in time is: completed external writes/DMA -> message/timer wakes -> core issue windows in coreID order -> fabric arbitration -> scheduling of new device events. Newly accepted requests respond no earlier than a later tick. A FIFO entry accepted only at the end of a tick cannot be read retroactively in the same core window.
All round-robin pointers start at source 0. For simultaneous events with equal priority, the stored (tick, event phase, source ID, local sequence) order decides. Quotas are nonnegative tick credits per class. At 100-tick boundaries, 4/12/64/20 credits are added; at most one full quota is additionally carried into the next section. Each transmitted beat consumes one credit. Borrowing free credits is allowed only when the original class has no waiting beat. For RAM service, five credits are reserved at the start; this also makes the service window, which cannot be interrupted beat by beat, deterministic.
Long idle phases are skipped up to the next event without changing guest cycles or deadlines. Host parallelization must not change the specified guest order. Every random/error generator has a saved seed. Host audio/image preview follows guest state and does not control its time.
30.3 Experiment Contents
A Scrapbook entry contains: unique experiment ID, title/note, source code, ISA/compiler/model version, 64-core target profile, resources and their SHA-256 hashes, initial state, seed, event/latency profile, expectations, execution result, measurements, and proof status. Comparable repetitions refer to the same input baseline; changed sources/profiles produce a new experiment version.
A canonical snapshot contains PC, registers, ACC, slots/roles, resources, FIFO/Tx/tags, timers, RAM, dispatcher, and all scheduled external events. Saving only PC/registers is insufficient for continuation or backward stepping.
The app may save snapshot deltas; it must reconstruct the same complete state from them. A checksum verifies it. Cancellation, fault, timeout, and PASS are separate results. Actual Mac runtime and simulated guest duration are displayed separately.
30.4 Observation and Control Contract
The 8x8 grid shows ID, state, clock level, task, and activity. Selection shows PC/instruction, registers, ACC, slot roles/ownership, messages, and open transfers. A timeline shows core sections, fabric occupancy, DMA, deadline, and present/audio events.
Instruction single-step means: run until the next completed instruction of the selected core; the other masters continue according to their guest time. System-tick single-step is a separate operation. A halt freezes the entire model state.
Breakpoints trigger before IF of the target instruction; watchpoints report after a valid completed write. A debug halt requires no guest FAULT and changes no queue. Backward stepping is performed through snapshot plus deterministic replay, not through invented inverse opcodes.
Experiments for RAM pauses, full FIFOs, USB stalls, and underflow are saved as profiles. The app then determines which budget assumption was violated instead of interpreting host stutter as a uRISC performance problem.
30.5 Limits and Future Delivery
The app may display graphics/audio output without physical PHYs. The display proves neither TMDS/USB/PDM compliance nor actual ASIC clock capability. Code signing protects the origin/integrity of the Mac app; it certifies no architectural claim. Developer ID, notarization, and bundle integration are implemented during app construction using the existing Mac packaging paths.
30.6 Reproducible Model Profiles
The standard profile is named t1-lab-nominal-1. It is a synthetic
working assumption for experiments, not a reproduction of a measured LPDDR4 PHY:
| Model component | Nominal contract |
|---|---|
| Core clocks | RUN every 2, IDLE every 10 system ticks |
| Start/frequency change | only at tick boundaries; first issue at the next suitable global 2-/10-tick grid point |
| Group/router link | 1 beat per tick; quotas from 27.8 |
| Local message path | same acceptance/ACK contract; earliest acceptance in the tick after SEND |
| RAM base latency | packet ready for service 64 system ticks after complete acceptance |
| RAM data service | one shared unit; 5 ticks per packet up to 64 bytes, reads and writes together |
| RAM order | arbitration by traffic class/quota, FIFO within the same class |
| Refresh/bank conflicts | disabled in the nominal profile and identified as such in the report |
| Service command | earliest processing 1 tick after complete acceptance |
| Display | present at the next rational 60-Hz VBlank boundary |
| Audio | sample consumption at the configured rational rate |
| USB storage | synthetic 35000000 bytes/s, startup latency 500000 ticks, no random stall pause |
| Seed | 0 unless a random profile is active |
RAM read bytes are read at the end of data service; write bytes are visibly written there. A service-ready packet waits for the shared unit. Read responses then wait for the response link; write ACK is generated only after the last visible packet. The 5-tick service permits at most 6.4 GB/s of aggregate array payload for full packets, never separate 6.4 GB/s for reads and writes. A short transfer also occupies a service section. A transaction data buffer is limited to eight requests per core plus bounded autonomous device buffers; unbounded guest queues are not assumed.
Frames and sample periods are rational numbers. The model uses a phase accumulator; an event at a noninteger tick falls on the first tick after the ideal time. This causes at most one tick of quantization per event, no cumulative drift. For example, VBlank intervals at 500 MHz/60 Hz follow a fixed pattern of 8333333/8333334 ticks.
The stress profile t1-lab-stress-1 uses the same rules and adds
every 3900 ticks a synthetic 100-tick RAM service pause.
A running service operation is paused and then resumed.
USB adds, after every 1048576 delivered bytes, 50000000 ticks of pause.
The numbers are deliberately saved test assumptions, not
guaranteed electrical DRAM/USB maxima.
Further profiles declare every deviation together with the seed.
A deadline met only in the nominal profile receives TIMED, never automatically BOUNDED. A worst-case profile must demonstrate upper bounds for its entire permitted input set. Endless disruptions are logged as TIMEOUT through a bounded experiment duration; they are not reinterpreted as successful runtime.
30.7 Research Experiments as the App's Baseline Set
| Experiment | Expected result / Question |
|---|---|
| ISA scalar | all boundary values and encodings; correct register/SRAM results |
| ISA lanes | no lane carries, correct masks and saturation |
| LD-use / branch | no stall; two flush windows only on a taken branch |
| 64-Core message ring | ordered tokens, no duplication, controlled FIFO load |
| Slot conflict | deliberate IF/EX/DMA conflict causes precise FAULT |
| Ping-Pong-DMA | read A while filling B; correct cookie and ownership |
| Shared RAM data | write completion before ownership transfer, then identical read |
| RAU stream | byte-identical PCM, rate/slot/deadline report |
| RFXL image | byte-identical RGBA for all existing decoder paths |
| RFXA canvas | ordered deltas and optional audio sprites |
| PMF0 roots | identical 1/4/60-worker output, correct history/root dependency |
| RPMC music | same event sequence, deterministic mixer join |
| LeXA / FormA / WORM0 | byte-identical resources, correct boundaries and nested codecs |
| Media combination | PMF0 plus audio plus input/USB; no claims that costs have disappeared |
| Stress / Underflow | defined error and violated budget assumption visible |
| Snapshot replay | identical final state and identical guest trace |
This is the binding future experiment catalog, not an app feature created or executed in this round.
31. Conformance, Open Physical Questions, and References
31.1 Binding Theoretical Release Scope
T1 version 1.0 is a theoretical specification with: 31 opcodes, fixed operand/mode bits, register/memory semantics, branch/WAIT contract, messages, resources, ownership and visibility, compiler target/ABI, all eight format profiles, and a reproducible lab contract. Its completeness means defined target rules, not an already working compiler or proven chip.
Mandatory future model checks:
| Area | Properties to check |
|---|---|
| Encoding | all 31 opcodes, reserved fields, register limits, immediate values |
| Arithmetic | signs, boundary values, lane isolation, saturation, MUL-high, ACC forwarding |
| Memory | alignment, slot boundaries, initialization, IF/EX/DMA conflict, no partial ST |
| Pipeline | RAW/WAW, LD-use, branch flush, precise FAULT, WAIT race |
| Communication | FIFO quotas, retry/dedup, ordering, completion reservation, delivery fault |
| Resources | rights/generation, overflow checking, cookies, visibility, cancellation |
| Compiler | types, alias/ownership plan, ABI, branch relocation, slot spills, UNPROVEN |
| Formats | active golden corpora, pixel/PCM equality, dependency, truncation |
| System | boot, 64-core time ordering, buffer switching, audio/display underflow |
| Replay | snapshot completeness, same seed, same hashes and guest traces |
This list is the future test contract. No smoke tests or codec programs are executed in this documentation round.
31.2 What Is Theoretically Decided and Physically Open
Decided for T1: no LOOP/CTZ/DIV, scalar MUL/MAC, atomic local descriptor acceptance, 0-byte CPUlet-SRAM in the standard profile, 8 Tx/tags/completion slots, 16-byte contact header, 64-core numbering, fixed ABI and model time.
Open for future manufacturing: process/PDK, SRAM/register-file/128-bit descriptor bank macros, 250-MHz timing, 500-MHz fabric, DRAM timing/refresh, PLL/clock domains, PHYs and signal integrity, power supply, package, area, power and temperature. Single-cycle MUL and LD/descriptor forwarding in particular require physical validation. An FPGA with a different clock frequency can confirm functional rules, but cannot prove the ASIC frequency.
The PPA figures in section 21 remain historical guidance with unknown accuracy. A theoretical simulation run will never turn them into watt/mm2 measurements.
31.3 Source Baseline and Responsibility
Baseline of this normative specification: 10 October 2026. Local active sources take precedence over old ZIP codec samples. The archives serve only as provenance of emulator behavior.
- Smallsome Code language and grammar.
- Format overview and measurement matrix.
- RAU/RFXL audit and PMF0 audit.
- Golden corpora:
Formats/ReleaseCorpus/,Formats/PMF0ReleaseCorpus/,Formats/RemainingFormatReleaseCorpus/. - Project architecture for memory/engine-area boundaries.
SHA-256 of the source baseline read during preparation:
| Source | SHA-256 |
|---|---|
| RAU module | 56a16198483c4a8a10c5e59dc1e086cebc2429a062f201b6b9bc73cb7e3b1331 |
| RFXL module | 7ab9534ab3101b992cfcc056f603fa2244a04dd45c8708b3c64d24047f7355b9 |
| PMF0 module | e6577ce03ecdc9b805c178d1d3fb574315fcf1fbf429c349be6d0203dcc34a4c |
| RFXA module | 3946992cc0198acbc9504ba56d1b3d36727447dee4bdcf79d3047cd56f32396f |
| RPMC module | 65bb2fa7ab6e3226fa2333e59111858e0b7b576c560a96e0f19db3003a4c90d7 |
| LeXA module | d85b26c21b1448c3e4db1727656bbf27d0076b5a7991c7a4e3f73823ead68057 |
| FormA module | 32a0d687ae9f548d413e00237f380f246e712c6193211cc014b9aa6f5db4d72b |
| WORM0 module | 9383fe7d40114c11b571d1402064bf7f860064168e5d4469553fdfb286191939 |
| Smallsome Code frontend (historical name: Cymple) | 233482e27a73233531234b8775934e5ba4211eb6023205ed83c35452a2b4cbc3 |
| Smallsome Code Ozon backend (historical name: Cymple) | 9d0adabf69a6c8729a636506ef67c3870e82ec13cee2f004de7335eb825e8dc0 |
| v4.1 emulator archive | 092b60b187990d5d49fc9444c48937ad40deb095e686fe80a91a51bc2dbdd1f7 |
Hashes document a reading baseline, not passed verification. When a module changes, its T1 compatibility is checked using the same integration contract and updated golden vectors.
31.4 Version Policy
Changes to instruction semantics or encoding increase the ISA version and require reassembly, a new compiler report, and new experiment snapshots. Language changes version the language; runtime/ownership changes version the runtime contract; mandatory manifest changes version the schema. These IDs are independent and all appear in an experiment. Revision 4 retains ISA 1.0 but uses program schema 2 and runtime contract 2. Version-1 manifests require an explicit migration/report, never silent acceptance as equivalent. Editorial clarification alone increases only the document revision.
A future app may display several old Scrapbook experiments, but execute them only with clear version association. T1 has exactly one reference core; historical archives are not incorporated as a second active architectural path.