- bits per weight
- 32
- multiplications
- one per weight
- memory per token
- grows with context
32bits per weight
A network is millions of decimals.
Every weight is a 32 bit number, and every step of inference multiplies them. That is why models need big GPUs.
16×smaller than fp32
Core-1 allows only three values.
Every weight is forced to −1, 0 or +1 during training. Packed at two bits each, the model gets sixteen times smaller.
0multiplications per weight
So multiplying becomes adding.
Zero means skip. Plus one means add. Minus one means subtract. Plain integer math that cheap processors are good at.
O(1)memory per new token
And memory stays flat.
Attention keeps a cache that grows with every token. Core-1 uses a state space backbone with a fixed size state instead.