Add a Backbone¶
A backbone reads tokens and returns a conditioning vector, plus one-token streaming decode with fixed state. Follow Mamba.
The contract¶
class MyBackbone {
MyBackbone(std::span<const TensorView> weights, Arena& scratch) noexcept;
bool valid() const noexcept;
const Config& config() const noexcept; // d_model, n_layers, vocab, ...
const float* embedding() const noexcept; // [vocab][d_model], the engine gathers tokens
std::size_t state_size() const noexcept; // floats of streaming state
void forward(std::span<const float> in, std::span<float> out, std::size_t seq_len) noexcept;
void decode(std::span<const float> x, std::span<float> state, std::span<float> out) noexcept;
};
The conditioning vector is the last row of forward, size d_model.
Steps¶
Create
src/core/models/<name>/. Derive config from checkpoint shapes. Store raw weight pointers.Implement
forwardfor the prefix. Carve per-layer scratch from the arena and rewind per layer.Implement
decodefor one token. Keep all state fixed-size and in the callerstatespan. It must equalforwardstep by step.Add kernels behind
kernels.h. Do not grow state with sequence length. A transformer sizes its KV-cache to the max prefix.Add a converter mapping. Select the backbone at load from the tensor names.
Verify against PyTorch. Check
decodematchesforward. Add tests and an ADR.
Reference: src/core/models/mamba/.