Production Training Networks
Two transformer topologies are currently used for real training runs on the network. Both are decoder-only, causal language models defined entirely in network YAML — the full files are in the compute-all/data directory.
They share the same building blocks:
- Token embedding (
embedder) with the embedding matrix published astoken_embeddingand tied to the LM head. - Pre-norm transformer blocks with distributed attention: each block broadcasts to 8 independent
attention_headVNodes (head dim 64, causal, RoPE position encoding, θ = 10000), stacks the head outputs, and applies an output projection — so every head can be placed on a different PNode. (At runtime, the orchestrator’s fusion pass may batch co-located heads into a single multi-head VNode for GPU efficiency.) - A GeLU feed-forward block (512 → 2048 → 512) with residual connections and dropout.
- A SwiGLU LM head: final layer norm, up-projection 512 → 4096, SwiGLU gate (4096 → 2048), down-projection 2048 → 512, then the tied
lm_headlinear producing vocabulary logits. ce_logitscross-entropy cost computed in float32 for numerical stability, while the network itself runs in fp16 (network-wideprecision: fp16, with fp32 master weights in the optimizers).- Adam as the per-node optimizer, with the learning rate exposed as a training metric.
One transformer block looks like this:
flowchart TD
IN[block input] --> BC[broadcast]
BC --> LN1[layernorm]
LN1 --> H[broadcast to 8 heads]
H --> A0[attention_head 0..7<br/>head_dim 64, causal, RoPE]
A0 --> ST[stack]
ST --> OP[linear output_proj]
OP --> DO1[dropout]
BC -- residual --> ADD1[add]
DO1 --> ADD1
ADD1 --> BC2[broadcast]
BC2 --> LN2[layernorm]
LN2 --> UP["linear 512→2048"]
UP --> G[gelu]
G --> DN["linear 2048→512"]
DN --> DO2[dropout]
BC2 -- residual --> ADD2[add]
DO2 --> ADD2
ADD2 --> OUT[next block]
Topology 1 — 6 layers, 10K vocab, 128 context
distributed_attention_example_6_swiglu_head.yml (distributed-multi-head-attention-6l-swiglu_ffn)
| Property | Value |
|---|---|
| Transformer blocks | 6 (8 attention heads each — 48 distributed heads) |
| Model dimension | 512 (head dim 64) |
| Context length | 128 tokens |
| Vocabulary | 10,000 BPE tokens |
| FFN | GeLU, 512 → 2048 → 512 |
| LM head | SwiGLU (512 → 4096 → 2048 → 512), tied embeddings |
| Precision | fp16 (fp32 loss and master weights) |
| Graph size | 143 VNodes |
| Parameters | ≈ 27M (with tied embeddings) |
This is the smaller iteration/testbed network: fast steps, quick end-to-end experiments on tokenizer, dataset, and training-loop changes.
Topology 2 — 8 layers, 32K vocab, 512 context
distributed_attention_example_8_swiglu_head_vocab32768_seq512.yml (distributed-multi-head-attention-8l-swiglu_ffn)
| Property | Value |
|---|---|
| Transformer blocks | 8 (8 attention heads each — 64 distributed heads) |
| Model dimension | 512 (head dim 64) |
| Context length | 512 tokens |
| Vocabulary | 32,768 BPE tokens |
| FFN | GeLU, 512 → 2048 → 512 |
| LM head | SwiGLU (512 → 4096 → 2048 → 512), tied embeddings |
| Precision | fp16 (fp32 loss and master weights) |
| Graph size | 183 VNodes |
| Parameters | ≈ 45M (with tied embeddings) |
This is the current production training network — the workload driving the ongoing performance push (FlashAttention-2-style tiled attention, fused LM-head cross-entropy, attention-head fusion). Live sessions train it at batch 128 with an average step time of roughly 175 seconds.
Placement
Every node in a block carries colocate_group: transformer_block_N with colocate_policy: required, and the embedding/LM-head group is colocated as token_embedding — so the orchestrator keeps each block together on one PNode (minimizing cross-node traffic within a block) while different blocks can be spread across the fleet, giving pipeline-style distribution with per-head parallelism available when more PNodes are present.