attention(q,k,v,mask,dropout) | verbatim | scaled dot-product, masked_fill(mask==0, -1e9), softmax, dropout |
MultiHeadedAttention | verbatim | 4 linears, view/transpose to (b,h,t,d_k), self-attention only |
PositionwiseFeedForward | verbatim | w_2(dropout(relu(w_1(x)))) |
LayerNorm | verbatim | a_2, b_2, eps 1e-6 |
SublayerConnection | verbatim | x + dropout(sublayer(norm(x))) |
Embeddings | verbatim | lut(x) * sqrt(d_model) |
PositionalEncoding | verbatim | sinusoidal, div_term = exp(arange(0,d,2) · −ln(10000)/d) |
subsequent_mask | verbatim | triu(ones, k=1) == 0 |
Generator | verbatim | linear + log_softmax |
NoamOpt | verbatim | factor · d_model^−0.5 · min(step^−0.5, step·warmup^−1.5) |
LabelSmoothing | verbatim | KLDivLoss, confidence/(V−2), padding zeroed |
| Xavier init | verbatim | xavier_uniform_ on dim > 1 |
Encoder, EncoderLayer, src-attn | dropped | no encoder in a causal LM |
EncoderDecoder | adapted → SolomonLM | embed → N decoder layers (causal mask) → final LayerNorm → Generator |
Batch, run_epoch, SimpleLossCompute | adapted | LM batches: y = x shifted by 1; ntokens excludes pad |
greedy_decode | adapted | + beam, temperature/top-k/top-p — decode-time only, no architecture change |