#5 LSTM + Mixture of Experts in JAX / Flax / Optax
· 24 min read
A hand-written LSTM with a sparse Mixture-of-Experts FFN between two recurrent layers, trained on Tiny Shakespeare. Anchor: Shazeer et al., Outrageously Large Neural Networks (2017) — the original LSTM-with-MoE paper.
Full runnable notebook: createcentury/jax-flax-optax-lab — 02-lstm-moe.
