RWKV

1 update on RWKV.

  1. RWKV trains like a transformer and runs with constant memory per token

    The paper releases pretrained RNN weights from 169M to 14B parameters and reports constant time and memory per token during inference.