
Training an LLM on 160GB of 1800s English Text
A researcher has compiled a 160GB dataset of 1800-1875 English texts to train specialized language models. A 500M parameter evaluation model is currently available, with plans for a larger 2B parameter model.






