Efficient training of language models to fill in the middle
We show that autoregressive language models can learn to infill text after we apply a straightforward transformation to the dataset, which simply moves a span of text from the middle of a document to its end. While this data augmentation has garnered much...
Topics: ModelsInfrastructure
Entities: ModelsInfrastructurePerplexity