GPT-2

About

This a GPT-2 fork. The idea for this repo is to be able to update and experinment with GPT-2 in a new way. For example, this repo offers a reengineered version of the orginal GPT (tensorflow) in PyTorch (pytorch).

Notes

Code and models from the paper "Language Models are Unsupervised Multitask Learners".

You can read about GPT-2 and its staged release in our original blog post, 6 month follow-up post, and final post.

We have also released a dataset for researchers to study their behaviors.

^* Note that our original parameter counts were wrong due to an error (in our previous blog posts and paper). Thus you may have seen small referred to as 117M and medium referred to as 345M.

Usage

This repository is meant to be a starting point for researchers and engineers to experiment with GPT-2.

For basic information, see our model card.

gpt-2-output-dataset

This dataset contains:

250K documents from the WebText test set
For each GPT-2 model (trained on the WebText training set), 250K random samples (temperature 1, no truncation) and 250K samples generated with Top-K 40 truncation

We look forward to the research produced using this data!

Download Dataset

For each model, we have a training split of 250K generated examples, as well as validation and test splits of 5K examples.

All data is located in Google Cloud Storage, under the directory gs://gpt-2/output-dataset/v1. (NOTE: everything has been migrated to Azure https://openaipublic.blob.core.windows.net/gpt-2/output-dataset/v1/)

Some caveats

GPT-2 models' robustness and worst case behaviors are not well-understood. As with any machine-learned model, carefully evaluate GPT-2 for your use case, especially if used without fine-tuning or in safety-critical applications where reliability is important.
The dataset our GPT-2 models were trained on contains many texts with biases and factual inaccuracies, and thus GPT-2 models are likely to be biased and inaccurate as well.
To avoid having samples mistaken as human-written, we recommend clearly labeling samples as synthetic before wide dissemination. Our models are often incoherent or inaccurate in subtle ways, which takes more than a quick read for a human to notice.
PyTorch version of GPT-2 is not fully completed

Development

See DEVELOPERS.md

Contributors

See CONTRIBUTORS.md

Citation

Please use the following bibtex entry:

@article{radford2019language,
  title={Language Models are Unsupervised Multitask Learners},
  author={Radford, Alec and Wu, Jeff and Child, Rewon and Luan, David and Amodei, Dario and Sutskever, Ilya},
  year={2019}
}

License

Modified MIT

Name		Name	Last commit message	Last commit date
Latest commit History 76 Commits
dataset		dataset
pytorch		pytorch
tensorflow		tensorflow
.gitattributes		.gitattributes
.gitignore		.gitignore
CONTRIBUTORS.md		CONTRIBUTORS.md
DEVELOPERS.md		DEVELOPERS.md
Dockerfile.cpu		Dockerfile.cpu
Dockerfile.gpu		Dockerfile.gpu
LICENSE		LICENSE
README.md		README.md
domains.txt		domains.txt
download_model.py		download_model.py
model_card.md		model_card.md
requirements.txt		requirements.txt

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

GPT-2

About

Notes

Usage

gpt-2-output-dataset

Download Dataset

Some caveats

Development

Contributors

Citation

License

About

Releases 7

Packages

Languages

License

MitchellShibilski-Unkel/GPT-2

Folders and files

Latest commit

History

Repository files navigation

GPT-2

About

Notes

Usage

gpt-2-output-dataset

Download Dataset

Some caveats

Development

Contributors

Citation

License

About

Topics

Resources

License

Stars

Watchers

Forks

Releases 7

Packages 0

Languages

Packages