| Dataset | URL | Description |
|---|---|---|
| SQuAD | SQuAD | Stanford Question Answering Dataset, used for training and evaluating question answering systems. |
| SuperGLUE | SuperGLUE | A benchmark for evaluating the performance of natural language understanding systems. |
| WebText | WebText | A dataset created by OpenAI from a variety of web pages, used to train GPT-2. |
| PILE | PILE | A large-scale, diverse, open-source language modeling dataset. |
| BIGQUERY | BIGQUERY | Google's serverless, highly scalable, and cost-effective multi-cloud data warehouse. |
| BIGPYTHON | BIGPYTHON | A dataset for training large-scale language models on Python code. |
Created
July 29, 2024 20:51
-
-
Save pydemo/07d127d347024a60677db7000b05bf1f to your computer and use it in GitHub Desktop.
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment