| language | ko | |||
|---|---|---|---|---|
| tags |
|
|||
| license | MIT | |||
| datasets |
|
|||
| metrics |
|
This model is no longer actively in training due to lack of computing resources. NIPA์ธก์ ์ ์ฑ ๋ณ๊ฒฝ์ผ๋ก ๊ณ ์ฑ๋ฅ์ปดํจํ ์ง์๋์์์ ์ ์ธ๋จ์ ๋ฐ๋ผ, ๋ชจ๋ธ ํ๋ จ์ ๋ ์ด์ ์งํํ ์ ์๊ฒ ๋์์ต๋๋ค.
KoGPT2
Demo NOT available at : ์๋ฌด๋ง ๋์์น
GPT2 and GPT3 trained on ~40GB of Korean datasets. see the included json files for hyperparameter details.
Available models (Training ATM):
- KoGPT2-base(117M)
- KoGPT2-medium(345M)
- KoGPT2-large(774M)
- KoGPT2-xlarge(1.5B)
- KoGPT2-2.7B TBA
Models are available as TF checkpoint files Training script or Huggingface transformers compatible ones
n_ctx available : 1024 2048 384
Intended for Korean text generation for ai-text-adventure(https://github.com/ksjae/ai-text-adventure) with PPLM.
Download files from links at the Releases tab. Alternatively, it is available from my own server(wget-friendly)
Try out on colab
or go to KoGPT2-train and use scripts/demo.py
v0.1 may have faulty tokenizers, producing bad outputs.
v0.2+ be GPT2 with n_ctx of 2048. True form of GPT-3 implementation(alternating layers) will not be available within the year.
If other limitations or errors are found, please open an issue.
Initialized with GPT(774M,https://github.com/openai/gpt-2/blob/master/model_card.md).
The following data was used, and is available for redistribution here:
- Namuwiki database dump, Early 2020
- KCC(Kookmin University Corpus)
- Dump of Korean Wikipedia
- NAVER movie reviews
- Korean news(about 1GB) from Leipzig(a German university)
- Context data from KorSQUAD questions
- Parsed Korean CommonCrawl data(WIP)
Please note the completed dataset includes <|endoftext|> tags.
The following data were used, but is unavailable for redistribution:
- Sejong Corpus
- '๋ชจ๋์ ๋ง๋ญ์น' from corpus.korean.go.kr
- A PRIVATE collection of korean novels
- Webcrawl of modern, uploaded text novels('ํ ๋ณธ' - If you want to prevent your novel from going in the training set, please contact me and I will blacklist it)
- Game storylines (with authors' approval)
All hyperparameters are the same as GPT2-large One paragraph per line(TextDataset)
Early models(GPT2-large v0.2 and prior) are trained on 2xTesla V100 for 3~4 weeks. Models up to XL size are trained on v3-8 TPUs.
prompt >>> ๋๋ ์ด๋์ด ์ฒ ์์ ๊ฑฐ๋๊ณ ์๋ค.
์ด๋์ ๋๋ ๊ทธ ์์ ์ฐ๋ค์ ํฅํด ๋ฐ๊ธฐ ์์ํ๋ค. ๊ทธ๋ฆฌ๊ณ ๋ด ์์ผ์๋ ์ด ๊ณจ์ง๊ธฐ์ ๋ํ ์ด๋ค ๋๊ฒฝ๋, ํน์ ๋๊ฒฝ๊ณผ ํํฌ์กฐ์ฐจ ์ฟ๋ณด์๋ค๊ฐ ์ฌ๋ผ์ก๋ค๊ฐ๋ ์ฌ๋ผ์ ธ ๋ฒ๋ฆฌ๊ณ ๋ง์ ๋ค. ๊ทธ๋ฌ๋ ๋ ์ญ์ ๊ทธ๊ฒ์ ๋ฏฟ์ง ์์๋ค. ์๋ ๊ทธ๊ฒ๋ ๋ชจ๋ฅธ๋คโฆโฆ. ๊ทธ๋ ๋ค๋ฉด ๊ทธ๊ฒ์ ๋ ๋ฌด์จ ๋ง์ธ๊ฐ? ๋ด๊ฐ ์ด๋ ๊ฒ ๋งํด๋ ์ข์ ํ ๋ฐโฆโฆ ํ์ง๋ง ์ด์จ๋ ์ด๊ณณ์ ์ ๊ทธ๋ฆฌ ๋จผ ๊ณณ์์๋ถํฐ ์ฐพ์์ฌ๊น ํ๊ณ ์๊ฐํด ๋ณด์๋ค. ๊ทธ๋์ ๊ทธ๋ ์ง๊ธ ์ด๋ ํ ๊ณณ๋ง์ ํค๋งค๋ฉด์ ์ ๊พธ๋ง ๊ฑธ์ด์ค๋ ๊ฒ์ผ๊น?
๋๋ ์ด๋์ด ์ฒ ์์ ๊ฑฐ๋๊ณ ์๋ค. #์จ๋ฆฌ๊ฑดํธํ ์ด๋ธ ์ #์ฝ์ฝ๋์ ค๋ฆฌ ์์ด ๋ ์์๊ณ ๋ง๋๋ค. #๊ทธ๋ฆญ์๊ฑฐํธ ์๋ฌผ์ด ๋ค์ด๊ฐ ๋ง์ง๋ง ๋ฒ ์ด์ปค๋ฆฌ #๋ง์ฐจ์ด์ฝ์ผ์ดํฌ ๋ ๋ง๋๊ณ ๊ฒ์ ์ด์ฝ๋ ๋๊ป๊ฒ ์นํ๊ณ ..โจ๐๐๐๐ป . . #์จ๋ฆฌ๊ฑดํธํ ์ด๋ธ @eleganttable_ #๊ฐ์ #์นดํ์ด๋์ธ๋ฒ ์ด์ปค๋ฆฌ๋ฒ ์ด์ปค #์ฝฉ์ฝฉ๋ณผ #์ฟ ํค #๋์ ํธ๋ง์ง #๋์ ํธ์นดํ #dessert #์นดํ์คํ๊ทธ๋จ #๋ง์คํ๊ทธ๋จ #๋จน์คํ ๊ทธ๋จ #๋จน๋ฐฉ #์ผ์ #daily
๋๋ ์ด๋์ด ์ฒ ์์ ๊ฑฐ๋๊ณ ์๋ค. ์ ๋ช ์ด ์ด์์์ ์ง์ ์ฑ ๊ณณ์ ๊ฑธ์ด๋์๋ค. ... ... ... ... ์ฌ๋งํ ๋์ด๊น์ง ์ฌ๋ผ๊ฐ์ผํ๋?
@unpublished{CitekeyUnpublished,
author = "Seungjae Kim",
title = "Introducing larger KoGPT2",
year = 2020
}