bert.txt
2018
Jacob Devlin & the BERT team
BERT
Showed that pre-training a Transformer on lots of text, then fine-tuning it, beats task-specific models across language benchmarks.
The people
- Jacob Devlin
- Ming-Wei Chang
- Kenton Lee
- Kristina Toutanova
Filed under
Key work
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (opens in new tab)
arXiv / NAACL, 2018