bert.txt

2018

Jacob Devlin & the BERT team

BERT

Showed that pre-training a Transformer on lots of text, then fine-tuning it, beats task-specific models across language benchmarks.

The people

  • Jacob Devlin
  • Ming-Wei Chang
  • Kenton Lee
  • Kristina Toutanova

Filed under

Key work

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (opens in new tab)

arXiv / NAACL, 2018

Sources