Subject
1 entry
Source Code
Bookmarks
Modeling Vocabulary for Big Code Machine Learning
An empirical study of vocabulary modeling decisions for machine learning systems on source code, evaluated across 14,436 projects. It matters because the choices made when tokenizing and preprocessing code vocabularies have an outsized impact on neural language model accuracy, yet were poorly documented before this work.
