spaCy/spacy/pipeline
Sofie Van Landeghem 311133e579
Train textcat with config (#5143)
* bring back default build_text_classifier method

* remove _set_dims_ hack in favor of proper dim inference

* add tok2vec initialize to unit test

* small fixes

* add unit test for various textcat config settings

* logistic output layer does not have nO

* fix window_size setting

* proper fix

* fix W initialization

* Update textcat training example

* Use ml_datasets
* Convert training data to `Example` format
* Use `n_texts` to set proportionate dev size

* fix _init renaming on latest thinc

* avoid setting a non-existing dim

* update to thinc==8.0.0a2

* add BOW and CNN defaults for easy testing

* various experiments with train_textcat script, fix softmax activation in textcat bow

* allow textcat train script to work on other datasets as well

* have dataset as a parameter

* train textcat from config, with example config

* add config for training textcat

* formatting

* fix exclusive_classes

* fixing BOW for GPU

* bump thinc to 8.0.0a3 (not published yet so CI will fail)

* add in link_vectors_to_models which got deleted

Co-authored-by: Adriane Boyd <adrianeboyd@gmail.com>
2020-03-29 19:40:36 +02:00
..
__init__.py Update spaCy for thinc 8.0.0 (#4920) 2020-01-29 17:06:46 +01:00
entityruler.py Tidy up and auto-format 2020-03-25 12:28:12 +01:00
functions.py Drop Python 2.7 and 3.5 (#4828) 2019-12-22 01:53:56 +01:00
hooks.py Default settings to configurations (#4995) 2020-02-27 18:42:27 +01:00
morphologizer.pyx Tidy up compiler flags and imports (#5071) 2020-03-02 11:48:10 +01:00
pipes.pyx Train textcat with config (#5143) 2020-03-29 19:40:36 +02:00
tok2vec.py Train textcat with config (#5143) 2020-03-29 19:40:36 +02:00