Vadim Mazaev
|
81314f8659
|
Fixed tokenizer: added char classes; added first lemmatizer and
tokenizer tests
|
2017-11-21 22:23:59 +03:00 |
|
Vadim Mazaev
|
52ee1f9bf9
|
Updated Russian Language, added lemmatizer, norm exceptions and lex
attrs
|
2017-11-21 11:44:46 +03:00 |
|
Vadim Mazaev
|
a0739a06d4
|
Returned russian support from v1.10 branch
|
2017-11-17 17:06:15 +03:00 |
|
yuukos
|
7401152289
|
updated Russian tokenizer
moved the trying to import pymorph into __init__
|
2017-11-17 17:04:50 +03:00 |
|
yuukos
|
3aad66cf00
|
added russian language support
|
2017-11-17 17:04:22 +03:00 |
|
Ines Montani
|
1e3068ec33
|
Merge pull request #1594 from pavillet/patch-1
Update _spacy.jade
|
2017-11-16 23:42:59 +00:00 |
|
pavillet
|
ad2935f0c3
|
Update _spacy.jade
Doc example gives 'object is not subscriptable' error.
Correcting as an attribuet
|
2017-11-17 00:02:20 +01:00 |
|
ines
|
a3d4dd1a5d
|
Test adding of lots of pipeline components (see #1585)
Just to make sure that there's no error now or in the future with adding a large number of pipeline components.
|
2017-11-15 17:28:06 +01:00 |
|
Matthew Honnibal
|
b60d92aca8
|
Increment version
|
2017-11-15 16:14:46 +01:00 |
|
Matthew Honnibal
|
cf0be62096
|
Increment version
|
2017-11-15 15:00:18 +01:00 |
|
Matthew Honnibal
|
716ccbb71e
|
Require thinc 6.10.1
|
2017-11-15 14:59:34 +01:00 |
|
ines
|
97a4f9362b
|
Merge branch 'master' of https://github.com/explosion/spaCy
|
2017-11-15 14:24:00 +01:00 |
|
ines
|
8e65247886
|
Fix lex.id if vectors is None
|
2017-11-15 14:23:58 +01:00 |
|
Matthew Honnibal
|
437ad1a852
|
Merge pull request #1570 from explosion/feature/fix-beam-leak
Fix memory leak in beam parser
|
2017-11-15 14:15:05 +01:00 |
|
Matthew Honnibal
|
2f169fdb0a
|
Set lex ID correctly for new tokens in Vocab
|
2017-11-15 13:58:03 +01:00 |
|
Matthew Honnibal
|
fe3c42a06b
|
Fix caching in tokenizer
|
2017-11-15 13:55:46 +01:00 |
|
Matthew Honnibal
|
8d692771f6
|
Improve profiling
|
2017-11-15 13:51:25 +01:00 |
|
Matthew Honnibal
|
b797dca977
|
Merge branch 'master' of https://github.com/explosion/spaCy
|
2017-11-15 13:11:43 +01:00 |
|
Ines Montani
|
9177c7d7aa
|
Merge pull request #1583 from yogendrasoni/master (resolves #1582)
Add rstrip after reading line from vec file #1582
|
2017-11-15 12:02:48 +00:00 |
|
ines
|
c9d72de0fb
|
Add dummy serialization methods for Japanese and missing lang getter (resolves #1557)
|
2017-11-15 12:44:02 +01:00 |
|
yogendrasoni
|
334ed433b2
|
rstrip line before rsplit
loading english fast text giving error because line contains new line at the end and rsplit is splitting it incorrectly
|
2017-11-15 13:55:08 +05:30 |
|
Matthew Honnibal
|
d274d3a3b9
|
Let beam forward use minibatches
|
2017-11-15 00:51:42 +01:00 |
|
Matthew Honnibal
|
855872f872
|
Remove state hashing
|
2017-11-14 23:36:46 +01:00 |
|
ines
|
40c4e8fc09
|
Remove "optional" from dev_data arg and add more info (see #1578)
|
2017-11-14 20:26:05 +01:00 |
|
Matthew Honnibal
|
2512ea9eeb
|
Fix memory leak in beam parser
|
2017-11-14 02:11:40 +01:00 |
|
Ines Montani
|
48b6cfe59e
|
Merge pull request #1569 from KMLDS/patch-1
trivial typo in docs
|
2017-11-14 01:46:34 +01:00 |
|
Matthew Honnibal
|
86ddf692a1
|
Fix bug in limit calculation on dev data
|
2017-11-14 01:37:10 +01:00 |
|
KMLDS
|
d5b20ac3b6
|
Update span.jade
|
2017-11-13 19:27:20 -05:00 |
|
Ines Montani
|
ea6c85c67a
|
Merge pull request #1566 from MathiasDesch/master (resolves #1248)
Add exceptions to tokenizer and norm
|
2017-11-13 19:05:22 +01:00 |
|
Matthew Honnibal
|
1b348389bb
|
Merge branch 'master' of https://github.com/explosion/spaCy
|
2017-11-13 18:18:48 +01:00 |
|
Matthew Honnibal
|
ca73d0d8fe
|
Cleanup states after beam parsing, explicitly
|
2017-11-13 18:18:26 +01:00 |
|
Matthew Honnibal
|
63ef9a2e73
|
Remove __dealloc__ from ParserBeam
|
2017-11-13 18:18:08 +01:00 |
|
Mathias Deschamps
|
d82f868e1c
|
Ignore pycharm project files
|
2017-11-13 17:46:05 +01:00 |
|
Mathias Deschamps
|
c0691b2ab4
|
Add tokenizer exceptions for ing verbs
Extend list of tokenizing exceptions introduced in 123810b
|
2017-11-13 17:46:05 +01:00 |
|
Mathias Deschamps
|
288298ead9
|
Add norm exception for ing verbs
Some ing verbs are sometimes written in or in'. Make the NORM form correct
|
2017-11-13 17:46:05 +01:00 |
|
ines
|
0e5642593e
|
Merge branch 'master' of https://github.com/explosion/spaCy
|
2017-11-13 17:00:07 +01:00 |
|
ines
|
bc79274706
|
Fix typo
|
2017-11-13 17:00:03 +01:00 |
|
Ines Montani
|
339675c9fb
|
Merge pull request #1565 from DuyguA/patch-2
added contributor agreement for DuyguA
|
2017-11-13 16:21:50 +01:00 |
|
Ines Montani
|
6ef702f79f
|
Merge pull request #1563 from abhi18av/patch-2
improved upon the list of included stop_words
|
2017-11-13 16:14:28 +01:00 |
|
Duygu Altinok
|
c263c3acce
|
added contributor agreement for DuyguA
|
2017-11-13 15:45:13 +01:00 |
|
Abhinav Sharma
|
4dd34058a2
|
Create abhi18av.md
|
2017-11-13 17:23:05 +05:30 |
|
Abhinav Sharma
|
59f5740ede
|
improved upon the list of included stop_words
|
2017-11-13 17:13:49 +05:30 |
|
ines
|
7a7b01feb1
|
Update links
|
2017-11-13 08:30:06 +01:00 |
|
ines
|
b3e502a076
|
Add videos section to resources
|
2017-11-13 08:29:57 +01:00 |
|
ines
|
f2b6b98b75
|
Fix typo in code example (resolves #1556)
|
2017-11-13 08:29:16 +01:00 |
|
Matthew Honnibal
|
f0e28e8ae5
|
Make fasttext reader accommodate whitespace
|
2017-11-12 12:07:13 +01:00 |
|
Ines Montani
|
94d8b711a3
|
Update CONTRIBUTING.md
|
2017-11-12 12:06:59 +01:00 |
|
Matthew Honnibal
|
6e641f46d4
|
Create a preprocess function that gets bigrams
|
2017-11-12 00:43:41 +01:00 |
|
Matthew Honnibal
|
86d37301c9
|
Merge pull request #1552 from ligser/master
Try to add ability to clean up StringStore in pipe
|
2017-11-11 18:39:48 +01:00 |
|
Matthew Honnibal
|
c9251d79e3
|
Edit comment
|
2017-11-11 18:38:32 +01:00 |
|