spaCy

mirror of https://github.com/explosion/spaCy.git synced 2025-11-15 07:15:52 +03:00

Author	SHA1	Message	Date
Matthew Honnibal	a27b23cc8f	* Have SBD return start/end indices	2015-01-22 22:24:44 +11:00
Matthew Honnibal	d460c28838	* Rename vec to repvec	2015-01-22 02:06:22 +11:00
Matthew Honnibal	8b9d913d97	* Rename vec to repvec	2015-01-22 02:05:58 +11:00
Matthew Honnibal	9cd0b6b3e9	* Various tweaks to Tokens class	2015-01-22 02:05:37 +11:00
Matthew Honnibal	5928d158ce	* Pass the string to Tokens	2015-01-22 02:04:58 +11:00
Matthew Honnibal	45264e356b	* Rename vec to repvec	2015-01-22 02:04:24 +11:00
Matthew Honnibal	5e63c606ad	* Rename vec to repvec	2015-01-22 02:03:54 +11:00
Matthew Honnibal	56e6cf0672	* Add _string attr to Tokens object	2015-01-21 18:57:09 +11:00
Matthew Honnibal	d6ac60e91c	* Bug fixes to sentences method, and improved vector transport for tokens	2015-01-21 18:56:32 +11:00
Matthew Honnibal	f2a229136c	* Fix data_dir=None argument to English class	2015-01-21 18:27:31 +11:00
Matthew Honnibal	ef49b8c179	* Add stop-word flag	2015-01-21 18:22:31 +11:00
Matthew Honnibal	6646bfc5df	* Add LOWER attr	2015-01-21 18:19:08 +11:00
Matthew Honnibal	f149259bf5	* Fix negative indices in tokens	2015-01-20 01:16:29 +11:00
Matthew Honnibal	b65b0c07bf	* Messily hook up vector in tokens	2015-01-19 19:59:55 +11:00
Matthew Honnibal	8ff5b8bd84	* Add attribute for POS scheme	2015-01-17 17:33:16 +11:00
Matthew Honnibal	6c7e44140b	* Work on word vectors, and other stuff	2015-01-17 16:21:17 +11:00
Matthew Honnibal	802867e96a	* Revise interface to Token. Strings now have attribute names like norm1_	2015-01-15 03:51:47 +11:00
Matthew Honnibal	7d3c40de7d	* Tests passing after refactor. API has obvious warts, particularly in Token and Lexeme	2015-01-15 00:33:16 +11:00
Matthew Honnibal	0930892fc1	* Tmp. Working on refactor. Compiles, must hook up lexical feats.	2015-01-14 00:03:48 +11:00
Matthew Honnibal	46da3d74d2	* Tmp. Refactoring, introducing a Lexeme PyObject.	2015-01-12 11:23:44 +11:00
Matthew Honnibal	ce2edd6312	* Tmp commit. Refactoring to create a Python Lexeme class.	2015-01-12 10:26:22 +11:00
Matthew Honnibal	aacaf1a0f0	* Fix parser	2015-01-08 01:19:23 +11:00
Matthew Honnibal	9a21127bf7	* Fix parser, which was importing the wrong model	2015-01-08 00:10:15 +11:00
Matthew Honnibal	6a3e39cdd1	* Add typedefs.pyx	2015-01-06 04:51:40 +11:00
Matthew Honnibal	a58920cc5e	* Import orth.word_shape as a C module	2015-01-06 03:18:22 +11:00
Matthew Honnibal	6b68f7ef75	* Finally get string types right for orth function	2015-01-06 03:17:39 +11:00
Matthew Honnibal	90c143bd85	* Fix orth import	2015-01-05 18:49:19 +11:00
Matthew Honnibal	7689dccd0f	* Remove unused import	2015-01-05 18:48:48 +11:00
Matthew Honnibal	3f1944d688	* Make PyPy work	2015-01-05 17:54:38 +11:00
Matthew Honnibal	a510d9f677	* Another assertion removed	2015-01-05 13:01:40 +11:00
Matthew Honnibal	2856946a66	* Remove assertion that doesn't work on Python 3	2015-01-05 12:51:16 +11:00
Matthew Honnibal	94034f1112	* Fix encoding in lemmatization	2015-01-05 11:54:29 +11:00
Matthew Honnibal	b132b3caa6	* Fix unicode error in lemmatizer	2015-01-05 11:53:54 +11:00
Matthew Honnibal	477e7fbffe	* Fix data reading for lemmatizer	2015-01-05 06:01:32 +11:00
Matthew Honnibal	58f75abaca	* Fix unicode error in orth	2015-01-05 05:53:08 +11:00
Matthew Honnibal	4e085d5166	* Fix lemmatizer for Python3	2015-01-05 05:51:26 +11:00
Matthew Honnibal	ae7c811fd1	* Use Exception instead of StandardError	2015-01-04 01:22:12 +11:00
Matthew Honnibal	0e4c2ba036	* Fix loading of special morph words	2015-01-03 23:13:00 +11:00
Matthew Honnibal	f5d41028b5	* Move around data files for test release	2015-01-03 01:59:22 +11:00
Matthew Honnibal	a24321b63a	* Add downloader	2015-01-02 21:44:41 +11:00
Matthew Honnibal	5d9a096e2f	* Some minor clean-up after HastyModel	2014-12-31 19:46:04 +11:00
Matthew Honnibal	aafaf58cbe	* Refactor _ml.Model, and finish implementing HastyModel so far not worthwhile.	2014-12-31 19:40:59 +11:00
Matthew Honnibal	bcd038e7b6	* Implement HastyModel	2014-12-31 01:16:47 +11:00
Matthew Honnibal	1a075f77ff	* Don't over-ride pre-loaded POS tags, if set by special-cases	2014-12-30 23:26:32 +11:00
Matthew Honnibal	785c7ba76a	* Embed signature on attrs	2014-12-30 23:25:31 +11:00
Matthew Honnibal	30e5805656	* Lazy-load tagger and parser	2014-12-30 23:25:09 +11:00
Matthew Honnibal	9976aa976e	* Messily fix morphology and POS tags on special tokens.	2014-12-30 23:24:37 +11:00
Matthew Honnibal	c1ef3febee	* Embedsignature in tokens.pyx	2014-12-30 21:22:00 +11:00
Matthew Honnibal	aac5028b6e	* Move tagger to _ml	2014-12-30 21:21:38 +11:00
Matthew Honnibal	1ffb0229ed	* Import tokens in parser.pxd	2014-12-30 21:21:17 +11:00
Matthew Honnibal	bb0b00f819	* Repurporse the Tagger class as a generic Model, wrapping thinc's interface	2014-12-30 21:20:15 +11:00
Matthew Honnibal	fe2a5e0370	* Work on docstrings	2014-12-27 21:46:04 +11:00
Matthew Honnibal	bb80937544	* Upd docstrings	2014-12-27 18:45:16 +11:00
Matthew Honnibal	b8b65903fc	* Tmp	2014-12-24 17:42:00 +11:00
Matthew Honnibal	ab61673edd	* Fix api of array method	2014-12-23 15:18:48 +11:00
Matthew Honnibal	7708d0e24a	* Move lemmatizer to en dir	2014-12-23 15:16:57 +11:00
Matthew Honnibal	98eb4c0426	* Fix path to parser model	2014-12-23 15:09:09 +11:00
Matthew Honnibal	b00bc01d8c	* All tests now passing for reorg	2014-12-23 13:18:59 +11:00
Matthew Honnibal	73f200436f	* Tests passing except for morphology/lemmatization stuff	2014-12-23 11:40:32 +11:00
Matthew Honnibal	cf8d26c3d2	* POS tagger training working after reorg	2014-12-22 08:54:47 +11:00
Matthew Honnibal	4c4aa2c5c9	* Work on train	2014-12-22 07:25:43 +11:00
Matthew Honnibal	61df50b598	* Add English-subclass POS tagger	2014-12-21 20:59:07 +11:00
Matthew Honnibal	9f3f07cab6	* Add attrs file for English	2014-12-21 11:29:11 +11:00
Matthew Honnibal	2a89d70429	* Add vocab.pyx to setup, and ensure we can import spacy.en.lang	2014-12-21 06:03:53 +11:00
Matthew Honnibal	b34a1325d3	* Everything compiling after reorg. About to start testing.	2014-12-21 05:42:23 +11:00
Matthew Honnibal	e1c1a4b868	* Tmp	2014-12-21 05:36:29 +11:00
Matthew Honnibal	d11c1edf8c	* Import slice_unicode from strings.pyx	2014-12-20 07:56:26 +11:00
Matthew Honnibal	be1bdcbd85	* Move lang.pyx to tokenizer.pyx	2014-12-20 07:55:40 +11:00
Matthew Honnibal	89a1cc1a48	* Move murmurhash to .pxd in strings file	2014-12-20 07:41:08 +11:00
Matthew Honnibal	d5a942c4a4	* Rename lang.pyx to tokenizer.pyx	2014-12-20 07:30:39 +11:00
Matthew Honnibal	a60ae261ae	* Move tokenizer to its own file, and refactor	2014-12-20 07:29:16 +11:00
Matthew Honnibal	867a4a000c	* Export set_morph_from_dict function	2014-12-20 07:28:27 +11:00
Matthew Honnibal	4e30195c6d	* Refactor morphology.pyx	2014-12-20 07:27:28 +11:00
Matthew Honnibal	4c6ce7ee84	* Update tokens.pyx as part of reorg	2014-12-20 07:03:26 +11:00
Matthew Honnibal	116f7f3bc1	* Rename Lexicon to Vocab, and move it to its own file	2014-12-20 06:54:03 +11:00
Matthew Honnibal	780cbd68b1	* Move all struct definitions to structs.pxd, to avoid circular dependencies	2014-12-20 06:51:33 +11:00
Matthew Honnibal	f6556d8e5d	* Refactor, move Lexeme struct to structs.pxd	2014-12-20 06:51:03 +11:00
Matthew Honnibal	7d48bba6c4	* Move StringStore class to its own file	2014-12-20 06:42:01 +11:00
Matthew Honnibal	b066102d2d	* Remove POS cache for now	2014-12-20 03:49:58 +11:00
Matthew Honnibal	ff252dd535	* Clean up 'guess_cache' idea, which didnt work well enough	2014-12-20 03:49:11 +11:00
Matthew Honnibal	9d3ca13909	* Start work on parse-tree iteration classes	2014-12-20 03:48:10 +11:00
Matthew Honnibal	bed680c632	* Remove commented-out features	2014-12-20 03:47:32 +11:00
Matthew Honnibal	3d178c03ae	* Prune the features a bit	2014-12-20 02:46:14 +11:00
Matthew Honnibal	a0408e1758	* Working DecisionMemory class	2014-12-20 01:43:26 +11:00
Matthew Honnibal	7920ea72b4	* Working parser with the decision memory idea. Disabling that for now, for simplicity	2014-12-20 01:43:15 +11:00
Matthew Honnibal	a2f2a48da9	* Add some extra features	2014-12-20 01:42:24 +11:00
Matthew Honnibal	8fd9762d91	* Start laying out parse tree iteration methods	2014-12-20 01:42:09 +11:00
Matthew Honnibal	53b8bc1f3c	* Work on implementing a trainable cache for the parser. So far, doesn't improve efficiency	2014-12-19 09:30:50 +11:00
Matthew Honnibal	033d6c9ac2	* Adapt POS tagger decision-memory for use in parser	2014-12-19 07:23:04 +11:00
Matthew Honnibal	809ddf7887	* Add index.pxd	2014-12-19 07:23:00 +11:00
Matthew Honnibal	1879abd16a	* Set const-correctness for tagger	2014-12-18 20:41:52 +11:00
Matthew Honnibal	f72243b156	* Set const-correctness for Feature* array	2014-12-18 20:41:32 +11:00
Matthew Honnibal	6ab7e40590	* Add non-monotonic parsing with cost-sensitive update. 92.26 on Y&M set	2014-12-18 11:33:25 +11:00
Matthew Honnibal	7e0c692daf	* Automatically push when the stack is empty	2014-12-18 09:16:10 +11:00
Matthew Honnibal	61142a8eff	* Tweak features	2014-12-18 09:15:03 +11:00
Matthew Honnibal	8446ebfbbb	* Work on parser. Up to 92 UAS on YM labels	2014-12-18 09:05:31 +11:00
Matthew Honnibal	55de747bfc	* Remove .cpp files	2014-12-18 02:43:13 +11:00
Matthew Honnibal	4448a840f7	* Work on greedy parsing. Scoring about 91.2	2014-12-18 02:42:55 +11:00
Matthew Honnibal	87e9487d76	* Work on parser	2014-12-17 21:10:12 +11:00
Matthew Honnibal	9d7d97978d	* Work on greedy parser	2014-12-17 21:09:29 +11:00

1 2 3 4 5 ...

422 Commits