spaCy

mirror of https://github.com/explosion/spaCy.git synced 2025-07-15 02:32:37 +03:00

Author	SHA1	Message	Date
Matthew Honnibal	a78ad4152d	* Broken version being refactored for docs	2014-08-20 13:39:39 +02:00
Matthew Honnibal	5fddb8d165	* Working refactor, with updated data model for Lexemes	2014-08-19 04:21:20 +02:00
Matthew Honnibal	1b71cbfe28	* Roll back to using unicode, and never Py_UNICODE. No dependence on murmurhash either.	2014-08-18 20:48:48 +02:00
Matthew Honnibal	bbf9a2c944	* Working version that uses arrays for chunks, which should be more memory efficient	2014-08-18 20:23:54 +02:00
Matthew Honnibal	8d3f6082be	* Working version, adding improvements	2014-08-18 19:59:59 +02:00
Matthew Honnibal	01469b0888	* Refactor spacy so that chunks return arrays of lexemes, so that there is properly one lexeme per word.	2014-08-18 19:14:00 +02:00
Matthew Honnibal	b94c9b72c9	* WordTree in use. Need to reform the way chunks are handled. Should be properly one Lexeme per word, with split points being the things that are cached.	2014-08-16 20:10:22 +02:00
Matthew Honnibal	34b68a18ab	* Progress to getting WordTree working. Tests pass, but so far it's slower.	2014-08-16 19:59:38 +02:00
Matthew Honnibal	515d41d325	* Restore string saving to spacy	2014-08-16 16:09:24 +02:00
Matthew Honnibal	a225ca5b0d	* Refactoring tokenizer	2014-08-16 03:22:03 +02:00
Matthew Honnibal	f11c8e22eb	* Remove happax stuff	2014-08-02 22:11:28 +01:00
Matthew Honnibal	d6e07aa922	* Switch to 32bit hash for strings	2014-08-02 21:51:52 +01:00
Matthew Honnibal	365a2af756	* Restore happax. commit uncommited work	2014-08-02 21:27:03 +01:00
Matthew Honnibal	18fb76b2c4	* Removed happax. Not sure if good idea.	2014-08-02 20:53:35 +01:00
Matthew Honnibal	fc7c10d7f8	* Ugly but seemingly working fix to the token memory leak	2014-08-01 09:43:19 +01:00
Matthew Honnibal	f39211b2b1	* Add FixedTable for hashing	2014-08-01 07:27:21 +01:00
Matthew Honnibal	5b81ee716f	* Use a sparse_hash_map to store happax vocab items, with a max size.	2014-07-31 17:40:43 +01:00
Matthew Honnibal	b9016c4633	* Switch to using sparsehash and murmurhash libraries out of pip	2014-07-25 15:47:27 +01:00
Matthew Honnibal	057c21969b	* Refactor for string view features. Working on setting up flags and enums.	2014-07-07 16:58:48 +02:00
Matthew Honnibal	f1bcbd4c4e	* Reorganized code to accomodate Tokens class. Need string views before group_by and count_by can be done well.	2014-07-07 12:47:21 +02:00
Matthew Honnibal	ff1869ff07	* Fixed major efficiency problem, from not quite grokking pass by reference in cython c++	2014-07-07 07:36:43 +02:00
Matthew Honnibal	d5bef02c72	* Reorganized, moving language-independent stuff to spacy. The functions in spacy ask for the dictionaries and split function on input, but the language-specific modules are curried versions that use the globals	2014-07-07 04:21:06 +02:00
Matthew Honnibal	556f6a18ca	* Initial commit. Tests passing for punctuation handling. Need contractions, file transport, tokenize function, etc.	2014-07-05 20:51:42 +02:00

23 Commits