Support for sparse tiledb arrays for the X matrix
1. cxgtool can now output sparse matrices
2. cxg_adaptor and diffexp_cxg updated to handle sparse matrices
3. added a test in test_diffexp to test sparse diffexp and get_X_array
6.7 does not have the `hidden` flag used in the code. Users building the
app with an older version of click within the current range specified by
requirements.txt may fail.
* hello world
* stuff
* successful build
* updates"
* maybe a basic example
* simplify
* reamde into dockerfile
* some more stuff
* Release procfile
* package.json at top levle
* don't release in procfile for now
* more package.json stuff
* copy assets
* merge master
* not in the relase phase
* revert not necessary
* pin gunicorn version
* reset common.mk
* Update package.json
Co-authored-by: Madison Dunitz <dunitzm@gmail.com>
* s3 region should have a single config param
The s3 region can also now be automatically determined to further
reduce errors.
This patch also fixes a bug with order of handling the config params.
The tiledb config needs to be fixed before attempting to load
(need to handle_adaptor before handle_single_dataset)
* Introduce a config file to cellxgene
The config file format is in yaml. The default config is located
in server/common/default_config.py. A user may create a yaml file
that contains a subset of these fields. It can be used during cellxgene
launch, or for hosted cellxgene.
The code has also been refactored. Much of the logic to check arguments
has moved from launch to app config.
It is now possible to set the tiledb context parameters using the config
file. Other feature will soon be handled in a similar way.
This PR contains a refactoring to make adding new features easier.
The new features include supporting the tiledb format, and the multi dataset application.
The refactoring includes
Simplifying the directory structure and files.
a class structure to handle annotations (currently one type: AnnotationsLocalFile).
a class to handle application configuration
a class structure to handle matrix data (currently AnndataAdaptor and CxgAdaptor). CxgAdaptor uses tiledb.
Algorithms that were previously dependent on the scanpy anndata object are now generalized to work with an abstract interface.
The multi dataset option is not fully supported yet, and so the option to use it is hidden.
Use "cli launch --dataroot ..."
To access this feature.
All combinations of app single dataset/ app multi dataset and AnndataAdaptor/CxgAdaptor work with all the features, such as annotations, ontologies, diffexp.
* revert MatrixProxy; replace with correct use of adata slicing
* work around 0.6 adata slicing bug
* fix incorrect var slice
* simplify slicing of X
* add warning about performance impact of anndata<=0.7
* lint and remove unused code
* improve comment
* lint
* correctly parse versions
* temp files should preserve file suffix if possible - anndata 0.7 compat
* update anndata dependency to 0.6.20
* resolve PR review comments
* add sample ontologies file
* add ontologies reducer
* Move select category to own component
* Dialog and Input factored out
* refactoring categorical, partway
* validationn
* anno
* suggest populates input
* frontend for ontology working
* initial implementation of back-end support for ontologies
* edit is now dialog again
* autosuggest working on edit
* part way through create arbitrary label
* handle choice in function
* pass duplicate cat prop
* editing works
* update test to match new CLI params
* fix occupancy alignment
* edit category as dialogue
* secondary button
* remove stubbed out ontologies
* add label setting upon new label creation
* Update legal characters for labels (#1119)
* Allow any term in the ontology (bypass legal name check)
* Add hyphens and parens to legal characters in names
* improve performance for large ontologies
* correctly handle case where ontologies are disabled
* fix logic error in CLI
Co-authored-by: Bruce Martin <bruce@chanzuckerberg.com>
* PR cleanup 1
* lint
* validate user generated labels
* finish hooking up connected suggest component
* protect against undefined callbacks
* Fix illegal characters error message
* break out npm run commands
* fix error detection on label edit
Co-authored-by: Bruce Martin <bruce@chanzuckerberg.com>
Co-authored-by: Sidney Bell <sidneymbell@users.noreply.github.com>
* initial cut at backed mode
* make flask multithreading conditional on debug flag
* update X access to support backed mode
* lint
* improve help message for backed mode
* fix tests
* add MatrixProxy to normalize supported matrix types
* add FAQ entry for --backed
* remove use of matrix.T
* clean up
* add ability to disable diffexp from CLI; add hueristic to detect likely slow diffexp calculation, and warn user
* fix tests
* do not print diffexp speed warning if diffexp is disabled
* tweak wording of diffexp speed messages
* add FAQ entry on --disable-diffexp
* revise heuristic for warning about slow diffexp
* use quick tooltip delay on diffexp button
* initial commit of URL support for launch
* lint
* modify tests to use new data locator
* add locator unit tests
* fix typo in faq
* more lint
* update faq per PR review
* dead code and route removal
* more dead code cleanup
* fix scanpy_engine tests
* lint
* add missing catch in filter parsing
* update scanpy NaN tests
* more fbs tests and dead test removal
* remove forced default for content type negotiation
* bit of cleanup
* more fbs test cleanup
* lint
* remove swagger
* swagger cleanup
* lint
* correctly handle lack of templates
* more dead code removal
* remove unused files
* fix dev build
* lint
* first flatbuffer schema
* do not lint auto-generated files
* add flatbuffers package
* add flatbuffer module
* wire up /data/X/T route
* use flatbuffers for matrix data fetc
* clarity and comments
* add flatbuffer layout route
* clean up obsolete code
* fix tests
* move flake8 config to setup.cfg
* add comments
* lint
* rework layout routes for fbs
* add more type support to fbs
* lint
* add flatbuffer support for annotations
* function name improvements
* fix botched merge with master
* remove unused import
* route cleanup for flatbuffers
* rename function for clarity
* add missing globals to Jest tests
* fix client JS tests
* fix routes for Python tests
* comments for clarity
* non-finite floating point hardening
* more non-finite number handling
* lint
* fix tests for summarizeAnnotations
* harden diffexp calculation against FP errors
* cleanup unused code
* lint
* add encoding tests for flatbuffers
* application type specified as strings
* fix spelling error
* improve variable names
* add note about documentation gap
* rename FBS DataFrame to Matrix
* fix anaconda build
* Added link for TKAgg
* Add matplotlib to requirements
We are pulling it in through scanpy, but since we are importing it directly we should include it explicitly
* remove memoization
* update to flash 1.0.2; turn on threading
* stand-alone helper routines for array slicing
* fix issue #405
* reset diffexp state when world changes
* performance work in dimension creation; fix world slicing bug
* update tests to match new state mgmt api
* update flask
* do not make dimensions for useless annotations
* update test to match optimizations
* add --obs-names and --var-names CLI params
* fix lint
* performance improvements in scanpy engine
* fix lint
* fix typo
* correctly handle sparse formats in diffexp
* fix diffexp and 1d slicing
* diffexp uses t-stat, not pval; clean up arg handling
* make _slice a static method
* revise scanpy tests to match new API
* add prepare cli
* fix handling of user path
* fixes for linter
* add flags and options for handling obs and var names
* add prepare cli
* fix handling of user path
* fixes for linter
* add flags and options for handling obs and var names
* address review requests
* range encode filter range lists
* speed up data load
* add comment on scanpy read params
* update to latest scanpy/anndata
* performance improvments in data loading
* fix typo
* work around scanpy bug
* remove debugging print statements
* Upgrade version of scanpy
* /data/var
This works for everything except the case where there is only one gene. Anndata flattens X when there is only one var thus causing the transpose to fail.
* Fix edge case when an axis (obs/var) only contains 1 element
Move to new REST v0.2 communication between front and back-end. This is a first cut implementation which is functional, but will need follow-up enhancements for performance, error checking, etc. Protocol spec is in docs directory.
* Add filtering via indexing
* Using new filter specs
Indexing working
* Added filtering by annotation value
* factor out common methods
* Documentation
* create enum for axis (obs/var)
* Better description for filter's return
* Add boolean to enumerated types
* Augmented enum for scanpy axis
* Create schema for annotations
Based on datatype within scanpy/anndata
+ tests
* remove obsolete schema parse script
* Update rest api to remove old routes and add schema route
* Separate development requirements
* Warning for unsupported datatypes
* include -r requirements.txt in dev
* Merged downcast warnings
* Fixed bug where names were NaNs
Needed to include the index too when creating the series
* Add config endpoint
* Generate app features from CLI selections
* Move features to driver
* Add tests for schema
* Clearer version wording
* python3 version of super
* version from engine to package level
* move features to driver
* Revise layout function to match the new spec
* GET for layout/obs
* PUT Layout (#211)
* PUT Layout
* Csweaver/annotations (#212)
* Update scanpy engine to support the rest v0.2 annotation requests
* GET endpoint for obs annotations + tests
* Documentation
* Test annotations in scanpy engine
* Description for annotation-keys param
* annotation->annotations
* clarified return for annotations
* Use URL query list for annotations fields
* parse_filter parses v0.2 GET filters (#215)
* parse_filter parses v0.2 GET filters
* Don't allow index filters from query params
* Better variable conversion
* Parse filter improvements
- uses default dict
- renamed filter -> query_filter
* Cleanup Tasks (#216)
* Add test_api back into travis build
* Do custom JSON encoding the correct way
* Run cellxgene server in test setup
* Cleanup new tests too
* Option to bind to all interfaces (#225)
app.run("0.0.0.0") instead of app.run("127.0.0.1") binds to all interfaces.
Note: There are comments on the internet that says that the flask server is not up to the task of production serving. I don't think that such scalability concerns apply here, but I was able to get cellxgene working with twistd relatively easily, and we could switch to that if there are scalability concerns.
Test plan: browsed to <ip>:5005/api/v0.2/config on a different host.
* Add filtering via indexing
* Using new filter specs
Indexing working
* Added filtering by annotation value
* factor out common methods
* Documentation
* create enum for axis (obs/var)
* Better description for filter's return
* Add boolean to enumerated types
* Augmented enum for scanpy axis
* Create schema for annotations
Based on datatype within scanpy/anndata
+ tests
* remove obsolete schema parse script
* Update rest api to remove old routes and add schema route
* Separate development requirements
* Warning for unsupported datatypes
* include -r requirements.txt in dev
* Merged downcast warnings
* Fixed bug where names were NaNs
Needed to include the index too when creating the series
* Add config endpoint
* Generate app features from CLI selections
* Move features to driver
* Add tests for schema
* Clearer version wording
* python3 version of super
* version from engine to package level
* move features to driver
* Revise layout function to match the new spec
* GET for layout/obs
* PUT Layout (#211)
* PUT Layout
* Csweaver/annotations (#212)
* Update scanpy engine to support the rest v0.2 annotation requests
* GET endpoint for obs annotations + tests
* Documentation
* Test annotations in scanpy engine
* Description for annotation-keys param
* annotation->annotations
* clarified return for annotations
* Use URL query list for annotations fields
* parse_filter parses v0.2 GET filters (#215)
* parse_filter parses v0.2 GET filters
* Don't allow index filters from query params
* Better variable conversion
* Parse filter improvements
- uses default dict
- renamed filter -> query_filter
* Cleanup Tasks (#216)
* Add test_api back into travis build
* Do custom JSON encoding the correct way
* Run cellxgene server in test setup
* Cleanup new tests too
* Option to bind to all interfaces (#225)
app.run("0.0.0.0") instead of app.run("127.0.0.1") binds to all interfaces.
Note: There are comments on the internet that says that the flask server is not up to the task of production serving. I don't think that such scalability concerns apply here, but I was able to get cellxgene working with twistd relatively easily, and we could switch to that if there are scalability concerns.
Test plan: browsed to <ip>:5005/api/v0.2/config on a different host.
* Fix merge errors
- import warnings was improperly deleted
- scanpy engine tests were totally wrong
* Fix merge error with driver
* PUT /annotations (#235)
* Add query param for annotation name
* fix descriptions, eliminate else clause
* first cut at initial data load on rest 0.2 api
* Annotation var (#248)
* Fix bug strings are always objects in pandas
* Add axis to annotation method
* Add /annotation/var to REST api
* Csweaver/expressiondata (#242)
* Refactor expression method for REST v2
* Add message to QueryStringError
* Fix range filters
* Add GET route for /data
* /data PUT route
* rename expression to data_frame
* clarification of error
* Improve accept type handling
* support all schema types for 0.2 REST API
* remove REST 0.1 code; connect var annotations loading
* config reducer; use config to set data set title; remove obsolete templating code for data set title
* REST 0.2 expression conversion support
* partial port of expression to REST 0.2
* diffexp (#273)
* Add diffexp method to scanpy
and test
* Minor tweaks to diffexp
Get a minimal working version to unblock FE development
* Fixing things git deleted
* cleanup print statements
* Add index test
* additional, partial REST 0.2 bring up of diffexp
* Ignore unstructured annotations for data (#275)
This is a temp hack, need to figure out how to include data.uns if there is only one gene
* diffexp REST 0.2 port finish
* ignore unstructured annotaitons on all routes except layout
* correctly use varDataCache; maintain state during world rebuild
* correct varDataCache use
* temporarily disable all memoization
* refinements to expression data caching
* clear cell sets upon regraph/reset
* update version of REST to 0.2
* Travis build fixes
- comment out cache import
- fix duplicate test name
* Remove dependency from travis
* clarify semantics of config variables
* move generic action helpers into util