* Add basic authentication in the server
A pattern for creating authentication methods is introduced, with three
authentication types defined:
none - no authentication
session - like the current session based auth used for user annotations
test - used to test the login/logout process end to end
The config endpoint now returns informations about the authentication, like if
the user is authenticated and their username. The redirect uri's for login and
logout are also returned if the authentication type requires login
This is the first a several PRs for authentication.
*. Update server tests to avoid hardcoded ports
test_api and test_nan_rest now use a common function for starting a test server,
than will initially choose a random port.
* Small fix for handling display versions
Making a distinction between __version__ and the version we display in the info panel (displayr_version).
The hosted cellxgene can overwrite the display_version using a plugin.
Improve version handling in the customized assets
This will give us the ability to specify different config options for
different dataroots.
the key of the dataroot dictionary is no longer the same as the dataroot_url.
Previously key==dataroot_url, and now those are separated.
Added an "is_multi_dataset" function to simplify logic where it branched on single vs multi.
Simplified the rest.py interface by no longer passing in the user annotations object, since
that can be retrieved from the dataset.
Many of our matrices are log normalized, which tends to eliminate
the number of non zero values (if there were any). This prevents
the matrix from being stored as a sparse matrix. The solution here
is to use a simple transformation to make it sparse again. The most
common value from each column is subtracted from that column. These
values that were subtracted are saved in an array called X_col_shift.
The cellxgene code needs to understand how to undo the transformation when
operating over the X matrix.
- added script to create a synthetic dataset for testing
- added a script to convert an existing CXG dataset to a sparse CXG dataset
Support for sparse tiledb arrays for the X matrix
1. cxgtool can now output sparse matrices
2. cxg_adaptor and diffexp_cxg updated to handle sparse matrices
3. added a test in test_diffexp to test sparse diffexp and get_X_array
* app_config, fix bug with list/tuple command line arguments.
There was a error caused by pyyaml using lists, and click using tuples.
Now tuples are automatically converted to lists when the config is
updated.
* hosted, update order to look for config file.
The app now uses a local config.yaml file bundled with the artifact
(if present), if it exists, then looks in the CXG_CONFIG_FILE
environment variable. This is the reverse of previous behavior.
The purpose of this change is to move away from using the
config file on s3, since that could lead to problem where an older
version of the app uses a newer version of the config.
Also in this PR:
1. Changed documentation around dataroot, to describe the posibility of using lustre.
2. Added a few improvements around the secret manager region name. If we use lustre for dataroot and a local config file, then we will no longer be able to
auto determine the region for the secret manager. I plan to start using the
environment variable option for hosted cellxgene.
* small edit to README
Co-authored-by: Severiano Badajoz <sbadajoz@chanzuckerberg.com>
* Improve diffexp for tiledb
- The rows from the A and B sets are gathered and processed at the same time. In this
way the matrix is only accessed once instead of twice for each tile.
- There is now a single thread queue that gets shared between all callers of the diffexp.
This will slow down work if diffexp gets too busy.
- There is a target_workunit amount of work given to each thread. Previously the
workunit was (rows selected * width of tile), which could be small. Now multiple
column tiles can be combined into one workunit. If the target is too small then
thread and other overheads may reduce performance. If target_workunit is too large
then the size of the gathered sub matrix may take up too much memory.
- add configuration parameters (max_workers, cpu_multiplier, and target_workunit)
* Specialize diffexp for tiledb
This patch adds a new diffexp algorithm which is tuned for tiledb.
This algorithm was written by Bruce and is adapted here to plug into the
current framework. The anndata_adaptor still calls the original
algotithm (which was move from diffexp.py to diffexp_generic.py).
The cxg_adaptor now calls the new diffexp_tiledb version. Some
code is shared between the two.
This is part 1 of the diffexp for tiledb. Further tuning and
global throttles are still needed.
A script to run and time diffexp with various options is also
added: test/run_diffexp.py.
* s3 region should have a single config param
The s3 region can also now be automatically determined to further
reduce errors.
This patch also fixes a bug with order of handling the config params.
The tiledb config needs to be fixed before attempting to load
(need to handle_adaptor before handle_single_dataset)
fixes an issue with "cellxgene launch" which had a bad interaction between
command line parameters and config file parameters.
Now, the config files are applied first, followed by the parameters that
were provided in the command line.
There is also now a check that each of the config attributes is type checked.
* Improvements to the matrix cache
- Add a timelimit for the matrix in the cache.
Once the timelimit is reached, the matrix can be removed.
- If a DatasetAccessError occurs, then remove the dataset
from the matrix cache.
Fixes#1322
There is a small chicken and egg problem.
The config file could be in s3, therefore when using the DataLocator to
download the config file, we don't yet have an app_config object.
Adding a check to handle this case.
Mostly this is just instructions for how to do this,
with a small addition to the makefile.
This enables support for serving the about_legal_tos and about_legal_privacy
from the cellxgene server.
* Added a config hook for secret key into the app.
the server first looks in an environment variable,
then looks in a config file.
For the cellxgene launch app, a default key is used if none is provided.
For the eb app, a secret key must be provided.
* Improved fix for matrix cache handling.
During the MatrixDataCacheItem acquire function there was a
time when the write lock was released and the read lock was taken.
During that time, the dataset could have been deleted, later
result in the MatrixDataCacheManageri data adaptor returning None.
The solution is to demote the writer lock to a reader lock instead
of unlocking and relocking.
Also, when a the cache needs to delete an entry, the delete
is done outside the MatrixDataCacheManager lock. This operation
only requires the write lock for the MatrixDataCacheItem.
Fixes#1255
* Introduce a config file to cellxgene
The config file format is in yaml. The default config is located
in server/common/default_config.py. A user may create a yaml file
that contains a subset of these fields. It can be used during cellxgene
launch, or for hosted cellxgene.
The code has also been refactored. Much of the logic to check arguments
has moved from launch to app config.
It is now possible to set the tiledb context parameters using the config
file. Other feature will soon be handled in a similar way.
* Improve hosted cellxgene
- option to turn off the test index page, or supply a page for redirect.
For EB, The default is to return 404. For cli launch, the default is the test page.
- option to select which matrix types are allowed for multi dataset servers.
For EB, The default is CXG only. For cli launch, the default is any matrix type.
- Return early with an error response if diffexp is requested when not configured
- Verified that reembedings and user annotations also return with an error response
if used when not enabled.
TODO: The new options cannot currently be set by the user.
I plan to add a configuration file where these and all other settings can be set.
Fixes#1210Fixes#1228Fixes#1229
* Fix for favicon with --dataroot
* fix static assets in hosted cxg
The web proxy at aws eb was not finding the static assets.
The solution here is very simple: just copy the directory
containing the static assets to the top level of the artifact.zip.
This is not really the ideal solution. According to the AWS
docs you can make a mapping to the correct location in an
an ebextentions config file. I tried this and many combinations but
was not able to get this to work following that pattern.
Since we control the construction of the zip file, the solution
here isn't bad, but it could probably be made better.
* early, non-working eb config
* hosted cellxgene
In this PR, contains scripts and instructions for deploying cellxgene
for AWS elastic beanstalk. It supports the multi-dataset option.
The Makefile in the server/eb directory creates an artifact.zip
file, which can be deploy at AWS EB.
The server/eb directory contains:
app.py - flask app to run the server
Makefile - which creates an artifact.zip file which can be deployed.
README.md - instructions for setting up and deploying the eb app.
* hosted cellxgene (#38)
In this PR, contains scripts and instructions for deploying cellxgene
for AWS elastic beanstalk. It supports the multi-dataset option.
The Makefile in the server/eb directory creates an artifact.zip
file, which can be deploy at AWS EB.
The server/eb directory contains:
app.py - flask app to run the server
Makefile - which creates an artifact.zip file which can be deployed.
README.md - instructions for setting up and deploying the eb app.
* Update how artifact.zip is created
prune the server/test and server/eb directories
* Remove debugging print statements
* fixes from review comments
* fix lint
Co-authored-by: bkmartinjr <bruce@chanzuckerberg.com>