187 Commits

Author SHA1 Message Date
Alok Saldanha
105dd8c493 Merge pull request #102 from Novartis/reenable-codecov
reenable-codecov
2025-12-01 21:47:23 -05:00
Alok Saldanha
c92fe18ee7 add some backend cache tests 2025-12-01 21:38:49 -05:00
Alok Saldanha
54aacab276 reenable-codecov 2025-12-01 21:13:16 -05:00
Alok Saldanha
7dd10f1f3d change module name to match package name (cellxgene_gateway) 2025-12-01 20:52:57 -05:00
Alok Saldanha
8f676f28d0 Prepare 0.4.1 release 2025-12-01 20:38:35 -05:00
Alok Saldanha
a74576ade5 Merge pull request #101 from Novartis/delay_itemsource_init
Delay itemsource init
2025-12-01 20:16:21 -05:00
Alok Saldanha
d747860118 made pruner a daemon thread 2025-11-08 18:40:33 -05:00
Alok Saldanha
79fef57010 updated start scripts to run in subshells 2025-11-08 17:51:06 -05:00
Alok Saldanha
cd9c0a3671 upload coverage reports as artifacts 2025-11-08 17:38:20 -05:00
Alok Saldanha
55b268125f blacken 2025-11-08 17:23:47 -05:00
Alok Saldanha
4df58f9ceb fixe bug in status.json 2025-11-08 17:21:37 -05:00
Alok Saldanha
8b4565e745 added status.json test 2025-11-08 17:21:37 -05:00
Alok Saldanha
c8056991b0 simplify tests 2025-11-08 11:00:07 -05:00
Alok Saldanha
1d1d8b4e59 Address linter warnings 2025-11-08 10:57:50 -05:00
Alok Saldanha
a35c6b9b1e added start scripts for flask, gunicorn and uwsgi 2025-11-05 22:02:56 -05:00
Alok Saldanha
f260180a76 set default_item_source and start pruner thread 2025-11-05 21:54:18 -05:00
Alok Saldanha
58ae41fe0c Delay itemsource initialization until first request is served 2025-11-05 21:41:10 -05:00
Alok Saldanha
0c5adc9fef Merge pull request #99 from andynu/gunicorn-support
Gunicorn support for now. Will revisit populating item sources on module load to improve reusability of module (specifically in tests for now).
2025-11-05 06:55:39 -05:00
Alok Saldanha
34ac73ba01 second attempt to skip codecov on PRs from forks 2025-11-05 06:47:33 -05:00
Alok Saldanha
d16a97906c Skip codecov for PRs from fork 2025-11-05 06:44:54 -05:00
Alok Saldanha
3e5accad65 fixed linting and tests 2025-11-04 06:57:43 -05:00
Andy
903d25763f Fix AttributeError by storing ItemSource objects in default_item_source
Previously, default_item_source was set to a string ("s3" or "local"),
but matching_source() tried to access default_item_source.name, causing:
  AttributeError: 'str' object has no attribute 'name'

This bug occurred when source_name=None (single data source configuration)
and has existed since the ItemSource interface was introduced in 2021.

Changes:
- Store the actual ItemSource object reference instead of string name
- Assign to intermediate variables (s3_source, file_source) for clarity

Fixes the error when viewing datasets with a single data source configured.
2025-10-29 14:08:56 -04:00
Andy
08c546f40a Fix WSGI server initialization by extracting data source setup
Addresses the issue where Gunicorn/uWSGI servers import the gateway
module but never call main(), leaving item_sources empty and causing
the file crawler to fail.

Changes:
- Extract data source initialization into initialize_data_sources()
- Call initialization at module import time for WSGI compatibility
- Add _initialized flag to prevent double initialization
- Simplify main() to delegate to initialize_data_sources()

This ensures data sources are populated when running under WSGI servers
(Gunicorn, uWSGI) while maintaining backward compatibility with the
Flask development server.

Related to GitHub issues #33 and #92
2025-10-29 14:08:56 -04:00
Andy
b9f4d35812 Fix UnicodeDecodeError when viewing compressed datasets
When viewing datasets through the gateway, requests would fail with:
  UnicodeDecodeError: 'utf-8' codec can't decode byte 0xb5 in position 1

The gateway was copying the accept-encoding header from browser requests
when proxying to cellxgene backend servers. When accept-encoding is
manually set, the Python requests library assumes the caller will handle
decompression and leaves response content compressed.

The cellxgene server responded with zstd-compressed content (magic bytes
28 b5 2f fd), but the gateway attempted to decode this compressed binary
data as UTF-8 text, causing the decode error.

Solution: Remove accept-encoding from the copied headers list in
cache_entry.py. This allows the requests library to automatically handle
compression negotiation and transparently decompress responses (gzip,
deflate, brotli, zstd, etc.).

This is the standard practice when proxying with requests and maintains
all other gateway functionality (URL rewriting, auth, caching, etc.).

Tested:
- Dataset viewing works with compressed responses
- File browser and static assets load correctly
- URL rewriting continues to function properly
2025-10-29 14:04:25 -04:00
Alok Saldanha
05a95c2ca9 Prepare 0.4.0 release 2024-03-10 09:33:50 -04:00
Alok Saldanha
5d29153544 #87 remove version pins for markupsafe, flask and werkzeug
also remove dependency on flask-api
2024-03-10 09:31:39 -04:00
Alok Saldanha
6a2bc409db Prepare for 0.3.12 release 2024-03-03 07:45:47 -05:00
Alok Saldanha
fa72481b66 Merge remote-tracking branch 'origin/dependabot/pip/werkzeug-2.3.8' 2024-03-03 07:33:03 -05:00
Alok Saldanha
3d0166904b Merge pull request #90 from Novartis/dependabot/pip/flask-2.2.5
Bump flask from 2.2.2 to 2.2.5
2024-03-03 07:31:42 -05:00
Alok Saldanha
4e63ff95a8 #87 blacken 2024-03-02 12:25:21 -05:00
Alok Saldanha
c1111e2cb4 #87 patch enable annotations 2024-03-02 12:23:11 -05:00
Alok Saldanha
c4b9084286 #87 Fix test 2024-03-02 12:08:57 -05:00
dependabot[bot]
0282da03c8 Bump werkzeug from 2.3.0 to 2.3.8
Bumps [werkzeug](https://github.com/pallets/werkzeug) from 2.3.0 to 2.3.8.
- [Release notes](https://github.com/pallets/werkzeug/releases)
- [Changelog](https://github.com/pallets/werkzeug/blob/main/CHANGES.rst)
- [Commits](https://github.com/pallets/werkzeug/compare/2.3.0...2.3.8)

---
updated-dependencies:
- dependency-name: werkzeug
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
2024-03-02 16:56:15 +00:00
dependabot[bot]
bec74bec45 Bump flask from 2.2.2 to 2.2.5
Bumps [flask](https://github.com/pallets/flask) from 2.2.2 to 2.2.5.
- [Release notes](https://github.com/pallets/flask/releases)
- [Changelog](https://github.com/pallets/flask/blob/main/CHANGES.rst)
- [Commits](https://github.com/pallets/flask/compare/2.2.2...2.2.5)

---
updated-dependencies:
- dependency-name: flask
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
2024-03-02 16:56:15 +00:00
Alok Saldanha
35c8e8180c #87 temporarily pin versions 2024-03-02 11:55:07 -05:00
Alok Saldanha
0000a60eb0 #73 hide annotation links when disabled 2024-03-02 11:51:20 -05:00
Alok Saldanha
1e02e0abb8 Merge pull request #74 from Novartis/73_reorder_filecrawl
#73 moved new annotation link to front
2024-03-02 11:44:49 -05:00
Alok Saldanha
66cca86b51 Merge remote-tracking branch 'ghall/just_gene_sets' 2024-03-02 11:09:00 -05:00
Alok Saldanha
f046e7c5d2 Merge pull request #88 from Mye-InfoBank/master
Fix dockerfile installation problems
2024-02-24 09:22:49 -05:00
Nico Trummer
a8f6e45f34 Implement pip upgrade to Dockerfile 2024-02-21 09:47:46 +01:00
george-hall-ucl
19caa6cb80 Sorry -- forgot to lint 2023-08-08 16:04:03 +01:00
george-hall-ucl
56bd079024 Save gene sets without cell annotations
This fixes a bug whereby new gene_sets csv files created without
accompanying cell-level annotations could not be detected by the
filecrawler.
2023-08-08 15:49:57 +01:00
Alok Saldanha
9d10932b06 prepare for 0.3.11 release 2023-07-09 19:27:32 -04:00
Alok Saldanha
c7c156b4cf Merge pull request #77 from aeisenbarth/filter-empty-folders
Filter directories without h5ad files
2023-07-09 07:04:06 -06:00
Alok Saldanha
624d1f8567 #78 Revert "Rename argument "filter" to "subpath""
This reverts commit fdd6cca297.
2023-07-09 08:25:27 -04:00
Alok Saldanha
79c3f588b6 Merge pull request #80 from Novartis/79_add_docker_example
79 add docker example
2023-07-09 06:05:50 -06:00
Alok Saldanha
d8fd07572c Merge pull request #83 from Novartis/81_gene_set_support
gene set support
2023-07-09 06:03:33 -06:00
Alok Saldanha
64d636a1c5 #81 added unit test for gene sets 2023-07-09 07:46:01 -04:00
Alok Saldanha
7b314d4457 #81 switch to latest ubuntu 2023-07-06 17:15:00 -06:00
Alok Saldanha
2754bc1ef1 #81 Combined GATEWAY_ENABLE_ANNOTATIONS and GATEWAY_ENABLE_GENE_SETS flags 2023-07-06 08:53:19 -06:00
Alok Saldanha
5a650334df #81 moved gene set check into fileitem_source 2023-07-06 08:18:50 -06:00
george-hall-ucl
81c8ce4219 #81 Add support for gene sets
This adds the flag `GATEWAY_ENABLE_GENE_SETS` to enable support for gene
sets.  To simplify implementation, activating this flag also activates
`GATEWAY_ENABLE_ANNOTATIONS`.  The gene sets are saved in a file that
has the same name as the annotations `csv` but with `_gene_sets`
appended to the file name (before the extension).  This file is hidden
in filecrawler, and the gene sets are loaded when the associated
annotations file is loaded.

If the annotations file is missing, then an Exception is raised.

I have updated one unit test to make it expect
`--disable-gene-sets-save` in the default case (i.e. if
`GATEWAY_ENABLE_ANNOTATIONS = 0`).  All units tests pass.

I have updated the README to document `GATEWAY_ENABLE_GENE_SETS`.
2023-07-06 07:59:12 -06:00
Alok Saldanha
f296efcc55 #79 added cellxgene-data directory so it actually works 2022-12-21 16:02:50 -05:00
Alok Saldanha
facfb27d5c #79 add simple example to customize cellxgene-gateway ui 2022-12-21 15:42:13 -05:00
Andreas Eisenbarth
6607b15085 Exclude directories having no h5ad files 2022-10-12 17:31:22 +02:00
Andreas Eisenbarth
0e92b73347 Add test case for dirs without h5ad 2022-10-12 17:31:22 +02:00
Andreas Eisenbarth
88b9b815c0 Adjust test case for dirs with h5ad 2022-10-12 17:31:22 +02:00
Andreas Eisenbarth
6ef82b36f1 For running individual tests, make sure flask_util.view_url is callable 2022-10-12 16:36:46 +02:00
Andreas Eisenbarth
fdd6cca297 Rename argument "filter" to "subpath" 2022-10-12 13:40:39 +02:00
Alok Saldanha
a00403c60e #73 moved new annotation link to front 2022-08-21 08:15:09 -04:00
Alok Saldanha
390fe24ea4 prepare for 0.3.10 release 2022-06-20 21:50:42 -04:00
Alok Saldanha
590565bea2 Merge pull request #69 from Novartis/68_read_from_subprocess
68 read from subprocess
2022-06-20 21:50:28 -04:00
Alok Saldanha
d32a31e855 #68 read process output until it exits 2022-06-20 21:46:19 -04:00
Alok Saldanha
0cd551382e #65 add environment variable to control how long cellxgene processes can remain idle 2022-06-20 21:46:19 -04:00
Alok Saldanha
36c0a4d3d7 #68 add param to set log level 2022-06-20 21:29:19 -04:00
Alok Saldanha
8d8a0a3483 #68 close responses 2022-06-20 21:29:19 -04:00
Alok Saldanha
eaa157079c Merge pull request #67 from Novartis/docker
Remove version pins to upgrade Flask
2022-06-07 12:06:49 -04:00
Alok Saldanha
977c50ce8c #66 switched from mocks to test request context 2022-06-07 07:22:30 -04:00
Alok Saldanha
833cad3bc2 #66 remove version pins 2022-06-07 06:17:11 -04:00
Alok Saldanha
c5f3c68740 Merge pull request #66 from romanhaa/docker
Dockerise cellxgene-gateway
2022-06-07 06:04:13 -04:00
Roman Hillje
d944a31d59 Dockerise cellxgene-gateway 2022-05-20 19:44:46 +02:00
Alok Saldanha
5814cb9943 clarified purpose of refresh query param 2022-03-14 23:17:57 -04:00
Alok Saldanha
9a91cdf795 prepare for 0.3.9 release 2022-03-14 23:12:01 -04:00
Alok Saldanha
25aff5c020 Merge pull request #60 from Novartis/59_s3_caching
#59 add refresh query param to force refresh of S3 cache
2022-03-14 23:08:45 -04:00
Alok Saldanha
6bcb594712 #59 add temporary workaround for jinja 2022-03-14 22:58:06 -04:00
Alok Saldanha
f9ed4c4047 #59 document S3_ENABLE_LISTINGS_CACHE 2022-03-14 22:36:42 -04:00
Alok Saldanha
a9753c4101 #59 change s3 cache variable from S3_DISABLE_LISTINGS_CACHE to S3_ENABLE_LISTINGS_CACHE 2022-03-14 22:36:42 -04:00
Alok Saldanha
fd48920c5b #59 add refresh query param to force refresh of S3 cache 2022-03-14 21:44:05 -04:00
Alok Saldanha
09db93b2b5 Merge pull request #62 from arogozhnikov/patch-1
Force reload of s3 file structure on every request
2022-03-14 20:45:20 -04:00
Alex Rogozhnikov
3e3bd22512 add environment variable S3_DISABLE_LISTINGS_CACHE per Alok's request 2022-03-14 09:59:55 -07:00
Alex Rogozhnikov
893b2f1af1 remove listing cache at the level of fs 2022-03-11 03:20:06 -08:00
Alex Rogozhnikov
757487b772 Force reload folder on every request 2022-03-11 02:43:29 -08:00
Alok Saldanha
87a8dbfa78 #57 Reverted incorrect change to unit test 2021-12-21 13:59:30 -05:00
Alok Saldanha
9c38e48c5c prepare for 0.3.8 release 2021-12-21 12:19:02 -05:00
Alok Saldanha
073f5f945c #57 changed logic to take last path element 2021-12-21 12:05:01 -05:00
Alok Saldanha
9dc4409f1a #57 added failing unit test 2021-12-21 12:03:39 -05:00
Alok Saldanha
551cb46af8 #42 add support for is_authorized hook 2021-11-14 17:28:10 -05:00
Alok Saldanha
620181ae4d prepare for 0.3.7 release 2021-08-12 14:02:26 -04:00
Alok Saldanha
73a7920cc8 add back ip_address endpoint 2021-08-12 13:58:19 -04:00
Alok Saldanha
ed3e999cd1 prepare for 0.3.6 release 2021-07-18 10:42:42 -04:00
Alok Saldanha
2ae2e53863 Pin version of workzeug
This is required by earlier flask-api versions

  File "/home/alokito/code/cellxgene-gateway/cellxgene_gateway/gateway.py", line 25, in <module>
    from flask_api import status
  File "/home/alokito/miniconda3/envs/cellxgene-gateway/lib/python3.7/site-packages/flask_api/__init__.py", line 1, in <module>
    from flask_api.app import FlaskAPI
  File "/home/alokito/miniconda3/envs/cellxgene-gateway/lib/python3.7/site-packages/flask_api/app.py", line 6, in <module>
    from flask_api.request import APIRequest
  File "/home/alokito/miniconda3/envs/cellxgene-gateway/lib/python3.7/site-packages/flask_api/request.py", line 9, in <module>
    from werkzeug._compat import to_unicode
ModuleNotFoundError: No module named 'werkzeug._compat'
2021-07-18 10:32:10 -04:00
Alok Saldanha
fd0e7d9c31 preparing for 0.3.5 release 2021-07-18 09:43:45 -04:00
Alok Saldanha
98ef6efd0c pinned version of flask, to match cellxgene 2021-07-18 09:39:52 -04:00
Alok Saldanha
82e43ff943 preparing for 0.3.4 release 2021-07-18 09:15:27 -04:00
Alok Saldanha
f8a77423eb Merge pull request #51 from Novartis/nested_subdirs
Enable listing nested subdirs
2021-07-18 09:14:16 -04:00
Alok Saldanha
0375a717c9 #50 fixed bug in listing subdirs 2021-07-18 09:04:36 -04:00
Alok Saldanha
4bf57832a0 #50 added failing test for listing subdirs 2021-07-18 08:51:37 -04:00
Alok Saldanha
b7d14dba6a preparing for 0.3.3 release 2021-07-12 20:03:14 -04:00
Alok Saldanha
26286f94b1 Merge pull request #49 from Novartis/fix_cache_pruning
Fix cache pruning
2021-07-12 19:26:53 -04:00
Alok Saldanha
264a324946 #48 fix bug in cache pruning 2021-07-12 19:19:29 -04:00
Alok Saldanha
520069a825 #48 added failing test for cache pruning 2021-07-12 19:12:52 -04:00
Alok Saldanha
3cb0e4d725 added test for is_port_in_use 2021-05-11 07:27:30 -04:00
Alok Saldanha
d2b508e371 Create SECURITY.md 2021-05-06 05:02:36 -04:00
Alok Saldanha
e74f6d01d1 added unit tests for dir_util 2021-04-23 07:14:09 -04:00
Alok Saldanha
7b799d0159 removed unused code 2021-04-23 07:00:53 -04:00
Alok Saldanha
1f0885afdd added pypi badges to Readme.md 2021-04-22 19:16:05 -04:00
Alok Saldanha
a3a1a2d095 Preparing for 0.3.2 release 2021-04-22 19:01:07 -04:00
Alok Saldanha
b4739b0cbb Merge pull request #46 from Novartis/nested_s3
This PR should address https://github.com/Novartis/cellxgene-gateway/issues/45
2021-04-22 07:06:17 -04:00
Alok Saldanha
c29e5c0d92 #45 implemented test__list_items__pass_filter_into_scan_directory 2021-04-22 06:59:08 -04:00
Alok Saldanha
a3e3b6cea8 #45 implemented test__scan_directory__properly_recurses_suburls 2021-04-21 08:21:46 -04:00
Alok Saldanha
687dc31b7a improved error message on invalid extra scripts 2021-04-21 06:21:26 -04:00
Alok Saldanha
68da42ae0e #45 implemented test_GIVEN_some_filter_THEN_includes_filterpart_in_heading 2021-04-20 13:32:28 -04:00
Alok Saldanha
ccf1ba58a8 #45 add extra_scripts to cache_status page 2021-04-20 13:17:37 -04:00
Alok Saldanha
d072034911 #45 pass filter into scan_directory
also include filterpart in heading
2021-04-20 06:03:16 -04:00
Alok Saldanha
1200629d74 #45 handle keys as full subpath rather than last path element 2021-04-20 05:57:48 -04:00
Alok Saldanha
dba13ccacd lowering threshold due to including of itemsource modules 2021-04-05 08:10:55 -04:00
Alok Saldanha
3239aefad4 Preparing for v0.3.1 release 2021-04-05 08:06:25 -04:00
Alok Saldanha
dd727a9469 added missing __init__.py 2021-04-05 08:04:48 -04:00
Alok Saldanha
ae5cf09d04 Preparing for 0.3.0 release 2021-04-05 07:55:20 -04:00
Alok Saldanha
ae35d24d80 only set wsgi.url_scheme when EXTERNAL_PROTOCOL is set 2021-04-05 07:55:20 -04:00
Alokito
ad3517f002 Create codeql-analysis.yml 2021-04-05 07:08:34 -04:00
Alok Saldanha
2cefc6d741 added token and path to codecov upload 2021-04-05 06:53:44 -04:00
Alok Saldanha
1586c9a3e8 added code coverage badge 2021-04-04 20:03:02 -04:00
Alok Saldanha
f16bd5e2b9 added test for SubprocessBackend.launch 2021-04-04 19:48:13 -04:00
Alok Saldanha
10105c43a8 added support for code coverage 2021-04-04 14:55:13 -04:00
Alokito
638923fb43 Merge pull request #44 from Novartis/itemsource
Itemsource
2021-04-03 07:15:35 -04:00
Alok Saldanha
169da934b1 updated annotation.js comment 2021-04-03 07:12:50 -04:00
Alok Saldanha
33331c9198 Template has correct relaunch_url. 2021-04-03 07:08:20 -04:00
Alok Saldanha
d2d22cecaa updated tests, formatting 2021-04-02 11:52:49 -04:00
Alok Saldanha
5c488d8d5e only include source path element in url when multiple sources 2021-04-02 11:40:11 -04:00
Alok Saldanha
b1ae9e8ce8 ensure s3 bucket does not include trailing slash 2021-04-02 11:37:11 -04:00
Alok Saldanha
76c1d9e80c removed .csv suffix from annotation name 2021-03-29 07:28:16 -04:00
Alok Saldanha
586dc6623a ensure annotation directory for new annotations 2021-03-29 07:03:05 -04:00
Alok Saldanha
76f3939fdd rebased "Introduction of ItemSource interface" patch 2021-03-28 22:05:08 -04:00
Alok Saldanha
5404a8b0f3 removed trailing slash from render_annotations 2021-03-28 21:59:03 -04:00
Alok Saldanha
6f2b372d32 Preparing for 0.2.3 release 2021-02-28 09:55:26 -05:00
Alokito
25755d3f69 Merge pull request #41 from Novartis/grst_master
Grst master
2021-02-22 08:59:05 -05:00
Alok Saldanha
893092f199 fixed unit tests 2021-02-13 17:30:21 -05:00
Alok Saldanha
d3ee04b6e5 blacken 2021-02-13 17:02:55 -05:00
Alok Saldanha
f1f4b0c0ca added environment variables for ProxyFix 2021-02-13 16:42:34 -05:00
Alok Saldanha
01dccef7db added back trailing slash to url_for change 2021-02-13 16:41:32 -05:00
Gregor Sturm
9c17e2ff32 Fix redirect in cache_entry 2021-01-18 22:08:25 +01:00
Gregor Sturm
fbc18fb636 Add ProxyFix to gateway.py 2021-01-18 20:01:53 +01:00
Gregor Sturm
8c5a635de9 Use url_for in all templates 2021-01-18 19:47:18 +01:00
Gregor Sturm
94062c2d64 apply proxy fix 2021-01-18 18:19:25 +01:00
Gregor Sturm
068e8f7633 Fix url_for 2021-01-18 18:09:45 +01:00
Gregor Sturm
16b54f9409 Use url_for to generate URLs 2021-01-18 16:59:15 +01:00
Alok Saldanha
942410bb44 added instructions on pre-commit installation to README.md 2020-12-31 17:12:11 -05:00
Alokito
0d084a405e Merge pull request #38 from Novartis/feature/action_push
Feature/action push
2020-12-31 16:33:30 -05:00
Alok Saldanha
49d679e779 try evaling the bash hook
per https://github.com/conda/conda/issues/7980
2020-12-31 13:23:25 -05:00
Alok Saldanha
5e6faa4b02 run pr checks on push 2020-12-31 13:06:56 -05:00
Alokito
9fe846c786 Merge pull request #37 from ericmjl/master
Migrated PR checks to GitHub actions
2020-12-28 11:38:26 -05:00
Eric Ma
1c7907aabd Change file extension 2020-12-27 21:26:32 -05:00
Eric Ma
30d2b07b1a Migrated PR checks to GitHub actions 2020-12-27 21:10:07 -05:00
Alokito
21ff56ea8b Merge pull request #35 from dfeinzeig/fix/redirect
add missing trailing slash in effort to avoid whatever is redirecting
2020-10-28 07:39:06 -04:00
David Feinzeig
0b866a46ec add missing trailing slash in effort to avoid whatever is redirecting 2020-10-23 17:41:05 -04:00
Alok Saldanha
d5246d4b4b prepared for 0.2.2 release 2020-09-08 21:11:59 -04:00
Alok Saldanha
465373ca7e #30 black formatting 2020-09-08 21:11:59 -04:00
Alok Saldanha
fc83523c84 #30 added missing js asset 2020-09-08 21:07:18 -04:00
Alok Saldanha
4e367b0ec9 update travis.yml to new conda env 2020-08-30 13:53:57 -04:00
Alok Saldanha
ce3eb0eedf Merge branch 'gmerge' into gmaster
# Conflicts:
#	cellxgene_gateway/backend_cache.py
#	cellxgene_gateway/cache_entry.py
#	cellxgene_gateway/env.py
#	cellxgene_gateway/gateway.py
#	cellxgene_gateway/subprocess_backend.py
#	tests/test_cache_entry.py
#	tests/test_dir_util.py
2020-08-30 13:47:22 -04:00
Alok Saldanha
9219336356 blackened code 2020-08-30 13:38:35 -04:00
Alok Saldanha
b689357f29 Applied "Formatting and new enumeration for Cellxgene-gateway" patch 2020-08-30 13:36:39 -04:00
Alok Saldanha
5349df36e8 prepared for 0.2.1 release
updated docs
made GATEWAY_IP optional
2020-08-18 20:02:24 -04:00
Alokito
13802b72b9 Merge pull request #26 from Novartis/feature/25
#25 added CELLXGENE_ARGS environment variable
2020-08-18 19:45:31 -04:00
Alok Saldanha
0f8092e07a #25 added CELLXGENE_ARGS environment variable 2020-08-16 17:08:06 -04:00
Alok Saldanha
61c7c43da7 blackened code 2020-08-16 17:07:23 -04:00
Alokito
0604a0ca9f Merge pull request #27 from Novartis/feature/24
Feature/24
2020-08-16 17:05:49 -04:00
Alok Saldanha
8cdd23a57b #24 switched to absolute paths
relative paths won't necessarily work in css files
2020-08-16 16:17:28 -04:00
Alok Saldanha
486206bd32 #24 make static paths relative 2020-08-16 16:06:34 -04:00
Alokito
4105fc32e1 Update README.md
added link to travis badge
2020-08-10 20:06:31 -04:00
Alok Saldanha
59cdac5f85 move black into environment.yml 2020-08-10 18:01:22 -04:00
Alokito
bd1b7e8da2 Merge pull request #23 from Novartis/feature/22_metadata_api
Feature/22 metadata api
2020-08-10 17:57:50 -04:00
Alok Saldanha
bfa80f7ef8 #22 blackened code 2020-08-10 17:56:32 -04:00
Alok Saldanha
63cdfe1953 #22 fixed failing tests 2020-08-10 17:56:24 -04:00
Alok Saldanha
578845b133 #22 Added /metadata/ip_address endpoint 2020-08-10 17:49:03 -04:00
Alokito
849c8a58f2 Rename travis.yml to .travis.yml 2020-08-10 16:54:44 -04:00
Alokito
021d981bef Create travis.yml 2020-08-10 16:52:32 -04:00
Alok Saldanha
df812f3203 ignore patch files 2020-06-11 13:44:10 -04:00
Alok Saldanha
804ea29f36 move version string into module 2020-06-11 13:43:06 -04:00
Alok Saldanha
da99abdd9f Increment version to 0.2.0 2020-06-11 10:53:14 -04:00
Alok Saldanha
bf221309d8 Merge branch for 0.1.0 release into master 2020-06-11 10:47:43 -04:00
Alokito
7c7fa5f12a Use all_entries in cellxgene_gateway/filecrawl.py 2020-05-17 14:57:20 -04:00
Gervaise H. Henry
eb4693f1fa Sort os.listdir in filecrawler.py 2020-05-17 14:57:20 -04:00
Alok Saldanha
84db11077a bump version, update environment.yml 2020-05-06 17:17:09 -04:00
Alokito
f8e619c53e Merge pull request #18 from fidelram/enable_backed_mode
updated annot arguments for cellxgene v 0.15 & added option for backed mode.
2020-05-06 17:16:22 -04:00
Fidel Ramírez
9a130a8eb7 updated arguments for cellxgene v 0.15. Added option for backed mode. 2020-05-04 13:49:09 +02:00
65 changed files with 3113 additions and 680 deletions

13
.coveragerc Normal file
View File

@@ -0,0 +1,13 @@
[run]
branch = True
source = cellxgene_gateway
[report]
exclude_lines =
if self.debug:
pragma: no cover
raise NotImplementedError
if __name__ == .__main__.:
ignore_errors = True
omit =
tests/*

67
.github/workflows/codeql-analysis.yml vendored Normal file
View File

@@ -0,0 +1,67 @@
# For most projects, this workflow file will not need changing; you simply need
# to commit it to your repository.
#
# You may wish to alter this file to override the set of languages analyzed,
# or to provide custom queries or build logic.
#
# ******** NOTE ********
# We have attempted to detect the languages in your repository. Please check
# the `language` matrix defined below to confirm you have the correct set of
# supported CodeQL languages.
#
name: "CodeQL"
on:
push:
branches: [ master ]
pull_request:
# The branches below must be a subset of the branches above
branches: [ master ]
schedule:
- cron: '18 6 * * 6'
jobs:
analyze:
name: Analyze
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
language: [ 'javascript', 'python' ]
# CodeQL supports [ 'cpp', 'csharp', 'go', 'java', 'javascript', 'python' ]
# Learn more:
# https://docs.github.com/en/free-pro-team@latest/github/finding-security-vulnerabilities-and-errors-in-your-code/configuring-code-scanning#changing-the-languages-that-are-analyzed
steps:
- name: Checkout repository
uses: actions/checkout@v2
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@v1
with:
languages: ${{ matrix.language }}
# If you wish to specify custom queries, you can do so here or in a config file.
# By default, queries listed here will override any specified in a config file.
# Prefix the list here with "+" to use these queries and those in the config file.
# queries: ./path/to/local/query, your-org/your-repo/queries@main
# Autobuild attempts to build any compiled languages (C/C++, C#, or Java).
# If this step fails, then you should remove it and run the build manually (see below)
- name: Autobuild
uses: github/codeql-action/autobuild@v1
# Command-line programs to run using the OS shell.
# 📚 https://git.io/JvXDl
# ✏️ If the Autobuild fails above, remove it and uncomment the following three lines
# and modify them (or add more) to build your code if your project
# uses a compiled language
#- run: |
# make bootstrap
# make release
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@v1

91
.github/workflows/pr-checks.yaml vendored Normal file
View File

@@ -0,0 +1,91 @@
# Tests that run on every PR
name: Pull Request Checks
on: [push, pull_request]
jobs:
black:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v2
name: Checkout repository
- uses: actions/setup-python@v2
name: Setup Python
with:
python-version: 3.9
- name: Install black
run: |
python -m pip install --upgrade pip
pip install black
- name: Run black
run: |
black . --check
# This job is copied over from `deploy.yaml`
run-tests:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v2
# See: https://github.com/marketplace/actions/setup-conda
- uses: s-weigand/setup-conda@v1
with:
conda-channels: "conda-forge"
- name: Build environment
run: |
conda env create -f environment.yml
eval "$(conda shell.bash hook)"
conda activate cellxgene-gateway
python setup.py install
pip install -r requirements-test.txt
- name: Run tests
run: |
eval "$(conda shell.bash hook)"
conda activate cellxgene-gateway
coverage run -m unittest discover tests
- name: Check coverage
run: |
eval "$(conda shell.bash hook)"
conda activate cellxgene-gateway
coverage report --fail-under 41
coverage report > coverage.txt
coverage html -i
coverage xml -i
- name: Upload coverage HTML report
uses: actions/upload-artifact@v4
with:
name: coverage-html
path: htmlcov/
retention-days: 30
- name: Upload coverage xml
uses: actions/upload-artifact@v4
with:
name: coverage-xml
path: coverage.xml
retention-days: 30
- name: Upload coverage summary
uses: actions/upload-artifact@v4
with:
name: coverage-summary
path: coverage.txt
retention-days: 30
- name: "Upload coverage to Codecov"
if: ${{ github.event_name == 'push' || (github.event_name == 'pull_request' && github.event.pull_request.head.repo.full_name == github.repository) }}
uses: codecov/codecov-action@v1
with:
token: ${{ secrets.CODECOV_TOKEN }}
files: ./coverage.xml
flags: unittests
env_vars: OS,PYTHON
name: codecov-umbrella
fail_ci_if_error: true
path_to_write_report: ./codecov_report.txt
verbose: true

4
.gitignore vendored
View File

@@ -55,6 +55,7 @@ htmlcov/
.nox/
.coverage
.coverage.*
htmlcov
.cache
nosetests.xml
coverage.xml
@@ -136,3 +137,6 @@ dmypy.json
.pyre/
# End of https://www.gitignore.io/api/python
*.patch
.vscode

View File

@@ -7,14 +7,7 @@ repos:
language: system
types: [python]
stages: [commit]
- id: flake8
name: flake8
language: system
entry: flake8
types: [python]
stages: [commit]
- id: black
language_version: python3.6+
name: black
language: system
entry: black

41
.travis.yml Normal file
View File

@@ -0,0 +1,41 @@
# This is necessary for nxviz as matplotlib is involved.
# before_script:
# - "export DISPLAY=:99.0"
# - "sh -e /etc/init.d/xvfb start"
# - sleep 5 # give xvfb some time to start
language: python
matrix:
include:
- python: 3.5 # we don't actually use this
env: PYTHON_VERSION=3.7
install:
# We do this conditionally because it saves us some downloading if the
# version is the same.
- wget https://repo.continuum.io/miniconda/Miniconda3-latest-Linux-x86_64.sh -O miniconda.sh;
- bash miniconda.sh -b -p $HOME/miniconda
- export PATH="$HOME/miniconda/bin:$PATH"
- hash -r
- conda config --set always_yes yes --set changeps1 no
- conda update -q conda
- conda config --add channels conda-forge
# Useful for debugging any issues with conda
- conda info -a
# Install Python, py.test, and required packages.
- conda env create -f environment.yml
- source activate cellxgene-gateway
- python setup.py install
script:
# Your test script goes here
- black -l 79 . --check
- python -m unittest discover tests
after_success:
- bash <(curl -s https://codecov.io/bash)
notifications:
email: true

108
Changelog.md Normal file
View File

@@ -0,0 +1,108 @@
# 0.4.2
* update package name
# 0.4.1
* Fix UnicodeDecodeError when viewing compressed datasets
* Fix WSGI server initialization by extracting data source setup
* Fix AttributeError by storing ItemSource objects in default_item_source
* Delay itemsource initialization until first request is served
* Set default_item_source and start pruner thread
* Added start scripts for flask, gunicorn and uwsgi
* Made pruner a daemon thread
* Updated start scripts to run in subshells
* Fixed bug in status.json
# 0.4.0
* Removed dependency on flask-api
* Updated dependencies (python 3.11, numpy, unpinned flask, werkzeug)
# 0.3.12
* #81 List gene set annotations when cell annotations not present
* #86 Upgrade pip within docker image
* #73 Moved new link to front
* #87 Temporarily pin versions of werkzeug and flask
# 0.3.11
* #81 added support for gene sets
* #79 added example for cellxgene-gateway customized docker image
* #78 prune directories that do not contain h5ad files
# 0.3.10
* #65 Added GATEWAY_EXPIRE_SECONDS to set how long cellxgene servers can remain idle before being terminated.
* Added GATEWAY_LOG_LEVEL to set the log level
* #68 Close connections after reading response
* #68 Background thread reads from output of cellxgene process until it exits
# 0.3.9
* Added S3_ENABLE_LISTINGS_CACHE variable (See README.md)
# 0.3.8
* Fixed bug #57 affecting deeply nested subdirectory listing
# 0.3.7
* added back /metadata/ip_address endpoint
# 0.3.6
* pinned version of werkzeug
# 0.3.5
* Pinned flask version to match cellxgene 0.17.0
# 0.3.4
* Fixed bug #50 affecting subdirectory listing
# 0.3.3
* Fixed bug #48 affecting cache pruning
# 0.3.2
* Fixed bug #45 affecting multi-level S3 folders
* Added extra_scripts to cache_status page
# 0.3.1
* Added missing __init__.py
# 0.3.0
* Added support for itemsource interface, allowing s3 hosting
* Removed support for http file uploads
* Only set wsgi.url_scheme when EXTERNAL_PROTOCOL is set (see issue #43)
* Dropped flake8 due to conflicts with black
* Added code coverage metrics
# 0.2.3
* Added support for ProxyFix
# 0.2.2
* Fixed bug with annotations (missing annotation.js asset)
# 0.2.1
* Minor fixes to enable cellxgene 0.16.0
* Added CELLXGENE_ARGS to enable passing additional arguments to cellxgene
* added metadata/ip_address endpoint
# 0.2.0
Incrementing minor version since the changes for 0.15 are breaking, and we may want to release bugfixes from 0.1.0 branch.
# 0.1.1
Added support for cellxgene 0.15

9
Dockerfile Normal file
View File

@@ -0,0 +1,9 @@
FROM python:3.11
RUN pip install --upgrade pip
RUN pip install "cellxgene-gateway>=0.4"
ENV CELLXGENE_DATA=/cellxgene-data
ENV CELLXGENE_LOCATION=/usr/local/bin/cellxgene
CMD ["cellxgene-gateway"]

117
README.md
View File

@@ -2,6 +2,8 @@
Cellxgene Gateway allows you to use the Cellxgene Server provided by the Chan Zuckerberg Institute (https://github.com/chanzuckerberg/cellxgene) with multiple datasets. It displays an index of available h5ad (anndata) files. When a user clicks on a file name, it launches a Cellxgene Server instance that loads that particular data file and once it is available proxies requests to that server.
[![codecov](https://codecov.io/gh/Novartis/cellxgene-gateway/branch/master/graph/badge.svg?token=ndEFSzRKJn)](https://codecov.io/gh/Novartis/cellxgene-gateway) [![PyPI](https://img.shields.io/pypi/v/cellxgene-gateway)](https://pypi.org/project/cellxgene-gateway/) [![PyPI - Downloads](https://img.shields.io/pypi/dm/cellxgene-gateway)](https://pypistats.org/packages/cellxgene-gateway)
# Running locally
## Prequisites
@@ -30,7 +32,7 @@ Note: you may need to downgrade h5py with `pip install h5py==2.9.0` due to an [i
### Option 2: Install from PyPI
```bash
# NOT YET DONE, COMING! STAY TUNED
pip install cellxgene-gateway
```
## Running cellxgene gateway
@@ -39,7 +41,7 @@ Note: you may need to downgrade h5py with `pip install h5py==2.9.0` due to an [i
```bash
mkdir ../cellxgene_data
wget https://github.com/chanzuckerberg/cellxgene/raw/master/example-dataset/pbmc3k.h5ad -O ../cellxgene_data/pbmc3k.h5ad
wget https://raw.githubusercontent.com/chanzuckerberg/cellxgene/master/example-dataset/pbmc3k.h5ad -O ../cellxgene_data/pbmc3k.h5ad
```
@@ -59,17 +61,76 @@ cellxgene-gateway
Here's what the environment variables mean:
* `CELLXGENE_LOCATION` - the location of the cellxgene executable, e.g. `~/anaconda2/envs/cellxgene/bin/cellxgene`
At least one of the following is required:
* `CELLXGENE_DATA` - a directory that can contain subdirectories with `.h5ad` data files, *without* trailing slash, e.g. `/mnt/cellxgene_data`
* `CELLXGENE_BUCKET` - an s3 bucket that can contain keys with `.h5ad` data files, e.g. `my-cellxgene-data-bucket`
Cellxgene Gateway is designed to make it easy to add additional data sources, please see the source code for gateway.py and the ItemSource interface in items/item_source.py
Optional environment variables:
* `CELLXGENE_ARGS` - catch-all variable that can be used to pass additional command line args to cellxgene server
* `EXTERNAL_HOST` - the hostname and port from the perspective of the web browser, typically `localhost:5005` if running locally. Defaults to "localhost:{GATEWAY_PORT}"
* `EXTERNAL_PROTOCOL` - typically http when running locally, can be https when deployed if the gateway is behind a load balancer or reverse proxy that performs https termination. Default value "http"
* `GATEWAY_IP` - ip addess of instance gateway is running on, mostly used to display SSH instructions. Defaults to `socket.gethostbyname(socket.gethostname())`
* `GATEWAY_PORT` - local port that the gateway should bind to, defaults to 5005
* `GATEWAY_EXPIRE_SECONDS` - time in seconds that a cellxgene process will remain idle before being terminated. Defaults to 3600 (one hour)
* `GATEWAY_EXTRA_SCRIPTS` - JSON array of script paths, will be embedded into each page and forwarded with `--scripts` to cellxgene server
* `GATEWAY_ENABLE_UPLOAD` - Set to `true` or `1` to enable HTTP uploads. This is not recommended for a public server.
* `GATEWAY_ENABLE_ANNOTATIONS` - Set to `true` or to `1` to enable cellxgene annotations and gene sets.
* `GATEWAY_ENABLE_BACKED_MODE` - Set to `true` or to `1` to load AnnData in file-backed mode. This saves memory and speeds up launch time but may reduce overall performance.
* `GATEWAY_LOG_LEVEL` - default is `INFO`. set to `DEBUG` to increase logging and to `WARNING` to decrease logging.
* `S3_ENABLE_LISTINGS_CACHE` - Set to `true` or to `1` to cache listings of S3 folders for performance. If the cache becomes stale, set `filecrawl.html?refresh=true` query parameter to refresh the cache.
If any of the following optional variables are set, [ProxyFix](https://werkzeug.palletsprojects.com/en/1.0.x/middleware/proxy_fix/) will be used.
* `PROXY_FIX_FOR` - Number of upstream proxies setting X-Forwarded-For
* `PROXY_FIX_PROTO` - Number of upstream proxies setting X-Forwarded-Proto
* `PROXY_FIX_HOST` - Number of upstream proxies setting X-Forwarded-Host
* `PROXY_FIX_PORT` - Number of upstream proxies setting X-Forwarded-Port
* `PROXY_FIX_PREFIX` - Number of upstream proxies setting X-Forwarded-Prefix
The defaults should be fine if you set up a venv and cellxgene_data folder as above.
## Running cellxgene-gateway with Docker
First, build Docker image:
```bash
docker build -t cellxgene-gateway .
```
Then, cellxgene-gateway can be launched as such:
```bash
docker run -it --rm \
-v <local_data_dir>:/cellxgene-data \
-p 5005:5005 \
cellxgene-gateway
```
Additional environment variables can be provided with the `-e` parameter:
```bash
docker run -it --rm \
-v ../cellxgene_data:/cellxgene-data \
-e GATEWAY_PORT=8080 \
-p 8080:8080 \
cellxgene-gateway
```
## Running cellxgene gateway with start scripts
For your convenience, we provide start scripts for flask, gunicorn and uwsgi.
First, set up a .env
```bash
cp env_example .env
# edit .env
open .env
```
Then run the scripts in a subshell
```bash
( ./start_flask.sh )
```
# Customization
The current paradigm for customization is to modify files during a build or deployment phase:
@@ -110,26 +171,70 @@ python setup.py develop
For convenience, the code repo includes a `run.sh.example` shell script to run the gateway.
4. Install pre-commit hooks
```bash
conda install -c conda-forge pre-commit
pre-commit install
```
## Running Tests
[![Build Status](https://travis-ci.org/Novartis/cellxgene-gateway.svg?branch=master)](https://travis-ci.org/Novartis/cellxgene-gateway)
```bash
python -m unittest discover tests
```
## Code Coverage
```bash
coverage run -m unittest discover tests
coverage html
```
## Running Linters
pip install isort flake8 black
```bash
isort -rc .
flake8 .
black -l 79 .
isort -rc . # rc means recursive, and was deprecated in dev version of isort
black .
```
# Getting Help
If you need help for any reason, please make a github ticket. One of the contributors should help you out.
# Releasing New Versions
## How to prepare for release
- Update Changelog.md and version number in __init__.py
- Cut a release on github
- Go to your project homepage on GitHub
- On right side, you will see [Releases](https://github.com/Novartis/cellxgene-gateway/releases) link. Click on it.
- Click on Draft a new release
- Fill in all the details
- Tag version should be the version number of your package release
- Release Title can be anything you want, but we use v0.3.11 (the same as the tag to be created on publish)
- Description should be changelog
- Click Publish release at the bottom of the page
- Now under Releases you can view all of your releases.
- Copy the download link (tar.gz) and save it somewhere
## How to publish to PyPI
Make sure your `.pypirc` is set up for testpypi and pypi index servers.
```bash
rm -rf dist
python setup.py sdist bdist_wheel
python -m twine upload --repository testpypi dist/*
python -m twine upload dist/*
```
# Contributors
* Niket Patel - https://github.com/NiketPatel9

15
SECURITY.md Normal file
View File

@@ -0,0 +1,15 @@
# Security Policy
## Supported Versions
Use this section to tell people about which versions of your project are
currently being supported with security updates.
| Version | Supported |
| ------- | ------------------ |
| 0.3.2 | :white_check_mark: |
| <= 0.3.1 | :x: |
## Reporting a Vulnerability
Please file a bug report issue.

2
cellxgene_gateway/__init__.py Executable file → Normal file
View File

@@ -6,3 +6,5 @@
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
__version__ = "0.4.2"

View File

@@ -8,12 +8,13 @@
# the specific language governing permissions and limitations under the License.
import time
from http import HTTPStatus
from threading import Thread
from flask_api import status
from typing import List
from cellxgene_gateway import env
from cellxgene_gateway.cache_entry import CacheEntry
from cellxgene_gateway.cache_entry import CacheEntry, CacheEntryStatus
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.cellxgene_exception import CellxgeneException
from cellxgene_gateway.subprocess_backend import SubprocessBackend
@@ -22,8 +23,10 @@ process_backend = SubprocessBackend()
def is_port_in_use(port):
import socket
with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as s:
return s.connect_ex(('localhost', port)) == 0
return s.connect_ex(("localhost", port)) == 0
class BackendCache:
def __init__(self):
@@ -33,12 +36,14 @@ class BackendCache:
contents = self.entry_list
return [c.port for c in contents]
def check_entry(self, key):
def check_path(self, source, path):
contents = self.entry_list
matches = [
c
for c in contents
if c.key.dataset == key.dataset and c.key.annotation_file == key.annotation_file and c.status != "terminated"
if c.key.source.name == source.name
and path.startswith(c.key.descriptor)
and c.status != CacheEntryStatus.terminated
]
if len(matches) == 0:
@@ -47,11 +52,29 @@ class BackendCache:
return matches[0]
else:
raise CellxgeneException(
status.HTTP_500_INTERNAL_SERVER_ERROR,
"Found " + str(len(matches)) + " for " + dataset,
HTTPStatus.INTERNAL_SERVER_ERROR,
"Found " + str(len(matches)) + " for " + path,
)
def create_entry(self, key, scripts):
def check_entry(self, key):
contents = self.entry_list
matches = [
c
for c in contents
if c.key.equals(key) and c.status != CacheEntryStatus.terminated
]
if len(matches) == 0:
return None
elif len(matches) == 1:
return matches[0]
else:
raise CellxgeneException(
HTTPStatus.INTERNAL_SERVER_ERROR,
"Found " + str(len(matches)) + " for " + key.dataset,
)
def create_entry(self, key: CacheKey, scripts: List[str]):
port = 8000
existing_ports = self.get_ports()

View File

@@ -6,17 +6,30 @@
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import psutil
import logging
import datetime
import logging
import re
from enum import Enum
from flask import make_response, request, render_template
import psutil
from flask import make_response, render_template, request
from flask.wrappers import Response
from requests import get, post, put
from cellxgene_gateway import env
from cellxgene_gateway.cellxgene_exception import CellxgeneException
from cellxgene_gateway.util import current_time_stamp
from cellxgene_gateway.flask_util import querystring
from cellxgene_gateway.util import current_time_stamp
logger = logging.getLogger(__name__)
class CacheEntryStatus(Enum):
loaded = "loaded"
loading = "loading"
error = "error"
terminated = "terminated"
class CacheEntry:
def __init__(
@@ -26,7 +39,7 @@ class CacheEntry:
port,
launchtime,
timestamp,
status,
status: CacheEntryStatus,
message,
all_output,
stderr,
@@ -45,29 +58,32 @@ class CacheEntry:
@classmethod
def for_key(cls, key, port):
return cls(
None,
key,
port,
current_time_stamp(),
current_time_stamp(),
"loading",
CacheEntryStatus.loading,
None,
None,
None,
None,
)
@property
def source_name(self):
return self.key.source_name
def set_loaded(self, pid):
self.pid = pid
self.status = "loaded"
self.status = CacheEntryStatus.loaded
def set_error(self, message, stderr, http_status):
self.message = message
self.stderr = stderr
self.http_status = http_status
self.status = "error"
self.status = CacheEntryStatus.error
def append_output(self, output):
if self.all_output == None:
@@ -77,101 +93,118 @@ class CacheEntry:
def terminate(self):
pid = self.pid
if pid != None and self.status != "terminated":
if pid != None and self.status != CacheEntryStatus.terminated:
terminated = []
def on_terminate(p):
terminated.append(p.pid)
p = psutil.Process(pid)
children = p.children()
for child in children:
child.terminate()
psutil.wait_procs(children, callback=on_terminate)
terminated.append(p.pid)
p.terminate()
psutil.wait_procs([p], callback=on_terminate)
logging.getLogger("cellxgene_gateway").info(f"terminated {terminated}")
self.status = "terminated"
# the parent process may automatically die once its children have --
try:
p.terminate()
psutil.wait_procs([p], callback=on_terminate)
except psutil.NoSuchProcess:
pass
logger.info(f"terminated {terminated}")
self.status = CacheEntryStatus.terminated
def rewrite_text_content(self, cellxgene_content):
# for v0.16.0 compatibility, see issue #24
gateway_content = (
re.sub(
'(="|\()/static/',
f"\\1{self.key.gateway_basepath()}static/",
cellxgene_content,
)
.replace("http://fonts.gstatic.com", "https://fonts.gstatic.com")
.replace(self.cellxgene_basepath(), self.key.gateway_basepath())
)
return gateway_content
def cellxgene_basepath(self):
return f"http://127.0.0.1:{self.port}"
def serve_content(self, path):
gateway_basepath = (
f"{env.external_protocol}://{env.external_host}/view/{self.key.pathpart}/"
)
subpath = path[len(self.key.pathpart) :] # noqa: E203
gateway_basepath = self.key.gateway_basepath()
subpath = path[len(self.key.descriptor) :] # noqa: E203
if len(subpath) == 0:
r = make_response(f"Redirect to {gateway_basepath}\n", 301)
r.headers["location"] = gateway_basepath+querystring()
r = make_response(f"Redirect to {gateway_basepath}\n", 302)
r.headers["location"] = gateway_basepath + querystring()
return r
elif self.status == "loading":
elif self.status == CacheEntryStatus.loading:
launch_time = datetime.datetime.fromtimestamp(self.launchtime)
return render_template(
"loading.html", launchtime=launch_time, all_output=self.all_output
"loading.html",
launchtime=launch_time,
all_output=self.all_output,
)
port = self.port
cellxgene_basepath = f"http://127.0.0.1:{port}"
headers = {}
copy_headers = [
'accept',
'accept-encoding',
'accept-language',
'cache-control',
'connection',
'content-length',
'content-type',
'cookie',
'host',
'origin',
'pragma',
'referer',
'sec-fetch-mode',
'sec-fetch-site',
'user-agent'
"accept",
# "accept-encoding" - removed: let requests library handle compression/decompression
"accept-language",
"cache-control",
"connection",
"content-length",
"content-type",
"cookie",
"host",
"origin",
"pragma",
"referer",
"sec-fetch-mode",
"sec-fetch-site",
"user-agent",
]
for h in copy_headers:
if h in request.headers:
headers[h] = request.headers[h]
full_path = cellxgene_basepath + subpath + querystring()
full_path = self.cellxgene_basepath() + subpath + querystring()
cellxgene_response = None
try:
if request.method in ["GET", "HEAD", "OPTIONS"]:
cellxgene_response = get(full_path, headers=headers)
elif request.method == "PUT":
cellxgene_response = put(
full_path,
headers=headers,
data=request.data,
)
elif request.method == "POST":
cellxgene_response = post(
full_path,
headers=headers,
data=request.data,
)
else:
raise CellxgeneException(f"Unexpected method {request.method}", 400)
content_type = cellxgene_response.headers["content-type"]
if "text" in content_type:
gateway_content = self.rewrite_text_content(
cellxgene_response.content.decode()
)
else:
gateway_content = cellxgene_response.content
if request.method in ["GET", "HEAD", "OPTIONS"]:
cellxgene_response = get(
full_path, headers=headers
)
elif request.method == "PUT":
cellxgene_response = put(
full_path,
headers=headers,
data=request.data,
)
elif request.method == "POST":
cellxgene_response = post(
full_path,
headers=headers,
data=request.data,
)
else:
raise CellxgeneException(
f"Unexpected method {request.method}", 400
)
content_type = cellxgene_response.headers["content-type"]
if "text" in content_type:
cellxgene_content = cellxgene_response.content.decode()
gateway_content = cellxgene_content.replace(
"http://fonts.gstatic.com", "https://fonts.gstatic.com"
).replace(cellxgene_basepath, gateway_basepath)
else:
gateway_content = cellxgene_response.content
resp_headers = {}
for h in copy_headers:
if h in cellxgene_response.headers:
resp_headers[h] = cellxgene_response.headers[h]
gateway_response = make_response(
gateway_content,
cellxgene_response.status_code,
resp_headers,
)
resp_headers = {}
for h in copy_headers:
if h in cellxgene_response.headers:
resp_headers[h] = cellxgene_response.headers[h]
gateway_response = make_response(
gateway_content,
cellxgene_response.status_code,
resp_headers,
)
finally:
if cellxgene_response is not None:
cellxgene_response.close()
return gateway_response

View File

@@ -7,23 +7,75 @@
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
from flask_api import status
from cellxgene_gateway import env
from cellxgene_gateway.cellxgene_exception import CellxgeneException
# There are three kinds of CacheKey:
# 1) somedir/dataset.h5ad: a dataset
# in this case, pathpart == dataset == 'somedir/dataset.h5ad'
# 2) somedir/dataset_annotations/saldaal1-T5HMVBNV.csv : an actual annotaitons file.
# in this case, pathpart == 'dataset_annotations/saldaal1-T5HMVBNV.csv', dataset == 'somedir/dataset.h5ad'
# in this case, descriptor == dataset == 'somedir/dataset.h5ad'
# 2) somedir/dataset_annotations/my_annotations.csv : an actual annotations file.
# in this case, descriptor == 'somedir/dataset_annotations/my_annotations.csv', dataset == 'somedir/dataset.h5ad'
# 3) somedir/dataset_annotations: an annotation directory. The corresponding h5ad must exist, but the directory may not.
# in this case, pathpart == 'dataset_annotations', dataset == 'somedir/dataset.h5ad'
# in this case, descriptor == 'somedir/dataset_annotations', dataset == 'somedir/dataset.h5ad'
from cellxgene_gateway import flask_util
from cellxgene_gateway.items.item import Item
from cellxgene_gateway.items.item_source import ItemSource, LookupResult
class CacheKey:
def __init__(self, pathpart, dataset, annotation_file):
self.pathpart = pathpart
self.dataset = dataset
self.annotation_file = annotation_file
@property
def descriptor(self):
if self.annotation_item is None:
return self.h5ad_item.descriptor
else:
return self.annotation_item.descriptor
@property
def file_path(self):
return self.source.get_local_path(self.h5ad_item)
@property
def annotation_file_path(self):
if self.annotation_item is None:
return None
else:
return self.source.get_local_path(self.annotation_item)
def relaunch_url(self):
return flask_util.relaunch_url(self.descriptor, self.source_name)
def gateway_basepath(self):
return self.view_url + "/"
@property
def view_url(self):
return flask_util.view_url(self.descriptor, self.source_name)
@property
def source_name(self):
return self.source.name
@property
def annotation_descriptor(self):
if self.annotation_item is None:
return None
else:
return self.annotation_item.descriptor
def equals(self, other):
return (
(self.source.name == other.source.name)
and (self.h5ad_item.descriptor == other.h5ad_item.descriptor)
and (self.annotation_descriptor == other.annotation_descriptor)
)
def __init__(
self, h5ad_item: Item, source: ItemSource, annotation_item: Item = None
):
assert h5ad_item is not None
assert source is not None
self.h5ad_item = h5ad_item
self.annotation_item = annotation_item
self.source = source
@classmethod
def for_lookup(cls, source: ItemSource, lookup: LookupResult):
return CacheKey(lookup.h5ad_item, source, lookup.annotation_item)

View File

@@ -9,50 +9,21 @@
import os
from flask_api import status
from cellxgene_gateway import env
from cellxgene_gateway.cellxgene_exception import CellxgeneException
def is_subdir(full_path, parent_path):
subdir = os.path.realpath(full_path)
parent = os.path.realpath(parent_path)
return subdir.startswith(parent)
annotations_suffix = "_annotations"
h5ad_suffix = ".h5ad"
def create_dir(parent_path, dir_name):
full_path = os.path.join(parent_path, dir_name)
if "/" in dir_name:
raise CellxgeneException(
"Please have no slashes in the intended directory.",
status.HTTP_400_BAD_REQUEST,
)
elif not os.path.exists(parent_path):
raise CellxgeneException(
"The selected User directory does not exist.",
status.HTTP_400_BAD_REQUEST,
)
elif os.path.exists(full_path):
raise CellxgeneException(
"The provided subdirectory already exists within Directory.",
status.HTTP_400_BAD_REQUEST,
)
elif not is_subdir(full_path, parent_path):
raise CellxgeneException(
"The directory must be a subdirectory of the parent path.",
status.HTTP_400_BAD_REQUEST,
)
elif not os.path.isdir(parent_path):
raise CellxgeneException(
"The parent is not a directory.", status.HTTP_400_BAD_REQUEST
)
else:
os.mkdir(full_path)
annotations_suffix = '_annotations'
def make_h5ad(el):
return el[:-len(annotations_suffix)]+'.h5ad'
return el[: -len(annotations_suffix)] + h5ad_suffix
def make_annotations(el):
return el[:-5]+annotations_suffix
return el[:-5] + annotations_suffix
def ensure_dir_exists(file_path):
if not os.path.exists(file_path):
os.makedirs(file_path)

View File

@@ -7,37 +7,65 @@
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
import logging
import socket
import os
cellxgene_location = os.environ.get("CELLXGENE_LOCATION")
cellxgene_data = os.environ.get("CELLXGENE_DATA")
cellxgene_data = os.environ.get("CELLXGENE_DATA", "")
cellxgene_args = os.environ.get("CELLXGENE_ARGS", None)
gateway_port = int(os.environ.get("GATEWAY_PORT", "5005"))
external_host = os.environ.get("EXTERNAL_HOST", os.environ.get("GATEWAY_HOST", f"localhost:{gateway_port}"))
external_protocol = os.environ.get("EXTERNAL_PROTOCOL", os.environ.get("GATEWAY_PROTOCOL", "http"))
external_host = os.environ.get(
"EXTERNAL_HOST",
os.environ.get("GATEWAY_HOST", f"localhost:{gateway_port}"),
)
external_protocol = os.environ.get(
"EXTERNAL_PROTOCOL", os.environ.get("GATEWAY_PROTOCOL", None)
)
ip = os.environ.get("GATEWAY_IP")
extra_scripts = os.environ.get("GATEWAY_EXTRA_SCRIPTS")
ttl = os.environ.get("GATEWAY_TTL")
enable_upload = os.environ.get("GATEWAY_ENABLE_UPLOAD", "").lower() in ['true', '1']
enable_annotations = os.environ.get("GATEWAY_ENABLE_ANNOTATIONS", "").lower() in ['true', '1']
expire_seconds = int(
os.environ.get("GATEWAY_EXPIRE_SECONDS", os.environ.get("GATEWAY_TTL", "3600"))
)
enable_annotations = os.environ.get("GATEWAY_ENABLE_ANNOTATIONS", "").lower() in [
"true",
"1",
]
enable_backed_mode = os.environ.get("GATEWAY_ENABLE_BACKED_MODE", "").lower() in [
"true",
"1",
]
log_level = logging.getLevelName(os.environ.get("GATEWAY_LOG_LEVEL", "INFO"))
env_vars = {
"CELLXGENE_LOCATION": cellxgene_location,
"CELLXGENE_DATA": cellxgene_data,
"GATEWAY_IP": ip,
}
proxy_fix_for = int(os.environ.get("PROXY_FIX_FOR", "0"))
proxy_fix_proto = int(os.environ.get("PROXY_FIX_PROTO", "0"))
proxy_fix_host = int(os.environ.get("PROXY_FIX_HOST", "0"))
proxy_fix_port = int(os.environ.get("PROXY_FIX_PORT", "0"))
proxy_fix_prefix = int(os.environ.get("PROXY_FIX_PREFIX", "0"))
optional_env_vars = {
"EXTERNAL_HOST": external_host,
"EXTERNAL_PROTOCOL": external_protocol,
"GATEWAY_IP": ip,
"GATEWAY_PORT": gateway_port,
"GATEWAY_EXTRA_SCRIPTS": extra_scripts,
"GATEWAY_TTL": ttl,
"GATEWAY_ENABLE_UPLOAD": enable_upload,
"GATEWAY_EXPIRE_SECONDS": expire_seconds,
"GATEWAY_ENABLE_ANNOTATIONS": enable_annotations,
"GATEWAY_ENABLE_BACKED_MODE": enable_backed_mode,
"GATEWAY_LOG_LEVEL": log_level,
"CELLXGENE_ARGS": cellxgene_args,
"CELLXGENE_DATA": cellxgene_data,
"PROXY_FIX_FOR": proxy_fix_for,
"PROXY_FIX_PROTO": proxy_fix_proto,
"PROXY_FIX_HOST": proxy_fix_host,
"PROXY_FIX_PORT": proxy_fix_port,
"PROXY_FIX_PREFIX": proxy_fix_prefix,
}
def validate():
if not all(env_vars.values()):
raise ValueError(
@@ -52,9 +80,12 @@ def validate():
export CELLXGENE_LOCATION=~/anaconda/envs/cellxgene-dev/bin/cellxgene
export CELLXGENE_DATA=../cellxgene_data
export GATEWAY_IP=127.0.0.1
"""
)
else:
logging.getLogger("cellxgene_gateway").info(f"Got required env: {env_vars}", )
logging.getLogger("cellxgene_gateway").info(f"Got optional env: {optional_env_vars}")
logging.getLogger("cellxgene_gateway").info(
f"Got required env: {env_vars}",
)
logging.getLogger("cellxgene_gateway").info(
f"Got optional env: {optional_env_vars}"
)

View File

@@ -7,8 +7,10 @@
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
from cellxgene_gateway import env
from json import loads
from json.decoder import JSONDecodeError
from cellxgene_gateway import env
def get_extra_scripts():
@@ -16,4 +18,9 @@ def get_extra_scripts():
# ['https://www.googletagmanager.com/gtag/js?id=UA-123456-2',
# f"{env.external_protocol}://{env.external_host}/static/js/google_ua.js"]
# where google_ua.js is a script you add to the static/js folder prior to deployment.
return [] if env.extra_scripts is None else loads(env.extra_scripts)
try:
return [] if env.extra_scripts is None else loads(env.extra_scripts)
except JSONDecodeError as exc:
raise Exception(
f'Error parsing GATEWAY_EXTRA_SCRIPTS, expected JSON array e.g. ["https://example.com/path/to/script.js"]'
) from exc

View File

@@ -7,78 +7,62 @@
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
from cellxgene_gateway import env
from cellxgene_gateway.dir_util import make_h5ad, make_annotations, annotations_suffix
import html
import urllib.parse
def recurse_dir(path):
if not os.path.exists(path):
raise CellxgeneException(
"The given path does not exist.", status.HTTP_400_BAD_REQUEST
)
all_entries = os.listdir(path)
def is_h5ad(el):
return el.endswith('.h5ad') and os.path.isfile(os.path.join(path, el))
h5ad_entries = [x for x in all_entries if is_h5ad(x)]
annotation_dir_entries = [x for x in all_entries if x.endswith(annotations_suffix) and make_h5ad(x) in h5ad_entries]
def list_annotations(el):
full_path = os.path.join(path, el)
if not os.path.isdir(full_path):
entries = []
else:
entries = [{
"name": x[:-13] if (len(x) > 13 and x[-13] in ['-','_']) else (
x[:-4] if x.endswith('.csv') else x),
"path": os.path.join(full_path, x).replace(env.cellxgene_data, ""),
} for x in os.listdir(full_path) if x.endswith('.csv') and os.path.isfile(os.path.join(full_path, x))]
return [{"name":'new', "class":'new', "path":full_path.replace(env.cellxgene_data, "")}] + entries
def make_entry(el):
full_path = os.path.join(path, el)
if el in h5ad_entries:
return {
"path": full_path.replace(env.cellxgene_data, ""),
"name": el,
"type": "file",
"annotations": list_annotations(make_annotations(el)),
}
elif os.path.isdir(full_path) and el not in annotation_dir_entries:
return {
"path": full_path.replace(env.cellxgene_data, ""),
"name": el,
"type": "directory",
"children": recurse_dir(full_path),
}
else:
return {
"path": full_path,
"name": el,
"type": "neither",
}
return [make_entry(x) for x in os.listdir(path)]
from cellxgene_gateway import flask_util
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.dir_util import annotations_suffix, make_annotations, make_h5ad
from cellxgene_gateway.env import enable_annotations
def render_entries(entries):
return "<ul>" + "\n".join([render_entry(e) for e in entries]) + "</ul>"
def get_url(entry):
return f"/view/{ entry['path'].lstrip('/') }"
def get_class(entry):
return f" class='{entry['class']}'" if 'class' in entry else ''
def render_annotations(entry):
if len(entry['annotations']) > 0:
return ' | annotations: ' + ", ".join([f"<a href='{get_url(a)}'{get_class(a)}>{a['name']}</a>" for a in entry['annotations']])
else:
return ''
def render_entry(entry):
if entry["type"] == "file":
return f"<li> <a href='{ get_url(entry) }'>{entry['name']}</a> {render_annotations(entry)}</li>"
elif entry["type"] == "directory":
url = f"/filecrawl/{entry['path'].lstrip('/')}"
return f"<li><a href='{url}'>{entry['name']}</a>{render_entries(entry['children'])}</li>"
else:
def render_annotations(item, item_source):
if not enable_annotations:
return ""
url = flask_util.view_url(
item_source.get_annotations_subpath(item), item_source.name
)
new_annotation = [f"<a class='new' href='{url}'>new</a>"]
annotations = (
[
f"<a href='{CacheKey(item, item_source, a).view_url}/'>{html.escape(a.name)}</a>"
for a in item.annotations
]
if item.annotations
else []
)
return "| annotations: " + ", ".join(new_annotation + annotations)
def render_item(item, item_source):
item_string = f"<li> <a href='{ CacheKey(item, item_source).view_url }/'>{item.name}</a> {render_annotations(item, item_source)}</li>"
return item_string
def render_item_tree(item_tree, item_source):
items = (
"\n".join([render_item(i, item_source) for i in item_tree.items])
if item_tree.items
else ""
)
branches = (
"\n".join([render_item_tree(b, item_source) for b in item_tree.branches])
if item_tree.branches
else ""
)
html = "<ul>" + items + branches + "</ul>"
if item_tree.descriptor:
descriptor = item_tree.descriptor.lstrip("/")
url = f"/filecrawl/{descriptor}?source={item_source.name}"
name = descriptor.rsplit("/", 1)[-1]
return f"<li><a href='{url}'>{name}</a>{html}</li>"
else:
return html
def render_item_source(item_source, filter=None):
item_tree = item_source.list_items(filter)
filterpart = "" if filter is None else ":" + filter
heading = f"<h6><a href='/filecrawl.html?source={urllib.parse.quote_plus(item_source.name)}'>{item_source.name}</a>{filterpart}</h6>"
return heading + render_item_tree(item_tree, item_source)

View File

@@ -7,8 +7,27 @@
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
from flask import request
from flask import request, url_for
def querystring():
qs = request.query_string.decode()
return f'?{qs}' if len(qs) > 0 else ''
return f"?{qs}" if len(qs) > 0 else ""
include_source_in_url = False
def url(endpoint, descriptor, source_name):
if include_source_in_url:
return url_for(endpoint, source_name=source_name, path=descriptor)
else:
return url_for(endpoint, path=descriptor)
def view_url(descriptor, source_name):
return url("do_view", descriptor, source_name)
def relaunch_url(descriptor, source_name):
return url("do_relaunch", descriptor, source_name)

View File

@@ -6,52 +6,164 @@
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
# import BaseHTTPServer
import os
import logging
from threading import Thread, Lock
import json
import logging
import os
import urllib.parse
from threading import Lock, Thread
from flask import (
Flask,
redirect,
make_response,
redirect,
render_template,
request,
send_from_directory,
url_for,
)
from flask_api import status
from werkzeug.utils import secure_filename
from werkzeug.middleware.proxy_fix import ProxyFix
from cellxgene_gateway import env
from cellxgene_gateway import env, flask_util
from cellxgene_gateway.backend_cache import BackendCache
from cellxgene_gateway.cache_entry import CacheEntryStatus
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.cellxgene_exception import CellxgeneException
from cellxgene_gateway.dir_util import create_dir, is_subdir
from cellxgene_gateway.filecrawl import recurse_dir, render_entries
from cellxgene_gateway.extra_scripts import get_extra_scripts
from cellxgene_gateway.filecrawl import render_item_source
from cellxgene_gateway.process_exception import ProcessException
from cellxgene_gateway.prune_process_cache import PruneProcessCache
from cellxgene_gateway.util import current_time_stamp
from cellxgene_gateway.path_util import get_key
app = Flask(__name__)
item_sources = []
default_item_source = None
# Guard for lazy initialization so tests can import this module without
# triggering environment-dependent side effects. initialize_data_sources()
# will set this to True when it has run.
data_sources_initialized = False
data_sources_init_lock = Lock()
def _force_https(app):
def wrapper(environ, start_response):
environ['wsgi.url_scheme'] = env.external_protocol
if env.external_protocol is not None:
environ["wsgi.url_scheme"] = env.external_protocol
return app(environ, start_response)
return wrapper
def set_no_cache(resp):
resp.headers["Cache-Control"] = "no-cache, no-store, must-revalidate"
resp.headers["Pragma"] = "no-cache"
resp.headers["Expires"] = "0"
resp.headers["Cache-Control"] = "public, max-age=0"
return resp
app.wsgi_app = _force_https(app.wsgi_app)
if (
env.proxy_fix_for > 0
or env.proxy_fix_proto > 0
or env.proxy_fix_host > 0
or env.proxy_fix_port > 0
or env.proxy_fix_prefix > 0
):
app.wsgi_app = ProxyFix(
app.wsgi_app,
x_for=env.proxy_fix_for,
x_proto=env.proxy_fix_proto,
x_host=env.proxy_fix_host,
x_port=env.proxy_fix_port,
x_prefix=env.proxy_fix_prefix,
)
# WSGI middleware to ensure data sources are initialized before the first
# WSGI request is handled. This guarantees initialization works under
# Gunicorn/uWSGI (which import the module but don't call main()). The
# initialize_data_sources() function is idempotent-protected by
# data_sources_initialized and data_sources_init_lock.
def _init_on_first_wsgi_request(wsgi_app):
def middleware(environ, start_response):
global data_sources_initialized
if not data_sources_initialized:
with data_sources_init_lock:
if not app.extensions.get("cellxgene_gateway", {}).get("launchtime"):
app.extensions.setdefault("cellxgene_gateway", {})[
"launchtime"
] = current_time_stamp()
if not data_sources_initialized:
initialize_data_sources()
env.validate()
if not item_sources or not len(item_sources):
raise Exception(
"No data sources specified for Cellxgene Gateway"
)
global default_item_source
if default_item_source is None:
default_item_source = item_sources[0]
data_sources_initialized = True
return wsgi_app(environ, start_response)
return middleware
# Wrap the WSGI app so Gunicorn/uWSGI will trigger initialization when the
# first request comes in. Tests that need initialization can call
# initialize_data_sources() directly.
app.wsgi_app = _init_on_first_wsgi_request(app.wsgi_app)
cache = BackendCache()
location = f"{env.external_protocol}://{env.external_host}"
# Initialize data sources - this is defined later in the file but called here
# to ensure initialization happens when WSGI servers (Gunicorn) import the module
def initialize_data_sources():
"""Initialize data sources from environment variables.
Called at module import time for WSGI server compatibility (Gunicorn).
Uses a guard flag to prevent double initialization within a process."""
global default_item_source
logging.basicConfig(
level=env.log_level,
format="%(asctime)s:%(name)s:%(levelname)s:%(message)s",
)
logger = logging.getLogger(__name__)
cellxgene_data = os.environ.get("CELLXGENE_DATA", None)
cellxgene_bucket = os.environ.get("CELLXGENE_BUCKET", None)
if cellxgene_bucket is not None:
from cellxgene_gateway.items.s3.s3item_source import S3ItemSource
s3_source = S3ItemSource(cellxgene_bucket, name="s3")
item_sources.append(s3_source)
default_item_source = s3_source
logger.info("Initialized S3 data source")
logger.debug(f"S3 bucket: {cellxgene_bucket}")
if cellxgene_data is not None:
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
file_source = FileItemSource(cellxgene_data, name="local")
item_sources.append(file_source)
default_item_source = file_source
logger.info("Initialized local file data source")
logger.debug(f"Data directory: {cellxgene_data}")
if len(item_sources) == 0:
raise Exception("Please specify CELLXGENE_DATA or CELLXGENE_BUCKET")
flask_util.include_source_in_url = len(item_sources) > 1
@app.errorhandler(CellxgeneException)
def handle_invalid_usage(error):
message = f"{error.http_status} Error : {error.message}"
return (
@@ -66,7 +178,6 @@ def handle_invalid_usage(error):
@app.errorhandler(ProcessException)
def handle_invalid_process(error):
message = []
message.append(error.message)
@@ -82,8 +193,8 @@ def handle_invalid_process(error):
http_status=error.http_status,
stdout=error.stdout,
stderr=error.stderr,
dataset=error.key.dataset,
annotation_file=error.key.annotation_file,
relaunch_url=error.key.relaunch_url(),
annotation_file=error.key.annotation_descriptor,
),
error.http_status,
)
@@ -100,167 +211,196 @@ def favicon():
@app.route("/")
def index():
users = [
name
for name in os.listdir(env.cellxgene_data)
if os.path.isdir(os.path.join(env.cellxgene_data, name))
]
return render_template(
"index.html",
ip=env.ip,
cellxgene_data=env.cellxgene_data,
extra_scripts=get_extra_scripts(),
users=users,
enable_upload=env.enable_upload,
)
def make_user():
dir_name = request.form["directory"]
create_dir(env.cellxgene_data, dir_name)
return redirect(location, code=302)
def make_subdir():
parent_path = os.path.join(env.cellxgene_data, request.form["usernames"])
dir_name = request.form["directory"]
create_dir(parent_path, dir_name)
return redirect(location, code=302)
def upload_file():
upload_dir = request.form["path"]
full_upload_path = os.path.join(env.cellxgene_data, upload_dir)
if is_subdir(full_upload_path, env.cellxgene_data) and os.path.isdir(full_upload_path):
if request.method == "POST":
if "file" in request.files:
f = request.files["file"]
if f and f.filename.endswith(".h5ad"):
f.save(
os.path.join(full_upload_path, secure_filename(f.filename))
)
return redirect("/filecrawl.html", code=302)
else:
raise CellxgeneException(
"Uploaded file must be in anndata (.h5ad) format.",
status.HTTP_400_BAD_REQUEST,
)
else:
raise CellxgeneException(
"A file must be chosen to upload.",
status.HTTP_400_BAD_REQUEST,
)
else:
raise CellxgeneException(
"Invalid directory.", status.HTTP_400_BAD_REQUEST
)
return redirect(location, code=302)
if env.enable_upload:
app.add_url_rule('/make_user', 'make_user', make_user, methods=["POST"])
app.add_url_rule('/make_subdir', 'make_subdir', make_subdir, methods=["POST"])
app.add_url_rule('/upload_file', 'upload_file', upload_file, methods=["POST"])
@app.route("/filecrawl.html")
def filecrawl():
entries = recurse_dir(env.cellxgene_data)
rendered_html = render_entries(entries)
resp = make_response(render_template(
"filecrawl.html",
extra_scripts=get_extra_scripts(),
rendered_html=rendered_html,
))
resp.headers["Cache-Control"] = "no-cache, no-store, must-revalidate"
resp.headers["Pragma"] = "no-cache"
resp.headers["Expires"] = "0"
resp.headers['Cache-Control'] = 'public, max-age=0'
@app.route("/filecrawl/<path:path>")
def filecrawl(path=None):
source_name = request.args.get("source")
sources = (
filter(
lambda x: x.name == urllib.parse.unquote_plus(source_name),
item_sources,
)
if source_name
else item_sources
)
# loop all data sources --
rendered_sources = [
render_item_source(item_source, path) for item_source in sources
] # will we need to make this async in the page???
rendered_html = "\n".join(rendered_sources)
resp = make_response(
render_template(
"filecrawl.html",
extra_scripts=get_extra_scripts(),
rendered_html=rendered_html,
path=path,
)
)
set_no_cache(resp)
return resp
@app.route("/filecrawl/<path:path>")
def do_filecrawl(path):
filecrawl_path = os.path.join(env.cellxgene_data, path)
if not os.path.isdir(filecrawl_path):
raise CellxgeneException(
"Path is not directory: " + filecrawl_path, status.HTTP_400_BAD_REQUEST
)
entries = recurse_dir(filecrawl_path)
rendered_html = render_entries(entries)
return render_template(
"filecrawl.html",
extra_scripts=get_extra_scripts(),
rendered_html=rendered_html,
path=path,
)
entry_lock = Lock()
def matching_source(source_name):
if source_name is None and default_item_source is not None:
source_name = default_item_source.name
matching = [i for i in item_sources if i.name == source_name]
if len(matching) != 1:
raise Exception(f"Could not find matching item source {source_name}")
source = matching[0]
return source
@app.route(
"/source/<path:source_name>/view/<path:path>",
methods=["GET", "PUT", "POST"],
)
@app.route("/view/<path:path>", methods=["GET", "PUT", "POST"])
def do_view(path):
key = get_key(path)
print(f"view path={path}, dataset={key.dataset}, annotation_file= {key.annotation_file}, key={key.pathpart}")
with entry_lock:
match = cache.check_entry(key)
if match is None:
uascripts = get_extra_scripts()
match = cache.create_entry(key, uascripts)
def do_view(path, source_name=None):
source = matching_source(source_name)
match = cache.check_path(source, path)
if match is None:
lookup = source.lookup(path)
if lookup is None:
raise CellxgeneException(
f"Could not find item for path {path} in source {source.name}",
404,
)
key = CacheKey.for_lookup(source, lookup)
print(
f"view path={path}, source_name={source_name}, dataset={key.file_path}, annotation_file= {key.annotation_file_path}, key={key.descriptor}, source={key.source_name}"
)
with entry_lock:
match = cache.check_entry(key)
if match is None:
uascripts = get_extra_scripts()
match = cache.create_entry(key, uascripts)
match.timestamp = current_time_stamp()
if match.status == "loaded" or match.status == "loading":
return match.serve_content(path)
elif match.status == "error":
if (
match.status == CacheEntryStatus.loaded
or match.status == CacheEntryStatus.loading
):
if source.is_authorized(match.key.descriptor):
return match.serve_content(path)
else:
raise CellxgeneException("User not authorized to access this data", 403)
elif match.status == CacheEntryStatus.error:
raise ProcessException.from_cache_entry(match)
else:
raise CellxgeneException(
f"Unexpected cache entry status {match.status} for key {match.key.descriptor}",
500,
)
@app.route("/cache_status", methods=["GET"])
def do_GET_status():
return render_template("cache_status.html", entry_list=cache.entry_list)
return render_template(
"cache_status.html",
entry_list=cache.entry_list,
extra_scripts=get_extra_scripts(),
)
@app.route("/cache_status.json", methods=["GET"])
def do_GET_status_json():
return json.dumps({'launchtime':app.launchtime,
'entry_list':[{
'dataset': entry.key.dataset,
'annotation_file': entry.key.annotation_file,
'launchtime': entry.launchtime,
'last_access': entry.timestamp,
'status': entry.status
} for entry in cache.entry_list]})
def map_entry(entry):
dataset = entry.key.h5ad_item.descriptor
annotation_file = entry.key.annotation_descriptor
return {
"dataset": dataset,
"annotation_file": annotation_file,
"launchtime": entry.launchtime,
"last_access": entry.timestamp,
"status": entry.status.name,
}
return json.dumps(
{
"launchtime": app.extensions.get("cellxgene_gateway", {}).get("launchtime"),
"entry_list": [map_entry(entry) for entry in cache.entry_list],
}
)
def get_cache_key(path):
if request.args.get("source_name"):
source_name = request.args.get("source_name")
elif default_item_source:
source_name = default_item_source.name
else:
source_name = None
source = matching_source(source_name)
key = CacheKey.for_lookup(source, source.lookup(path))
return key
@app.route("/relaunch/<path:path>", methods=["GET"])
def do_relaunch(path):
key = get_key(path)
key = get_cache_key(path)
match = cache.check_entry(key)
if not match is None:
match.terminate()
qs = request.query_string.decode()
return redirect(url_for("do_view", path=path) + (f'?{qs}' if len(qs) > 0 else ''), code=302)
return redirect(
key.view_url,
code=302,
)
@app.route("/terminate/<path:path>", methods=["GET"])
def do_terminate(path):
key = get_key(path)
key = get_cache_key(path)
match = cache.check_entry(key)
if not match is None:
match.terminate()
return redirect(url_for("do_GET_status"), code=302)
def main():
logging.basicConfig(level=logging.INFO, format='%(asctime)s:%(name)s:%(levelname)s:%(message)s')
env.validate()
pruner = PruneProcessCache(cache)
@app.route("/metadata/ip_address", methods=["GET"])
def ip_address():
resp = make_response(env.ip)
return set_no_cache(resp)
background_thread = Thread(target=pruner)
def start_pruner_thread():
pruner = PruneProcessCache(cache)
# Run the pruner as a daemon thread so it won't block interpreter
# shutdown (for example when Ctrl-C is used in the main thread).
# This avoids "Exception ignored in: <module 'threading'...>" at exit.
background_thread = Thread(target=pruner, daemon=True)
background_thread.start()
app.launchtime = current_time_stamp()
def launch():
start_pruner_thread()
app.extensions.setdefault("cellxgene_gateway", {})[
"launchtime"
] = current_time_stamp()
app.run(host="0.0.0.0", port=env.gateway_port, debug=False)
app.extensions.setdefault("cellxgene_gateway", {})["launchtime"] = None
def main():
"""CLI entry point for Flask development server."""
launch()
if __name__ == "__main__":
main()

View File

View File

View File

@@ -0,0 +1,28 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
from cellxgene_gateway.items.item import Item
class FileItem(Item):
"""e.g. FileItem(subpath = subpath, name = filename, type = ItemType.h5ad)
The Item superclass expects a 'name' and 'type'.
"""
def __init__(self, subpath: str, ext: str = "", *args, **kwargs):
super().__init__(*args, **kwargs)
self.subpath = subpath
self.ext = ext
@property
def descriptor(self) -> str:
return os.path.join(self.subpath, self.name + self.ext).strip("/")

View File

@@ -0,0 +1,214 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
from typing import List
from cellxgene_gateway import dir_util
from cellxgene_gateway.items.file.fileitem import FileItem
from cellxgene_gateway.items.item import ItemTree, ItemType
from cellxgene_gateway.items.item_source import ItemSource, LookupResult
class FileItemSource(ItemSource):
def __init__(
self,
base_path,
name=None,
h5ad_suffix=dir_util.h5ad_suffix,
annotation_dir_suffix=dir_util.annotations_suffix,
annotation_file_suffix=".csv",
gene_set_file_suffix="_gene_sets.csv",
):
self._name = name
self.base_path = base_path
self.h5ad_suffix = h5ad_suffix
self.annotation_dir_suffix = annotation_dir_suffix
self.annotation_file_suffix = annotation_file_suffix
self.gene_set_file_suffix = gene_set_file_suffix
@property
def name(self):
return self._name or f"Files:{self.base_path}"
def is_gene_set(self, path: str) -> bool:
return path.endswith(self.gene_set_file_suffix)
def is_h5ad_file(self, path: str) -> bool:
return path.endswith(self.h5ad_suffix) and os.path.isfile(path)
def convert_annotation_path_to_h5ad(self, path):
return path[: -len(self.annotation_dir_suffix)] + self.h5ad_suffix
def convert_h5ad_path_to_annotation(self, path):
return path[: -len(self.h5ad_suffix)] + self.annotation_dir_suffix
def get_local_path(self, item: FileItem) -> str:
return os.path.join(self.base_path, item.descriptor)
def get_annotations_subpath(self, item) -> str:
return self.convert_h5ad_path_to_annotation(item.descriptor)
def list_items(self, filter: str = None) -> ItemTree:
item_tree = self.scan_directory("" if filter is None else filter)
"""def get_items(dir):
if dir.branches:
return [*dir.items, *[item for subdir in dir.branches for item in get_items(subdir)]]
else:
return dir.items
return get_items(self.item_tree)"""
return item_tree
def scan_directory(self, subpath: str = "") -> ItemTree:
base_path = os.path.join(self.base_path, subpath)
if not os.path.exists(base_path):
raise Exception(f"Path for local files '{base_path}' does not exist.")
filepath_map = dict(
(filepath, os.path.join(base_path, filepath))
for filepath in sorted(os.listdir(base_path))
)
def is_annotation_dir(dir):
return (
dir.endswith(self.annotation_dir_suffix)
and self.convert_annotation_path_to_h5ad(dir) in h5ad_paths
)
h5ad_paths = [
filepath
for filepath, full_path in filepath_map.items()
if self.is_h5ad_file(full_path)
]
subdirs = [
filepath
for filepath, full_path in filepath_map.items()
if os.path.isdir(full_path) and not is_annotation_dir(filepath)
]
items = [
self.make_fileitem_from_path(filename, subpath) for filename in h5ad_paths
]
branches = None
if len(subdirs) > 0:
branches = [
self.scan_directory(os.path.join(subpath, subdir)) for subdir in subdirs
]
# Exclude branches without files as leaves. Since traversal is applied pre-order,
# branch.branches has already been processed and we don't need to check deeper nesting.
branches = [
branch for branch in branches if branch.items or branch.branches
]
return ItemTree(subpath, items, branches)
def create_annotation(self, item: FileItem, name: str) -> FileItem:
annotation = self.make_fileitem_from_path(
name, self.get_annotations_subpath(item), is_annotation=True
)
item.annotations = (item.annotations or []).append(annotation)
return annotation
def update(self, item: FileItem) -> None:
pass
def full_path(self, p):
return os.path.join(self.base_path, p)
def lookup_item(self, descriptor):
full_path = self.full_path(descriptor)
if self.is_h5ad_file(full_path):
return self.shallowitem_from_descriptor(descriptor)
def is_authorized(self, descriptor):
return True
def lookup(self, indescriptor: str) -> LookupResult:
descriptor = indescriptor.strip("/")
if descriptor.endswith(self.annotation_file_suffix):
annotation_item = self.shallowitem_from_descriptor(descriptor, True)
h5ad_descriptor = self.convert_annotation_path_to_h5ad(
annotation_item.subpath
)
item = self.lookup_item(h5ad_descriptor)
if item is not None:
dir_util.ensure_dir_exists(self.full_path(annotation_item.subpath))
return LookupResult(item, annotation_item)
else:
item = self.lookup_item(descriptor)
if item is not None:
return LookupResult(item)
def shallowitem_from_descriptor(self, descriptor, is_annotation=False):
filename = os.path.basename(descriptor)
subpath = os.path.dirname(descriptor)
return self.make_fileitem_from_path(
filename,
subpath,
is_annotation,
True,
)
def make_fileitem_from_path(
self, filename, subpath, is_annotation=False, is_shallow=False
) -> FileItem:
if is_annotation and filename.endswith(self.annotation_file_suffix):
name = filename[: -len(self.annotation_file_suffix)]
ext = self.annotation_file_suffix
else:
name = filename
ext = ""
item = FileItem(
subpath=subpath,
name=name,
ext=ext,
type=ItemType.annotation if is_annotation else ItemType.h5ad,
)
if not is_annotation and not is_shallow:
annotations = self.make_annotations_for_fileitem(item)
item.annotations = annotations
return item
def make_annotations_for_fileitem(self, item: FileItem) -> List[FileItem]:
annotations_subpath = self.get_annotations_subpath(item)
annotations_fullpath = self.full_path(annotations_subpath)
if os.path.isdir(annotations_fullpath):
sorted_files = sorted(os.listdir(annotations_fullpath))
annotation_files = [
self.make_fileitem_from_path(annotation, annotations_subpath, True)
for annotation in sorted_files
if annotation.endswith(self.annotation_file_suffix)
and not self.is_gene_set(annotation)
and os.path.isfile(os.path.join(annotations_fullpath, annotation))
]
# Catch gene sets without accompanying [annotations].csv
gene_sets_files = [
self.make_fileitem_from_path(
annotation[: -len(self.gene_set_file_suffix)] + ".csv",
annotations_subpath,
True,
)
for annotation in sorted_files
if self.is_gene_set(annotation)
and annotation[: -len(self.gene_set_file_suffix)]
not in [a.name for a in annotation_files]
and os.path.isfile(os.path.join(annotations_fullpath, annotation))
]
return sorted(annotation_files + gene_sets_files, key=lambda x: x.name)
else:
return None

View File

@@ -0,0 +1,41 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
from abc import ABC, abstractmethod
from enum import Enum
from typing import List
class ItemType(Enum):
annotation = "annotation"
h5ad = "h5ad"
class Item(ABC):
def __init__(self, name: str, type: ItemType, annotations: List["Item"] = None):
self.name = name
self.type = type
self.annotations = annotations
@property
@abstractmethod
def descriptor(self):
raise Exception('"descriptor" not implemented')
class ItemTree:
def __init__(
self,
descriptor: str,
items: List[Item] = None,
branches: List["ItemTree"] = None,
):
self.descriptor = descriptor
self.items = items
self.branches = branches

View File

@@ -0,0 +1,54 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
from abc import ABC, abstractmethod
from typing import List
from cellxgene_gateway.items.item import Item
class LookupResult:
def __init__(self, h5ad_item: Item, annotation_item: Item = None):
self.h5ad_item = h5ad_item
self.annotation_item = annotation_item
class ItemSource(ABC):
@abstractmethod
def list_items(self, filter: str = None) -> List[Item]:
raise Exception('"list_items" unimplemented')
@abstractmethod
def get_local_path(self, item: Item) -> str:
raise Exception('"local_path" unimplemented')
@abstractmethod
def get_annotations_subpath(self, item) -> str:
raise Exception('"annotations_path" unimplemented')
@abstractmethod
def create_annotation(self, item: Item, name: str) -> Item:
raise Exception('"annotation" unimplemented')
@abstractmethod
def update(self, item: Item) -> None:
raise Exception('"update" unimplemented')
@abstractmethod
def is_authorized(self, descriptor: str) -> bool:
raise Exception('"is_authorized" unimplemented')
@abstractmethod
def lookup(self, descriptor: str) -> LookupResult:
raise Exception('"lookup" unimplemented')
@property
@abstractmethod
def name(self):
pass

View File

View File

@@ -0,0 +1,27 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
from cellxgene_gateway.items.item import Item
class S3Item(Item):
"""e.g. FileItem(subpath = subpath, name = filename, type = ItemType.h5ad)
The Item superclass expects a 'name' and 'type'.
"""
def __init__(self, s3key: str, *args, **kwargs):
super().__init__(*args, **kwargs)
self.s3key = s3key
@property
def descriptor(self) -> str:
return self.s3key

View File

@@ -0,0 +1,195 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
from os.path import basename, dirname
from typing import List
import flask
import s3fs
from cellxgene_gateway import dir_util
from cellxgene_gateway.items.item import ItemTree, ItemType
from cellxgene_gateway.items.item_source import ItemSource, LookupResult
from cellxgene_gateway.items.s3.s3item import S3Item
def truthy(val: str):
return val.lower() in ["true", "1"]
class S3ItemSource(ItemSource):
def __init__(
self,
bucket,
name=None,
h5ad_suffix=dir_util.h5ad_suffix,
annotation_dir_suffix=dir_util.annotations_suffix,
annotation_file_suffix=".csv",
):
self._name = name
enable_cache = os.environ.get("S3_ENABLE_LISTINGS_CACHE", "false").lower()
assert enable_cache in ["0", "1", "false", "true"]
self.use_listings_cache = truthy(enable_cache)
self.s3 = s3fs.S3FileSystem(use_listings_cache=self.use_listings_cache)
if bucket.startswith("s3://"):
raise Exception(
f"Bucket name should not include s3:// prefix, got {bucket}"
)
self.bucket = bucket
self.h5ad_suffix = h5ad_suffix
self.annotation_dir_suffix = annotation_dir_suffix
self.annotation_file_suffix = annotation_file_suffix
def url(self, key):
return "s3://" + self.bucket + "/" + key
def remove_bucket(self, filepath):
return filepath[len(self.bucket) :].lstrip("/")
@property
def name(self):
return self._name or f"Items:{self.url('')}"
def is_h5ad_url(self, s3url: str) -> bool:
return s3url.endswith(self.h5ad_suffix) and self.s3.exists(s3url)
def convert_annotation_key_to_h5ad(self, s3key):
return s3key[: -len(self.annotation_dir_suffix)] + self.h5ad_suffix
def convert_h5ad_key_to_annotation(self, s3key):
return s3key[: -len(self.h5ad_suffix)] + self.annotation_dir_suffix
def get_local_path(self, item: S3Item) -> str:
return self.url(item.descriptor)
def get_annotations_subpath(self, item) -> str:
return self.convert_h5ad_key_to_annotation(item.descriptor)
def list_items(self, filter: str = None) -> ItemTree:
item_tree = self.scan_directory("" if filter is None else filter)
return item_tree
@property
def refresh(self):
return (
truthy(flask.request.args.get("refresh", default="false"))
or not self.use_listings_cache
)
def scan_directory(self, directory_key="") -> dict:
url = self.url(directory_key)
if not self.s3.exists(url):
raise Exception(f"S3 url '{url}' does not exist.")
s3key_map = dict(
(self.remove_bucket(filepath), "s3://" + filepath)
for filepath in sorted(self.s3.ls(url, refresh=self.refresh))
)
def is_annotation_dir(dir_s3key):
return (
dir_s3key.endswith(self.annotation_dir_suffix)
and self.convert_annotation_key_to_h5ad(dir_s3key) in h5ad_keys
)
h5ad_keys = [
filepath
for filepath, item_url in s3key_map.items()
if self.is_h5ad_url(item_url)
]
subdir_keys = [
filepath
for filepath, item_url in s3key_map.items()
if self.s3.isdir(item_url) and not is_annotation_dir(filepath)
]
items = [self.make_s3item_from_key(basename(key), key) for key in h5ad_keys]
branches = None
if len(subdir_keys) > 0:
branches = [self.scan_directory(key) for key in subdir_keys]
branches = [
branch for branch in branches if branch.items or branch.branches
]
return ItemTree(directory_key, items, branches)
def create_annotation(self, item: S3Item, name: str) -> S3Item:
annotation = self.make_s3item_from_key(
name, self.get_annotations_subpath(item), is_annotation=True
)
item.annotations = (item.annotations or []).append(annotation)
return annotation
def update(self, item: S3Item) -> None:
pass
def is_authorized(self, descriptor):
return True
def lookup_item(self, descriptor):
full_path = self.url(descriptor)
if self.is_h5ad_url(full_path):
return self.shallowitem_from_descriptor(descriptor)
def lookup(self, indescriptor: str) -> LookupResult:
descriptor = indescriptor.strip("/")
if descriptor.endswith(self.annotation_file_suffix):
annotation_item = self.shallowitem_from_descriptor(descriptor, True)
if not self.s3.exists(self.url(annotation_item.s3key)):
with self.s3.open(self.url(annotation_item.s3key), "w") as f:
f.write("")
h5ad_descriptor = self.convert_annotation_key_to_h5ad(
dirname(annotation_item.s3key)
)
item = self.shallowitem_from_descriptor(h5ad_descriptor)
return LookupResult(item, annotation_item)
else:
item = self.lookup_item(descriptor)
if item is not None:
return LookupResult(item)
def shallowitem_from_descriptor(self, descriptor, is_annotation=False):
return self.make_s3item_from_key(
basename(descriptor), descriptor, is_annotation, True
)
def make_s3item_from_key(
self, name, s3key, is_annotation=False, is_shallow=False
) -> S3Item:
item = S3Item(
s3key=s3key,
name=name,
type=ItemType.annotation if is_annotation else ItemType.h5ad,
)
if not is_annotation and not is_shallow:
annotations = self.make_annotations_for_fileitem(item)
item.annotations = annotations
return item
def make_annotations_for_fileitem(self, item: S3Item) -> List[S3Item]:
annotations_subpath = self.get_annotations_subpath(item)
annotations_fullpath = self.url(annotations_subpath)
if self.s3.isdir(annotations_fullpath):
return [
self.make_s3item_from_key(
basename(annotation), self.remove_bucket(annotation), True
)
for annotation in sorted(
self.s3.ls(annotations_fullpath, refresh=self.refresh)
)
if annotation.endswith(self.annotation_file_suffix)
and self.s3.isfile("s3://" + annotation)
]
else:
return None

View File

@@ -1,95 +0,0 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
from flask_api import status
from cellxgene_gateway import env
from cellxgene_gateway.cellxgene_exception import CellxgeneException
from cellxgene_gateway.dir_util import make_h5ad
from cellxgene_gateway.cache_key import CacheKey
def get_key(path):
if path == "/" or path == "":
raise CellxgeneException(
"No matching dataset found.", status.HTTP_404_NOT_FOUND
)
trimmed = path[:-1] if path[-1] == "/" else path
try:
# valid paths come in three forms:
if trimmed.endswith('.h5ad') and data_file_exists(trimmed):
# 1) somedir/dataset.h5ad: a dataset
return CacheKey(trimmed, trimmed, None)
elif trimmed.endswith('.csv'):
# 2) somedir/dataset_annotations/saldaal1-T5HMVBNV.csv : an actual annotations file.
annotations_dir = os.path.split(trimmed)[0]
dataset = make_h5ad(annotations_dir)
if data_file_exists(dataset):
data_dir_ensure(annotations_dir)
return CacheKey(trimmed, dataset, trimmed)
elif trimmed.endswith('_annotations') and data_dir_exists(trimmed):
# 3) somedir/dataset_annotations: an annotation directory. The corresponding h5ad must exist, but the directory may not.
dataset = make_h5ad(trimmed)
if data_file_exists(dataset):
return CacheKey(trimmed, dataset, '')
except CellxgeneException:
pass
split = os.path.split(trimmed)
return get_key(split[0])
def validate_exists(file_path):
if not os.path.exists(file_path):
raise CellxgeneException(
"File does not exist: " + file_path, status.HTTP_400_BAD_REQUEST
)
def validate_is_file(file_path):
validate_exists(file_path)
if not os.path.isfile(file_path):
raise CellxgeneException(
"Path is not file: " + file_path, status.HTTP_400_BAD_REQUEST
)
return
def validate_is_dir(file_path):
validate_exists(file_path)
if not os.path.isdir(file_path):
raise CellxgeneException(
"Path is not dir: " + file_path, status.HTTP_400_BAD_REQUEST
)
return
def data_file_exists(dataset):
file_path = os.path.join(env.cellxgene_data, dataset)
validate_is_file(file_path)
return True
def data_dir_exists(dataset):
file_path = os.path.join(env.cellxgene_data, dataset)
validate_is_dir(file_path)
return True
def data_dir_ensure(dataset):
file_path = os.path.join(env.cellxgene_data, dataset)
if not os.path.exists(file_path):
os.makedirs(file_path)
def get_file_path(key):
dataset = key.dataset
file_path = os.path.join(env.cellxgene_data, dataset)
validate_is_file(file_path)
return file_path
def get_annotation_file_path(key):
if key.annotation_file is None:
return None
if key.annotation_file == '':
return ''
file_path = os.path.join(env.cellxgene_data, key.annotation_file)
return file_path

View File

@@ -7,17 +7,18 @@
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import time
import logging
import time
from cellxgene_gateway.util import current_time_stamp
from cellxgene_gateway.env import ttl
from cellxgene_gateway import env, util
logger = logging.getLogger(__name__)
class PruneProcessCache:
def __init__(self, cache):
self.cache = cache
self.expire_seconds = (3600 if ttl is None else int(ttl))
self.expire_seconds = env.expire_seconds
def __call__(self):
while True:
@@ -25,16 +26,22 @@ class PruneProcessCache:
self.prune()
def prune(self):
timestamp = current_time_stamp()
timestamp = util.current_time_stamp()
cutoff = timestamp - self.expire_seconds
processes_to_delete = [p for p in self.cache.entry_list if p.timestamp < cutoff]
processes_to_keep = [p for p in self.cache.entry_list if not p.timestamp < cutoff]
logger = logging.getLogger("cellxgene_gateway")
logger.debug(f"Cutoff {cutoff} = timestamp {timestamp} - expire seconds {self.expire_seconds} , keeping {processes_to_keep}")
processes_to_delete = [p for p in self.cache.entry_list if p.timestamp < cutoff]
processes_to_keep = [
p for p in self.cache.entry_list if not p.timestamp < cutoff
]
logger.debug(
f"Cutoff {cutoff} = timestamp {timestamp} - expire seconds {self.expire_seconds} , keeping {processes_to_keep}, pruning {processes_to_delete}"
)
for process in processes_to_delete:
try:
logger.info(f"pruning process {process.pid} ({process.key.dataset})")
logger.info(f"pruning process {process.pid} ({process.key.descriptor})")
self.cache.prune(process)
except Exception:
logger.exception("failed to prune process {process.pid} ({process.dataset})")
logger.exception(
"failed to prune process {process.pid} ({process.key.descriptor})"
)

View File

@@ -1,4 +1,5 @@
// neandertal javascript
// Annotations only work with file itemsources at the moment. If they work with others in the future we may need to revisit this.
const new_annotation_callback = (() =>{
const suffix = `.csv`;
return (e) => {

View File

@@ -9,12 +9,15 @@
import logging
import subprocess
from http import HTTPStatus
from flask_api import status
from cellxgene_gateway.env import enable_annotations
from cellxgene_gateway.process_exception import ProcessException
from cellxgene_gateway.cache_entry import CacheEntryStatus
from cellxgene_gateway.dir_util import make_annotations
from cellxgene_gateway.path_util import get_file_path, get_annotation_file_path
from cellxgene_gateway.env import cellxgene_args, enable_annotations, enable_backed_mode
from cellxgene_gateway.process_exception import ProcessException
logger = logging.getLogger(__name__)
class SubprocessBackend:
def __init__(self):
@@ -22,19 +25,25 @@ class SubprocessBackend:
def create_cmd(self, cellxgene_loc, file_path, port, scripts, annotation_file_path):
if enable_annotations and not annotation_file_path is None:
annotation_args_prefix = " --experimental-annotations"
if annotation_file_path == "":
annotation_args = f"{annotation_args_prefix} --experimental-annotations-output-dir {make_annotations(file_path)}"
extra_args = f" --annotations-dir {make_annotations(file_path)}"
else:
annotation_args = f"{annotation_args_prefix} --experimental-annotations-file {annotation_file_path}"
extra_args = f" --annotations-file {annotation_file_path}"
gene_sets_file_path = annotation_file_path[:-4] + "_gene_sets.csv"
extra_args += f" --gene-sets-file {gene_sets_file_path}"
else:
annotation_args = ""
extra_args = " --disable-annotations"
extra_args += " --disable-gene-sets-save"
if enable_backed_mode:
extra_args += " --backed"
if not cellxgene_args is None:
extra_args += f" {cellxgene_args}"
cmd = (
f"yes | {cellxgene_loc} launch {file_path}"
+ " --port "
+ str(port)
+ f" --port {port}"
+ " --host 127.0.0.1"
+ annotation_args
+ extra_args
)
for s in scripts:
@@ -43,11 +52,14 @@ class SubprocessBackend:
return cmd
def launch(self, cellxgene_loc, scripts, cache_entry):
cmd = self.create_cmd(
cellxgene_loc, get_file_path(cache_entry.key), cache_entry.port, scripts, get_annotation_file_path(cache_entry.key)
cellxgene_loc,
cache_entry.key.file_path,
cache_entry.port,
scripts,
cache_entry.key.annotation_file_path,
)
logging.getLogger("cellxgene_gateway").info(f"launching {cmd}")
logger.info(f"launching {cmd}")
process = subprocess.Popen(
[cmd], stdout=subprocess.PIPE, stderr=subprocess.PIPE, shell=True
)
@@ -63,12 +75,12 @@ class SubprocessBackend:
or "Could not open file" in stderr
):
message = "File was invalid."
http_status = status.HTTP_400_BAD_REQUEST
http_status = HTTPStatus.BAD_REQUEST
else:
message = "Cellxgene failed to launch dataset."
http_status = status.HTTP_500_INTERNAL_SERVER_ERROR
http_status = HTTPStatus.INTERNAL_SERVER_ERROR
cache_entry.status = "error"
cache_entry.status = CacheEntryStatus.error
cache_entry.set_error(message, stderr, http_status)
raise ProcessException.from_cache_entry(cache_entry)
@@ -76,5 +88,6 @@ class SubprocessBackend:
cache_entry.append_output(output)
cache_entry.set_loaded(process.pid)
return
for output in process.communicate():
logger.debug(f"cellxgene:{output}")
logger.info(f"exiting {cmd}")

View File

@@ -10,65 +10,77 @@
-->
<html>
<head>
<title>Cellxgene Gateway - FILE CRAWLER</title>
<script src="https://ajax.googleapis.com/ajax/libs/jquery/3.3.1/jquery.min.js"></script>
<link rel="icon" type="image/png" href="{{ url_for('static', filename='nibr.ico') }}">
{% for script in extra_scripts %}
<script src="{{ script }}"></script>
{% endfor %}
<link rel="stylesheet" href="https://stackpath.bootstrapcdn.com/bootstrap/4.1.3/css/bootstrap.min.css" integrity="sha384-MCw98/SFnGE8fJT3GXwEOngsV7Zt27NXFoaoApmYm81iuXoPkFOJwJ8ERdknLPMO" crossorigin="anonymous">
{% for script in extra_scripts %}
<script src="{{ script }}"></script>
{% endfor %}
<link rel="stylesheet" href="https://stackpath.bootstrapcdn.com/bootstrap/4.1.3/css/bootstrap.min.css"
integrity="sha384-MCw98/SFnGE8fJT3GXwEOngsV7Zt27NXFoaoApmYm81iuXoPkFOJwJ8ERdknLPMO" crossorigin="anonymous">
</head>
<body>
<header class="navbar navbar-expand navbar-dark flex-column flex-md-row bd-navbar">
<h3>Cellxgene Gateway - Cache Status</h3>
<h3>Cellxgene Gateway - Cache Status</h3>
</header>
<br>
<table class="table">
<thead>
<tr>
<th>PID</th>
<th>dataset</th>
<th>annotation_file</th>
<th>port</th>
<th>launchtime</th>
<th>last access</th>
<th>status</th>
<th>message</th>
<th>http_status</th>
<th>actions</th>
</tr>
</thead>
<tbody>
{% for entry in entry_list %}
<tr>
<td>{{ entry.pid }}</td>
<td><a href="{{ url_for('do_view', path=entry.key.pathpart) }}">{{ entry.key.dataset }}</a></td>
<td>{{ entry.key.annotation_file }}</td>
<td>{{ entry.port }}</td>
<td class="timestamp">{{ entry.launchtime }}</td>
<td class="timestamp">{{ entry.timestamp }}</td>
<td>{{ entry.status }}</td>
<td>{{ entry.message }}</td>
<td>{{ entry.http_status }}</td>
<td>
{% if entry.status == 'loaded' %}
<a href="{{ url_for('do_terminate', path=entry.key.pathpart) }}"> terminate </a>
{% endif %}
</td>
</tr>
{% endfor %}
</tbody>
<tr>
<th>PID</th>
<th>dataset</th>
<th>annotation_file</th>
<th>source</th>
<th>port</th>
<th>launchtime</th>
<th>last access</th>
<th>status</th>
<th>message</th>
<th>http_status</th>
<th>actions</th>
</tr>
</thead>
<tbody>
{% for entry in entry_list %}
<tr>
<td>{{ entry.pid }}</td>
<td><a
href="{{ entry.key.view_url }}">{{ entry.key.h5ad_item.descriptor }}</a>
</td>
<td>{{ entry.key.annotation_descriptor }}</td>
<td>{{ entry.source_name }}</td>
<td>{{ entry.port }}</td>
<td class="timestamp">{{ entry.launchtime }}</td>
<td class="timestamp">{{ entry.timestamp }}</td>
<td>{{ entry.status.name }}</td>
<td>{{ entry.message }}</td>
<td>{{ entry.http_status }}</td>
<td>
{% if entry.status.name == 'loaded' %}
<a
href="{{ url_for('do_terminate', path=entry.key.descriptor, source_name=entry.key.source_name) }}">
terminate </a>
{% endif %}
</td>
</tr>
{% endfor %}
</tbody>
</table>
<script>
$(() => {
$(".timestamp").each(function(){
$(".timestamp").each(function () {
const el = $(this);
const ts = el.text();
const dt = new Date(parseInt(ts * 1000));
el.html(`${dt.toISOString()}<br>(${ts})`);
el.prepend(`${dt.toISOString()}<br>(`);
el.append(')');
});
})
</script>
</body>
</html>
</html>

View File

@@ -28,11 +28,11 @@
<h4>{{ message }}</h4>
<a href="/filecrawl.html">
<a href="{{ url_for('filecrawl') }}">
Please click here to be redirected to the file directory.
</a>
<br>
<a href="/">
<a href="{{ url_for('index') }}">
Please click here to return to the homepage.
</a>
</div>

View File

@@ -36,10 +36,10 @@
Navigation:
<ul>
{% if path %}
<li><a href="/filecrawl.html">top level</a></li>
<li><a href="{{ url_for('filecrawl') }}">top level</a></li>
{% else %}
{% endif %}
<li><a href="/">homepage</a></li>
<li><a href="{{ url_for('index') }}">homepage</a></li>
</ul>
</p>
<script>

View File

@@ -35,56 +35,15 @@
Links:
</h1>
<div class="list-group" style="width:50%;padding-left:65px">
<a href="/filecrawl.html" class="list-group-item list-group-item-action">
<a href="{{ url_for('filecrawl') }}" class="list-group-item list-group-item-action">
<u>File Crawler: Allows you to view all uploaded data.</u></a>
</div>
<div class="list-group" style="width:50%;padding-left:65px">
<a href="/cache_status" class="list-group-item list-group-item-action">
<a href="{{ url_for('do_GET_status') }}" class="list-group-item list-group-item-action">
<u>Cache Status: view status of launched cellxgene servers.</u></a>
</div>
{% if enable_upload %}
<br>
<h1 style="padding-left:35px">
How To Upload Data via HTTP:
</h1>
<ol style="padding-left:85px;">
<li>
Create a folder for your Username:
</li>
<br>
<form action="{{ url_for('make_user') }}" method="post">
Username <input type="text" name="directory">
<input type="submit" value="Create">
</form>
<li>
Create a subdirectory under the selected Folder:
</li>
<br>
<form action="{{ url_for('make_subdir') }}" method="post">
<select name="usernames" id="usernames">
{% for user in users %}
<option value="{{ user }}">{{ user }}</option>
{% endfor %}
</select>
<br>
Subdirectory Name <input type="text" name="directory">
<input type="submit" value="Create">
</form>
<li>Choose a folder to copy your data to, then upload your data file (must be in .h5ad format).</li>
<br>
<form action="{{ url_for('upload_file') }}" method="post" enctype="multipart/form-data">
Type in the name of the directory and subdirectory you wish to upload to, i.e. "USER/cells". <input type="text" name="path">
<br>
File: <input type="file" name="file"><br>
<input style="position:relative; top:10px;" type="submit" value="Upload">
</form>
<br>
<li>Take a look at your data using the file crawler link above</li>
</ol>
{% endif %}
<br>
<h1 style="padding-left:35px">

View File

@@ -35,11 +35,11 @@
The page will refresh shortly.
</p>
<a href="/filecrawl.html">
<a href="{{ url_for('filecrawl') }}">
Please click here to be redirected to the file directory.
</a>
<br>
<a href="/">
<a href="{{ url_for('index') }}">
Please click here to return to the homepage.
</a>
</div>

View File

@@ -36,13 +36,13 @@
<h4>Options</h4>
Please choose one of the following, or use the back button:
<ul>
<li><a href="{{url_for('do_relaunch', path=dataset)}}">
<li><a href="{{ relaunch_url }}">
Attempt to relaunch the cellxgene server.
</a></li>
<li><a href="/filecrawl.html">
<li><a href="{{ url_for('filecrawl') }}">
Return to the file directory.
</a></li>
<li><a href="/">
<li><a href="{{ url_for('index') }}">
Return to the homepage.
</a></li>
</ul>

3
env_example Normal file
View File

@@ -0,0 +1,3 @@
export CELLXGENE_LOCATION=$(pwd)/.venv/bin/cellxgene
export CELLXGENE_DATA=../cellxgene_data
export GATEWAY_IP=127.0.0.1

View File

@@ -1,11 +0,0 @@
name: cellxgene-dev
channels:
- conda-forge
dependencies:
- python=3.7
- requests
- flask
- psutil
- pip:
- flask-api
- cellxgene==0.14.1

17
environment.yml Normal file
View File

@@ -0,0 +1,17 @@
name: cellxgene-gateway
channels:
- conda-forge
dependencies:
- python=3.11
- requests
- flask
- psutil
- black
- twine
- isort
- coverage
- pip
- pip:
- pre_commit
- werkzeug
- cellxgene

View File

@@ -0,0 +1,14 @@
FROM python:3.9
RUN pip install cellxgene-gateway 'MarkupSafe<2.1'
COPY customize_ui.sh customize_ui.sh
RUN CELLXGENE_GATEWAY_DIR=/usr/local/lib/python3.9/site-packages/cellxgene_gateway . ./customize_ui.sh
ENV CELLXGENE_DATA=/cellxgene-data
ENV CELLXGENE_LOCATION=/usr/local/bin/cellxgene
EXPOSE 5005
RUN mkdir /cellxgene-data
CMD ["cellxgene-gateway"]

View File

@@ -0,0 +1,14 @@
# Purpose
This is a simple example of how to make a small script to customize the UI of cellxgene-gateway. The script that does the customization is `customize_ui.sh`, it simply makes the main header green using CSS but you could do anything you want there (including adding more script tags, etc).
# Usage
```
docker build -t cellxgene_custom .
CELLXGENE_DATA=`pwd`/../../../cellxgene_data
docker run -p 5005:5005 --mount src=$CELLXGENE_DATA,target=/cellxgene-data,type=bind cellxgene_custom
```
If you now open http://localhost:5005 you should see a green cellxgene gateway header.

View File

@@ -0,0 +1,3 @@
# make the header bright green
find "${CELLXGENE_GATEWAY_DIR}/templates" -name index.html -exec sed -i -e 's/<head>/<head>\
> <style> header h3 {color: #0F0;} <\/style>/g' {} \;

1
requirements-test.txt Normal file
View File

@@ -0,0 +1 @@
freezegun~=1.5.5

View File

@@ -1,5 +1,5 @@
cellxgene==0.14.1
cellxgene
flask
flask_api
werkzeug
psutil
requests

View File

@@ -1,6 +0,0 @@
export CELLXGENE_LOCATION=$(pwd)/.cellxgene-gateway/bin/cellxgene
export CELLXGENE_DATA=../cellxgene_data
export GATEWAY_IP=127.0.0.1
#Once these are set, you run like a normal Flask app
cellxgene-gateway

View File

@@ -1,13 +1,36 @@
import codecs
import os
from setuptools import setup
import sys
from setuptools import find_packages, setup
if sys.version_info < (3, 6):
sys.exit("Sorry, Python < 3.6 is not supported")
def read(rel_path):
here = os.path.abspath(os.path.dirname(__file__))
with codecs.open(os.path.join(here, rel_path), "r") as fp:
return fp.read()
def get_version(rel_path):
for line in read(rel_path).splitlines():
if line.startswith("__version__"):
delim = '"' if '"' in line else "'"
return line.split(delim)[1]
else:
raise RuntimeError("Unable to find version string.")
def parse_requirements():
reqs = []
with open("requirements.txt", "r") as f:
for l in f.readlines():
reqs.append(l.strip("\n"))
for line in f.readlines():
reqs.append(line.strip("\n"))
return reqs
with open("README.md", "r") as fh:
long_description = fh.read()
@@ -15,9 +38,9 @@ install_reqs = parse_requirements()
setup(
# mandatory
name="cellxgene-gateway",
name="cellxgene_gateway",
# mandatory
version="0.1.0",
version=get_version("cellxgene_gateway/__init__.py"),
# mandatory
author="Niket Patel, Yohann Potier, Alok Saldanha",
author_email="alok.saldanha@novartis.com",
@@ -27,18 +50,20 @@ setup(
license="MIT",
keywords="visualization, genomics",
url="http://github.com/Novartis/cellxgene-gateway",
packages=["cellxgene_gateway"],
packages=find_packages(),
package_data={
"cellxgene_gateway": [
"static/css/homepagestyle.css",
"static/js/annotation.js",
"static/nibr.ico",
"templates/*.html"
]},
data_files=[('', ['README.md', 'LICENSE'])],
"templates/*.html",
]
},
data_files=[("", ["README.md", "LICENSE"])],
install_requires=install_reqs,
entry_points={
"console_scripts": ["cellxgene-gateway=cellxgene_gateway.gateway:main"]
},
classifiers=["Topic :: Scientific/Engineering :: Visualization"],
python_requires='>=3.6',
python_requires=">=3.6",
)

41
start_flask.sh Executable file
View File

@@ -0,0 +1,41 @@
#!/bin/bash
# start_gunicorn.sh - Start Cellxgene Gateway with Gunicorn
#
# PREREQUISITES:
# - Gunicorn installed (included with cellxgene 1.3.0, or: pip install gunicorn)
# - Virtual environment activated
# - .env file with CELLXGENE_LOCATION and CELLXGENE_DATA (or CELLXGENE_BUCKET)
#
# USAGE:
# ./start_gunicorn.sh
# Exit on error
set -e
# Get the directory where this script is located
SCRIPT_DIR="$( cd "$( dirname "${BASH_SOURCE[0]}" )" && pwd )"
# Source environment variables
echo "Loading environment variables..."
if [ -f "$SCRIPT_DIR/.env" ]; then
source "$SCRIPT_DIR/.env"
else
echo "Error: .env file not found at $SCRIPT_DIR/.env"
echo "Please create it with required environment variables"
exit 1
fi
# Verify required environment variables
if [ -z "$CELLXGENE_LOCATION" ]; then
echo "Error: CELLXGENE_LOCATION not set"
exit 1
fi
if [ -z "$CELLXGENE_DATA" ] && [ -z "$CELLXGENE_BUCKET" ]; then
echo "Error: Either CELLXGENE_DATA or CELLXGENE_BUCKET must be set"
exit 1
fi
exec cellxgene-gateway

91
start_gunicorn.sh Executable file
View File

@@ -0,0 +1,91 @@
#!/bin/bash
# start_gunicorn.sh - Start Cellxgene Gateway with Gunicorn
#
# PREREQUISITES:
# - Gunicorn installed (included with cellxgene 1.3.0, or: pip install gunicorn)
# - Virtual environment activated
# - .env file with CELLXGENE_LOCATION and CELLXGENE_DATA (or CELLXGENE_BUCKET)
#
# USAGE:
# ./start_gunicorn.sh
# Exit on error
set -e
# Get the directory where this script is located
SCRIPT_DIR="$( cd "$( dirname "${BASH_SOURCE[0]}" )" && pwd )"
# Source environment variables
echo "Loading environment variables..."
if [ -f "$SCRIPT_DIR/.env" ]; then
source "$SCRIPT_DIR/.env"
else
echo "Error: .env file not found at $SCRIPT_DIR/.env"
echo "Please create it with required environment variables"
exit 1
fi
# Verify required environment variables
if [ -z "$CELLXGENE_LOCATION" ]; then
echo "Error: CELLXGENE_LOCATION not set"
exit 1
fi
if [ -z "$CELLXGENE_DATA" ] && [ -z "$CELLXGENE_BUCKET" ]; then
echo "Error: Either CELLXGENE_DATA or CELLXGENE_BUCKET must be set"
exit 1
fi
# Gunicorn configuration
# WARNING: Multi-worker mode has cache synchronization issues (see plans/002-shared-cache-implementation.md)
# Each worker maintains its own in-memory cache, causing 404s for static assets when different
# workers handle requests for the same dataset. Use GUNICORN_WORKERS=1 until shared cache is implemented.
WORKERS=${GUNICORN_WORKERS:-1}
BIND=${GATEWAY_IP:-0.0.0.0}:${GATEWAY_PORT:-5005}
TIMEOUT=${GUNICORN_TIMEOUT:-120}
WORKER_CLASS=${GUNICORN_WORKER_CLASS:-sync}
KEEPALIVE=${GUNICORN_KEEPALIVE:-5}
LOG_LEVEL=${GUNICORN_LOG_LEVEL:-info}
# Production optimization: enable backed mode to reduce memory usage
export GATEWAY_ENABLE_BACKED_MODE=${GATEWAY_ENABLE_BACKED_MODE:-true}
# Check if gunicorn is installed
if ! command -v gunicorn &> /dev/null; then
echo "Error: gunicorn not found. Install with: pip install gunicorn"
exit 1
fi
# Display configuration
echo "Starting Cellxgene Gateway with Gunicorn..."
echo "Configuration:"
echo " Data source: ${CELLXGENE_DATA:-$CELLXGENE_BUCKET}"
echo " Binding to: $BIND"
echo " Workers: $WORKERS"
echo " Worker class: $WORKER_CLASS"
echo " Timeout: ${TIMEOUT}s"
echo " Keepalive: ${KEEPALIVE}s"
echo " Log level: $LOG_LEVEL"
echo " Backed mode: ${GATEWAY_ENABLE_BACKED_MODE}"
echo ""
cd "$SCRIPT_DIR"
# Start Gunicorn with optimized settings
# Additional options you can add via environment variables:
# - GUNICORN_MAX_REQUESTS: Restart worker after N requests (prevents memory leaks)
# - GUNICORN_MAX_REQUESTS_JITTER: Add randomness to max-requests
exec gunicorn cellxgene_gateway.gateway:app \
--workers "$WORKERS" \
--worker-class "$WORKER_CLASS" \
--bind "$BIND" \
--timeout "$TIMEOUT" \
--keep-alive "$KEEPALIVE" \
--access-logfile - \
--error-logfile - \
--log-level "$LOG_LEVEL" \
--preload \
${GUNICORN_MAX_REQUESTS:+--max-requests "$GUNICORN_MAX_REQUESTS"} \
${GUNICORN_MAX_REQUESTS_JITTER:+--max-requests-jitter "$GUNICORN_MAX_REQUESTS_JITTER"} \
"$@"

89
start_uwsgi.sh Executable file
View File

@@ -0,0 +1,89 @@
#!/bin/bash
# start_uwsgi.sh - Start Cellxgene Gateway with uWSGI
#
# PREREQUISITES:
# - uWSGI installed (pip install uwsgi)
# - Virtual environment activated
# - .env file with CELLXGENE_LOCATION and CELLXGENE_DATA (or CELLXGENE_BUCKET)
#
# USAGE:
# ./start_uwsgi.sh
# Exit on error
set -e
# Get the directory where this script is located
SCRIPT_DIR="$( cd "$( dirname "${BASH_SOURCE[0]}" )" && pwd )"
# Source environment variables
echo "Loading environment variables..."
if [ -f "$SCRIPT_DIR/.env" ]; then
source "$SCRIPT_DIR/.env"
else
echo "Error: .env file not found at $SCRIPT_DIR/.env"
echo "Please create it with required environment variables"
exit 1
fi
# Verify required environment variables
if [ -z "$CELLXGENE_LOCATION" ]; then
echo "Error: CELLXGENE_LOCATION not set"
exit 1
fi
if [ -z "$CELLXGENE_DATA" ] && [ -z "$CELLXGENE_BUCKET" ]; then
echo "Error: Either CELLXGENE_DATA or CELLXGENE_BUCKET must be set"
exit 1
fi
# uWSGI configuration
# WARNING: Multi-worker mode has cache synchronization issues (see plans/002-shared-cache-implementation.md)
# Each worker maintains its own in-memory cache, causing 404s for static assets when different
# workers handle requests for the same dataset. Use UWSGI_WORKERS=1 until shared cache is implemented.
WORKERS=${UWSGI_WORKERS:-1}
HOST=${GATEWAY_IP:-0.0.0.0}
PORT=${GATEWAY_PORT:-5005}
TIMEOUT=${UWSGI_TIMEOUT:-120}
THREADS=${UWSGI_THREADS:-1}
# Production optimization: enable backed mode to reduce memory usage
export GATEWAY_ENABLE_BACKED_MODE=${GATEWAY_ENABLE_BACKED_MODE:-true}
# Check if uwsgi is installed
if ! command -v uwsgi &> /dev/null; then
echo "Error: uwsgi not found. Install with: pip install uwsgi"
exit 1
else
# Display configuration
echo "Starting Cellxgene Gateway with uWSGI..."
echo "Configuration:"
echo " Data source: ${CELLXGENE_DATA:-$CELLXGENE_BUCKET}"
echo " Binding to: $HOST:$PORT"
echo " Workers: $WORKERS"
echo " Threads: $THREADS"
echo " Timeout: ${TIMEOUT}s"
echo " Backed mode: ${GATEWAY_ENABLE_BACKED_MODE}"
echo ""
cd "$SCRIPT_DIR"
# Start uWSGI with optimized settings
# Additional options you can add via environment variables:
# - UWSGI_MAX_REQUESTS: Restart worker after N requests (prevents memory leaks)
exec uwsgi \
--http "$HOST:$PORT" \
--module cellxgene_gateway.gateway:app \
--workers "$WORKERS" \
--threads "$THREADS" \
--harakiri "$TIMEOUT" \
--master \
--enable-threads \
--single-interpreter \
--need-app \
--die-on-term \
--log-x-forwarded-for \
${UWSGI_MAX_REQUESTS:+--max-requests "$UWSGI_MAX_REQUESTS"} \
"$@"
fi

0
tests/items/__init__.py Normal file
View File

View File

View File

@@ -0,0 +1,43 @@
import tempfile
import unittest
from unittest.mock import patch
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
def stub_join(path):
path.join = lambda x, y: x + "/" + y
class TestFileItemSource(unittest.TestCase):
@patch("os.path")
@patch("os.listdir")
def test_list_items_GIVEN_no_subpath_THEN_checks_dir(self, listdir, path):
stub_join(path)
source = FileItemSource("/tmp/unittest", "local")
source.list_items()
path.exists.assert_called_once_with("/tmp/unittest/")
@patch("os.path")
@patch("os.listdir")
def test_list_items_GIVEN_subpath_THEN_checks_subpath(self, listdir, path):
stub_join(path)
source = FileItemSource("/tmp/unittest", "local")
source.list_items("foo")
path.exists.assert_called_once_with("/tmp/unittest/foo")
def test_make_fileitem_from_path_GIVEN_annotation_file_THEN_name_lacks_csv(
self,
):
source = FileItemSource(tempfile.gettempdir(), "local")
item = source.make_fileitem_from_path(
"customanno.csv", "someh5ad_annotations", True
)
self.assertEqual(item.name, "customanno")
self.assertEqual(item.descriptor, "someh5ad_annotations/customanno.csv")
def test_make_fileitem_from_path_GIVEN_h5ad_file_THEN_returns_name(self):
source = FileItemSource(tempfile.gettempdir(), "local")
item = source.make_fileitem_from_path("someanalysis.h5ad", "studydir")
self.assertEqual(item.name, "someanalysis.h5ad")
self.assertEqual(item.descriptor, "studydir/someanalysis.h5ad")

View File

View File

@@ -0,0 +1,184 @@
import unittest
import os
import shutil
import tempfile
from unittest.mock import MagicMock, Mock, patch
from cellxgene_gateway.items.item import ItemType
from cellxgene_gateway.items.s3.s3item import S3Item
from cellxgene_gateway.items.s3.s3item_source import S3ItemSource
from cellxgene_gateway.gateway import app
class TestScanDirectory(unittest.TestCase):
def setUp(self):
self.app = app
def tearDown(self):
pass
@patch("s3fs.S3FileSystem")
def test_GIVEN_invalid_bucket_THEN_throws_error(self, s3func):
class S3Mock:
def exists(path):
if path in ["s3://my-bucket/"]:
return False
s3func.return_value = S3Mock
source = S3ItemSource("my-bucket")
with self.assertRaises(Exception) as context:
source.scan_directory()
self.assertEqual(
"S3 url 's3://my-bucket/' does not exist.",
str(context.exception),
)
@patch("s3fs.S3FileSystem")
def test__GIVEN_multilevel_bucket_THEN_properly_recurses_suburls(self, s3func):
class S3Mock:
def exists(path):
if path in [
"s3://my-bucket/",
"s3://my-bucket/pbmc3k.h5ad",
"s3://my-bucket/lvl1",
"s3://my-bucket/lvl1/pbmc3k_l1.h5ad",
"s3://my-bucket/lvl1/lvl2",
"s3://my-bucket/lvl1/lvl2/pbmc3k_l2.h5ad",
]:
return True
raise Exception("exists called with " + path)
def ls(path, refresh):
assert refresh == True
if path == "s3://my-bucket/":
return [
"my-bucket/lvl1",
"my-bucket/pbmc3k.h5ad",
"my-bucket/pbmc3k_annotations",
]
elif path == "s3://my-bucket/pbmc3k_annotations":
return ["my-bucket/pbmc3k_annotations/annot.csv"]
elif path == "s3://my-bucket/lvl1":
return ["my-bucket/lvl1/lvl2", "my-bucket/lvl1/pbmc3k_l1.h5ad"]
elif path == "s3://my-bucket/lvl1/lvl2":
return ["my-bucket/lvl1/lvl2/pbmc3k_l2.h5ad"]
raise Exception("ls called with " + path)
def isdir(path):
if path in [
"s3://my-bucket/lvl1",
"s3://my-bucket/pbmc3k_annotations",
"s3://my-bucket/lvl1/lvl2",
]:
return True
if path in [
"s3://my-bucket/pbmc3k.h5ad",
"s3://my-bucket/lvl1/pbmc3k_l1.h5ad",
"s3://my-bucket/lvl1/pbmc3k_l1_annotations",
"s3://my-bucket/lvl1/lvl2/pbmc3k_l2.h5ad",
"s3://my-bucket/lvl1/lvl2/pbmc3k_l2_annotations",
]:
return False
raise Exception("isdir called with " + path)
def isfile(path):
if path in ["s3://my-bucket/pbmc3k_annotations/annot.csv"]:
return True
if path in ["s3://my-bucket/pbmc3k_annotations"]:
return False
raise Exception("isfile called with " + path)
s3func.return_value = S3Mock
source = S3ItemSource("my-bucket")
with self.app.test_request_context(query_string="refresh=true") as test_context:
tree = source.scan_directory()
def s3item_compare(i1, i2, msg=""):
self.assertEqual(i1.name, i2.name, "name equals")
self.assertEqual(i1.type, i2.type, "type equals")
self.assertEqual(i1.s3key, i2.s3key, "s3key equals")
if i1.annotations is None:
self.assertEqual(i1.annotations, i2.annotations, "annotations equals")
else:
self.assertEqual(
len(i1.annotations),
len(i2.annotations),
"annotations length equals",
)
for a1, a2 in zip(i1.annotations, i2.annotations):
self.assertEqual(a1, a2)
return True
self.addTypeEqualityFunc(S3Item, s3item_compare)
def assertTree(t, descriptor, items):
self.assertEqual(t.descriptor, descriptor)
self.assertEqual(len(t.items), len(items))
for i1, i2 in zip(t.items, items):
self.assertEqual(i1, i2)
assertTree(
tree,
"",
[
S3Item(
"pbmc3k.h5ad",
name="pbmc3k.h5ad",
type=ItemType.h5ad,
annotations=[
S3Item(
"pbmc3k_annotations/annot.csv",
name="annot.csv",
type=ItemType.annotation,
)
],
)
],
)
self.assertEqual(len(tree.branches), 1)
lvl1 = tree.branches[0]
assertTree(
lvl1,
"lvl1",
[
S3Item(
"lvl1/pbmc3k_l1.h5ad",
name="pbmc3k_l1.h5ad",
type=ItemType.h5ad,
annotations=None,
)
],
)
self.assertEqual(len(lvl1.branches), 1)
lvl2 = lvl1.branches[0]
assertTree(
lvl2,
"lvl1/lvl2",
[
S3Item(
"lvl1/lvl2/pbmc3k_l2.h5ad",
name="pbmc3k_l2.h5ad",
type=ItemType.h5ad,
annotations=None,
)
],
)
self.assertEqual(lvl2.branches, None)
class TestListItems(unittest.TestCase):
def test_GIVEN_filter_THEN_pass_filter_into_scan_directory(self):
source = S3ItemSource("my-bucket")
source.scan_directory = MagicMock()
tree = source.list_items("some-filter")
source.scan_directory.assert_called_once_with("some-filter")
def test_GIVEN_no_filter_THEN_pass_empty_string_into_scan_directory(self):
source = S3ItemSource("my-bucket")
source.scan_directory = MagicMock()
tree = source.list_items()
source.scan_directory.assert_called_once_with("")

362
tests/test_backend_cache.py Normal file
View File

@@ -0,0 +1,362 @@
import unittest
from http import HTTPStatus
from unittest.mock import MagicMock, Mock, patch, call
from freezegun import freeze_time
from cellxgene_gateway.backend_cache import BackendCache, is_port_in_use
from cellxgene_gateway.cache_entry import CacheEntry, CacheEntryStatus
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.cellxgene_exception import CellxgeneException
class TestIsPortInUse(unittest.TestCase):
@patch("socket.socket")
def test_GIVEN_free_port_THEN_returns_true(self, socketMock):
connectMock = socketMock()
connectMock.connect_ex.return_value = 0
connectMock.__enter__.return_value = connectMock
self.assertEqual(is_port_in_use(123), True)
self.assertTrue(connectMock.__enter__.calledOnce)
self.assertTrue(connectMock.__exit__.calledOnce)
self.assertTrue(connectMock.connect_ex.calledOnceWith("a"))
self.assertTrue(socketMock.calledOnceWith("a"))
@patch("socket.socket")
def test_GIVEN_used_port_THEN_returns_false(self, socketMock):
connectMock = socketMock()
connectMock.__enter__.return_value = connectMock
connectMock.connect_ex.return_value = 1
self.assertTrue(connectMock.__enter__.calledOnce)
self.assertTrue(connectMock.__exit__.calledOnce)
self.assertTrue(connectMock.connect_ex.calledOnceWith("a"))
self.assertTrue(socketMock.calledOnceWith("a"))
self.assertEqual(is_port_in_use(123), False)
class TestBackendCacheInit(unittest.TestCase):
def test_GIVEN_new_backend_cache_THEN_entry_list_is_empty(self):
"""BackendCache should initialize with an empty entry list."""
cache = BackendCache()
self.assertEqual(cache.entry_list, [])
class TestBackendCacheGetPorts(unittest.TestCase):
def test_GIVEN_empty_cache_THEN_get_ports_returns_empty_list(self):
"""get_ports should return empty list when no entries exist."""
cache = BackendCache()
ports = cache.get_ports()
self.assertEqual(ports, [])
def test_GIVEN_cache_with_entries_THEN_get_ports_returns_all_ports(self):
"""get_ports should return list of all ports from entries."""
cache = BackendCache()
# Create mock entries with different ports
entry1 = Mock()
entry1.port = 8000
entry2 = Mock()
entry2.port = 8001
entry3 = Mock()
entry3.port = 8002
cache.entry_list = [entry1, entry2, entry3]
ports = cache.get_ports()
self.assertEqual(ports, [8000, 8001, 8002])
def test_GIVEN_single_entry_THEN_get_ports_returns_single_port(self):
"""get_ports should correctly handle a single entry."""
cache = BackendCache()
entry = Mock()
entry.port = 9000
cache.entry_list = [entry]
ports = cache.get_ports()
self.assertEqual(ports, [9000])
class TestBackendCacheCheckPath(unittest.TestCase):
def setUp(self):
self.cache = BackendCache()
def test_GIVEN_no_matching_path_THEN_check_path_returns_none(self):
"""check_path should return None when no entries match the path."""
source = Mock()
source.name = "test_source"
cache_entry = Mock()
cache_entry.status = CacheEntryStatus.loaded
cache_entry.key = Mock()
cache_entry.key.source.name = "other_source"
cache_entry.key.descriptor = "/some/path"
self.cache.entry_list = [cache_entry]
result = self.cache.check_path(source, "/test/path")
self.assertIsNone(result)
def test_GIVEN_terminated_entry_THEN_check_path_ignores_it(self):
"""check_path should ignore entries with terminated status."""
source = Mock()
source.name = "test_source"
cache_entry = Mock()
cache_entry.status = CacheEntryStatus.terminated
cache_entry.key = Mock()
cache_entry.key.source.name = "test_source"
cache_entry.key.descriptor = "/data"
self.cache.entry_list = [cache_entry]
result = self.cache.check_path(source, "/data/file.txt")
self.assertIsNone(result)
def test_GIVEN_single_matching_entry_THEN_check_path_returns_it(self):
"""check_path should return the matching entry when exactly one matches."""
source = Mock()
source.name = "test_source"
cache_entry = Mock()
cache_entry.status = CacheEntryStatus.loaded
cache_entry.key = Mock()
cache_entry.key.source.name = "test_source"
cache_entry.key.descriptor = "/data"
self.cache.entry_list = [cache_entry]
result = self.cache.check_path(source, "/data/file.txt")
self.assertIs(result, cache_entry)
def test_GIVEN_path_that_does_not_start_with_descriptor_THEN_check_path_returns_none(
self,
):
"""check_path should return None if path doesn't start with descriptor."""
source = Mock()
source.name = "test_source"
cache_entry = Mock()
cache_entry.status = CacheEntryStatus.loaded
cache_entry.key = Mock()
cache_entry.key.source.name = "test_source"
cache_entry.key.descriptor = "/data"
self.cache.entry_list = [cache_entry]
result = self.cache.check_path(source, "/other/file.txt")
self.assertIsNone(result)
def test_GIVEN_multiple_matching_entries_THEN_check_path_raises_exception(self):
"""check_path should raise exception when multiple entries match."""
source = Mock()
source.name = "test_source"
cache_entry1 = Mock()
cache_entry1.status = CacheEntryStatus.loaded
cache_entry1.key = Mock()
cache_entry1.key.source.name = "test_source"
cache_entry1.key.descriptor = "/data"
cache_entry2 = Mock()
cache_entry2.status = CacheEntryStatus.loaded
cache_entry2.key = Mock()
cache_entry2.key.source.name = "test_source"
cache_entry2.key.descriptor = "/data"
self.cache.entry_list = [cache_entry1, cache_entry2]
with self.assertRaises(CellxgeneException) as context:
self.cache.check_path(source, "/data/file.txt")
# The CellxgeneException is raised with HTTPStatus as first arg and message as second
self.assertEqual(context.exception.message, HTTPStatus.INTERNAL_SERVER_ERROR)
self.assertIn("Found 2", context.exception.http_status)
def test_GIVEN_mixed_entries_THEN_check_path_returns_only_matching_active_entry(
self,
):
"""check_path should correctly filter by source, path, and status."""
source = Mock()
source.name = "target_source"
# Terminated entry - should be ignored
terminated_entry = Mock()
terminated_entry.status = CacheEntryStatus.terminated
terminated_entry.key = Mock()
terminated_entry.key.source.name = "target_source"
terminated_entry.key.descriptor = "/data"
# Different source - should be ignored
other_source_entry = Mock()
other_source_entry.status = CacheEntryStatus.loaded
other_source_entry.key = Mock()
other_source_entry.key.source.name = "other_source"
other_source_entry.key.descriptor = "/data"
# Matching entry - should be returned
matching_entry = Mock()
matching_entry.status = CacheEntryStatus.loaded
matching_entry.key = Mock()
matching_entry.key.source.name = "target_source"
matching_entry.key.descriptor = "/data"
self.cache.entry_list = [terminated_entry, other_source_entry, matching_entry]
result = self.cache.check_path(source, "/data/file.txt")
self.assertIs(result, matching_entry)
class TestBackendCacheCheckEntry(unittest.TestCase):
def setUp(self):
self.cache = BackendCache()
def test_GIVEN_no_matching_entry_THEN_check_entry_returns_none(self):
"""check_entry should return None when no entries match the key."""
key = Mock()
cache_entry = Mock()
cache_entry.status = CacheEntryStatus.loaded
cache_entry.key = Mock()
cache_entry.key.equals.return_value = False
self.cache.entry_list = [cache_entry]
result = self.cache.check_entry(key)
self.assertIsNone(result)
def test_GIVEN_terminated_entry_THEN_check_entry_ignores_it(self):
"""check_entry should ignore entries with terminated status."""
key = Mock()
cache_entry = Mock()
cache_entry.status = CacheEntryStatus.terminated
cache_entry.key = Mock()
cache_entry.key.equals.return_value = True
self.cache.entry_list = [cache_entry]
result = self.cache.check_entry(key)
self.assertIsNone(result)
def test_GIVEN_single_matching_entry_THEN_check_entry_returns_it(self):
"""check_entry should return the matching entry when exactly one matches."""
key = Mock()
cache_entry = Mock()
cache_entry.status = CacheEntryStatus.loaded
cache_entry.key = Mock()
cache_entry.key.equals.return_value = True
self.cache.entry_list = [cache_entry]
result = self.cache.check_entry(key)
self.assertIs(result, cache_entry)
def test_GIVEN_multiple_matching_entries_THEN_check_entry_raises_exception(self):
"""check_entry should raise exception when multiple entries match."""
key = Mock()
key.dataset = "test_dataset"
cache_entry1 = Mock()
cache_entry1.status = CacheEntryStatus.loaded
cache_entry1.key = Mock()
cache_entry1.key.equals.return_value = True
cache_entry2 = Mock()
cache_entry2.status = CacheEntryStatus.loaded
cache_entry2.key = Mock()
cache_entry2.key.equals.return_value = True
self.cache.entry_list = [cache_entry1, cache_entry2]
with self.assertRaises(CellxgeneException) as context:
self.cache.check_entry(key)
# The CellxgeneException is raised with HTTPStatus as first arg and message as second
self.assertEqual(context.exception.message, HTTPStatus.INTERNAL_SERVER_ERROR)
self.assertIn("Found 2", context.exception.http_status)
def test_GIVEN_mixed_entries_THEN_check_entry_returns_only_matching_active_entry(
self,
):
"""check_entry should correctly filter by key equality and status."""
key = Mock()
# Terminated entry - should be ignored
terminated_entry = Mock()
terminated_entry.status = CacheEntryStatus.terminated
terminated_entry.key = Mock()
terminated_entry.key.equals.return_value = True
# Non-matching entry - should be ignored
non_matching_entry = Mock()
non_matching_entry.status = CacheEntryStatus.loaded
non_matching_entry.key = Mock()
non_matching_entry.key.equals.return_value = False
# Matching entry - should be returned
matching_entry = Mock()
matching_entry.status = CacheEntryStatus.loaded
matching_entry.key = Mock()
matching_entry.key.equals.return_value = True
self.cache.entry_list = [terminated_entry, non_matching_entry, matching_entry]
result = self.cache.check_entry(key)
self.assertIs(result, matching_entry)
class TestBackendCachePrune(unittest.TestCase):
def test_GIVEN_entry_in_cache_THEN_prune_removes_it(self):
"""prune should remove the entry from the cache."""
cache = BackendCache()
entry_mock = Mock()
cache.entry_list = [entry_mock]
cache.prune(entry_mock)
self.assertEqual(len(cache.entry_list), 0)
self.assertNotIn(entry_mock, cache.entry_list)
def test_GIVEN_entry_in_cache_THEN_prune_terminates_it(self):
"""prune should call terminate on the entry."""
cache = BackendCache()
entry_mock = Mock()
cache.entry_list = [entry_mock]
cache.prune(entry_mock)
entry_mock.terminate.assert_called_once()
def test_GIVEN_multiple_entries_THEN_prune_removes_only_target_entry(self):
"""prune should only remove the specified entry, not others."""
cache = BackendCache()
entry1 = Mock()
entry2 = Mock()
entry3 = Mock()
cache.entry_list = [entry1, entry2, entry3]
cache.prune(entry2)
self.assertEqual(len(cache.entry_list), 2)
self.assertIn(entry1, cache.entry_list)
self.assertNotIn(entry2, cache.entry_list)
self.assertIn(entry3, cache.entry_list)
# Verify only entry2 was terminated
entry1.terminate.assert_not_called()
entry2.terminate.assert_called_once()
entry3.terminate.assert_not_called()
def test_GIVEN_empty_cache_THEN_prune_raises_value_error(self):
"""prune should raise ValueError if entry is not in cache."""
cache = BackendCache()
entry_mock = Mock()
with self.assertRaises(ValueError):
cache.prune(entry_mock)

69
tests/test_cache_entry.py Normal file
View File

@@ -0,0 +1,69 @@
import unittest
import tempfile
import os
import shutil
from flask import Flask
from cellxgene_gateway import flask_util
from cellxgene_gateway.cache_entry import CacheEntry, CacheEntryStatus
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.items.item import ItemType
from cellxgene_gateway.items.file.fileitem import FileItem
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
from cellxgene_gateway.gateway import app
key = CacheKey(
FileItem("/czi/", name="pbmc3k.h5ad", type=ItemType.h5ad),
FileItemSource("/tmp", "local"),
)
class TestRenderEntry(unittest.TestCase):
def setUp(self):
self.app = app
self.app_context = self.app.test_request_context()
self.app_context.push()
self.client = self.app.test_client()
def tearDown(self):
self.app_context.pop()
def test_GIVEN_key_and_port_THEN_returns_loading_CacheEntry(self):
entry = CacheEntry.for_key("some-key", 1)
self.assertEqual(entry.status, CacheEntryStatus.loading)
def test_GIVEN_absolute_static_url_THEN_include_path(self):
flask_util.include_source_in_url = False
actual = CacheEntry.for_key(key, 8000).rewrite_text_content(
"src:url(/static/assets/"
)
expected = "src:url(/view/czi/pbmc3k.h5ad/static/assets/"
self.assertEqual(actual, expected)
def test_GIVEN_absolute_src_THEN_include_path(self):
flask_util.include_source_in_url = False
actual = CacheEntry.for_key(key, 8000).rewrite_text_content(
'<link rel="shortcut icon" href="/static/assets/favicon.ico">'
)
expected = '<link rel="shortcut icon" href="/view/czi/pbmc3k.h5ad/static/assets/favicon.ico">'
self.assertEqual(actual, expected)
def test_GIVEN_absolute_static_url_include_source_THEN_include_path(self):
flask_util.include_source_in_url = True
actual = CacheEntry.for_key(key, 8000).rewrite_text_content(
"src:url(/static/assets/"
)
expected = "src:url(/source/local/view/czi/pbmc3k.h5ad/static/assets/"
self.assertEqual(actual, expected)
def test_GIVEN_absolute_src_include_source_THEN_include_path(self):
flask_util.include_source_in_url = True
actual = CacheEntry.for_key(key, 8000).rewrite_text_content(
'<link rel="shortcut icon" href="/static/assets/favicon.ico">'
)
expected = '<link rel="shortcut icon" href="/source/local/view/czi/pbmc3k.h5ad/static/assets/favicon.ico">'
self.assertEqual(actual, expected)
if __name__ == "__main__":
unittest.main()

View File

@@ -1,38 +1,35 @@
import unittest
from unittest.mock import MagicMock, patch
from cellxgene_gateway.dir_util import render_entry
class TestRenderEntry(unittest.TestCase):
def test_GIVEN_path_both_slash_THEN_view_has_single_slash(self):
entry = {
"path": "/somepath/",
"name": "entry",
"type": "file",
}
rendered = render_entry(entry)
self.assertIn('view/somepath', rendered)
def test_GIVEN_path_starts_slash_THEN_view_has_single_slash(self):
entry = {
"path": "/somepath",
"name": "entry",
"type": "file",
}
rendered = render_entry(entry)
self.assertIn('view/somepath', rendered)
def test_GIVEN_path_ends_slash_THEN_view_has_single_slash(self):
entry = {
"path": "somepath/",
"name": "entry",
"type": "file",
}
rendered = render_entry(entry)
self.assertIn('view/somepath', rendered)
def test_GIVEN_path_no_slash_THEN_view_has_single_slash(self):
entry = {
"path": "somepath",
"name": "entry",
"type": "file",
}
rendered = render_entry(entry)
self.assertIn('view/somepath', rendered)
from cellxgene_gateway.dir_util import ensure_dir_exists, make_annotations, make_h5ad
class TestMakeH5ad(unittest.TestCase):
def test_GIVEN_annotation_dir_THEN_returns_h5ad(self):
self.assertEqual(make_h5ad("pbmc_annotations"), "pbmc.h5ad")
class TestMakeAnnotations(unittest.TestCase):
def test_GIVEN_h5ad_THEN_returns_annotations(self):
self.assertEqual(make_annotations("pbmc.h5ad"), "pbmc_annotations")
class TestMakeAnnotations(unittest.TestCase):
def test_GIVEN_h5ad_THEN_returns_annotations(self):
self.assertEqual(make_annotations("pbmc.h5ad"), "pbmc_annotations")
class TestEnsureDirExists(unittest.TestCase):
@patch("os.path.exists")
@patch("os.makedirs")
def test_GIVEN_existing_THEN_does_not_call_makedir(self, makedirsMock, existsMock):
existsMock.return_value = True
ensure_dir_exists("/foo")
makedirsMock.assert_not_called()
@patch("os.path.exists")
@patch("os.makedirs")
def test_GIVEN_not_existing_THEN_calls_makedir(self, makedirsMock, existsMock):
existsMock.return_value = False
ensure_dir_exists("/foo")
makedirsMock.assert_called_once_with("/foo")

View File

@@ -1,23 +1,35 @@
import unittest
from unittest.mock import MagicMock, patch
from cellxgene_gateway.extra_scripts import get_extra_scripts
class TestExtraScripts(unittest.TestCase):
@patch('cellxgene_gateway.env.extra_scripts', new='["abc","def"]')
@patch("cellxgene_gateway.env.extra_scripts", new='["abc","def"]')
def test_GIVEN_two_scripts_THEN_returns_two_strings(self):
self.assertEqual(get_extra_scripts(), ['abc', 'def'])
self.assertEqual(get_extra_scripts(), ["abc", "def"])
@patch('cellxgene_gateway.env.extra_scripts', new='["abc", "def"]')
@patch("cellxgene_gateway.env.extra_scripts", new='["abc", "def"]')
def test_GIVEN_two_scripts_space_THEN_returns_two_strings(self):
self.assertEqual(get_extra_scripts(), ['abc', 'def'])
self.assertEqual(get_extra_scripts(), ["abc", "def"])
@patch('cellxgene_gateway.env.extra_scripts', new=None)
@patch("cellxgene_gateway.env.extra_scripts", new=None)
def test_GIVEN_none_THEN_returns_empty_array(self):
self.assertEqual(get_extra_scripts(), [])
@patch('cellxgene_gateway.env.extra_scripts', new='[]')
@patch("cellxgene_gateway.env.extra_scripts", new="[]")
def test_GIVEN_empty_string_THEN_returns_empty_array(self):
self.assertEqual(get_extra_scripts(), [])
if __name__ == '__main__':
unittest.main()
@patch("cellxgene_gateway.env.extra_scripts", new="'asdf'")
def test_GIVEN_bare_string_THEN_throws_Exception(self):
with self.assertRaises(Exception) as context:
self.assertEqual(get_extra_scripts(), [])
self.assertEqual(
'Error parsing GATEWAY_EXTRA_SCRIPTS, expected JSON array e.g. ["https://example.com/path/to/script.js"]',
str(context.exception),
)
if __name__ == "__main__":
unittest.main()

170
tests/test_filecrawl.py Normal file
View File

@@ -0,0 +1,170 @@
import os
import shutil
import tempfile
import unittest
from collections import defaultdict
from unittest.mock import patch
from cellxgene_gateway.filecrawl import (
render_item,
render_item_source,
render_item_tree,
)
from cellxgene_gateway.items.file.fileitem import FileItem
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
from cellxgene_gateway.items.item import ItemTree, ItemType
from cellxgene_gateway.gateway import app
source = FileItemSource("/tmp")
def make_entry(subpath="somepath", annotations=None):
return FileItem(
subpath=subpath,
name="entry",
ext=".h5ad",
type=ItemType.h5ad,
annotations=annotations,
)
class TestRenderEntry(unittest.TestCase):
def setUp(self):
self.app = app
self.app_context = self.app.test_request_context()
self.app_context.push()
def tearDown(self):
self.app_context.pop()
def test_GIVEN_path_both_slash_THEN_view_has_single_slash(self):
entry = make_entry(subpath="/somepath/")
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry.h5ad/'", rendered)
def test_GIVEN_path_starts_slash_THEN_view_has_single_slash(self):
entry = make_entry(subpath="/somepath")
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry.h5ad/'", rendered)
def test_GIVEN_path_ends_slash_THEN_view_has_single_slash(self):
entry = make_entry(subpath="somepath/")
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry.h5ad/'", rendered)
def test_GIVEN_path_no_slash_THEN_view_has_single_slash(self):
entry = make_entry(subpath="somepath")
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry.h5ad/'", rendered)
class TestRenderAnnotation(unittest.TestCase):
def setUp(self):
self.app = app
self.app_context = self.app.test_request_context()
self.app_context.push()
def tearDown(self):
self.app_context.pop()
@patch("cellxgene_gateway.filecrawl.enable_annotations", new=True)
def test_GIVEN_no_annotation_THEN_new_alone(self):
entry = make_entry(annotations=None)
rendered = render_item(entry, source)
self.assertIn(
"> | annotations: <a class='new' href='/source/Files:/tmp/view/somepath/entry_annotations'>new</a></li>",
rendered,
)
@patch("cellxgene_gateway.filecrawl.enable_annotations", new=True)
def test_GIVEN_annotation_THEN_new_before(self):
annotation = FileItem(
subpath="somepath/entry_annotations",
name="annot",
ext=".csv",
type=ItemType.annotation,
)
entry = make_entry(annotations=[annotation])
rendered = render_item(entry, source)
self.assertIn(
"> | annotations: <a class='new' href='/source/Files:/tmp/view/somepath/entry_annotations'>new</a>,"
" <a href='/source/Files:/tmp/view/somepath/entry_annotations/annot.csv/'>annot</a></li>",
rendered,
)
@patch("cellxgene_gateway.filecrawl.enable_annotations", new=True)
def test_GIVEN_annotation_THEN_escaped(self):
annotation = FileItem(
subpath="somepath/entry_annotations",
name="hot&cold",
ext=".csv",
type=ItemType.annotation,
)
entry = make_entry(annotations=[annotation])
rendered = render_item(entry, source)
self.assertIn(
"> | annotations: <a class='new' href='/source/Files:/tmp/view/somepath/entry_annotations'>new</a>,"
" <a href='/source/Files:/tmp/view/somepath/entry_annotations/hot&cold.csv/'>hot&amp;cold</a></li>",
rendered,
)
class TestRenderItemSource(unittest.TestCase):
@patch("cellxgene_gateway.items.file.fileitem_source.FileItemSource")
def test_GIVEN_some_filter_THEN_includes_filterpart_in_heading(self, item_source):
item_source.name = "FakeSource"
item_source.list_items.return_value = ItemTree("rootdir", [], [])
rendered = render_item_source(item_source, "some_filter")
self.assertEqual(
rendered,
"<h6><a href='/filecrawl.html?source=FakeSource'>FakeSource</a>:some_filter</h6><li><a href='/filecrawl/rootdir?source=FakeSource'>rootdir</a><ul></ul></li>",
)
class TestRenderItemTree(unittest.TestCase):
def setUp(self):
self.app = app
self.app_context = self.app.test_request_context()
self.app_context.push()
def tearDown(self):
self.app_context.pop()
@patch("cellxgene_gateway.items.file.fileitem_source.FileItemSource")
def test_GIVEN_deep_nested_dirs_THEN_includes_dirs_in_output(self, item_source):
item_source.name = "FakeSource"
item_source.get_annotations_subpath = lambda _: "FakeAnnotations"
file_item = FileItem(
subpath="foo/bar/baz", name="file.h5ad", type=ItemType.h5ad
)
item_tree = ItemTree("foo/bar/baz", [file_item], [])
rendered = render_item_tree(item_tree, item_source)
self.assertEqual(
rendered,
"<li><a href='/filecrawl/foo/bar/baz?source=FakeSource'>baz</a><ul>"
"<li> <a href='/source/FakeSource/view/foo/bar/baz/file.h5ad/'>file.h5ad</a>"
" </li></ul></li>",
)
@patch(
"os.listdir",
side_effect=lambda parent: defaultdict(
list, {"tmp": ["foo"], "tmp/foo": ["bar"]}
)[parent],
)
@patch("os.path.exists", return_value=True)
def test_GIVEN_dirs_without_h5ad_THEN_excludes_dirs_in_output(
self, listdir, exists
):
# Directories:
# - tmp
# - foo
# - bar (no h5ad files)
item_source = FileItemSource("tmp", name="local")
item_tree = item_source.list_items("foo")
rendered = render_item_tree(item_tree, item_source)
self.assertEqual(
rendered,
"<li><a href='/filecrawl/foo?source=local'>foo</a><ul></ul></li>",
)

View File

@@ -0,0 +1,54 @@
import json
import unittest
from types import SimpleNamespace
from cellxgene_gateway.gateway import do_GET_status_json, app, cache
from cellxgene_gateway.cache_entry import CacheEntry, CacheEntryStatus
class TestGatewayStatusJson(unittest.TestCase):
def test_do_GET_status_json_returns_expected_structure(self):
# Create a minimal fake key with required attributes
h5ad_item = SimpleNamespace(descriptor="somedir/dataset.h5ad")
key = SimpleNamespace(
h5ad_item=h5ad_item,
annotation_descriptor="somedir/dataset_annotations/foo.csv",
)
# Create a CacheEntry with known launchtime/timestamp/status
entry = CacheEntry(
None,
key,
8000,
111,
222,
CacheEntryStatus.loaded,
None,
None,
None,
None,
)
# Install into the gateway cache and set app launchtime
cache.entry_list = [entry]
app.extensions.setdefault("cellxgene_gateway", {})["launchtime"] = "LAUNCH_TIME"
rv = do_GET_status_json()
data = json.loads(rv)
# top-level launchtime comes from app.extensions
self.assertEqual("LAUNCH_TIME", data["launchtime"])
self.assertIn("entry_list", data)
self.assertEqual(1, len(data["entry_list"]))
e = data["entry_list"][0]
self.assertEqual("somedir/dataset.h5ad", e["dataset"])
self.assertEqual("somedir/dataset_annotations/foo.csv", e["annotation_file"])
self.assertEqual("loaded", e["status"])
self.assertEqual(111, e["launchtime"])
self.assertEqual(222, e["last_access"])
if __name__ == "__main__":
unittest.main()

View File

@@ -1,26 +1,46 @@
import unittest
from unittest.mock import MagicMock, patch
from cellxgene_gateway.cache_entry import CacheEntry
from unittest.mock import patch, seal
from cellxgene_gateway.backend_cache import BackendCache
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.items.file.fileitem import FileItem
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
from cellxgene_gateway.items.item import ItemType
key = CacheKey(
FileItem("/czi/", name="pbmc3k.h5ad", type=ItemType.h5ad),
FileItemSource("/tmp", "local"),
)
class TestPruneProcessCache(unittest.TestCase):
@patch('cellxgene_gateway.util.current_time_stamp', new=lambda:0)
@patch('cellxgene_gateway.env.ttl', new='10')
@patch('cellxgene_gateway.cache_entry.CacheEntry')
@patch('cellxgene_gateway.cache_entry.CacheEntry')
@patch("cellxgene_gateway.util.current_time_stamp", new=lambda: 0)
@patch("cellxgene_gateway.env.expire_seconds", new=10)
@patch("cellxgene_gateway.cache_entry.CacheEntry")
@patch("cellxgene_gateway.cache_entry.CacheEntry")
def test_GIVEN_one_old_one_new_THEN_prune_old(self, old, new):
from cellxgene_gateway.prune_process_cache import PruneProcessCache
cache = BackendCache()
old.timestamp = -100
old.foo = 12
old.pid = 1
old.key = key
old.terminate.return_value = None
seal(old)
new.key = key
cache.entry_list.append(old)
new.timestamp = -5
seal(new)
cache.entry_list.append(new)
self.assertEqual(len(cache.entry_list), 2)
ppc = PruneProcessCache(cache)
ppc.prune()
self.assertEqual(len(cache.entry_list), 1)
self.assertEqual(cache.entry_list[0], new)
self.assertEqual(cache.entry_list[0], new)
self.assertTrue(old.terminate.called)
if __name__ == '__main__':
unittest.main()
if __name__ == "__main__":
unittest.main()

View File

@@ -0,0 +1,77 @@
import unittest
from unittest.mock import MagicMock, patch
from cellxgene_gateway.cache_entry import CacheEntry
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.items.file.fileitem import FileItem
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
from cellxgene_gateway.items.item import ItemType
from cellxgene_gateway.process_exception import ProcessException
class TestSubprocessBackend(unittest.TestCase):
@patch("subprocess.Popen")
def test_launch_GIVEN_no_stdout_THEN_throw_ProcessException(self, popen):
subprocess = MagicMock()
subprocess.stdout.readline().decode.return_value = ""
subprocess.stderr.read().decode.return_value = "An unexpected error"
popen.return_value = subprocess
key = CacheKey(
FileItem("/czi/", name="pbmc3k.h5ad", type=ItemType.h5ad),
FileItemSource("/tmp", "local"),
)
entry = CacheEntry.for_key(key, 8000)
from cellxgene_gateway.subprocess_backend import SubprocessBackend
backend = SubprocessBackend()
cellxgene_loc = "/some/cellxgene"
scripts = ["http://example.com/script.js", "http://example.com/script2.js"]
with self.assertRaises(ProcessException) as context:
backend.launch(cellxgene_loc, scripts, entry)
popen.assert_called_once_with(
[
"yes | /some/cellxgene launch /tmp/czi/pbmc3k.h5ad --port 8000 --host 127.0.0.1 --disable-annotations --disable-gene-sets-save --scripts http://example.com/script.js --scripts http://example.com/script2.js"
],
shell=True,
stderr=-1,
stdout=-1,
)
self.assertEqual("An unexpected error", context.exception.stderr)
@patch("subprocess.Popen")
def test_launch_GIVEN_annotations_enabled_THEN_set_flags(self, popen):
subprocess = MagicMock()
subprocess.stdout.readline().decode.return_value = (
"[cellxgene] Type CTRL-C at any time to exit.\n"
)
subprocess.stderr.read().decode.return_value = ""
popen.return_value = subprocess
key = CacheKey(
FileItem("/czi/", name="pbmc3k.h5ad", type=ItemType.h5ad),
FileItemSource("/tmp", "local"),
FileItem(
"/czi/pbmc3k_annotations/", name="foo.csv", type=ItemType.annotation
),
)
entry = CacheEntry.for_key(key, 8000)
import cellxgene_gateway.subprocess_backend
cellxgene_gateway.subprocess_backend.enable_annotations = True
try:
backend = cellxgene_gateway.subprocess_backend.SubprocessBackend()
cellxgene_loc = "/some/cellxgene"
backend.launch(cellxgene_loc, [], entry)
finally:
cellxgene_gateway.subprocess_backend.enable_annotations = False
popen.assert_called_once_with(
[
"yes | /some/cellxgene launch /tmp/czi/pbmc3k.h5ad --port 8000 --host 127.0.0.1 --annotations-file /tmp/czi/pbmc3k_annotations/foo.csv --gene-sets-file /tmp/czi/pbmc3k_annotations/foo_gene_sets.csv"
],
shell=True,
stderr=-1,
stdout=-1,
)