56 Commits

Author SHA1 Message Date
Alok Saldanha
8f676f28d0 Prepare 0.4.1 release 2025-12-01 20:38:35 -05:00
Alok Saldanha
a74576ade5 Merge pull request #101 from Novartis/delay_itemsource_init
Delay itemsource init
2025-12-01 20:16:21 -05:00
Alok Saldanha
d747860118 made pruner a daemon thread 2025-11-08 18:40:33 -05:00
Alok Saldanha
79fef57010 updated start scripts to run in subshells 2025-11-08 17:51:06 -05:00
Alok Saldanha
cd9c0a3671 upload coverage reports as artifacts 2025-11-08 17:38:20 -05:00
Alok Saldanha
55b268125f blacken 2025-11-08 17:23:47 -05:00
Alok Saldanha
4df58f9ceb fixe bug in status.json 2025-11-08 17:21:37 -05:00
Alok Saldanha
8b4565e745 added status.json test 2025-11-08 17:21:37 -05:00
Alok Saldanha
c8056991b0 simplify tests 2025-11-08 11:00:07 -05:00
Alok Saldanha
1d1d8b4e59 Address linter warnings 2025-11-08 10:57:50 -05:00
Alok Saldanha
a35c6b9b1e added start scripts for flask, gunicorn and uwsgi 2025-11-05 22:02:56 -05:00
Alok Saldanha
f260180a76 set default_item_source and start pruner thread 2025-11-05 21:54:18 -05:00
Alok Saldanha
58ae41fe0c Delay itemsource initialization until first request is served 2025-11-05 21:41:10 -05:00
Alok Saldanha
0c5adc9fef Merge pull request #99 from andynu/gunicorn-support
Gunicorn support for now. Will revisit populating item sources on module load to improve reusability of module (specifically in tests for now).
2025-11-05 06:55:39 -05:00
Alok Saldanha
34ac73ba01 second attempt to skip codecov on PRs from forks 2025-11-05 06:47:33 -05:00
Alok Saldanha
d16a97906c Skip codecov for PRs from fork 2025-11-05 06:44:54 -05:00
Alok Saldanha
3e5accad65 fixed linting and tests 2025-11-04 06:57:43 -05:00
Andy
903d25763f Fix AttributeError by storing ItemSource objects in default_item_source
Previously, default_item_source was set to a string ("s3" or "local"),
but matching_source() tried to access default_item_source.name, causing:
  AttributeError: 'str' object has no attribute 'name'

This bug occurred when source_name=None (single data source configuration)
and has existed since the ItemSource interface was introduced in 2021.

Changes:
- Store the actual ItemSource object reference instead of string name
- Assign to intermediate variables (s3_source, file_source) for clarity

Fixes the error when viewing datasets with a single data source configured.
2025-10-29 14:08:56 -04:00
Andy
08c546f40a Fix WSGI server initialization by extracting data source setup
Addresses the issue where Gunicorn/uWSGI servers import the gateway
module but never call main(), leaving item_sources empty and causing
the file crawler to fail.

Changes:
- Extract data source initialization into initialize_data_sources()
- Call initialization at module import time for WSGI compatibility
- Add _initialized flag to prevent double initialization
- Simplify main() to delegate to initialize_data_sources()

This ensures data sources are populated when running under WSGI servers
(Gunicorn, uWSGI) while maintaining backward compatibility with the
Flask development server.

Related to GitHub issues #33 and #92
2025-10-29 14:08:56 -04:00
Andy
b9f4d35812 Fix UnicodeDecodeError when viewing compressed datasets
When viewing datasets through the gateway, requests would fail with:
  UnicodeDecodeError: 'utf-8' codec can't decode byte 0xb5 in position 1

The gateway was copying the accept-encoding header from browser requests
when proxying to cellxgene backend servers. When accept-encoding is
manually set, the Python requests library assumes the caller will handle
decompression and leaves response content compressed.

The cellxgene server responded with zstd-compressed content (magic bytes
28 b5 2f fd), but the gateway attempted to decode this compressed binary
data as UTF-8 text, causing the decode error.

Solution: Remove accept-encoding from the copied headers list in
cache_entry.py. This allows the requests library to automatically handle
compression negotiation and transparently decompress responses (gzip,
deflate, brotli, zstd, etc.).

This is the standard practice when proxying with requests and maintains
all other gateway functionality (URL rewriting, auth, caching, etc.).

Tested:
- Dataset viewing works with compressed responses
- File browser and static assets load correctly
- URL rewriting continues to function properly
2025-10-29 14:04:25 -04:00
Alok Saldanha
05a95c2ca9 Prepare 0.4.0 release 2024-03-10 09:33:50 -04:00
Alok Saldanha
5d29153544 #87 remove version pins for markupsafe, flask and werkzeug
also remove dependency on flask-api
2024-03-10 09:31:39 -04:00
Alok Saldanha
6a2bc409db Prepare for 0.3.12 release 2024-03-03 07:45:47 -05:00
Alok Saldanha
fa72481b66 Merge remote-tracking branch 'origin/dependabot/pip/werkzeug-2.3.8' 2024-03-03 07:33:03 -05:00
Alok Saldanha
3d0166904b Merge pull request #90 from Novartis/dependabot/pip/flask-2.2.5
Bump flask from 2.2.2 to 2.2.5
2024-03-03 07:31:42 -05:00
Alok Saldanha
4e63ff95a8 #87 blacken 2024-03-02 12:25:21 -05:00
Alok Saldanha
c1111e2cb4 #87 patch enable annotations 2024-03-02 12:23:11 -05:00
Alok Saldanha
c4b9084286 #87 Fix test 2024-03-02 12:08:57 -05:00
dependabot[bot]
0282da03c8 Bump werkzeug from 2.3.0 to 2.3.8
Bumps [werkzeug](https://github.com/pallets/werkzeug) from 2.3.0 to 2.3.8.
- [Release notes](https://github.com/pallets/werkzeug/releases)
- [Changelog](https://github.com/pallets/werkzeug/blob/main/CHANGES.rst)
- [Commits](https://github.com/pallets/werkzeug/compare/2.3.0...2.3.8)

---
updated-dependencies:
- dependency-name: werkzeug
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
2024-03-02 16:56:15 +00:00
dependabot[bot]
bec74bec45 Bump flask from 2.2.2 to 2.2.5
Bumps [flask](https://github.com/pallets/flask) from 2.2.2 to 2.2.5.
- [Release notes](https://github.com/pallets/flask/releases)
- [Changelog](https://github.com/pallets/flask/blob/main/CHANGES.rst)
- [Commits](https://github.com/pallets/flask/compare/2.2.2...2.2.5)

---
updated-dependencies:
- dependency-name: flask
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
2024-03-02 16:56:15 +00:00
Alok Saldanha
35c8e8180c #87 temporarily pin versions 2024-03-02 11:55:07 -05:00
Alok Saldanha
0000a60eb0 #73 hide annotation links when disabled 2024-03-02 11:51:20 -05:00
Alok Saldanha
1e02e0abb8 Merge pull request #74 from Novartis/73_reorder_filecrawl
#73 moved new annotation link to front
2024-03-02 11:44:49 -05:00
Alok Saldanha
66cca86b51 Merge remote-tracking branch 'ghall/just_gene_sets' 2024-03-02 11:09:00 -05:00
Alok Saldanha
f046e7c5d2 Merge pull request #88 from Mye-InfoBank/master
Fix dockerfile installation problems
2024-02-24 09:22:49 -05:00
Nico Trummer
a8f6e45f34 Implement pip upgrade to Dockerfile 2024-02-21 09:47:46 +01:00
george-hall-ucl
19caa6cb80 Sorry -- forgot to lint 2023-08-08 16:04:03 +01:00
george-hall-ucl
56bd079024 Save gene sets without cell annotations
This fixes a bug whereby new gene_sets csv files created without
accompanying cell-level annotations could not be detected by the
filecrawler.
2023-08-08 15:49:57 +01:00
Alok Saldanha
9d10932b06 prepare for 0.3.11 release 2023-07-09 19:27:32 -04:00
Alok Saldanha
c7c156b4cf Merge pull request #77 from aeisenbarth/filter-empty-folders
Filter directories without h5ad files
2023-07-09 07:04:06 -06:00
Alok Saldanha
624d1f8567 #78 Revert "Rename argument "filter" to "subpath""
This reverts commit fdd6cca297.
2023-07-09 08:25:27 -04:00
Alok Saldanha
79c3f588b6 Merge pull request #80 from Novartis/79_add_docker_example
79 add docker example
2023-07-09 06:05:50 -06:00
Alok Saldanha
d8fd07572c Merge pull request #83 from Novartis/81_gene_set_support
gene set support
2023-07-09 06:03:33 -06:00
Alok Saldanha
64d636a1c5 #81 added unit test for gene sets 2023-07-09 07:46:01 -04:00
Alok Saldanha
7b314d4457 #81 switch to latest ubuntu 2023-07-06 17:15:00 -06:00
Alok Saldanha
2754bc1ef1 #81 Combined GATEWAY_ENABLE_ANNOTATIONS and GATEWAY_ENABLE_GENE_SETS flags 2023-07-06 08:53:19 -06:00
Alok Saldanha
5a650334df #81 moved gene set check into fileitem_source 2023-07-06 08:18:50 -06:00
george-hall-ucl
81c8ce4219 #81 Add support for gene sets
This adds the flag `GATEWAY_ENABLE_GENE_SETS` to enable support for gene
sets.  To simplify implementation, activating this flag also activates
`GATEWAY_ENABLE_ANNOTATIONS`.  The gene sets are saved in a file that
has the same name as the annotations `csv` but with `_gene_sets`
appended to the file name (before the extension).  This file is hidden
in filecrawler, and the gene sets are loaded when the associated
annotations file is loaded.

If the annotations file is missing, then an Exception is raised.

I have updated one unit test to make it expect
`--disable-gene-sets-save` in the default case (i.e. if
`GATEWAY_ENABLE_ANNOTATIONS = 0`).  All units tests pass.

I have updated the README to document `GATEWAY_ENABLE_GENE_SETS`.
2023-07-06 07:59:12 -06:00
Alok Saldanha
f296efcc55 #79 added cellxgene-data directory so it actually works 2022-12-21 16:02:50 -05:00
Alok Saldanha
facfb27d5c #79 add simple example to customize cellxgene-gateway ui 2022-12-21 15:42:13 -05:00
Andreas Eisenbarth
6607b15085 Exclude directories having no h5ad files 2022-10-12 17:31:22 +02:00
Andreas Eisenbarth
0e92b73347 Add test case for dirs without h5ad 2022-10-12 17:31:22 +02:00
Andreas Eisenbarth
88b9b815c0 Adjust test case for dirs with h5ad 2022-10-12 17:31:22 +02:00
Andreas Eisenbarth
6ef82b36f1 For running individual tests, make sure flask_util.view_url is callable 2022-10-12 16:36:46 +02:00
Andreas Eisenbarth
fdd6cca297 Rename argument "filter" to "subpath" 2022-10-12 13:40:39 +02:00
Alok Saldanha
a00403c60e #73 moved new annotation link to front 2022-08-21 08:15:09 -04:00
28 changed files with 798 additions and 127 deletions

View File

@@ -6,7 +6,7 @@ on: [push, pull_request]
jobs:
black:
runs-on: ubuntu-18.04
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v2
name: Checkout repository
@@ -25,7 +25,7 @@ jobs:
black . --check
# This job is copied over from `deploy.yaml`
run-tests:
runs-on: ubuntu-18.04
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v2
@@ -39,7 +39,6 @@ jobs:
conda env create -f environment.yml
eval "$(conda shell.bash hook)"
conda activate cellxgene-gateway
pip install markupsafe==2.0.1 # temporary workaround for jinja2-2.11.3 calling soft_unicode in markupsafe
python setup.py install
- name: Run tests
@@ -53,16 +52,39 @@ jobs:
eval "$(conda shell.bash hook)"
conda activate cellxgene-gateway
coverage report --fail-under 41
coverage report > coverage.txt
coverage html -i
coverage xml -i
- name: "Upload coverage to Codecov"
uses: codecov/codecov-action@v1
- name: Upload coverage HTML report
uses: actions/upload-artifact@v4
with:
token: ${{ secrets.CODECOV_TOKEN }}
files: ./coverage.xml
flags: unittests
env_vars: OS,PYTHON
name: codecov-umbrella
fail_ci_if_error: true
path_to_write_report: ./codecov_report.txt
verbose: true
name: coverage-html
path: htmlcov/
retention-days: 30
- name: Upload coverage xml
uses: actions/upload-artifact@v4
with:
name: coverage-xml
path: coverage.xml
retention-days: 30
- name: Upload coverage summary
uses: actions/upload-artifact@v4
with:
name: coverage-summary
path: coverage.txt
retention-days: 30
# - name: "Upload coverage to Codecov"
# if: ${{ github.event_name == 'push' || (github.event_name == 'pull_request' && github.event.pull_request.head.repo.full_name == github.repository) }}
# uses: codecov/codecov-action@v1
# with:
# token: ${{ secrets.CODECOV_TOKEN }}
# files: ./coverage.xml
# flags: unittests
# env_vars: OS,PYTHON
# name: codecov-umbrella
# fail_ci_if_error: true
# path_to_write_report: ./codecov_report.txt
# verbose: true

View File

@@ -1,3 +1,34 @@
# 0.4.1
* Fix UnicodeDecodeError when viewing compressed datasets
* Fix WSGI server initialization by extracting data source setup
* Fix AttributeError by storing ItemSource objects in default_item_source
* Delay itemsource initialization until first request is served
* Set default_item_source and start pruner thread
* Added start scripts for flask, gunicorn and uwsgi
* Made pruner a daemon thread
* Updated start scripts to run in subshells
* Fixed bug in status.json
# 0.4.0
* Removed dependency on flask-api
* Updated dependencies (python 3.11, numpy, unpinned flask, werkzeug)
# 0.3.12
* #81 List gene set annotations when cell annotations not present
* #86 Upgrade pip within docker image
* #73 Moved new link to front
* #87 Temporarily pin versions of werkzeug and flask
# 0.3.11
* #81 added support for gene sets
* #79 added example for cellxgene-gateway customized docker image
* #78 prune directories that do not contain h5ad files
# 0.3.10
* #65 Added GATEWAY_EXPIRE_SECONDS to set how long cellxgene servers can remain idle before being terminated.

View File

@@ -1,6 +1,7 @@
FROM python:3.9
FROM python:3.11
RUN pip install cellxgene-gateway 'MarkupSafe<2.1'
RUN pip install --upgrade pip
RUN pip install "cellxgene-gateway>=0.4"
ENV CELLXGENE_DATA=/cellxgene-data
ENV CELLXGENE_LOCATION=/usr/local/bin/cellxgene

View File

@@ -75,7 +75,7 @@ Optional environment variables:
* `GATEWAY_PORT` - local port that the gateway should bind to, defaults to 5005
* `GATEWAY_EXPIRE_SECONDS` - time in seconds that a cellxgene process will remain idle before being terminated. Defaults to 3600 (one hour)
* `GATEWAY_EXTRA_SCRIPTS` - JSON array of script paths, will be embedded into each page and forwarded with `--scripts` to cellxgene server
* `GATEWAY_ENABLE_ANNOTATIONS` - Set to `true` or to `1` to enable cellxgene annotations.
* `GATEWAY_ENABLE_ANNOTATIONS` - Set to `true` or to `1` to enable cellxgene annotations and gene sets.
* `GATEWAY_ENABLE_BACKED_MODE` - Set to `true` or to `1` to load AnnData in file-backed mode. This saves memory and speeds up launch time but may reduce overall performance.
* `GATEWAY_LOG_LEVEL` - default is `INFO`. set to `DEBUG` to increase logging and to `WARNING` to decrease logging.
* `S3_ENABLE_LISTINGS_CACHE` - Set to `true` or to `1` to cache listings of S3 folders for performance. If the cache becomes stale, set `filecrawl.html?refresh=true` query parameter to refresh the cache.
@@ -110,11 +110,26 @@ Additional environment variables can be provided with the `-e` parameter:
```bash
docker run -it --rm \
-v <local_data_dir>:/cellxgene-data \
-v ../cellxgene_data:/cellxgene-data \
-e GATEWAY_PORT=8080 \
-p 8080:8080 \
cellxgene-gateway
```
## Running cellxgene gateway with start scripts
For your convenience, we provide start scripts for flask, gunicorn and uwsgi.
First, set up a .env
```bash
cp env_example .env
# edit .env
open .env
```
Then run the scripts in a subshell
```bash
( ./start_flask.sh )
```
# Customization
@@ -191,6 +206,35 @@ black .
If you need help for any reason, please make a github ticket. One of the contributors should help you out.
# Releasing New Versions
## How to prepare for release
- Update Changelog.md and version number in __init__.py
- Cut a release on github
- Go to your project homepage on GitHub
- On right side, you will see [Releases](https://github.com/Novartis/cellxgene-gateway/releases) link. Click on it.
- Click on Draft a new release
- Fill in all the details
- Tag version should be the version number of your package release
- Release Title can be anything you want, but we use v0.3.11 (the same as the tag to be created on publish)
- Description should be changelog
- Click Publish release at the bottom of the page
- Now under Releases you can view all of your releases.
- Copy the download link (tar.gz) and save it somewhere
## How to publish to PyPI
Make sure your `.pypirc` is set up for testpypi and pypi index servers.
```bash
rm -rf dist
python setup.py sdist bdist_wheel
python -m twine upload --repository testpypi dist/*
python -m twine upload dist/*
```
# Contributors
* Niket Patel - https://github.com/NiketPatel9

View File

@@ -7,4 +7,4 @@
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
__version__ = "0.3.10"
__version__ = "0.4.1"

View File

@@ -8,11 +8,10 @@
# the specific language governing permissions and limitations under the License.
import time
from http import HTTPStatus
from threading import Thread
from typing import List
from flask_api import status
from cellxgene_gateway import env
from cellxgene_gateway.cache_entry import CacheEntry, CacheEntryStatus
from cellxgene_gateway.cache_key import CacheKey
@@ -53,7 +52,7 @@ class BackendCache:
return matches[0]
else:
raise CellxgeneException(
status.HTTP_500_INTERNAL_SERVER_ERROR,
HTTPStatus.INTERNAL_SERVER_ERROR,
"Found " + str(len(matches)) + " for " + path,
)
@@ -71,7 +70,7 @@ class BackendCache:
return matches[0]
else:
raise CellxgeneException(
status.HTTP_500_INTERNAL_SERVER_ERROR,
HTTPStatus.INTERNAL_SERVER_ERROR,
"Found " + str(len(matches)) + " for " + key.dataset,
)

View File

@@ -58,7 +58,6 @@ class CacheEntry:
@classmethod
def for_key(cls, key, port):
return cls(
None,
key,
@@ -149,7 +148,7 @@ class CacheEntry:
headers = {}
copy_headers = [
"accept",
"accept-encoding",
# "accept-encoding" - removed: let requests library handle compression/decompression
"accept-language",
"cache-control",
"connection",
@@ -169,9 +168,8 @@ class CacheEntry:
headers[h] = request.headers[h]
full_path = self.cellxgene_basepath() + subpath + querystring()
cellxgene_response = None
try:
cellxgene_response = None
if request.method in ["GET", "HEAD", "OPTIONS"]:
cellxgene_response = get(full_path, headers=headers)
elif request.method == "PUT":

View File

@@ -9,8 +9,6 @@
import os
from flask_api import status
from cellxgene_gateway import env
from cellxgene_gateway.cellxgene_exception import CellxgeneException

View File

@@ -7,31 +7,32 @@
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
import html
import urllib.parse
from cellxgene_gateway import env, flask_util
from cellxgene_gateway import flask_util
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.dir_util import annotations_suffix, make_annotations, make_h5ad
from cellxgene_gateway.env import enable_annotations
def render_annotations(item, item_source):
if not enable_annotations:
return ""
url = flask_util.view_url(
item_source.get_annotations_subpath(item), item_source.name
)
new_annotation = f"<a class='new' href='{url}'>new</a>"
new_annotation = [f"<a class='new' href='{url}'>new</a>"]
annotations = (
", ".join(
[
f"<a href='{CacheKey(item, item_source, a).view_url}/'>{a.name}</a>"
for a in item.annotations
]
)
+ ", "
[
f"<a href='{CacheKey(item, item_source, a).view_url}/'>{html.escape(a.name)}</a>"
for a in item.annotations
]
if item.annotations
else ""
else []
)
return " | annotations: " + annotations + new_annotation
return "| annotations: " + ", ".join(new_annotation + annotations)
def render_item(item, item_source):

View File

@@ -40,6 +40,12 @@ app = Flask(__name__)
item_sources = []
default_item_source = None
# Guard for lazy initialization so tests can import this module without
# triggering environment-dependent side effects. initialize_data_sources()
# will set this to True when it has run.
data_sources_initialized = False
data_sources_init_lock = Lock()
def _force_https(app):
def wrapper(environ, start_response):
@@ -75,12 +81,89 @@ if (
x_prefix=env.proxy_fix_prefix,
)
# WSGI middleware to ensure data sources are initialized before the first
# WSGI request is handled. This guarantees initialization works under
# Gunicorn/uWSGI (which import the module but don't call main()). The
# initialize_data_sources() function is idempotent-protected by
# data_sources_initialized and data_sources_init_lock.
def _init_on_first_wsgi_request(wsgi_app):
def middleware(environ, start_response):
global data_sources_initialized
if not data_sources_initialized:
with data_sources_init_lock:
if not app.extensions.get("cellxgene_gateway", {}).get("launchtime"):
app.extensions.setdefault("cellxgene_gateway", {})[
"launchtime"
] = current_time_stamp()
if not data_sources_initialized:
initialize_data_sources()
env.validate()
if not item_sources or not len(item_sources):
raise Exception(
"No data sources specified for Cellxgene Gateway"
)
global default_item_source
if default_item_source is None:
default_item_source = item_sources[0]
data_sources_initialized = True
return wsgi_app(environ, start_response)
return middleware
# Wrap the WSGI app so Gunicorn/uWSGI will trigger initialization when the
# first request comes in. Tests that need initialization can call
# initialize_data_sources() directly.
app.wsgi_app = _init_on_first_wsgi_request(app.wsgi_app)
cache = BackendCache()
# Initialize data sources - this is defined later in the file but called here
# to ensure initialization happens when WSGI servers (Gunicorn) import the module
def initialize_data_sources():
"""Initialize data sources from environment variables.
Called at module import time for WSGI server compatibility (Gunicorn).
Uses a guard flag to prevent double initialization within a process."""
global default_item_source
logging.basicConfig(
level=env.log_level,
format="%(asctime)s:%(name)s:%(levelname)s:%(message)s",
)
logger = logging.getLogger(__name__)
cellxgene_data = os.environ.get("CELLXGENE_DATA", None)
cellxgene_bucket = os.environ.get("CELLXGENE_BUCKET", None)
if cellxgene_bucket is not None:
from cellxgene_gateway.items.s3.s3item_source import S3ItemSource
s3_source = S3ItemSource(cellxgene_bucket, name="s3")
item_sources.append(s3_source)
default_item_source = s3_source
logger.info("Initialized S3 data source")
logger.debug(f"S3 bucket: {cellxgene_bucket}")
if cellxgene_data is not None:
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
file_source = FileItemSource(cellxgene_data, name="local")
item_sources.append(file_source)
default_item_source = file_source
logger.info("Initialized local file data source")
logger.debug(f"Data directory: {cellxgene_data}")
if len(item_sources) == 0:
raise Exception("Please specify CELLXGENE_DATA or CELLXGENE_BUCKET")
flask_util.include_source_in_url = len(item_sources) > 1
@app.errorhandler(CellxgeneException)
def handle_invalid_usage(error):
message = f"{error.http_status} Error : {error.message}"
return (
@@ -95,7 +178,6 @@ def handle_invalid_usage(error):
@app.errorhandler(ProcessException)
def handle_invalid_process(error):
message = []
message.append(error.message)
@@ -171,7 +253,7 @@ entry_lock = Lock()
def matching_source(source_name):
if source_name is None:
if source_name is None and default_item_source is not None:
source_name = default_item_source.name
matching = [i for i in item_sources if i.name == source_name]
if len(matching) != 1:
@@ -218,6 +300,11 @@ def do_view(path, source_name=None):
raise CellxgeneException("User not authorized to access this data", 403)
elif match.status == CacheEntryStatus.error:
raise ProcessException.from_cache_entry(match)
else:
raise CellxgeneException(
f"Unexpected cache entry status {match.status} for key {match.key.descriptor}",
500,
)
@app.route("/cache_status", methods=["GET"])
@@ -231,28 +318,40 @@ def do_GET_status():
@app.route("/cache_status.json", methods=["GET"])
def do_GET_status_json():
def map_entry(entry):
dataset = entry.key.h5ad_item.descriptor
annotation_file = entry.key.annotation_descriptor
return {
"dataset": dataset,
"annotation_file": annotation_file,
"launchtime": entry.launchtime,
"last_access": entry.timestamp,
"status": entry.status.name,
}
return json.dumps(
{
"launchtime": app.launchtime,
"entry_list": [
{
"dataset": entry.key.dataset,
"annotation_file": entry.key.annotation_file,
"launchtime": entry.launchtime,
"last_access": entry.timestamp,
"status": entry.status,
}
for entry in cache.entry_list
],
"launchtime": app.extensions.get("cellxgene_gateway", {}).get("launchtime"),
"entry_list": [map_entry(entry) for entry in cache.entry_list],
}
)
@app.route("/relaunch/<path:path>", methods=["GET"])
def do_relaunch(path):
source_name = request.args.get("source_name") or default_item_source.name
def get_cache_key(path):
if request.args.get("source_name"):
source_name = request.args.get("source_name")
elif default_item_source:
source_name = default_item_source.name
else:
source_name = None
source = matching_source(source_name)
key = CacheKey.for_lookup(source, source.lookup(path))
return key
@app.route("/relaunch/<path:path>", methods=["GET"])
def do_relaunch(path):
key = get_cache_key(path)
match = cache.check_entry(key)
if not match is None:
match.terminate()
@@ -264,9 +363,7 @@ def do_relaunch(path):
@app.route("/terminate/<path:path>", methods=["GET"])
def do_terminate(path):
source_name = request.args.get("source_name") or default_item_source.name
source = matching_source(source_name)
key = CacheKey.for_lookup(source, source.lookup(path))
key = get_cache_key(path)
match = cache.check_entry(key)
if not match is None:
match.terminate()
@@ -279,46 +376,29 @@ def ip_address():
return set_no_cache(resp)
def launch():
env.validate()
if not item_sources or not len(item_sources):
raise Exception("No data sources specified for Cellxgene Gateway")
global default_item_source
if default_item_source is None:
default_item_source = item_sources[0]
def start_pruner_thread():
pruner = PruneProcessCache(cache)
background_thread = Thread(target=pruner)
# Run the pruner as a daemon thread so it won't block interpreter
# shutdown (for example when Ctrl-C is used in the main thread).
# This avoids "Exception ignored in: <module 'threading'...>" at exit.
background_thread = Thread(target=pruner, daemon=True)
background_thread.start()
app.launchtime = current_time_stamp()
def launch():
start_pruner_thread()
app.extensions.setdefault("cellxgene_gateway", {})[
"launchtime"
] = current_time_stamp()
app.run(host="0.0.0.0", port=env.gateway_port, debug=False)
app.extensions.setdefault("cellxgene_gateway", {})["launchtime"] = None
def main():
logging.basicConfig(
level=env.log_level,
format="%(asctime)s:%(name)s:%(levelname)s:%(message)s",
)
cellxgene_data = os.environ.get("CELLXGENE_DATA", None)
cellxgene_bucket = os.environ.get("CELLXGENE_BUCKET", None)
if cellxgene_bucket is not None:
from cellxgene_gateway.items.s3.s3item_source import S3ItemSource
item_sources.append(S3ItemSource(cellxgene_bucket, name="s3"))
default_item_source = "s3"
if cellxgene_data is not None:
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
item_sources.append(FileItemSource(cellxgene_data, name="local"))
default_item_source = "local"
if len(item_sources) == 0:
raise Exception("Please specify CELLXGENE_DATA or CELLXGENE_BUCKET")
flask_util.include_source_in_url = len(item_sources) > 1
"""CLI entry point for Flask development server."""
launch()

View File

@@ -24,17 +24,22 @@ class FileItemSource(ItemSource):
h5ad_suffix=dir_util.h5ad_suffix,
annotation_dir_suffix=dir_util.annotations_suffix,
annotation_file_suffix=".csv",
gene_set_file_suffix="_gene_sets.csv",
):
self._name = name
self.base_path = base_path
self.h5ad_suffix = h5ad_suffix
self.annotation_dir_suffix = annotation_dir_suffix
self.annotation_file_suffix = annotation_file_suffix
self.gene_set_file_suffix = gene_set_file_suffix
@property
def name(self):
return self._name or f"Files:{self.base_path}"
def is_gene_set(self, path: str) -> bool:
return path.endswith(self.gene_set_file_suffix)
def is_h5ad_file(self, path: str) -> bool:
return path.endswith(self.h5ad_suffix) and os.path.isfile(path)
@@ -63,7 +68,7 @@ class FileItemSource(ItemSource):
return item_tree
def scan_directory(self, subpath="") -> dict:
def scan_directory(self, subpath: str = "") -> ItemTree:
base_path = os.path.join(self.base_path, subpath)
if not os.path.exists(base_path):
@@ -100,6 +105,11 @@ class FileItemSource(ItemSource):
branches = [
self.scan_directory(os.path.join(subpath, subdir)) for subdir in subdirs
]
# Exclude branches without files as leaves. Since traversal is applied pre-order,
# branch.branches has already been processed and we don't need to check deeper nesting.
branches = [
branch for branch in branches if branch.items or branch.branches
]
return ItemTree(subpath, items, branches)
@@ -176,11 +186,29 @@ class FileItemSource(ItemSource):
annotations_subpath = self.get_annotations_subpath(item)
annotations_fullpath = self.full_path(annotations_subpath)
if os.path.isdir(annotations_fullpath):
return [
sorted_files = sorted(os.listdir(annotations_fullpath))
annotation_files = [
self.make_fileitem_from_path(annotation, annotations_subpath, True)
for annotation in sorted(os.listdir(annotations_fullpath))
for annotation in sorted_files
if annotation.endswith(self.annotation_file_suffix)
and not self.is_gene_set(annotation)
and os.path.isfile(os.path.join(annotations_fullpath, annotation))
]
# Catch gene sets without accompanying [annotations].csv
gene_sets_files = [
self.make_fileitem_from_path(
annotation[: -len(self.gene_set_file_suffix)] + ".csv",
annotations_subpath,
True,
)
for annotation in sorted_files
if self.is_gene_set(annotation)
and annotation[: -len(self.gene_set_file_suffix)]
not in [a.name for a in annotation_files]
and os.path.isfile(os.path.join(annotations_fullpath, annotation))
]
return sorted(annotation_files + gene_sets_files, key=lambda x: x.name)
else:
return None

View File

@@ -116,6 +116,9 @@ class S3ItemSource(ItemSource):
branches = None
if len(subdir_keys) > 0:
branches = [self.scan_directory(key) for key in subdir_keys]
branches = [
branch for branch in branches if branch.items or branch.branches
]
return ItemTree(directory_key, items, branches)

View File

@@ -9,8 +9,7 @@
import logging
import subprocess
from flask_api import status
from http import HTTPStatus
from cellxgene_gateway.cache_entry import CacheEntryStatus
from cellxgene_gateway.dir_util import make_annotations
@@ -30,8 +29,11 @@ class SubprocessBackend:
extra_args = f" --annotations-dir {make_annotations(file_path)}"
else:
extra_args = f" --annotations-file {annotation_file_path}"
gene_sets_file_path = annotation_file_path[:-4] + "_gene_sets.csv"
extra_args += f" --gene-sets-file {gene_sets_file_path}"
else:
extra_args = " --disable-annotations"
extra_args += " --disable-gene-sets-save"
if enable_backed_mode:
extra_args += " --backed"
if not cellxgene_args is None:
@@ -73,10 +75,10 @@ class SubprocessBackend:
or "Could not open file" in stderr
):
message = "File was invalid."
http_status = status.HTTP_400_BAD_REQUEST
http_status = HTTPStatus.BAD_REQUEST
else:
message = "Cellxgene failed to launch dataset."
http_status = status.HTTP_500_INTERNAL_SERVER_ERROR
http_status = HTTPStatus.INTERNAL_SERVER_ERROR
cache_entry.status = CacheEntryStatus.error
cache_entry.set_error(message, stderr, http_status)

3
env_example Normal file
View File

@@ -0,0 +1,3 @@
export CELLXGENE_LOCATION=$(pwd)/.venv/bin/cellxgene
export CELLXGENE_DATA=../cellxgene_data
export GATEWAY_IP=127.0.0.1

View File

@@ -2,7 +2,7 @@ name: cellxgene-gateway
channels:
- conda-forge
dependencies:
- python=3.9
- python=3.11
- requests
- flask
- psutil
@@ -13,6 +13,5 @@ dependencies:
- pip
- pip:
- pre_commit
- flask-api
- werkzeug
- cellxgene

View File

@@ -0,0 +1,14 @@
FROM python:3.9
RUN pip install cellxgene-gateway 'MarkupSafe<2.1'
COPY customize_ui.sh customize_ui.sh
RUN CELLXGENE_GATEWAY_DIR=/usr/local/lib/python3.9/site-packages/cellxgene_gateway . ./customize_ui.sh
ENV CELLXGENE_DATA=/cellxgene-data
ENV CELLXGENE_LOCATION=/usr/local/bin/cellxgene
EXPOSE 5005
RUN mkdir /cellxgene-data
CMD ["cellxgene-gateway"]

View File

@@ -0,0 +1,14 @@
# Purpose
This is a simple example of how to make a small script to customize the UI of cellxgene-gateway. The script that does the customization is `customize_ui.sh`, it simply makes the main header green using CSS but you could do anything you want there (including adding more script tags, etc).
# Usage
```
docker build -t cellxgene_custom .
CELLXGENE_DATA=`pwd`/../../../cellxgene_data
docker run -p 5005:5005 --mount src=$CELLXGENE_DATA,target=/cellxgene-data,type=bind cellxgene_custom
```
If you now open http://localhost:5005 you should see a green cellxgene gateway header.

View File

@@ -0,0 +1,3 @@
# make the header bright green
find "${CELLXGENE_GATEWAY_DIR}/templates" -name index.html -exec sed -i -e 's/<head>/<head>\
> <style> header h3 {color: #0F0;} <\/style>/g' {} \;

View File

@@ -1,6 +1,5 @@
cellxgene
flask
flask-api
werkzeug
psutil
requests

View File

@@ -1,6 +0,0 @@
export CELLXGENE_LOCATION=$(pwd)/.cellxgene-gateway/bin/cellxgene
export CELLXGENE_DATA=../cellxgene_data
export GATEWAY_IP=127.0.0.1
#Once these are set, you run like a normal Flask app
cellxgene-gateway

41
start_flask.sh Executable file
View File

@@ -0,0 +1,41 @@
#!/bin/bash
# start_gunicorn.sh - Start Cellxgene Gateway with Gunicorn
#
# PREREQUISITES:
# - Gunicorn installed (included with cellxgene 1.3.0, or: pip install gunicorn)
# - Virtual environment activated
# - .env file with CELLXGENE_LOCATION and CELLXGENE_DATA (or CELLXGENE_BUCKET)
#
# USAGE:
# ./start_gunicorn.sh
# Exit on error
set -e
# Get the directory where this script is located
SCRIPT_DIR="$( cd "$( dirname "${BASH_SOURCE[0]}" )" && pwd )"
# Source environment variables
echo "Loading environment variables..."
if [ -f "$SCRIPT_DIR/.env" ]; then
source "$SCRIPT_DIR/.env"
else
echo "Error: .env file not found at $SCRIPT_DIR/.env"
echo "Please create it with required environment variables"
exit 1
fi
# Verify required environment variables
if [ -z "$CELLXGENE_LOCATION" ]; then
echo "Error: CELLXGENE_LOCATION not set"
exit 1
fi
if [ -z "$CELLXGENE_DATA" ] && [ -z "$CELLXGENE_BUCKET" ]; then
echo "Error: Either CELLXGENE_DATA or CELLXGENE_BUCKET must be set"
exit 1
fi
exec cellxgene-gateway

91
start_gunicorn.sh Executable file
View File

@@ -0,0 +1,91 @@
#!/bin/bash
# start_gunicorn.sh - Start Cellxgene Gateway with Gunicorn
#
# PREREQUISITES:
# - Gunicorn installed (included with cellxgene 1.3.0, or: pip install gunicorn)
# - Virtual environment activated
# - .env file with CELLXGENE_LOCATION and CELLXGENE_DATA (or CELLXGENE_BUCKET)
#
# USAGE:
# ./start_gunicorn.sh
# Exit on error
set -e
# Get the directory where this script is located
SCRIPT_DIR="$( cd "$( dirname "${BASH_SOURCE[0]}" )" && pwd )"
# Source environment variables
echo "Loading environment variables..."
if [ -f "$SCRIPT_DIR/.env" ]; then
source "$SCRIPT_DIR/.env"
else
echo "Error: .env file not found at $SCRIPT_DIR/.env"
echo "Please create it with required environment variables"
exit 1
fi
# Verify required environment variables
if [ -z "$CELLXGENE_LOCATION" ]; then
echo "Error: CELLXGENE_LOCATION not set"
exit 1
fi
if [ -z "$CELLXGENE_DATA" ] && [ -z "$CELLXGENE_BUCKET" ]; then
echo "Error: Either CELLXGENE_DATA or CELLXGENE_BUCKET must be set"
exit 1
fi
# Gunicorn configuration
# WARNING: Multi-worker mode has cache synchronization issues (see plans/002-shared-cache-implementation.md)
# Each worker maintains its own in-memory cache, causing 404s for static assets when different
# workers handle requests for the same dataset. Use GUNICORN_WORKERS=1 until shared cache is implemented.
WORKERS=${GUNICORN_WORKERS:-1}
BIND=${GATEWAY_IP:-0.0.0.0}:${GATEWAY_PORT:-5005}
TIMEOUT=${GUNICORN_TIMEOUT:-120}
WORKER_CLASS=${GUNICORN_WORKER_CLASS:-sync}
KEEPALIVE=${GUNICORN_KEEPALIVE:-5}
LOG_LEVEL=${GUNICORN_LOG_LEVEL:-info}
# Production optimization: enable backed mode to reduce memory usage
export GATEWAY_ENABLE_BACKED_MODE=${GATEWAY_ENABLE_BACKED_MODE:-true}
# Check if gunicorn is installed
if ! command -v gunicorn &> /dev/null; then
echo "Error: gunicorn not found. Install with: pip install gunicorn"
exit 1
fi
# Display configuration
echo "Starting Cellxgene Gateway with Gunicorn..."
echo "Configuration:"
echo " Data source: ${CELLXGENE_DATA:-$CELLXGENE_BUCKET}"
echo " Binding to: $BIND"
echo " Workers: $WORKERS"
echo " Worker class: $WORKER_CLASS"
echo " Timeout: ${TIMEOUT}s"
echo " Keepalive: ${KEEPALIVE}s"
echo " Log level: $LOG_LEVEL"
echo " Backed mode: ${GATEWAY_ENABLE_BACKED_MODE}"
echo ""
cd "$SCRIPT_DIR"
# Start Gunicorn with optimized settings
# Additional options you can add via environment variables:
# - GUNICORN_MAX_REQUESTS: Restart worker after N requests (prevents memory leaks)
# - GUNICORN_MAX_REQUESTS_JITTER: Add randomness to max-requests
exec gunicorn cellxgene_gateway.gateway:app \
--workers "$WORKERS" \
--worker-class "$WORKER_CLASS" \
--bind "$BIND" \
--timeout "$TIMEOUT" \
--keep-alive "$KEEPALIVE" \
--access-logfile - \
--error-logfile - \
--log-level "$LOG_LEVEL" \
--preload \
${GUNICORN_MAX_REQUESTS:+--max-requests "$GUNICORN_MAX_REQUESTS"} \
${GUNICORN_MAX_REQUESTS_JITTER:+--max-requests-jitter "$GUNICORN_MAX_REQUESTS_JITTER"} \
"$@"

89
start_uwsgi.sh Executable file
View File

@@ -0,0 +1,89 @@
#!/bin/bash
# start_uwsgi.sh - Start Cellxgene Gateway with uWSGI
#
# PREREQUISITES:
# - uWSGI installed (pip install uwsgi)
# - Virtual environment activated
# - .env file with CELLXGENE_LOCATION and CELLXGENE_DATA (or CELLXGENE_BUCKET)
#
# USAGE:
# ./start_uwsgi.sh
# Exit on error
set -e
# Get the directory where this script is located
SCRIPT_DIR="$( cd "$( dirname "${BASH_SOURCE[0]}" )" && pwd )"
# Source environment variables
echo "Loading environment variables..."
if [ -f "$SCRIPT_DIR/.env" ]; then
source "$SCRIPT_DIR/.env"
else
echo "Error: .env file not found at $SCRIPT_DIR/.env"
echo "Please create it with required environment variables"
exit 1
fi
# Verify required environment variables
if [ -z "$CELLXGENE_LOCATION" ]; then
echo "Error: CELLXGENE_LOCATION not set"
exit 1
fi
if [ -z "$CELLXGENE_DATA" ] && [ -z "$CELLXGENE_BUCKET" ]; then
echo "Error: Either CELLXGENE_DATA or CELLXGENE_BUCKET must be set"
exit 1
fi
# uWSGI configuration
# WARNING: Multi-worker mode has cache synchronization issues (see plans/002-shared-cache-implementation.md)
# Each worker maintains its own in-memory cache, causing 404s for static assets when different
# workers handle requests for the same dataset. Use UWSGI_WORKERS=1 until shared cache is implemented.
WORKERS=${UWSGI_WORKERS:-1}
HOST=${GATEWAY_IP:-0.0.0.0}
PORT=${GATEWAY_PORT:-5005}
TIMEOUT=${UWSGI_TIMEOUT:-120}
THREADS=${UWSGI_THREADS:-1}
# Production optimization: enable backed mode to reduce memory usage
export GATEWAY_ENABLE_BACKED_MODE=${GATEWAY_ENABLE_BACKED_MODE:-true}
# Check if uwsgi is installed
if ! command -v uwsgi &> /dev/null; then
echo "Error: uwsgi not found. Install with: pip install uwsgi"
exit 1
else
# Display configuration
echo "Starting Cellxgene Gateway with uWSGI..."
echo "Configuration:"
echo " Data source: ${CELLXGENE_DATA:-$CELLXGENE_BUCKET}"
echo " Binding to: $HOST:$PORT"
echo " Workers: $WORKERS"
echo " Threads: $THREADS"
echo " Timeout: ${TIMEOUT}s"
echo " Backed mode: ${GATEWAY_ENABLE_BACKED_MODE}"
echo ""
cd "$SCRIPT_DIR"
# Start uWSGI with optimized settings
# Additional options you can add via environment variables:
# - UWSGI_MAX_REQUESTS: Restart worker after N requests (prevents memory leaks)
exec uwsgi \
--http "$HOST:$PORT" \
--module cellxgene_gateway.gateway:app \
--workers "$WORKERS" \
--threads "$THREADS" \
--harakiri "$TIMEOUT" \
--master \
--enable-threads \
--single-interpreter \
--need-app \
--die-on-term \
--log-x-forwarded-for \
${UWSGI_MAX_REQUESTS:+--max-requests "$UWSGI_MAX_REQUESTS"} \
"$@"
fi

View File

@@ -1,13 +1,24 @@
import unittest
import os
import shutil
import tempfile
from unittest.mock import MagicMock, Mock, patch
from cellxgene_gateway.gateway import app
from cellxgene_gateway.items.item import ItemType
from cellxgene_gateway.items.s3.s3item import S3Item
from cellxgene_gateway.items.s3.s3item_source import S3ItemSource
from cellxgene_gateway.gateway import app
class TestScanDirectory(unittest.TestCase):
def setUp(self):
self.app = app
def tearDown(self):
pass
@patch("s3fs.S3FileSystem")
def test_GIVEN_invalid_bucket_THEN_throws_error(self, s3func):
class S3Mock:
@@ -26,6 +37,7 @@ class TestScanDirectory(unittest.TestCase):
@patch("s3fs.S3FileSystem")
def test__GIVEN_multilevel_bucket_THEN_properly_recurses_suburls(self, s3func):
class S3Mock:
def exists(path):
if path in [
@@ -82,7 +94,7 @@ class TestScanDirectory(unittest.TestCase):
s3func.return_value = S3Mock
source = S3ItemSource("my-bucket")
with app.test_request_context(query_string="refresh=true") as test_context:
with self.app.test_request_context(query_string="refresh=true") as test_context:
tree = source.scan_directory()
def s3item_compare(i1, i2, msg=""):

View File

@@ -1,14 +1,16 @@
import unittest
import tempfile
import os
import shutil
from flask import Flask
from cellxgene_gateway import flask_util
from cellxgene_gateway.cache_entry import CacheEntry, CacheEntryStatus
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.gateway import app
from cellxgene_gateway.items.item import ItemType
from cellxgene_gateway.items.file.fileitem import FileItem
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
from cellxgene_gateway.items.item import ItemType
from cellxgene_gateway.gateway import app
key = CacheKey(
FileItem("/czi/", name="pbmc3k.h5ad", type=ItemType.h5ad),
@@ -23,6 +25,9 @@ class TestRenderEntry(unittest.TestCase):
self.app_context.push()
self.client = self.app.test_client()
def tearDown(self):
self.app_context.pop()
def test_GIVEN_key_and_port_THEN_returns_loading_CacheEntry(self):
entry = CacheEntry.for_key("some-key", 1)
self.assertEqual(entry.status, CacheEntryStatus.loading)

View File

@@ -1,5 +1,9 @@
import os
import shutil
import tempfile
import unittest
from unittest.mock import MagicMock, patch
from collections import defaultdict
from unittest.mock import patch
from cellxgene_gateway.filecrawl import (
render_item,
@@ -9,30 +13,101 @@ from cellxgene_gateway.filecrawl import (
from cellxgene_gateway.items.file.fileitem import FileItem
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
from cellxgene_gateway.items.item import ItemTree, ItemType
from cellxgene_gateway.gateway import app
source = FileItemSource("/tmp")
def make_entry(subpath="somepath", annotations=None):
return FileItem(
subpath=subpath,
name="entry",
ext=".h5ad",
type=ItemType.h5ad,
annotations=annotations,
)
class TestRenderEntry(unittest.TestCase):
def setUp(self):
self.app = app
self.app_context = self.app.test_request_context()
self.app_context.push()
def tearDown(self):
self.app_context.pop()
def test_GIVEN_path_both_slash_THEN_view_has_single_slash(self):
entry = FileItem(subpath="/somepath/", name="entry", type=ItemType.h5ad)
entry = make_entry(subpath="/somepath/")
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry/'", rendered)
self.assertIn("view/somepath/entry.h5ad/'", rendered)
def test_GIVEN_path_starts_slash_THEN_view_has_single_slash(self):
entry = FileItem(subpath="/somepath", name="entry", type=ItemType.h5ad)
entry = make_entry(subpath="/somepath")
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry/'", rendered)
self.assertIn("view/somepath/entry.h5ad/'", rendered)
def test_GIVEN_path_ends_slash_THEN_view_has_single_slash(self):
entry = FileItem(subpath="somepath/", name="entry", type=ItemType.h5ad)
entry = make_entry(subpath="somepath/")
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry/'", rendered)
self.assertIn("view/somepath/entry.h5ad/'", rendered)
def test_GIVEN_path_no_slash_THEN_view_has_single_slash(self):
entry = FileItem(subpath="somepath", name="entry", type=ItemType.h5ad)
entry = make_entry(subpath="somepath")
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry/'", rendered)
self.assertIn("view/somepath/entry.h5ad/'", rendered)
class TestRenderAnnotation(unittest.TestCase):
def setUp(self):
self.app = app
self.app_context = self.app.test_request_context()
self.app_context.push()
def tearDown(self):
self.app_context.pop()
@patch("cellxgene_gateway.filecrawl.enable_annotations", new=True)
def test_GIVEN_no_annotation_THEN_new_alone(self):
entry = make_entry(annotations=None)
rendered = render_item(entry, source)
self.assertIn(
"> | annotations: <a class='new' href='/source/Files:/tmp/view/somepath/entry_annotations'>new</a></li>",
rendered,
)
@patch("cellxgene_gateway.filecrawl.enable_annotations", new=True)
def test_GIVEN_annotation_THEN_new_before(self):
annotation = FileItem(
subpath="somepath/entry_annotations",
name="annot",
ext=".csv",
type=ItemType.annotation,
)
entry = make_entry(annotations=[annotation])
rendered = render_item(entry, source)
self.assertIn(
"> | annotations: <a class='new' href='/source/Files:/tmp/view/somepath/entry_annotations'>new</a>,"
" <a href='/source/Files:/tmp/view/somepath/entry_annotations/annot.csv/'>annot</a></li>",
rendered,
)
@patch("cellxgene_gateway.filecrawl.enable_annotations", new=True)
def test_GIVEN_annotation_THEN_escaped(self):
annotation = FileItem(
subpath="somepath/entry_annotations",
name="hot&cold",
ext=".csv",
type=ItemType.annotation,
)
entry = make_entry(annotations=[annotation])
rendered = render_item(entry, source)
self.assertIn(
"> | annotations: <a class='new' href='/source/Files:/tmp/view/somepath/entry_annotations'>new</a>,"
" <a href='/source/Files:/tmp/view/somepath/entry_annotations/hot&cold.csv/'>hot&amp;cold</a></li>",
rendered,
)
class TestRenderItemSource(unittest.TestCase):
@@ -48,12 +123,48 @@ class TestRenderItemSource(unittest.TestCase):
class TestRenderItemTree(unittest.TestCase):
def setUp(self):
self.app = app
self.app_context = self.app.test_request_context()
self.app_context.push()
def tearDown(self):
self.app_context.pop()
@patch("cellxgene_gateway.items.file.fileitem_source.FileItemSource")
def test_GIVEN_deep_nested_dirs_THEN_includes_dirs_in_output(self, item_source):
item_source.name = "FakeSource"
item_tree = ItemTree("foo/bar/baz", [], [])
item_source.get_annotations_subpath = lambda _: "FakeAnnotations"
file_item = FileItem(
subpath="foo/bar/baz", name="file.h5ad", type=ItemType.h5ad
)
item_tree = ItemTree("foo/bar/baz", [file_item], [])
rendered = render_item_tree(item_tree, item_source)
self.assertEqual(
rendered,
"<li><a href='/filecrawl/foo/bar/baz?source=FakeSource'>baz</a><ul></ul></li>",
"<li><a href='/filecrawl/foo/bar/baz?source=FakeSource'>baz</a><ul>"
"<li> <a href='/source/FakeSource/view/foo/bar/baz/file.h5ad/'>file.h5ad</a>"
" </li></ul></li>",
)
@patch(
"os.listdir",
side_effect=lambda parent: defaultdict(
list, {"tmp": ["foo"], "tmp/foo": ["bar"]}
)[parent],
)
@patch("os.path.exists", return_value=True)
def test_GIVEN_dirs_without_h5ad_THEN_excludes_dirs_in_output(
self, listdir, exists
):
# Directories:
# - tmp
# - foo
# - bar (no h5ad files)
item_source = FileItemSource("tmp", name="local")
item_tree = item_source.list_items("foo")
rendered = render_item_tree(item_tree, item_source)
self.assertEqual(
rendered,
"<li><a href='/filecrawl/foo?source=local'>foo</a><ul></ul></li>",
)

View File

@@ -0,0 +1,54 @@
import json
import unittest
from types import SimpleNamespace
from cellxgene_gateway.gateway import do_GET_status_json, app, cache
from cellxgene_gateway.cache_entry import CacheEntry, CacheEntryStatus
class TestGatewayStatusJson(unittest.TestCase):
def test_do_GET_status_json_returns_expected_structure(self):
# Create a minimal fake key with required attributes
h5ad_item = SimpleNamespace(descriptor="somedir/dataset.h5ad")
key = SimpleNamespace(
h5ad_item=h5ad_item,
annotation_descriptor="somedir/dataset_annotations/foo.csv",
)
# Create a CacheEntry with known launchtime/timestamp/status
entry = CacheEntry(
None,
key,
8000,
111,
222,
CacheEntryStatus.loaded,
None,
None,
None,
None,
)
# Install into the gateway cache and set app launchtime
cache.entry_list = [entry]
app.extensions.setdefault("cellxgene_gateway", {})["launchtime"] = "LAUNCH_TIME"
rv = do_GET_status_json()
data = json.loads(rv)
# top-level launchtime comes from app.extensions
self.assertEqual("LAUNCH_TIME", data["launchtime"])
self.assertIn("entry_list", data)
self.assertEqual(1, len(data["entry_list"]))
e = data["entry_list"][0]
self.assertEqual("somedir/dataset.h5ad", e["dataset"])
self.assertEqual("somedir/dataset_annotations/foo.csv", e["annotation_file"])
self.assertEqual("loaded", e["status"])
self.assertEqual(111, e["launchtime"])
self.assertEqual(222, e["last_access"])
if __name__ == "__main__":
unittest.main()

View File

@@ -1,7 +1,6 @@
import unittest
from unittest.mock import MagicMock, patch
from cellxgene_gateway.backend_cache import BackendCache
from cellxgene_gateway.cache_entry import CacheEntry
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.items.file.fileitem import FileItem
@@ -33,10 +32,46 @@ class TestSubprocessBackend(unittest.TestCase):
backend.launch(cellxgene_loc, scripts, entry)
popen.assert_called_once_with(
[
"yes | /some/cellxgene launch /tmp/czi/pbmc3k.h5ad --port 8000 --host 127.0.0.1 --disable-annotations --scripts http://example.com/script.js --scripts http://example.com/script2.js"
"yes | /some/cellxgene launch /tmp/czi/pbmc3k.h5ad --port 8000 --host 127.0.0.1 --disable-annotations --disable-gene-sets-save --scripts http://example.com/script.js --scripts http://example.com/script2.js"
],
shell=True,
stderr=-1,
stdout=-1,
)
self.assertEqual("An unexpected error", context.exception.stderr)
@patch("subprocess.Popen")
def test_launch_GIVEN_annotations_enabled_THEN_set_flags(self, popen):
subprocess = MagicMock()
subprocess.stdout.readline().decode.return_value = (
"[cellxgene] Type CTRL-C at any time to exit.\n"
)
subprocess.stderr.read().decode.return_value = ""
popen.return_value = subprocess
key = CacheKey(
FileItem("/czi/", name="pbmc3k.h5ad", type=ItemType.h5ad),
FileItemSource("/tmp", "local"),
FileItem(
"/czi/pbmc3k_annotations/", name="foo.csv", type=ItemType.annotation
),
)
entry = CacheEntry.for_key(key, 8000)
import cellxgene_gateway.subprocess_backend
cellxgene_gateway.subprocess_backend.enable_annotations = True
try:
backend = cellxgene_gateway.subprocess_backend.SubprocessBackend()
cellxgene_loc = "/some/cellxgene"
backend.launch(cellxgene_loc, [], entry)
finally:
cellxgene_gateway.subprocess_backend.enable_annotations = False
popen.assert_called_once_with(
[
"yes | /some/cellxgene launch /tmp/czi/pbmc3k.h5ad --port 8000 --host 127.0.0.1 --annotations-file /tmp/czi/pbmc3k_annotations/foo.csv --gene-sets-file /tmp/czi/pbmc3k_annotations/foo_gene_sets.csv"
],
shell=True,
stderr=-1,
stdout=-1,
)