32 Commits

Author SHA1 Message Date
Alok Saldanha
10105c43a8 added support for code coverage 2021-04-04 14:55:13 -04:00
Alokito
638923fb43 Merge pull request #44 from Novartis/itemsource
Itemsource
2021-04-03 07:15:35 -04:00
Alok Saldanha
169da934b1 updated annotation.js comment 2021-04-03 07:12:50 -04:00
Alok Saldanha
33331c9198 Template has correct relaunch_url. 2021-04-03 07:08:20 -04:00
Alok Saldanha
d2d22cecaa updated tests, formatting 2021-04-02 11:52:49 -04:00
Alok Saldanha
5c488d8d5e only include source path element in url when multiple sources 2021-04-02 11:40:11 -04:00
Alok Saldanha
b1ae9e8ce8 ensure s3 bucket does not include trailing slash 2021-04-02 11:37:11 -04:00
Alok Saldanha
76c1d9e80c removed .csv suffix from annotation name 2021-03-29 07:28:16 -04:00
Alok Saldanha
586dc6623a ensure annotation directory for new annotations 2021-03-29 07:03:05 -04:00
Alok Saldanha
76f3939fdd rebased "Introduction of ItemSource interface" patch 2021-03-28 22:05:08 -04:00
Alok Saldanha
5404a8b0f3 removed trailing slash from render_annotations 2021-03-28 21:59:03 -04:00
Alok Saldanha
6f2b372d32 Preparing for 0.2.3 release 2021-02-28 09:55:26 -05:00
Alokito
25755d3f69 Merge pull request #41 from Novartis/grst_master
Grst master
2021-02-22 08:59:05 -05:00
Alok Saldanha
893092f199 fixed unit tests 2021-02-13 17:30:21 -05:00
Alok Saldanha
d3ee04b6e5 blacken 2021-02-13 17:02:55 -05:00
Alok Saldanha
f1f4b0c0ca added environment variables for ProxyFix 2021-02-13 16:42:34 -05:00
Alok Saldanha
01dccef7db added back trailing slash to url_for change 2021-02-13 16:41:32 -05:00
Gregor Sturm
9c17e2ff32 Fix redirect in cache_entry 2021-01-18 22:08:25 +01:00
Gregor Sturm
fbc18fb636 Add ProxyFix to gateway.py 2021-01-18 20:01:53 +01:00
Gregor Sturm
8c5a635de9 Use url_for in all templates 2021-01-18 19:47:18 +01:00
Gregor Sturm
94062c2d64 apply proxy fix 2021-01-18 18:19:25 +01:00
Gregor Sturm
068e8f7633 Fix url_for 2021-01-18 18:09:45 +01:00
Gregor Sturm
16b54f9409 Use url_for to generate URLs 2021-01-18 16:59:15 +01:00
Alok Saldanha
942410bb44 added instructions on pre-commit installation to README.md 2020-12-31 17:12:11 -05:00
Alokito
0d084a405e Merge pull request #38 from Novartis/feature/action_push
Feature/action push
2020-12-31 16:33:30 -05:00
Alok Saldanha
49d679e779 try evaling the bash hook
per https://github.com/conda/conda/issues/7980
2020-12-31 13:23:25 -05:00
Alok Saldanha
5e6faa4b02 run pr checks on push 2020-12-31 13:06:56 -05:00
Alokito
9fe846c786 Merge pull request #37 from ericmjl/master
Migrated PR checks to GitHub actions
2020-12-28 11:38:26 -05:00
Eric Ma
1c7907aabd Change file extension 2020-12-27 21:26:32 -05:00
Eric Ma
30d2b07b1a Migrated PR checks to GitHub actions 2020-12-27 21:10:07 -05:00
Alokito
21ff56ea8b Merge pull request #35 from dfeinzeig/fix/redirect
add missing trailing slash in effort to avoid whatever is redirecting
2020-10-28 07:39:06 -04:00
David Feinzeig
0b866a46ec add missing trailing slash in effort to avoid whatever is redirecting 2020-10-23 17:41:05 -04:00
39 changed files with 1099 additions and 576 deletions

13
.coveragerc Normal file
View File

@@ -0,0 +1,13 @@
[run]
branch = True
source = cellxgene_gateway
[report]
exclude_lines =
if self.debug:
pragma: no cover
raise NotImplementedError
if __name__ == .__main__.:
ignore_errors = True
omit =
tests/*

60
.github/workflows/pr-checks.yaml vendored Normal file
View File

@@ -0,0 +1,60 @@
# Tests that run on every PR
name: Pull Request Checks
on: [push, pull_request]
jobs:
black:
runs-on: ubuntu-18.04
steps:
- uses: actions/checkout@v2
name: Checkout repository
- uses: actions/setup-python@v2
name: Setup Python
with:
python-version: 3.9
- name: Install black
run: |
python -m pip install --upgrade pip
pip install black
- name: Run black
run: |
black . --check
# This job is copied over from `deploy.yaml`
run-tests:
runs-on: ubuntu-18.04
steps:
- uses: actions/checkout@v2
# See: https://github.com/marketplace/actions/setup-conda
- uses: s-weigand/setup-conda@v1
with:
conda-channels: "conda-forge"
- name: Build environment
run: |
conda env create -f environment.yml
eval "$(conda shell.bash hook)"
conda activate cellxgene-gateway
python setup.py install
- name: Run tests
run: |
eval "$(conda shell.bash hook)"
conda activate cellxgene-gateway
coverage run -m unittest discover tests
- name: Check coverage
run: |
eval "$(conda shell.bash hook)"
conda activate cellxgene-gateway
coverage report --fail-under 41
coverage xml -i
- name: "Upload coverage to Codecov"
uses: codecov/codecov-action@v1
with:
fail_ci_if_error: true

1
.gitignore vendored
View File

@@ -55,6 +55,7 @@ htmlcov/
.nox/
.coverage
.coverage.*
htmlcov
.cache
nosetests.xml
coverage.xml

View File

@@ -7,12 +7,6 @@ repos:
language: system
types: [python]
stages: [commit]
- id: flake8
name: flake8
language: system
entry: flake8
types: [python]
stages: [commit]
- id: black
language_version: python3.6+
name: black

View File

@@ -1,3 +1,7 @@
# 0.2.3
* Added support for ProxyFix
# 0.2.2
* Fixed bug with annotations (missing annotation.js asset)

View File

@@ -59,7 +59,12 @@ cellxgene-gateway
Here's what the environment variables mean:
* `CELLXGENE_LOCATION` - the location of the cellxgene executable, e.g. `~/anaconda2/envs/cellxgene/bin/cellxgene`
At least one of the following is required:
* `CELLXGENE_DATA` - a directory that can contain subdirectories with `.h5ad` data files, *without* trailing slash, e.g. `/mnt/cellxgene_data`
* `CELLXGENE_BUCKET` - an s3 bucket that can contain keys with `.h5ad` data files, e.g. `my-cellxgene-data-bucket`
Cellxgene Gateway is designed to make it easy to add additional data sources, please see the source code for gateway.py and the ItemSource interface in items/item_source.py
Optional environment variables:
* `CELLXGENE_ARGS` - catch-all variable that can be used to pass additional command line args to cellxgene server
* `EXTERNAL_HOST` - the hostname and port from the perspective of the web browser, typically `localhost:5005` if running locally. Defaults to "localhost:{GATEWAY_PORT}"
@@ -67,9 +72,15 @@ Optional environment variables:
* `GATEWAY_IP` - ip addess of instance gateway is running on, mostly used to display SSH instructions. Defaults to `socket.gethostbyname(socket.gethostname())`
* `GATEWAY_PORT` - local port that the gateway should bind to, defaults to 5005
* `GATEWAY_EXTRA_SCRIPTS` - JSON array of script paths, will be embedded into each page and forwarded with `--scripts` to cellxgene server
* `GATEWAY_ENABLE_UPLOAD` - Set to `true` or `1` to enable HTTP uploads. This is not recommended for a public server.
* `GATEWAY_ENABLE_ANNOTATIONS` - Set to `true` or to `1` to enable cellxgene annotations.
* `GATEWAY_ENABLE_BACKED_MODE` - Set to `true` or to `1` to load AnnData in file-backed mode. This saves memory and speeds up launch time but may reduce overall performance.
* `GATEWAY_ENABLE_BACKED_MODE` - Set to `true` or to `1` to load AnnData in file-backed mode. This saves memory and speeds up launch time but may reduce overall performance.
If any of the following optional variables are set, [ProxyFix](https://werkzeug.palletsprojects.com/en/1.0.x/middleware/proxy_fix/) will be used.
* `PROXY_FIX_FOR` - Number of upstream proxies setting X-Forwarded-For
* `PROXY_FIX_PROTO` - Number of upstream proxies setting X-Forwarded-Proto
* `PROXY_FIX_HOST` - Number of upstream proxies setting X-Forwarded-Host
* `PROXY_FIX_PORT` - Number of upstream proxies setting X-Forwarded-Port
* `PROXY_FIX_PREFIX` - Number of upstream proxies setting X-Forwarded-Prefix
The defaults should be fine if you set up a venv and cellxgene_data folder as above.
@@ -113,6 +124,14 @@ python setup.py develop
For convenience, the code repo includes a `run.sh.example` shell script to run the gateway.
4. Install pre-commit hooks
```bash
conda install -c conda-forge pre-commit
pre-commit install
```
## Running Tests
[![Build Status](https://travis-ci.org/Novartis/cellxgene-gateway.svg?branch=master)](https://travis-ci.org/Novartis/cellxgene-gateway)
@@ -121,14 +140,19 @@ For convenience, the code repo includes a `run.sh.example` shell script to run t
python -m unittest discover tests
```
## Code Coverage
```bash
coverage run -m unittest discover tests
coverage html
```
## Running Linters
pip install isort flake8 black
```bash
isort -rc .
flake8 .
black -l 79 .
isort -rc . # rc means recursive, and was deprecated in dev version of isort
black .
```
# Getting Help

2
cellxgene_gateway/__init__.py Executable file → Normal file
View File

@@ -7,4 +7,4 @@
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
__version__ = "0.2.2"
__version__ = "0.2.3"

View File

@@ -9,11 +9,13 @@
import time
from threading import Thread
from typing import List
from flask_api import status
from cellxgene_gateway import env
from cellxgene_gateway.cache_entry import CacheEntry, CacheEntryStatus
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.cellxgene_exception import CellxgeneException
from cellxgene_gateway.subprocess_backend import SubprocessBackend
@@ -35,13 +37,13 @@ class BackendCache:
contents = self.entry_list
return [c.port for c in contents]
def check_entry(self, key):
def check_path(self, source, path):
contents = self.entry_list
matches = [
c
for c in contents
if c.key.dataset == key.dataset
and c.key.annotation_file == key.annotation_file
if c.key.source.name == source.name
and path.startswith(c.key.descriptor)
and c.status != CacheEntryStatus.terminated
]
@@ -52,10 +54,28 @@ class BackendCache:
else:
raise CellxgeneException(
status.HTTP_500_INTERNAL_SERVER_ERROR,
"Found " + str(len(matches)) + " for " + dataset,
"Found " + str(len(matches)) + " for " + path,
)
def create_entry(self, key, scripts):
def check_entry(self, key):
contents = self.entry_list
matches = [
c
for c in contents
if c.key.equals(key) and c.status != CacheEntryStatus.terminated
]
if len(matches) == 0:
return None
elif len(matches) == 1:
return matches[0]
else:
raise CellxgeneException(
status.HTTP_500_INTERNAL_SERVER_ERROR,
"Found " + str(len(matches)) + " for " + key.dataset,
)
def create_entry(self, key: CacheKey, scripts: List[str]):
port = 8000
existing_ports = self.get_ports()

View File

@@ -8,12 +8,14 @@
# the specific language governing permissions and limitations under the License.
import datetime
import logging
import re
import urllib.parse
from enum import Enum
import psutil
from enum import Enum
from flask import make_response, render_template, request
from flask.wrappers import Response
from requests import get, post, put
import re
from cellxgene_gateway import env
from cellxgene_gateway.cellxgene_exception import CellxgeneException
@@ -69,6 +71,10 @@ class CacheEntry:
None,
)
@property
def source_name(self):
return self.key.source_name
def set_loaded(self, pid):
self.pid = pid
self.status = CacheEntryStatus.loaded
@@ -98,12 +104,14 @@ class CacheEntry:
for child in children:
child.terminate()
psutil.wait_procs(children, callback=on_terminate)
terminated.append(p.pid)
p.terminate()
psutil.wait_procs([p], callback=on_terminate)
logging.getLogger("cellxgene_gateway").info(
f"terminated {terminated}"
)
# the parent process may automatically die once its children have --
try:
p.terminate()
psutil.wait_procs([p], callback=on_terminate)
except psutil.NoSuchProcess:
pass
logging.getLogger("cellxgene_gateway").info(f"terminated {terminated}")
self.status = CacheEntryStatus.terminated
def rewrite_text_content(self, cellxgene_content):
@@ -111,26 +119,22 @@ class CacheEntry:
gateway_content = (
re.sub(
'(="|\()/static/',
f"\\1{self.gateway_basepath()}static/",
f"\\1{self.key.gateway_basepath()}static/",
cellxgene_content,
)
.replace("http://fonts.gstatic.com", "https://fonts.gstatic.com")
.replace(self.cellxgene_basepath(), self.gateway_basepath())
.replace(self.cellxgene_basepath(), self.key.gateway_basepath())
)
return gateway_content
def gateway_basepath(self):
return f"{env.external_protocol}://{env.external_host}/view/{self.key.pathpart}/"
def cellxgene_basepath(self):
return f"http://127.0.0.1:{self.port}"
def serve_content(self, path):
gateway_basepath = self.gateway_basepath()
subpath = path[len(self.key.pathpart) :] # noqa: E203
gateway_basepath = self.key.gateway_basepath()
subpath = path[len(self.key.descriptor) :] # noqa: E203
if len(subpath) == 0:
r = make_response(f"Redirect to {gateway_basepath}\n", 301)
r = make_response(f"Redirect to {gateway_basepath}\n", 302)
r.headers["location"] = gateway_basepath + querystring()
return r
elif self.status == CacheEntryStatus.loading:
@@ -180,9 +184,7 @@ class CacheEntry:
data=request.data,
)
else:
raise CellxgeneException(
f"Unexpected method {request.method}", 400
)
raise CellxgeneException(f"Unexpected method {request.method}", 400)
content_type = cellxgene_response.headers["content-type"]
if "text" in content_type:
gateway_content = self.rewrite_text_content(
@@ -201,5 +203,4 @@ class CacheEntry:
cellxgene_response.status_code,
resp_headers,
)
return gateway_response

View File

@@ -9,15 +9,73 @@
# There are three kinds of CacheKey:
# 1) somedir/dataset.h5ad: a dataset
# in this case, pathpart == dataset == 'somedir/dataset.h5ad'
# 2) somedir/dataset_annotations/my_annotations.csv : an actual annotaitons file.
# in this case, pathpart == 'dataset_annotations/my_annotations.csv', dataset == 'somedir/dataset.h5ad'
# in this case, descriptor == dataset == 'somedir/dataset.h5ad'
# 2) somedir/dataset_annotations/my_annotations.csv : an actual annotations file.
# in this case, descriptor == 'somedir/dataset_annotations/my_annotations.csv', dataset == 'somedir/dataset.h5ad'
# 3) somedir/dataset_annotations: an annotation directory. The corresponding h5ad must exist, but the directory may not.
# in this case, pathpart == 'dataset_annotations', dataset == 'somedir/dataset.h5ad'
# in this case, descriptor == 'somedir/dataset_annotations', dataset == 'somedir/dataset.h5ad'
from cellxgene_gateway import flask_util
from cellxgene_gateway.items.item import Item
from cellxgene_gateway.items.item_source import ItemSource, LookupResult
class CacheKey:
def __init__(self, pathpart, dataset, annotation_file):
self.pathpart = pathpart
self.dataset = dataset
self.annotation_file = annotation_file
@property
def descriptor(self):
if self.annotation_item is None:
return self.h5ad_item.descriptor
else:
return self.annotation_item.descriptor
@property
def file_path(self):
return self.source.get_local_path(self.h5ad_item)
@property
def annotation_file_path(self):
if self.annotation_item is None:
return None
else:
return self.source.get_local_path(self.annotation_item)
def relaunch_url(self):
return flask_util.relaunch_url(self.descriptor, self.source_name)
def gateway_basepath(self):
return self.view_url + "/"
@property
def view_url(self):
return flask_util.view_url(self.descriptor, self.source_name)
@property
def source_name(self):
return self.source.name
@property
def annotation_descriptor(self):
if self.annotation_item is None:
return None
else:
return self.annotation_item.descriptor
def equals(self, other):
return (
(self.source.name == other.source.name)
and (self.h5ad_item.descriptor == other.h5ad_item.descriptor)
and (self.annotation_descriptor == other.annotation_descriptor)
)
def __init__(
self, h5ad_item: Item, source: ItemSource, annotation_item: Item = None
):
assert h5ad_item is not None
assert source is not None
self.h5ad_item = h5ad_item
self.annotation_item = annotation_item
self.source = source
@classmethod
def for_lookup(cls, source: ItemSource, lookup: LookupResult):
return CacheKey(lookup.h5ad_item, source, lookup.annotation_item)

View File

@@ -53,11 +53,17 @@ def create_dir(parent_path, dir_name):
annotations_suffix = "_annotations"
h5ad_suffix = ".h5ad"
def make_h5ad(el):
return el[: -len(annotations_suffix)] + ".h5ad"
return el[: -len(annotations_suffix)] + h5ad_suffix
def make_annotations(el):
return el[:-5] + annotations_suffix
def ensure_dir_exists(file_path):
if not os.path.exists(file_path):
os.makedirs(file_path)

View File

@@ -12,7 +12,7 @@ import os
import socket
cellxgene_location = os.environ.get("CELLXGENE_LOCATION")
cellxgene_data = os.environ.get("CELLXGENE_DATA")
cellxgene_data = os.environ.get("CELLXGENE_DATA", "")
cellxgene_args = os.environ.get("CELLXGENE_ARGS", None)
gateway_port = int(os.environ.get("GATEWAY_PORT", "5005"))
external_host = os.environ.get(
@@ -22,31 +22,28 @@ external_host = os.environ.get(
external_protocol = os.environ.get(
"EXTERNAL_PROTOCOL", os.environ.get("GATEWAY_PROTOCOL", "http")
)
ip = os.environ.get("GATEWAY_IP", "127.0.0.1")
ip = os.environ.get("GATEWAY_IP")
extra_scripts = os.environ.get("GATEWAY_EXTRA_SCRIPTS")
ttl = os.environ.get("GATEWAY_TTL")
enable_upload = os.environ.get("GATEWAY_ENABLE_UPLOAD", "").lower() in [
enable_annotations = os.environ.get("GATEWAY_ENABLE_ANNOTATIONS", "").lower() in [
"true",
"1",
]
enable_annotations = os.environ.get(
"GATEWAY_ENABLE_ANNOTATIONS", ""
).lower() in [
"true",
"1",
]
enable_backed_mode = os.environ.get(
"GATEWAY_ENABLE_BACKED_MODE", ""
).lower() in [
enable_backed_mode = os.environ.get("GATEWAY_ENABLE_BACKED_MODE", "").lower() in [
"true",
"1",
]
env_vars = {
"CELLXGENE_LOCATION": cellxgene_location,
"CELLXGENE_DATA": cellxgene_data,
}
proxy_fix_for = int(os.environ.get("PROXY_FIX_FOR", "0"))
proxy_fix_proto = int(os.environ.get("PROXY_FIX_PROTO", "0"))
proxy_fix_host = int(os.environ.get("PROXY_FIX_HOST", "0"))
proxy_fix_port = int(os.environ.get("PROXY_FIX_PORT", "0"))
proxy_fix_prefix = int(os.environ.get("PROXY_FIX_PREFIX", "0"))
optional_env_vars = {
"EXTERNAL_HOST": external_host,
"EXTERNAL_PROTOCOL": external_protocol,
@@ -54,10 +51,15 @@ optional_env_vars = {
"GATEWAY_PORT": gateway_port,
"GATEWAY_EXTRA_SCRIPTS": extra_scripts,
"GATEWAY_TTL": ttl,
"GATEWAY_ENABLE_UPLOAD": enable_upload,
"GATEWAY_ENABLE_ANNOTATIONS": enable_annotations,
"GATEWAY_ENABLE_BACKED_MODE": enable_backed_mode,
"CELLXGENE_ARGS": cellxgene_args,
"CELLXGENE_DATA": cellxgene_data,
"PROXY_FIX_FOR": proxy_fix_for,
"PROXY_FIX_PROTO": proxy_fix_proto,
"PROXY_FIX_HOST": proxy_fix_host,
"PROXY_FIX_PORT": proxy_fix_port,
"PROXY_FIX_PREFIX": proxy_fix_prefix,
}

View File

@@ -8,113 +8,59 @@
# the specific language governing permissions and limitations under the License.
import os
from cellxgene_gateway import env
from cellxgene_gateway.dir_util import (
make_h5ad,
make_annotations,
annotations_suffix,
)
import urllib.parse
from cellxgene_gateway import env, flask_util
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.dir_util import annotations_suffix, make_annotations, make_h5ad
def recurse_dir(path):
if not os.path.exists(path):
raise CellxgeneException(
"The given path does not exist.", status.HTTP_400_BAD_REQUEST
)
all_entries = sorted(os.listdir(path))
def is_h5ad(el):
return el.endswith(".h5ad") and os.path.isfile(os.path.join(path, el))
h5ad_entries = [x for x in all_entries if is_h5ad(x)]
annotation_dir_entries = [
x
for x in all_entries
if x.endswith(annotations_suffix) and make_h5ad(x) in h5ad_entries
]
def list_annotations(el):
full_path = os.path.join(path, el)
if not os.path.isdir(full_path):
entries = []
else:
entries = [
{
"name": x[:-13]
if (len(x) > 13 and x[-13] in ["-", "_"])
else (x[:-4] if x.endswith(".csv") else x),
"path": os.path.join(full_path, x).replace(
env.cellxgene_data, ""
),
}
for x in sorted(os.listdir(full_path))
if x.endswith(".csv")
and os.path.isfile(os.path.join(full_path, x))
]
return [
{
"name": "new",
"class": "new",
"path": full_path.replace(env.cellxgene_data, ""),
}
] + entries
def make_entry(el):
full_path = os.path.join(path, el)
if el in h5ad_entries:
return {
"path": full_path.replace(env.cellxgene_data, ""),
"name": el,
"type": "file",
"annotations": list_annotations(make_annotations(el)),
}
elif os.path.isdir(full_path) and el not in annotation_dir_entries:
return {
"path": full_path.replace(env.cellxgene_data, ""),
"name": el,
"type": "directory",
"children": recurse_dir(full_path),
}
else:
return {
"path": full_path,
"name": el,
"type": "neither",
}
return [make_entry(x) for x in all_entries]
def render_entries(entries):
return "<ul>" + "\n".join([render_entry(e) for e in entries]) + "</ul>"
def get_url(entry):
return f"/view/{ entry['path'].lstrip('/') }"
def get_class(entry):
return f" class='{entry['class']}'" if "class" in entry else ""
def render_annotations(entry):
if len(entry["annotations"]) > 0:
return " | annotations: " + ", ".join(
def render_annotations(item, item_source):
url = flask_util.view_url(
item_source.get_annotations_subpath(item), item_source.name
)
new_annotation = f"<a class='new' href='{url}'>new</a>"
annotations = (
", ".join(
[
f"<a href='{get_url(a)}'{get_class(a)}>{a['name']}</a>"
for a in entry["annotations"]
f"<a href='{CacheKey(item, item_source, a).view_url}/'>{a.name}</a>"
for a in item.annotations
]
)
else:
return ""
+ ", "
if item.annotations
else ""
)
return " | annotations: " + annotations + new_annotation
def render_entry(entry):
if entry["type"] == "file":
return f"<li> <a href='{ get_url(entry) }'>{entry['name']}</a> {render_annotations(entry)}</li>"
elif entry["type"] == "directory":
url = f"/filecrawl/{entry['path'].lstrip('/')}"
return f"<li><a href='{url}'>{entry['name']}</a>{render_entries(entry['children'])}</li>"
def render_item(item, item_source):
item_string = f"<li> <a href='{ CacheKey(item, item_source).view_url }/'>{item.name}</a> {render_annotations(item, item_source)}</li>"
return item_string
def render_item_tree(item_tree, item_source):
items = (
"\n".join([render_item(i, item_source) for i in item_tree.items])
if item_tree.items
else ""
)
branches = (
"\n".join([render_item_tree(b, item_source) for b in item_tree.branches])
if item_tree.branches
else ""
)
html = "<ul>" + items + branches + "</ul>"
if item_tree.descriptor:
descriptor = item_tree.descriptor.lstrip("/")
url = f"/filecrawl/{descriptor}?source={item_source.name}"
name = descriptor.rsplit("/")[1] if descriptor.find("/") >= 0 else descriptor
return f"<li><a href='{url}'>{name}</a>{html}</li>"
else:
return ""
return html
def render_item_source(item_source, filter=None):
item_tree = item_source.list_items(filter)
heading = f"<h6><a href='/filecrawl.html?source={urllib.parse.quote_plus(item_source.name)}'>{item_source.name}</a></h6>"
return heading + render_item_tree(item_tree, item_source)

View File

@@ -7,9 +7,27 @@
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
from flask import request
from flask import request, url_for
def querystring():
qs = request.query_string.decode()
return f"?{qs}" if len(qs) > 0 else ""
include_source_in_url = False
def url(endpoint, descriptor, source_name):
if include_source_in_url:
return url_for(endpoint, source_name=source_name, path=descriptor)
else:
return url_for(endpoint, path=descriptor)
def view_url(descriptor, source_name):
return url("do_view", descriptor, source_name)
def relaunch_url(descriptor, source_name):
return url("do_relaunch", descriptor, source_name)

View File

@@ -6,12 +6,11 @@
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
# import BaseHTTPServer
import json
import logging
# import BaseHTTPServer
import os
import urllib.parse
from threading import Lock, Thread
from flask import (
@@ -24,22 +23,26 @@ from flask import (
url_for,
)
from flask_api import status
from werkzeug.middleware.proxy_fix import ProxyFix
from werkzeug.utils import secure_filename
from cellxgene_gateway import env
from cellxgene_gateway import env, flask_util
from cellxgene_gateway.backend_cache import BackendCache
from cellxgene_gateway.cache_entry import CacheEntryStatus
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.cellxgene_exception import CellxgeneException
from cellxgene_gateway.dir_util import create_dir, is_subdir
from cellxgene_gateway.extra_scripts import get_extra_scripts
from cellxgene_gateway.filecrawl import recurse_dir, render_entries
from cellxgene_gateway.path_util import get_key
from cellxgene_gateway.filecrawl import render_item_source
from cellxgene_gateway.process_exception import ProcessException
from cellxgene_gateway.prune_process_cache import PruneProcessCache
from cellxgene_gateway.util import current_time_stamp
app = Flask(__name__)
item_sources = []
default_item_source = None
def _force_https(app):
def wrapper(environ, start_response):
@@ -50,9 +53,23 @@ def _force_https(app):
app.wsgi_app = _force_https(app.wsgi_app)
if (
env.proxy_fix_for > 0
or env.proxy_fix_proto > 0
or env.proxy_fix_host > 0
or env.proxy_fix_port > 0
or env.proxy_fix_prefix > 0
):
app.wsgi_app = ProxyFix(
app.wsgi_app,
x_for=env.proxy_fix_for,
x_proto=env.proxy_fix_proto,
x_host=env.proxy_fix_host,
x_port=env.proxy_fix_port,
x_prefix=env.proxy_fix_prefix,
)
cache = BackendCache()
location = f"{env.external_protocol}://{env.external_host}"
@app.errorhandler(CellxgeneException)
@@ -88,8 +105,8 @@ def handle_invalid_process(error):
http_status=error.http_status,
stdout=error.stdout,
stderr=error.stderr,
dataset=error.key.dataset,
annotation_file=error.key.annotation_file,
relaunch_url=error.key.relaunch_url(),
annotation_file=error.key.annotation_descriptor,
),
error.http_status,
)
@@ -106,84 +123,40 @@ def favicon():
@app.route("/")
def index():
users = [
name
for name in os.listdir(env.cellxgene_data)
if os.path.isdir(os.path.join(env.cellxgene_data, name))
]
return render_template(
"index.html",
ip=env.ip,
cellxgene_data=env.cellxgene_data,
extra_scripts=get_extra_scripts(),
users=users,
enable_upload=env.enable_upload,
)
def make_user():
dir_name = request.form["directory"]
create_dir(env.cellxgene_data, dir_name)
return redirect(location, code=302)
def make_subdir():
parent_path = os.path.join(env.cellxgene_data, request.form["usernames"])
dir_name = request.form["directory"]
create_dir(parent_path, dir_name)
return redirect(location, code=302)
def upload_file():
upload_dir = request.form["path"]
full_upload_path = os.path.join(env.cellxgene_data, upload_dir)
if is_subdir(full_upload_path, env.cellxgene_data) and os.path.isdir(
full_upload_path
):
if request.method == "POST":
if "file" in request.files:
f = request.files["file"]
if f and f.filename.endswith(".h5ad"):
f.save(
os.path.join(
full_upload_path, secure_filename(f.filename)
)
)
return redirect("/filecrawl.html", code=302)
else:
raise CellxgeneException(
"Uploaded file must be in anndata (.h5ad) format.",
status.HTTP_400_BAD_REQUEST,
)
else:
raise CellxgeneException(
"A file must be chosen to upload.",
status.HTTP_400_BAD_REQUEST,
)
else:
raise CellxgeneException(
"Invalid directory.", status.HTTP_400_BAD_REQUEST
@app.route("/filecrawl.html")
@app.route("/filecrawl/<path:path>")
def filecrawl(path=None):
source_name = request.args.get("source")
sources = (
filter(
lambda x: x.name == urllib.parse.unquote_plus(source_name),
item_sources,
)
return redirect(location, code=302)
if env.enable_upload:
app.add_url_rule("/make_user", "make_user", make_user, methods=["POST"])
app.add_url_rule(
"/make_subdir", "make_subdir", make_subdir, methods=["POST"]
if source_name
else item_sources
)
app.add_url_rule(
"/upload_file", "upload_file", upload_file, methods=["POST"]
# loop all data sources --
rendered_sources = [
render_item_source(item_source, path) for item_source in sources
] # will we need to make this async in the page???
rendered_html = "\n".join(rendered_sources)
resp = make_response(
render_template(
"filecrawl.html",
extra_scripts=get_extra_scripts(),
rendered_html=rendered_html,
path=path,
)
)
def set_no_cache(resp):
resp.headers["Cache-Control"] = "no-cache, no-store, must-revalidate"
resp.headers["Pragma"] = "no-cache"
resp.headers["Expires"] = "0"
@@ -191,52 +164,44 @@ def set_no_cache(resp):
return resp
@app.route("/filecrawl.html")
def filecrawl():
entries = recurse_dir(env.cellxgene_data)
rendered_html = render_entries(entries)
resp = make_response(
render_template(
"filecrawl.html",
extra_scripts=get_extra_scripts(),
rendered_html=rendered_html,
)
)
return set_no_cache(resp)
@app.route("/filecrawl/<path:path>")
def do_filecrawl(path):
filecrawl_path = os.path.join(env.cellxgene_data, path)
if not os.path.isdir(filecrawl_path):
raise CellxgeneException(
"Path is not directory: " + filecrawl_path,
status.HTTP_400_BAD_REQUEST,
)
entries = recurse_dir(filecrawl_path)
rendered_html = render_entries(entries)
return render_template(
"filecrawl.html",
extra_scripts=get_extra_scripts(),
rendered_html=rendered_html,
path=path,
)
entry_lock = Lock()
def matching_source(source_name):
if source_name is None:
source_name = default_item_source.name
matching = [i for i in item_sources if i.name == source_name]
if len(matching) != 1:
raise Exception(f"Could not find matching item source {source_name}")
source = matching[0]
return source
@app.route(
"/source/<path:source_name>/view/<path:path>",
methods=["GET", "PUT", "POST"],
)
@app.route("/view/<path:path>", methods=["GET", "PUT", "POST"])
def do_view(path):
key = get_key(path)
print(
f"view path={path}, dataset={key.dataset}, annotation_file= {key.annotation_file}, key={key.pathpart}"
)
with entry_lock:
match = cache.check_entry(key)
if match is None:
uascripts = get_extra_scripts()
match = cache.create_entry(key, uascripts)
def do_view(path, source_name=None):
source = matching_source(source_name)
match = cache.check_path(source, path)
if match is None:
lookup = source.lookup(path)
if lookup is None:
raise CellxgeneException(
f"Could not find item for path {path} in source {source.name}",
404,
)
key = CacheKey.for_lookup(source, lookup)
print(
f"view path={path}, source_name={source_name}, dataset={key.file_path}, annotation_file= {key.annotation_file_path}, key={key.descriptor}, source={key.source_name}"
)
with entry_lock:
match = cache.check_entry(key)
if match is None:
uascripts = get_extra_scripts()
match = cache.create_entry(key, uascripts)
match.timestamp = current_time_stamp()
@@ -275,38 +240,38 @@ def do_GET_status_json():
@app.route("/relaunch/<path:path>", methods=["GET"])
def do_relaunch(path):
key = get_key(path)
source_name = request.args.get("source_name") or default_item_source.name
source = matching_source(source_name)
key = CacheKey.for_lookup(source, source.lookup(path))
match = cache.check_entry(key)
if not match is None:
match.terminate()
qs = request.query_string.decode()
return redirect(
url_for("do_view", path=path) + (f"?{qs}" if len(qs) > 0 else ""),
key.view_url,
code=302,
)
@app.route("/terminate/<path:path>", methods=["GET"])
def do_terminate(path):
key = get_key(path)
source_name = request.args.get("source_name") or default_item_source.name
source = matching_source(source_name)
key = CacheKey.for_lookup(source, source.lookup(path))
match = cache.check_entry(key)
if not match is None:
match.terminate()
return redirect(url_for("do_GET_status"), code=302)
@app.route("/metadata/ip_address", methods=["GET"])
def ip_address():
resp = make_response(env.ip)
return set_no_cache(resp)
def main():
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s:%(name)s:%(levelname)s:%(message)s",
)
def launch():
env.validate()
if not item_sources or not len(item_sources):
raise Exception("No data sources specified for Cellxgene Gateway")
global default_item_source
if default_item_source is None:
default_item_source = item_sources[0]
pruner = PruneProcessCache(cache)
background_thread = Thread(target=pruner)
@@ -316,5 +281,30 @@ def main():
app.run(host="0.0.0.0", port=env.gateway_port, debug=False)
def main():
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s:%(name)s:%(levelname)s:%(message)s",
)
cellxgene_data = os.environ.get("CELLXGENE_DATA", None)
cellxgene_bucket = os.environ.get("CELLXGENE_BUCKET", None)
if cellxgene_bucket is not None:
from cellxgene_gateway.items.s3.s3item_source import S3ItemSource
item_sources.append(S3ItemSource(cellxgene_bucket, name="s3"))
default_item_source = "s3"
if cellxgene_data is not None:
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
item_sources.append(FileItemSource(cellxgene_data, name="local"))
default_item_source = "local"
if len(item_sources) == 0:
raise Exception("Please specify CELLXGENE_DATA or CELLXGENE_BUCKET")
flask_util.include_source_in_url = len(item_sources) > 1
launch()
if __name__ == "__main__":
main()

View File

@@ -0,0 +1,28 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
from cellxgene_gateway.items.item import Item
class FileItem(Item):
"""e.g. FileItem(subpath = subpath, name = filename, type = ItemType.h5ad)
The Item superclass expects a 'name' and 'type'.
"""
def __init__(self, subpath: str, ext: str = "", *args, **kwargs):
super().__init__(*args, **kwargs)
self.subpath = subpath
self.ext = ext
@property
def descriptor(self) -> str:
return os.path.join(self.subpath, self.name + self.ext).strip("/")

View File

@@ -0,0 +1,183 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
from typing import List
from cellxgene_gateway import dir_util
from cellxgene_gateway.items.file.fileitem import FileItem
from cellxgene_gateway.items.item import ItemTree, ItemType
from cellxgene_gateway.items.item_source import ItemSource, LookupResult
class FileItemSource(ItemSource):
def __init__(
self,
base_path,
name=None,
h5ad_suffix=dir_util.h5ad_suffix,
annotation_dir_suffix=dir_util.annotations_suffix,
annotation_file_suffix=".csv",
):
self._name = name
self.base_path = base_path
self.h5ad_suffix = h5ad_suffix
self.annotation_dir_suffix = annotation_dir_suffix
self.annotation_file_suffix = annotation_file_suffix
@property
def name(self):
return self._name or f"Files:{self.base_path}"
def is_h5ad_file(self, path: str) -> bool:
return path.endswith(self.h5ad_suffix) and os.path.isfile(path)
def convert_annotation_path_to_h5ad(self, path):
return path[: -len(self.annotation_dir_suffix)] + self.h5ad_suffix
def convert_h5ad_path_to_annotation(self, path):
return path[: -len(self.h5ad_suffix)] + self.annotation_dir_suffix
def get_local_path(self, item: FileItem) -> str:
return os.path.join(self.base_path, item.descriptor)
def get_annotations_subpath(self, item) -> str:
return self.convert_h5ad_path_to_annotation(item.descriptor)
def list_items(self, filter: str = None) -> ItemTree:
item_tree = self.scan_directory()
"""def get_items(dir):
if dir.branches:
return [*dir.items, *[item for subdir in dir.branches for item in get_items(subdir)]]
else:
return dir.items
return get_items(self.item_tree)"""
return item_tree
def scan_directory(self, subpath="") -> dict:
base_path = os.path.join(self.base_path, subpath)
if not os.path.exists(base_path):
raise Exception(f"Path for local files '{base_path}' does not exist.")
filepath_map = dict(
(filepath, os.path.join(base_path, filepath))
for filepath in sorted(os.listdir(base_path))
)
def is_annotation_dir(dir):
return (
dir.endswith(self.annotation_dir_suffix)
and self.convert_annotation_path_to_h5ad(dir) in h5ad_paths
)
h5ad_paths = [
filepath
for filepath, full_path in filepath_map.items()
if self.is_h5ad_file(full_path)
]
subdirs = [
filepath
for filepath, full_path in filepath_map.items()
if os.path.isdir(full_path) and not is_annotation_dir(filepath)
]
items = [
self.make_fileitem_from_path(filename, subpath) for filename in h5ad_paths
]
branches = None
if len(subdirs) > 0:
branches = [
self.scan_directory(os.path.join(subpath, subdir)) for subdir in subdirs
]
return ItemTree(subpath, items, branches)
def create_annotation(self, item: FileItem, name: str) -> FileItem:
annotation = self.make_fileitem_from_path(
name, self.get_annotations_subpath(item), is_annotation=True
)
item.annotations = (item.annotations or []).append(annotation)
return annotation
def update(self, item: FileItem) -> None:
pass
def full_path(self, p):
return os.path.join(self.base_path, p)
def lookup_item(self, descriptor):
full_path = self.full_path(descriptor)
if self.is_h5ad_file(full_path):
return self.shallowitem_from_descriptor(descriptor)
def lookup(self, indescriptor: str) -> LookupResult:
descriptor = indescriptor.strip("/")
if descriptor.endswith(self.annotation_file_suffix):
annotation_item = self.shallowitem_from_descriptor(descriptor, True)
h5ad_descriptor = self.convert_annotation_path_to_h5ad(
annotation_item.subpath
)
item = self.lookup_item(h5ad_descriptor)
if item is not None:
dir_util.ensure_dir_exists(self.full_path(annotation_item.subpath))
return LookupResult(item, annotation_item)
else:
item = self.lookup_item(descriptor)
if item is not None:
return LookupResult(item)
def shallowitem_from_descriptor(self, descriptor, is_annotation=False):
filename = os.path.basename(descriptor)
subpath = os.path.dirname(descriptor)
return self.make_fileitem_from_path(
filename,
subpath,
is_annotation,
True,
)
def make_fileitem_from_path(
self, filename, subpath, is_annotation=False, is_shallow=False
) -> FileItem:
if is_annotation and filename.endswith(self.annotation_file_suffix):
name = filename[: -len(self.annotation_file_suffix)]
ext = self.annotation_file_suffix
else:
name = filename
ext = ""
item = FileItem(
subpath=subpath,
name=name,
ext=ext,
type=ItemType.annotation if is_annotation else ItemType.h5ad,
)
if not is_annotation and not is_shallow:
annotations = self.make_annotations_for_fileitem(item)
item.annotations = annotations
return item
def make_annotations_for_fileitem(self, item: FileItem) -> List[FileItem]:
annotations_subpath = self.get_annotations_subpath(item)
annotations_fullpath = self.full_path(annotations_subpath)
if os.path.isdir(annotations_fullpath):
return [
self.make_fileitem_from_path(annotation, annotations_subpath, True)
for annotation in sorted(os.listdir(annotations_fullpath))
if annotation.endswith(self.annotation_file_suffix)
and os.path.isfile(os.path.join(annotations_fullpath, annotation))
]
else:
return None

View File

@@ -0,0 +1,41 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
from abc import ABC, abstractmethod
from enum import Enum
from typing import List
class ItemType(Enum):
annotation = "annotation"
h5ad = "h5ad"
class Item(ABC):
def __init__(self, name: str, type: ItemType, annotations: List["Item"] = None):
self.name = name
self.type = type
self.annotations = annotations
@property
@abstractmethod
def descriptor(self):
raise Exception('"descriptor" not implemented')
class ItemTree:
def __init__(
self,
descriptor: str,
items: List[Item] = None,
branches: List["ItemTree"] = None,
):
self.descriptor = descriptor
self.items = items
self.branches = branches

View File

@@ -0,0 +1,50 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
from abc import ABC, abstractmethod
from typing import List
from cellxgene_gateway.items.item import Item
class LookupResult:
def __init__(self, h5ad_item: Item, annotation_item: Item = None):
self.h5ad_item = h5ad_item
self.annotation_item = annotation_item
class ItemSource(ABC):
@abstractmethod
def list_items(self, filter: str = None) -> List[Item]:
raise Exception('"list_items" unimplemented')
@abstractmethod
def get_local_path(self, item: Item) -> str:
raise Exception('"local_path" unimplemented')
@abstractmethod
def get_annotations_subpath(self, item) -> str:
raise Exception('"annotations_path" unimplemented')
@abstractmethod
def create_annotation(self, item: Item, name: str) -> Item:
raise Exception('"annotation" unimplemented')
@abstractmethod
def update(self, item: Item) -> None:
raise Exception('"update" unimplemented')
@abstractmethod
def lookup(self, descriptor: str) -> LookupResult:
raise Exception('"lookup" unimplemented')
@property
@abstractmethod
def name(self):
pass

View File

@@ -0,0 +1,27 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
from cellxgene_gateway.items.item import Item
class S3Item(Item):
"""e.g. FileItem(subpath = subpath, name = filename, type = ItemType.h5ad)
The Item superclass expects a 'name' and 'type'.
"""
def __init__(self, s3key: str, *args, **kwargs):
super().__init__(*args, **kwargs)
self.s3key = s3key
@property
def descriptor(self) -> str:
return self.s3key

View File

@@ -0,0 +1,173 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
from os.path import basename, dirname, join
from typing import List
import s3fs
from cellxgene_gateway import dir_util
from cellxgene_gateway.items.item import ItemTree, ItemType
from cellxgene_gateway.items.item_source import ItemSource, LookupResult
from cellxgene_gateway.items.s3.s3item import S3Item
class S3ItemSource(ItemSource):
def __init__(
self,
bucket,
name=None,
h5ad_suffix=dir_util.h5ad_suffix,
annotation_dir_suffix=dir_util.annotations_suffix,
annotation_file_suffix=".csv",
):
self._name = name
self.s3 = s3fs.S3FileSystem()
if bucket.startswith("s3://"):
raise Exception(
f"Bucket name should not include s3:// prefix, got {bucket}"
)
self.bucket = bucket
self.h5ad_suffix = h5ad_suffix
self.annotation_dir_suffix = annotation_dir_suffix
self.annotation_file_suffix = annotation_file_suffix
def url(self, path):
return "s3://" + join(self.bucket, path)
@property
def name(self):
return self._name or f"Items:{self.url('')}"
def is_h5ad_url(self, s3url: str) -> bool:
return s3url.endswith(self.h5ad_suffix) and self.s3.exists(s3url)
def convert_annotation_key_to_h5ad(self, s3key):
return s3key[: -len(self.annotation_dir_suffix)] + self.h5ad_suffix
def convert_h5ad_key_to_annotation(self, s3key):
return s3key[: -len(self.h5ad_suffix)] + self.annotation_dir_suffix
def get_local_path(self, item: S3Item) -> str:
return self.url(item.descriptor)
def get_annotations_subpath(self, item) -> str:
return self.convert_h5ad_key_to_annotation(item.descriptor)
def list_items(self, filter: str = None) -> ItemTree:
item_tree = self.scan_directory()
return item_tree
def scan_directory(self, subpath="") -> dict:
url = self.url(subpath)
if not self.s3.exists(url):
raise Exception(f"S3 url '{url}' does not exist.")
s3key_map = dict(
(filepath[len(self.bucket) :].lstrip("/"), "s3://" + filepath)
for filepath in sorted(self.s3.ls(url))
)
def is_annotation_dir(dir_s3key):
return (
dir_s3key.endswith(self.annotation_dir_suffix)
and self.convert_annotation_key_to_h5ad(dir_s3key) in h5ad_paths
)
h5ad_paths = [
filepath
for filepath, item_url in s3key_map.items()
if self.is_h5ad_url(item_url)
]
subdirs = [
filepath
for filepath, item_url in s3key_map.items()
if self.s3.isdir(item_url) and not is_annotation_dir(filepath)
]
items = [
self.make_s3item_from_key(filename, join(subpath, filename))
for filename in h5ad_paths
]
branches = None
if len(subdirs) > 0:
branches = [
self.scan_directory(join(subpath, subdir)) for subdir in subdirs
]
return ItemTree(subpath, items, branches)
def create_annotation(self, item: S3Item, name: str) -> S3Item:
annotation = self.make_s3item_from_key(
name, self.get_annotations_subpath(item), is_annotation=True
)
item.annotations = (item.annotations or []).append(annotation)
return annotation
def update(self, item: S3Item) -> None:
pass
def lookup_item(self, descriptor):
full_path = self.url(descriptor)
if self.is_h5ad_url(full_path):
return self.shallowitem_from_descriptor(descriptor)
def lookup(self, indescriptor: str) -> LookupResult:
descriptor = indescriptor.strip("/")
if descriptor.endswith(self.annotation_file_suffix):
annotation_item = self.shallowitem_from_descriptor(descriptor, True)
if not self.s3.exists(self.url(annotation_item.s3key)):
with self.s3.open(self.url(annotation_item.s3key), "w") as f:
f.write("")
h5ad_descriptor = self.convert_annotation_key_to_h5ad(
dirname(annotation_item.s3key)
)
item = self.shallowitem_from_descriptor(h5ad_descriptor)
return LookupResult(item, annotation_item)
else:
item = self.lookup_item(descriptor)
if item is not None:
return LookupResult(item)
def shallowitem_from_descriptor(self, descriptor, is_annotation=False):
return self.make_s3item_from_key(
basename(descriptor), descriptor, is_annotation, True
)
def make_s3item_from_key(
self, name, s3key, is_annotation=False, is_shallow=False
) -> S3Item:
item = S3Item(
s3key=s3key,
name=name,
type=ItemType.annotation if is_annotation else ItemType.h5ad,
)
if not is_annotation and not is_shallow:
annotations = self.make_annotations_for_fileitem(item)
item.annotations = annotations
return item
def make_annotations_for_fileitem(self, item: S3Item) -> List[S3Item]:
annotations_subpath = self.get_annotations_subpath(item)
annotations_fullpath = self.url(annotations_subpath)
if self.s3.isdir(annotations_fullpath):
return [
self.make_s3item_from_key(
annotation, join(annotations_subpath, annotation), True
)
for annotation in sorted(self.s3.ls(annotations_fullpath))
if annotation.endswith(self.annotation_file_suffix)
and self.s3.isfile(join(annotations_fullpath, annotation))
]
else:
return None

View File

@@ -11,43 +11,11 @@ import os
from flask_api import status
from cellxgene_gateway import env
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.cellxgene_exception import CellxgeneException
from cellxgene_gateway.dir_util import make_h5ad
def get_key(path):
if path == "/" or path == "":
raise CellxgeneException(
"No matching dataset found.", status.HTTP_404_NOT_FOUND
)
trimmed = path[:-1] if path[-1] == "/" else path
try:
# valid paths come in three forms:
if trimmed.endswith(".h5ad") and data_file_exists(trimmed):
# 1) somedir/dataset.h5ad: a dataset
return CacheKey(trimmed, trimmed, None)
elif trimmed.endswith(".csv"):
# 2) somedir/dataset_annotations/my_annotations.csv : an actual annotations file.
annotations_dir = os.path.split(trimmed)[0]
dataset = make_h5ad(annotations_dir)
if data_file_exists(dataset):
data_dir_ensure(annotations_dir)
return CacheKey(trimmed, dataset, trimmed)
elif trimmed.endswith("_annotations") and data_dir_exists(trimmed):
# 3) somedir/dataset_annotations: an annotation directory. The corresponding h5ad must exist, but the directory may not.
dataset = make_h5ad(trimmed)
if data_file_exists(dataset):
return CacheKey(trimmed, dataset, "")
except CellxgeneException:
pass
split = os.path.split(trimmed)
return get_key(split[0])
def validate_exists(file_path):
if not os.path.exists(file_path):
raise CellxgeneException(
@@ -71,37 +39,3 @@ def validate_is_dir(file_path):
"Path is not dir: " + file_path, status.HTTP_400_BAD_REQUEST
)
return
def data_file_exists(dataset):
file_path = os.path.join(env.cellxgene_data, dataset)
validate_is_file(file_path)
return True
def data_dir_exists(dataset):
file_path = os.path.join(env.cellxgene_data, dataset)
validate_is_dir(file_path)
return True
def data_dir_ensure(dataset):
file_path = os.path.join(env.cellxgene_data, dataset)
if not os.path.exists(file_path):
os.makedirs(file_path)
def get_file_path(key):
dataset = key.dataset
file_path = os.path.join(env.cellxgene_data, dataset)
validate_is_file(file_path)
return file_path
def get_annotation_file_path(key):
if key.annotation_file is None:
return None
if key.annotation_file == "":
return ""
file_path = os.path.join(env.cellxgene_data, key.annotation_file)
return file_path

View File

@@ -10,14 +10,15 @@
import logging
import time
from cellxgene_gateway.env import ttl
from cellxgene_gateway.util import current_time_stamp
from cellxgene_gateway import env, util
logger = logging.getLogger(__name__)
class PruneProcessCache:
def __init__(self, cache):
self.cache = cache
self.expire_seconds = 3600 if ttl is None else int(ttl)
self.expire_seconds = 3600 if env.ttl is None else int(env.ttl)
def __call__(self):
while True:
@@ -25,24 +26,20 @@ class PruneProcessCache:
self.prune()
def prune(self):
timestamp = current_time_stamp()
timestamp = util.current_time_stamp()
cutoff = timestamp - self.expire_seconds
processes_to_delete = [
p for p in self.cache.entry_list if p.timestamp < cutoff
]
processes_to_delete = [p for p in self.cache.entry_list if p.timestamp < cutoff]
processes_to_keep = [
p for p in self.cache.entry_list if not p.timestamp < cutoff
]
logger = logging.getLogger("cellxgene_gateway")
logger.debug(
f"Cutoff {cutoff} = timestamp {timestamp} - expire seconds {self.expire_seconds} , keeping {processes_to_keep}"
f"Cutoff {cutoff} = timestamp {timestamp} - expire seconds {self.expire_seconds} , keeping {processes_to_keep}, pruning {processes_to_delete}"
)
for process in processes_to_delete:
try:
logger.info(
f"pruning process {process.pid} ({process.key.dataset})"
)
logger.info(f"pruning process {process.pid} ({process.key.dataset})")
self.cache.prune(process)
except Exception:
logger.exception(

View File

@@ -1,4 +1,5 @@
// neandertal javascript
// Annotations only work with file itemsources at the moment. If they work with others in the future we may need to revisit this.
const new_annotation_callback = (() =>{
const suffix = `.csv`;
return (e) => {

View File

@@ -11,14 +11,10 @@ import logging
import subprocess
from flask_api import status
from cellxgene_gateway.cache_entry import CacheEntryStatus
from cellxgene_gateway.dir_util import make_annotations
from cellxgene_gateway.path_util import get_annotation_file_path, get_file_path
from cellxgene_gateway.env import (
enable_annotations,
enable_backed_mode,
cellxgene_args,
)
from cellxgene_gateway.env import cellxgene_args, enable_annotations, enable_backed_mode
from cellxgene_gateway.process_exception import ProcessException
@@ -26,14 +22,10 @@ class SubprocessBackend:
def __init__(self):
pass
def create_cmd(
self, cellxgene_loc, file_path, port, scripts, annotation_file_path
):
def create_cmd(self, cellxgene_loc, file_path, port, scripts, annotation_file_path):
if enable_annotations and not annotation_file_path is None:
if annotation_file_path == "":
extra_args = (
f" --annotations-dir {make_annotations(file_path)}"
)
extra_args = f" --annotations-dir {make_annotations(file_path)}"
else:
extra_args = f" --annotations-file {annotation_file_path}"
else:
@@ -45,8 +37,7 @@ class SubprocessBackend:
cmd = (
f"yes | {cellxgene_loc} launch {file_path}"
+ " --port "
+ str(port)
+ f" --port {port}"
+ " --host 127.0.0.1"
+ extra_args
)
@@ -57,13 +48,12 @@ class SubprocessBackend:
return cmd
def launch(self, cellxgene_loc, scripts, cache_entry):
cmd = self.create_cmd(
cellxgene_loc,
get_file_path(cache_entry.key),
cache_entry.key.file_path,
cache_entry.port,
scripts,
get_annotation_file_path(cache_entry.key),
cache_entry.key.annotation_file_path,
)
logging.getLogger("cellxgene_gateway").info(f"launching {cmd}")
process = subprocess.Popen(

View File

@@ -10,59 +10,68 @@
-->
<html>
<head>
<title>Cellxgene Gateway - FILE CRAWLER</title>
<script src="https://ajax.googleapis.com/ajax/libs/jquery/3.3.1/jquery.min.js"></script>
<link rel="icon" type="image/png" href="{{ url_for('static', filename='nibr.ico') }}">
{% for script in extra_scripts %}
<script src="{{ script }}"></script>
{% endfor %}
<link rel="stylesheet" href="https://stackpath.bootstrapcdn.com/bootstrap/4.1.3/css/bootstrap.min.css" integrity="sha384-MCw98/SFnGE8fJT3GXwEOngsV7Zt27NXFoaoApmYm81iuXoPkFOJwJ8ERdknLPMO" crossorigin="anonymous">
{% for script in extra_scripts %}
<script src="{{ script }}"></script>
{% endfor %}
<link rel="stylesheet" href="https://stackpath.bootstrapcdn.com/bootstrap/4.1.3/css/bootstrap.min.css"
integrity="sha384-MCw98/SFnGE8fJT3GXwEOngsV7Zt27NXFoaoApmYm81iuXoPkFOJwJ8ERdknLPMO" crossorigin="anonymous">
</head>
<body>
<header class="navbar navbar-expand navbar-dark flex-column flex-md-row bd-navbar">
<h3>Cellxgene Gateway - Cache Status</h3>
<h3>Cellxgene Gateway - Cache Status</h3>
</header>
<br>
<table class="table">
<thead>
<tr>
<th>PID</th>
<th>dataset</th>
<th>annotation_file</th>
<th>port</th>
<th>launchtime</th>
<th>last access</th>
<th>status</th>
<th>message</th>
<th>http_status</th>
<th>actions</th>
</tr>
</thead>
<tbody>
{% for entry in entry_list %}
<tr>
<td>{{ entry.pid }}</td>
<td><a href="{{ url_for('do_view', path=entry.key.pathpart) }}">{{ entry.key.dataset }}</a></td>
<td>{{ entry.key.annotation_file }}</td>
<td>{{ entry.port }}</td>
<td class="timestamp">{{ entry.launchtime }}</td>
<td class="timestamp">{{ entry.timestamp }}</td>
<td>{{ entry.status.name }}</td>
<td>{{ entry.message }}</td>
<td>{{ entry.http_status }}</td>
<td>
{% if entry.status.name == 'loaded' %}
<a href="{{ url_for('do_terminate', path=entry.key.pathpart) }}"> terminate </a>
{% endif %}
</td>
</tr>
{% endfor %}
</tbody>
<tr>
<th>PID</th>
<th>dataset</th>
<th>annotation_file</th>
<th>source</th>
<th>port</th>
<th>launchtime</th>
<th>last access</th>
<th>status</th>
<th>message</th>
<th>http_status</th>
<th>actions</th>
</tr>
</thead>
<tbody>
{% for entry in entry_list %}
<tr>
<td>{{ entry.pid }}</td>
<td><a
href="{{ entry.key.view_url }}">{{ entry.key.h5ad_item.descriptor }}</a>
</td>
<td>{{ entry.key.annotation_descriptor }}</td>
<td>{{ entry.source_name }}</td>
<td>{{ entry.port }}</td>
<td class="timestamp">{{ entry.launchtime }}</td>
<td class="timestamp">{{ entry.timestamp }}</td>
<td>{{ entry.status.name }}</td>
<td>{{ entry.message }}</td>
<td>{{ entry.http_status }}</td>
<td>
{% if entry.status.name == 'loaded' %}
<a
href="{{ url_for('do_terminate', path=entry.key.descriptor, source_name=entry.key.source_name) }}">
terminate </a>
{% endif %}
</td>
</tr>
{% endfor %}
</tbody>
</table>
<script>
$(() => {
$(".timestamp").each(function(){
$(".timestamp").each(function () {
const el = $(this);
const ts = el.text();
const dt = new Date(parseInt(ts * 1000));
@@ -71,4 +80,5 @@
})
</script>
</body>
</html>
</html>

View File

@@ -28,11 +28,11 @@
<h4>{{ message }}</h4>
<a href="/filecrawl.html">
<a href="{{ url_for('filecrawl') }}">
Please click here to be redirected to the file directory.
</a>
<br>
<a href="/">
<a href="{{ url_for('index') }}">
Please click here to return to the homepage.
</a>
</div>

View File

@@ -36,10 +36,10 @@
Navigation:
<ul>
{% if path %}
<li><a href="/filecrawl.html">top level</a></li>
<li><a href="{{ url_for('filecrawl') }}">top level</a></li>
{% else %}
{% endif %}
<li><a href="/">homepage</a></li>
<li><a href="{{ url_for('index') }}">homepage</a></li>
</ul>
</p>
<script>

View File

@@ -35,56 +35,15 @@
Links:
</h1>
<div class="list-group" style="width:50%;padding-left:65px">
<a href="/filecrawl.html" class="list-group-item list-group-item-action">
<a href="{{ url_for('filecrawl') }}" class="list-group-item list-group-item-action">
<u>File Crawler: Allows you to view all uploaded data.</u></a>
</div>
<div class="list-group" style="width:50%;padding-left:65px">
<a href="/cache_status" class="list-group-item list-group-item-action">
<a href="{{ url_for('do_GET_status') }}" class="list-group-item list-group-item-action">
<u>Cache Status: view status of launched cellxgene servers.</u></a>
</div>
{% if enable_upload %}
<br>
<h1 style="padding-left:35px">
How To Upload Data via HTTP:
</h1>
<ol style="padding-left:85px;">
<li>
Create a folder for your Username:
</li>
<br>
<form action="{{ url_for('make_user') }}" method="post">
Username <input type="text" name="directory">
<input type="submit" value="Create">
</form>
<li>
Create a subdirectory under the selected Folder:
</li>
<br>
<form action="{{ url_for('make_subdir') }}" method="post">
<select name="usernames" id="usernames">
{% for user in users %}
<option value="{{ user }}">{{ user }}</option>
{% endfor %}
</select>
<br>
Subdirectory Name <input type="text" name="directory">
<input type="submit" value="Create">
</form>
<li>Choose a folder to copy your data to, then upload your data file (must be in .h5ad format).</li>
<br>
<form action="{{ url_for('upload_file') }}" method="post" enctype="multipart/form-data">
Type in the name of the directory and subdirectory you wish to upload to, i.e. "USER/cells". <input type="text" name="path">
<br>
File: <input type="file" name="file"><br>
<input style="position:relative; top:10px;" type="submit" value="Upload">
</form>
<br>
<li>Take a look at your data using the file crawler link above</li>
</ol>
{% endif %}
<br>
<h1 style="padding-left:35px">

View File

@@ -35,11 +35,11 @@
The page will refresh shortly.
</p>
<a href="/filecrawl.html">
<a href="{{ url_for('filecrawl') }}">
Please click here to be redirected to the file directory.
</a>
<br>
<a href="/">
<a href="{{ url_for('index') }}">
Please click here to return to the homepage.
</a>
</div>

View File

@@ -36,13 +36,13 @@
<h4>Options</h4>
Please choose one of the following, or use the back button:
<ul>
<li><a href="{{url_for('do_relaunch', path=dataset)}}">
<li><a href="{{ relaunch_url }}">
Attempt to relaunch the cellxgene server.
</a></li>
<li><a href="/filecrawl.html">
<li><a href="{{ url_for('filecrawl') }}">
Return to the file directory.
</a></li>
<li><a href="/">
<li><a href="{{ url_for('index') }}">
Return to the homepage.
</a></li>
</ul>

View File

@@ -7,6 +7,7 @@ dependencies:
- flask
- psutil
- black
- coverage
- pip
- pip:
- flask-api

View File

@@ -1,8 +1,9 @@
import os
import codecs
from setuptools import find_packages, setup
import os
import sys
from setuptools import find_packages, setup
if sys.version_info < (3, 6):
sys.exit("Sorry, Python < 3.6 is not supported")
@@ -25,8 +26,8 @@ def get_version(rel_path):
def parse_requirements():
reqs = []
with open("requirements.txt", "r") as f:
for l in f.readlines():
reqs.append(l.strip("\n"))
for line in f.readlines():
reqs.append(line.strip("\n"))
return reqs
@@ -49,7 +50,7 @@ setup(
license="MIT",
keywords="visualization, genomics",
url="http://github.com/Novartis/cellxgene-gateway",
packages=["cellxgene_gateway"],
packages=find_packages(),
package_data={
"cellxgene_gateway": [
"static/css/homepagestyle.css",

0
tests/items/__init__.py Normal file
View File

View File

View File

@@ -0,0 +1,22 @@
import tempfile
import unittest
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
class TestFileItemSource(unittest.TestCase):
def test_make_fileitem_from_path_GIVEN_annotation_file_THEN_name_lacks_csv(
self,
):
source = FileItemSource(tempfile.gettempdir(), "local")
item = source.make_fileitem_from_path(
"customanno.csv", "someh5ad_annotations", True
)
self.assertEqual(item.name, "customanno")
self.assertEqual(item.descriptor, "someh5ad_annotations/customanno.csv")
def test_make_fileitem_from_path_GIVEN_h5ad_file_THEN_returns_name(self):
source = FileItemSource(tempfile.gettempdir(), "local")
item = source.make_fileitem_from_path("someanalysis.h5ad", "studydir")
self.assertEqual(item.name, "someanalysis.h5ad")
self.assertEqual(item.descriptor, "studydir/someanalysis.h5ad")

View File

@@ -1,29 +1,62 @@
import unittest
from flask import Flask
from cellxgene_gateway import flask_util
from cellxgene_gateway.cache_entry import CacheEntry, CacheEntryStatus
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.gateway import app
from cellxgene_gateway.items.file.fileitem import FileItem
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
from cellxgene_gateway.items.item import ItemType
key = CacheKey("czi/pbmc3k.h5ad", "pbmc3k.h5ad", "tmp.csv")
key = CacheKey(
FileItem("/czi/", name="pbmc3k.h5ad", type=ItemType.h5ad),
FileItemSource("/tmp", "local"),
)
class TestRenderEntry(unittest.TestCase):
def setUp(self):
self.app = app
self.app_context = self.app.test_request_context()
self.app_context.push()
self.client = self.app.test_client()
def test_GIVEN_key_and_port_THEN_returns_loading_CacheEntry(self):
entry = CacheEntry.for_key("some-key", 1)
self.assertEqual(entry.status, CacheEntryStatus.loading)
def test_GIVEN_absolute_static_url_THEN_include_path(self):
flask_util.include_source_in_url = False
actual = CacheEntry.for_key(key, 8000).rewrite_text_content(
"src:url(/static/assets/"
)
expected = (
"src:url(http://localhost:5005/view/czi/pbmc3k.h5ad/static/assets/"
)
expected = "src:url(/view/czi/pbmc3k.h5ad/static/assets/"
self.assertEqual(actual, expected)
def test_GIVEN_absolute_src_THEN_include_path(self):
flask_util.include_source_in_url = False
actual = CacheEntry.for_key(key, 8000).rewrite_text_content(
'<link rel="shortcut icon" href="/static/assets/favicon.ico">'
)
expected = '<link rel="shortcut icon" href="http://localhost:5005/view/czi/pbmc3k.h5ad/static/assets/favicon.ico">'
expected = '<link rel="shortcut icon" href="/view/czi/pbmc3k.h5ad/static/assets/favicon.ico">'
self.assertEqual(actual, expected)
def test_GIVEN_absolute_static_url_include_source_THEN_include_path(self):
flask_util.include_source_in_url = True
actual = CacheEntry.for_key(key, 8000).rewrite_text_content(
"src:url(/static/assets/"
)
expected = "src:url(/source/local/view/czi/pbmc3k.h5ad/static/assets/"
self.assertEqual(actual, expected)
def test_GIVEN_absolute_src_include_source_THEN_include_path(self):
flask_util.include_source_in_url = True
actual = CacheEntry.for_key(key, 8000).rewrite_text_content(
'<link rel="shortcut icon" href="/static/assets/favicon.ico">'
)
expected = '<link rel="shortcut icon" href="/source/local/view/czi/pbmc3k.h5ad/static/assets/favicon.ico">'
self.assertEqual(actual, expected)

View File

@@ -1,49 +0,0 @@
import unittest
from unittest.mock import MagicMock, patch
from cellxgene_gateway.filecrawl import render_entry
class TestRenderEntry(unittest.TestCase):
def test_GIVEN_path_both_slash_THEN_view_has_single_slash(self):
entry = {
"path": "/somepath/",
"name": "entry",
"type": "file",
"annotations": [],
"children": [],
}
rendered = render_entry(entry)
self.assertIn("view/somepath", rendered)
def test_GIVEN_path_starts_slash_THEN_view_has_single_slash(self):
entry = {
"path": "/somepath",
"name": "entry",
"type": "file",
"annotations": [],
"children": [],
}
rendered = render_entry(entry)
self.assertIn("view/somepath", rendered)
def test_GIVEN_path_ends_slash_THEN_view_has_single_slash(self):
entry = {
"path": "somepath/",
"name": "entry",
"type": "file",
"annotations": [],
"children": [],
}
rendered = render_entry(entry)
self.assertIn("view/somepath", rendered)
def test_GIVEN_path_no_slash_THEN_view_has_single_slash(self):
entry = {
"path": "somepath",
"name": "entry",
"type": "file",
"annotations": [],
"children": [],
}
rendered = render_entry(entry)
self.assertIn("view/somepath", rendered)

View File

@@ -1,46 +1,31 @@
import unittest
from unittest.mock import MagicMock, patch
from cellxgene_gateway.filecrawl import render_entry
from cellxgene_gateway.filecrawl import render_item
from cellxgene_gateway.items.file.fileitem import FileItem
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
from cellxgene_gateway.items.item import ItemType
source = FileItemSource("/tmp")
class TestRenderEntry(unittest.TestCase):
def test_GIVEN_path_both_slash_THEN_view_has_single_slash(self):
entry = {
"path": "/somepath/",
"name": "entry",
"type": "file",
"annotations": [],
}
rendered = render_entry(entry)
self.assertIn("view/somepath", rendered)
entry = FileItem(subpath="/somepath/", name="entry", type=ItemType.h5ad)
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry/'", rendered)
def test_GIVEN_path_starts_slash_THEN_view_has_single_slash(self):
entry = {
"path": "/somepath",
"name": "entry",
"type": "file",
"annotations": [],
}
rendered = render_entry(entry)
self.assertIn("view/somepath", rendered)
entry = FileItem(subpath="/somepath", name="entry", type=ItemType.h5ad)
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry/'", rendered)
def test_GIVEN_path_ends_slash_THEN_view_has_single_slash(self):
entry = {
"path": "somepath/",
"name": "entry",
"type": "file",
"annotations": [],
}
rendered = render_entry(entry)
self.assertIn("view/somepath", rendered)
entry = FileItem(subpath="somepath/", name="entry", type=ItemType.h5ad)
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry/'", rendered)
def test_GIVEN_path_no_slash_THEN_view_has_single_slash(self):
entry = {
"path": "somepath",
"name": "entry",
"type": "file",
"annotations": [],
}
rendered = render_entry(entry)
self.assertIn("view/somepath", rendered)
entry = FileItem(subpath="somepath", name="entry", type=ItemType.h5ad)
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry/'", rendered)