1 Commits

Author SHA1 Message Date
Alok Saldanha
2d6d868c92 #9 only prune entries that have a pid 2019-10-29 21:20:24 -04:00
40 changed files with 420 additions and 1387 deletions

3
.gitignore vendored
View File

@@ -136,6 +136,3 @@ dmypy.json
.pyre/ .pyre/
# End of https://www.gitignore.io/api/python # End of https://www.gitignore.io/api/python
*.patch
.vscode

View File

@@ -1,41 +0,0 @@
# This is necessary for nxviz as matplotlib is involved.
# before_script:
# - "export DISPLAY=:99.0"
# - "sh -e /etc/init.d/xvfb start"
# - sleep 5 # give xvfb some time to start
language: python
matrix:
include:
- python: 3.5 # we don't actually use this
env: PYTHON_VERSION=3.7
install:
# We do this conditionally because it saves us some downloading if the
# version is the same.
- wget https://repo.continuum.io/miniconda/Miniconda3-latest-Linux-x86_64.sh -O miniconda.sh;
- bash miniconda.sh -b -p $HOME/miniconda
- export PATH="$HOME/miniconda/bin:$PATH"
- hash -r
- conda config --set always_yes yes --set changeps1 no
- conda update -q conda
- conda config --add channels conda-forge
# Useful for debugging any issues with conda
- conda info -a
# Install Python, py.test, and required packages.
- conda env create -f environment.yml
- source activate cellxgene-gateway
- python setup.py install
script:
# Your test script goes here
- black -l 79 . --check
- python -m unittest discover tests
after_success:
- bash <(curl -s https://codecov.io/bash)
notifications:
email: true

View File

@@ -1,17 +0,0 @@
# 0.2.2
* Fixed bug with annotations (missing annotation.js asset)
# 0.2.1
* Minor fixes to enable cellxgene 0.16.0
* Added CELLXGENE_ARGS to enable passing additional arguments to cellxgene
* added metadata/ip_address endpoint
# 0.2.0
Incrementing minor version since the changes for 0.15 are breaking, and we may want to release bugfixes from 0.1.0 branch.
# 0.1.1
Added support for cellxgene 0.15

View File

@@ -30,7 +30,7 @@ Note: you may need to downgrade h5py with `pip install h5py==2.9.0` due to an [i
### Option 2: Install from PyPI ### Option 2: Install from PyPI
```bash ```bash
pip install cellxgene-gateway # NOT YET DONE, COMING! STAY TUNED
``` ```
## Running cellxgene gateway ## Running cellxgene gateway
@@ -39,7 +39,7 @@ pip install cellxgene-gateway
```bash ```bash
mkdir ../cellxgene_data mkdir ../cellxgene_data
wget https://raw.githubusercontent.com/chanzuckerberg/cellxgene/master/example-dataset/pbmc3k.h5ad -O ../cellxgene_data/pbmc3k.h5ad wget https://github.com/chanzuckerberg/cellxgene/raw/master/example-dataset/pbmc3k.h5ad -O ../cellxgene_data/pbmc3k.h5ad
``` ```
@@ -48,6 +48,9 @@ wget https://raw.githubusercontent.com/chanzuckerberg/cellxgene/master/example-d
```bash ```bash
export CELLXGENE_DATA=../cellxgene_data # change this directory if you put data in a different place. export CELLXGENE_DATA=../cellxgene_data # change this directory if you put data in a different place.
export CELLXGENE_LOCATION=`which cellxgene` export CELLXGENE_LOCATION=`which cellxgene`
export GATEWAY_HOST=localhost:5005
export GATEWAY_PROTOCOL=http
export GATEWAY_IP=127.0.0.1
``` ```
3. Now, execute the cellxgene gateway: 3. Now, execute the cellxgene gateway:
@@ -59,21 +62,13 @@ cellxgene-gateway
Here's what the environment variables mean: Here's what the environment variables mean:
* `CELLXGENE_LOCATION` - the location of the cellxgene executable, e.g. `~/anaconda2/envs/cellxgene/bin/cellxgene` * `CELLXGENE_LOCATION` - the location of the cellxgene executable, e.g. `~/anaconda2/envs/cellxgene/bin/cellxgene`
At least one of the following is required:
* `CELLXGENE_DATA` - a directory that can contain subdirectories with `.h5ad` data files, *without* trailing slash, e.g. `/mnt/cellxgene_data` * `CELLXGENE_DATA` - a directory that can contain subdirectories with `.h5ad` data files, *without* trailing slash, e.g. `/mnt/cellxgene_data`
* `CELLXGENE_BUCKET` - an s3 bucket that can contain keys with `.h5ad` data files, e.g. `my-cellxgene-data-bucket` * `GATEWAY_HOST` - the hostname and port that the gateway will run on, typically `localhost:5005` if running locally
Cellxgene Gateway is designed to make it easy to add additional data sources, please see the source code for gateway.py and the ItemSource interface in items/item_source.py * `GATEWAY_PROTOCOL` - typically http when running locally, can be https when deployed if the gateway is behind a load balancer or reverse proxy.
* `GATEWAY_IP` - ip addess of instance gateway is running on, mostly used to display SSH instructions
Optional environment variables: Optional environment variables:
* `CELLXGENE_ARGS` - catch-all variable that can be used to pass additional command line args to cellxgene server
* `EXTERNAL_HOST` - the hostname and port from the perspective of the web browser, typically `localhost:5005` if running locally. Defaults to "localhost:{GATEWAY_PORT}"
* `EXTERNAL_PROTOCOL` - typically http when running locally, can be https when deployed if the gateway is behind a load balancer or reverse proxy that performs https termination. Default value "http"
* `GATEWAY_IP` - ip addess of instance gateway is running on, mostly used to display SSH instructions. Defaults to `socket.gethostbyname(socket.gethostname())`
* `GATEWAY_PORT` - local port that the gateway should bind to, defaults to 5005
* `GATEWAY_EXTRA_SCRIPTS` - JSON array of script paths, will be embedded into each page and forwarded with `--scripts` to cellxgene server * `GATEWAY_EXTRA_SCRIPTS` - JSON array of script paths, will be embedded into each page and forwarded with `--scripts` to cellxgene server
* `GATEWAY_ENABLE_ANNOTATIONS` - Set to `true` or to `1` to enable cellxgene annotations. * `GATEWAY_ENABLE_UPLOAD` - Set to `true` or `1` to enable HTTP uploads. This is not recommended for a public server.
* `GATEWAY_ENABLE_BACKED_MODE` - Set to `true` or to `1` to load AnnData in file-backed mode. This saves memory and speeds up launch time but may reduce overall performance.
The defaults should be fine if you set up a venv and cellxgene_data folder as above. The defaults should be fine if you set up a venv and cellxgene_data folder as above.
@@ -119,8 +114,6 @@ For convenience, the code repo includes a `run.sh.example` shell script to run t
## Running Tests ## Running Tests
[![Build Status](https://travis-ci.org/Novartis/cellxgene-gateway.svg?branch=master)](https://travis-ci.org/Novartis/cellxgene-gateway)
```bash ```bash
python -m unittest discover tests python -m unittest discover tests
``` ```
@@ -130,8 +123,9 @@ For convenience, the code repo includes a `run.sh.example` shell script to run t
pip install isort flake8 black pip install isort flake8 black
```bash ```bash
isort -rc . # rc means recursive, and was deprecated in dev version of isort isort -rc .
black . flake8 .
black -l 79 .
``` ```
# Getting Help # Getting Help

View File

@@ -6,5 +6,3 @@
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES # under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for # OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License. # the specific language governing permissions and limitations under the License.
__version__ = "0.2.2"

View File

@@ -13,21 +13,17 @@ from threading import Thread
from flask_api import status from flask_api import status
from cellxgene_gateway import env from cellxgene_gateway import env
from cellxgene_gateway.cache_entry import CacheEntry, CacheEntryStatus from cellxgene_gateway.cache_entry import CacheEntry
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.cellxgene_exception import CellxgeneException from cellxgene_gateway.cellxgene_exception import CellxgeneException
from cellxgene_gateway.subprocess_backend import SubprocessBackend from cellxgene_gateway.subprocess_backend import SubprocessBackend
from typing import List
process_backend = SubprocessBackend() process_backend = SubprocessBackend()
def is_port_in_use(port): def is_port_in_use(port):
import socket import socket
with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as s: with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as s:
return s.connect_ex(("localhost", port)) == 0 return s.connect_ex(('localhost', port)) == 0
class BackendCache: class BackendCache:
def __init__(self): def __init__(self):
@@ -37,14 +33,12 @@ class BackendCache:
contents = self.entry_list contents = self.entry_list
return [c.port for c in contents] return [c.port for c in contents]
def check_path(self, source, path): def check_entry(self, dataset):
contents = self.entry_list contents = self.entry_list
matches = [ matches = [
c c
for c in contents for c in contents
if c.key.source.name == source.name if c.dataset == dataset and c.status != "terminated"
and path.startswith(c.key.descriptor)
and c.status != CacheEntryStatus.terminated
] ]
if len(matches) == 0: if len(matches) == 0:
@@ -54,35 +48,16 @@ class BackendCache:
else: else:
raise CellxgeneException( raise CellxgeneException(
status.HTTP_500_INTERNAL_SERVER_ERROR, status.HTTP_500_INTERNAL_SERVER_ERROR,
"Found " + str(len(matches)) + " for " + path, "Found " + str(len(matches)) + " for " + dataset,
) )
def check_entry(self, key): def create_entry(self, dataset, file_path, scripts):
contents = self.entry_list
matches = [
c
for c in contents
if c.key.equals(key) and c.status != CacheEntryStatus.terminated
]
if len(matches) == 0:
return None
elif len(matches) == 1:
return matches[0]
else:
raise CellxgeneException(
status.HTTP_500_INTERNAL_SERVER_ERROR,
"Found " + str(len(matches)) + " for " + key.dataset,
)
def create_entry(self, key: CacheKey, scripts: List[str]):
port = 8000 port = 8000
existing_ports = self.get_ports() existing_ports = self.get_ports()
while (port in existing_ports) or is_port_in_use(port): while (port in existing_ports) or is_port_in_use(port):
port += 1 port += 1
entry = CacheEntry.for_key(key, port) entry = CacheEntry.for_dataset(dataset, file_path, port)
background_thread = Thread( background_thread = Thread(
target=process_backend.launch, target=process_backend.launch,

View File

@@ -6,45 +6,34 @@
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES # under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for # OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License. # the specific language governing permissions and limitations under the License.
import datetime
import logging
import urllib.parse
from enum import Enum
import psutil import psutil
from flask import make_response, render_template, request import logging
from flask import make_response, request
from requests import get, post, put from requests import get, post, put
import re
from cellxgene_gateway import env from cellxgene_gateway import env
from cellxgene_gateway.cellxgene_exception import CellxgeneException from cellxgene_gateway.cellxgene_exception import CellxgeneException
from cellxgene_gateway.flask_util import querystring
from cellxgene_gateway.util import current_time_stamp from cellxgene_gateway.util import current_time_stamp
class CacheEntryStatus(Enum):
loaded = "loaded"
loading = "loading"
error = "error"
terminated = "terminated"
class CacheEntry: class CacheEntry:
def __init__( def __init__(
self, self,
pid, pid,
key, dataset,
file_path,
port, port,
launchtime, launchtime,
timestamp, timestamp,
status: CacheEntryStatus, status,
message, message,
all_output, all_output,
stderr, stderr,
http_status, http_status,
): ):
self.pid = pid self.pid = pid
self.key = key self.dataset = dataset
self.file_path = file_path
self.port = port self.port = port
self.launchtime = launchtime self.launchtime = launchtime
self.timestamp = timestamp self.timestamp = timestamp
@@ -55,34 +44,30 @@ class CacheEntry:
self.http_status = http_status self.http_status = http_status
@classmethod @classmethod
def for_key(cls, key, port): def for_dataset(cls, dataset, file_path, port):
return cls( return cls(
None, None,
key, dataset,
file_path,
port, port,
current_time_stamp(), current_time_stamp(),
current_time_stamp(), current_time_stamp(),
CacheEntryStatus.loading, "loading",
None, None,
None, None,
None, None,
None, None,
) )
@property
def source_name(self):
return self.key.source_name
def set_loaded(self, pid): def set_loaded(self, pid):
self.pid = pid self.pid = pid
self.status = CacheEntryStatus.loaded self.status = "loaded"
def set_error(self, message, stderr, http_status): def set_error(self, message, stderr, http_status):
self.message = message self.message = message
self.stderr = stderr self.stderr = stderr
self.http_status = http_status self.http_status = http_status
self.status = CacheEntryStatus.error self.status = "error"
def append_output(self, output): def append_output(self, output):
if self.all_output == None: if self.all_output == None:
@@ -92,105 +77,61 @@ class CacheEntry:
def terminate(self): def terminate(self):
pid = self.pid pid = self.pid
if pid != None and self.status != CacheEntryStatus.terminated: if pid != None and self.status != "terminated":
terminated = [] terminated = []
def on_terminate(p): def on_terminate(p):
terminated.append(p.pid) terminated.append(p.pid)
p = psutil.Process(pid) p = psutil.Process(pid)
children = p.children() children = p.children()
for child in children: for child in children:
child.terminate() child.terminate()
psutil.wait_procs(children, callback=on_terminate) psutil.wait_procs(children, callback=on_terminate)
# the parent process may automatically die once its children have -- terminated.append(p.pid)
try: p.terminate()
p.terminate() psutil.wait_procs([p], callback=on_terminate)
psutil.wait_procs([p], callback=on_terminate) logging.getLogger("cellxgene_gateway").info(f"terminated {terminated}")
except psutil.NoSuchProcess: self.status = "terminated"
pass
logging.getLogger("cellxgene_gateway").info(
f"terminated {terminated}"
)
self.status = CacheEntryStatus.terminated
def rewrite_text_content(self, cellxgene_content):
# for v0.16.0 compatibility, see issue #24
gateway_content = (
re.sub(
'(="|\()/static/',
f"\\1{self.gateway_basepath()}static/",
cellxgene_content,
)
.replace("http://fonts.gstatic.com", "https://fonts.gstatic.com")
.replace(self.cellxgene_basepath(), self.gateway_basepath())
)
return gateway_content
def gateway_basepath(self):
source_path = (
f"/source/{urllib.parse.quote_plus(self.source_name)}"
if self.source_name
else ""
)
return f"{env.external_protocol}://{env.external_host}{source_path}/view/{self.key.descriptor}/"
def cellxgene_basepath(self):
return f"http://127.0.0.1:{self.port}"
def serve_content(self, path): def serve_content(self, path):
gateway_basepath = self.gateway_basepath() dataset = self.dataset
subpath = path[len(self.key.descriptor) :] # noqa: E203
gateway_basepath = (
f"{env.gateway_protocol}://{env.gateway_host}/view/{dataset}/"
)
subpath = path[len(dataset) :] # noqa: E203
if len(subpath) == 0: if len(subpath) == 0:
r = make_response(f"Redirect to {gateway_basepath}\n", 301) r = make_response(f"Redirect to {gateway_basepath}\n", 301)
r.headers["location"] = gateway_basepath + querystring() r.headers["location"] = gateway_basepath
return r return r
elif self.status == CacheEntryStatus.loading:
launch_time = datetime.datetime.fromtimestamp(self.launchtime) port = self.port
return render_template( cellxgene_basepath = f"http://127.0.0.1:{port}"
"loading.html",
launchtime=launch_time,
all_output=self.all_output,
)
headers = {} headers = {}
copy_headers = [
"accept",
"accept-encoding",
"accept-language",
"cache-control",
"connection",
"content-length",
"content-type",
"cookie",
"host",
"origin",
"pragma",
"referer",
"sec-fetch-mode",
"sec-fetch-site",
"user-agent",
]
for h in copy_headers:
if h in request.headers:
headers[h] = request.headers[h]
full_path = self.cellxgene_basepath() + subpath + querystring() if "accept" in request.headers:
headers["accept"] = request.headers["accept"]
if "user-agent" in request.headers:
headers["user-agent"] = request.headers["user-agent"]
if "content-type" in request.headers:
headers["content-type"] = request.headers["content-type"]
if request.method in ["GET", "HEAD", "OPTIONS"]: if request.method in ["GET", "HEAD", "OPTIONS"]:
cellxgene_response = get(full_path, headers=headers) cellxgene_response = get(
cellxgene_basepath + subpath, headers=headers
)
elif request.method == "PUT": elif request.method == "PUT":
cellxgene_response = put( cellxgene_response = put(
full_path, cellxgene_basepath + subpath,
headers=headers, headers=headers,
data=request.data, data=request.data.decode(),
) )
elif request.method == "POST": elif request.method == "POST":
cellxgene_response = post( cellxgene_response = post(
full_path, cellxgene_basepath + subpath,
headers=headers, headers=headers,
data=request.data, data=request.data.decode(),
) )
else: else:
raise CellxgeneException( raise CellxgeneException(
@@ -198,20 +139,17 @@ class CacheEntry:
) )
content_type = cellxgene_response.headers["content-type"] content_type = cellxgene_response.headers["content-type"]
if "text" in content_type: if "text" in content_type:
gateway_content = self.rewrite_text_content( cellxgene_content = cellxgene_response.content.decode()
cellxgene_response.content.decode() gateway_content = cellxgene_content.replace(
) "http://fonts.gstatic.com", "https://fonts.gstatic.com"
).replace(cellxgene_basepath, gateway_basepath)
else: else:
gateway_content = cellxgene_response.content gateway_content = cellxgene_response.content
resp_headers = {}
for h in copy_headers:
if h in cellxgene_response.headers:
resp_headers[h] = cellxgene_response.headers[h]
gateway_response = make_response( gateway_response = make_response(
gateway_content, gateway_content,
cellxgene_response.status_code, cellxgene_response.status_code,
resp_headers, {"Content-Type": content_type},
) )
return gateway_response return gateway_response

View File

@@ -1,70 +0,0 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
# There are three kinds of CacheKey:
# 1) somedir/dataset.h5ad: a dataset
# in this case, descriptor == dataset == 'somedir/dataset.h5ad'
# 2) somedir/dataset_annotations/my_annotations.csv : an actual annotations file.
# in this case, descriptor == 'somedir/dataset_annotations/my_annotations.csv', dataset == 'somedir/dataset.h5ad'
# 3) somedir/dataset_annotations: an annotation directory. The corresponding h5ad must exist, but the directory may not.
# in this case, descriptor == 'somedir/dataset_annotations', dataset == 'somedir/dataset.h5ad'
from cellxgene_gateway.items.item import Item
from cellxgene_gateway.items.item_source import ItemSource, LookupResult
class CacheKey:
@property
def descriptor(self):
if self.annotation_item is None:
return self.h5ad_item.descriptor
else:
return self.annotation_item.descriptor
@property
def file_path(self):
return self.source.get_local_path(self.h5ad_item)
@property
def annotation_file_path(self):
if self.annotation_item is None:
return None
else:
return self.source.get_local_path(self.annotation_item)
@property
def source_name(self):
return self.source.name
@property
def annotation_descriptor(self):
if self.annotation_item is None:
return None
else:
return self.annotation_item.descriptor
def equals(self, other):
return (
(self.source.name == other.source.name)
and (self.h5ad_item.descriptor == other.h5ad_item.descriptor)
and (self.annotation_descriptor == other.annotation_descriptor)
)
def __init__(
self, h5ad_item: Item, source: ItemSource, annotation_item: Item = None
):
assert h5ad_item is not None
assert source is not None
self.h5ad_item = h5ad_item
self.annotation_item = annotation_item
self.source = source
@classmethod
def for_lookup(cls, source: ItemSource, lookup: LookupResult):
return CacheKey(lookup.h5ad_item, source, lookup.annotation_item)

View File

@@ -52,13 +52,43 @@ def create_dir(parent_path, dir_name):
os.mkdir(full_path) os.mkdir(full_path)
annotations_suffix = "_annotations" def recurse_dir(path):
h5ad_suffix = ".h5ad" if not os.path.exists(path):
raise CellxgeneException(
"The given path does not exist.", status.HTTP_400_BAD_REQUEST
)
def make_entry(el):
full_path = os.path.join(path, el)
if os.path.isfile(full_path):
return {
"path": full_path.replace(env.cellxgene_data, ""),
"name": el,
"type": "file",
}
elif os.path.isdir(full_path):
return {
"path": full_path,
"name": el,
"type": "directory",
"children": recurse_dir(full_path),
}
else:
raise CellxgeneException(
"Given path is neither file nor directory.",
status.HTTP_400_BAD_REQUEST,
)
return [make_entry(x) for x in os.listdir(path)]
def make_h5ad(el): def render_entries(entries):
return el[: -len(annotations_suffix)] + h5ad_suffix return "<ul>" + "\n".join([render_entry(e) for e in entries]) + "</ul>"
def make_annotations(el): def render_entry(entry):
return el[:-5] + annotations_suffix if entry["type"] == "file":
url = 'view' + '/' + entry['path'].lstrip("/")
return f"<li> <a href='{ url}'>{entry['name']}</a></li>"
elif entry["type"] == "directory":
return f"<li>{entry['name']}{render_entries(entry['children'])}</li>"

View File

@@ -7,55 +7,32 @@
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for # OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License. # the specific language governing permissions and limitations under the License.
import logging
import os import os
import socket import logging
cellxgene_location = os.environ.get("CELLXGENE_LOCATION") cellxgene_location = os.environ.get("CELLXGENE_LOCATION")
cellxgene_data = os.environ.get("CELLXGENE_DATA") cellxgene_data = os.environ.get("CELLXGENE_DATA")
cellxgene_args = os.environ.get("CELLXGENE_ARGS", None) gateway_host = os.environ.get("GATEWAY_HOST")
gateway_port = int(os.environ.get("GATEWAY_PORT", "5005")) gateway_protocol = os.environ.get("GATEWAY_PROTOCOL")
external_host = os.environ.get(
"EXTERNAL_HOST",
os.environ.get("GATEWAY_HOST", f"localhost:{gateway_port}"),
)
external_protocol = os.environ.get(
"EXTERNAL_PROTOCOL", os.environ.get("GATEWAY_PROTOCOL", "http")
)
ip = os.environ.get("GATEWAY_IP") ip = os.environ.get("GATEWAY_IP")
extra_scripts = os.environ.get("GATEWAY_EXTRA_SCRIPTS") extra_scripts = os.environ.get("GATEWAY_EXTRA_SCRIPTS")
ttl = os.environ.get("GATEWAY_TTL") ttl = os.environ.get("GATEWAY_TTL")
enable_annotations = os.environ.get( enable_upload = os.environ.get("GATEWAY_ENABLE_UPLOAD", "").lower() in ['true', '1']
"GATEWAY_ENABLE_ANNOTATIONS", ""
).lower() in [
"true",
"1",
]
enable_backed_mode = os.environ.get(
"GATEWAY_ENABLE_BACKED_MODE", ""
).lower() in [
"true",
"1",
]
env_vars = { env_vars = {
"CELLXGENE_LOCATION": cellxgene_location, "CELLXGENE_LOCATION": cellxgene_location,
"CELLXGENE_DATA": cellxgene_data,
"GATEWAY_HOST": gateway_host,
"GATEWAY_PROTOCOL": gateway_protocol,
"GATEWAY_IP": ip,
} }
optional_env_vars = { optional_env_vars = {
"EXTERNAL_HOST": external_host,
"EXTERNAL_PROTOCOL": external_protocol,
"GATEWAY_IP": ip,
"GATEWAY_PORT": gateway_port,
"GATEWAY_EXTRA_SCRIPTS": extra_scripts, "GATEWAY_EXTRA_SCRIPTS": extra_scripts,
"GATEWAY_TTL": ttl, "GATEWAY_TTL": ttl,
"GATEWAY_ENABLE_ANNOTATIONS": enable_annotations, "GATEWAY_ENABLE_UPLOAD": enable_upload,
"GATEWAY_ENABLE_BACKED_MODE": enable_backed_mode,
"CELLXGENE_ARGS": cellxgene_args,
"CELLXGENE_DATA": cellxgene_data,
} }
def validate(): def validate():
if not all(env_vars.values()): if not all(env_vars.values()):
raise ValueError( raise ValueError(
@@ -70,12 +47,11 @@ def validate():
export CELLXGENE_LOCATION=~/anaconda/envs/cellxgene-dev/bin/cellxgene export CELLXGENE_LOCATION=~/anaconda/envs/cellxgene-dev/bin/cellxgene
export CELLXGENE_DATA=../cellxgene_data export CELLXGENE_DATA=../cellxgene_data
export GATEWAY_HOST=localhost:5005
export GATEWAY_PROTOCOL=http
export GATEWAY_IP=127.0.0.1
""" """
) )
else: else:
logging.getLogger("cellxgene_gateway").info( logging.getLogger("cellxgene_gateway").info(f"Got required env: {env_vars}", )
f"Got required env: {env_vars}", logging.getLogger("cellxgene_gateway").info(f"Got optional env: {optional_env_vars}")
)
logging.getLogger("cellxgene_gateway").info(
f"Got optional env: {optional_env_vars}"
)

View File

@@ -7,14 +7,13 @@
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for # OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License. # the specific language governing permissions and limitations under the License.
from json import loads
from cellxgene_gateway import env from cellxgene_gateway import env
from json import loads
def get_extra_scripts(): def get_extra_scripts():
# can be array of script tags to inject on every page, e.g. for google analytics could be # can be array of script tags to inject on every page, e.g. for google analytics could be
# ['https://www.googletagmanager.com/gtag/js?id=UA-123456-2', # ['https://www.googletagmanager.com/gtag/js?id=UA-123456-2',
# f"{env.external_protocol}://{env.external_host}/static/js/google_ua.js"] # f"{env.gateway_protocol}://{env.gateway_host}/static/js/google_ua.js"]
# where google_ua.js is a script you add to the static/js folder prior to deployment. # where google_ua.js is a script you add to the static/js folder prior to deployment.
return [] if env.extra_scripts is None else loads(env.extra_scripts) return [] if env.extra_scripts is None else loads(env.extra_scripts)

View File

@@ -1,68 +0,0 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
import urllib.parse
def render_annotations(item, item_source):
subpath = f"/source/{urllib.parse.quote_plus(item_source.name)}/view/"
new_annotation = f"<a class='new' href='{subpath}{item_source.get_annotations_subpath(item)}'>new</a>"
annotations = (
", ".join(
[
f"<a href='{subpath}{a.descriptor}/'>{a.name}</a>"
for a in item.annotations
]
)
+ ", "
if item.annotations
else ""
)
return " | annotations: " + annotations + new_annotation
def render_item(item, item_source):
url = f"/source/{urllib.parse.quote_plus(item_source.name)}/view/{item.descriptor}/"
item_string = f"<li> <a href='{ url }'>{item.name}</a> {render_annotations(item, item_source)}</li>"
return item_string
def render_item_tree(item_tree, item_source):
items = (
"\n".join([render_item(i, item_source) for i in item_tree.items])
if item_tree.items
else ""
)
branches = (
"\n".join(
[render_item_tree(b, item_source) for b in item_tree.branches]
)
if item_tree.branches
else ""
)
html = "<ul>" + items + branches + "</ul>"
if item_tree.descriptor:
descriptor = item_tree.descriptor.lstrip("/")
url = f"/filecrawl/{descriptor}?source={item_source.name}"
name = (
descriptor.rsplit("/")[1]
if descriptor.find("/") >= 0
else descriptor
)
return f"<li><a href='{url}'>{name}</a>{html}</li>"
else:
return html
def render_item_source(item_source, filter=None):
item_tree = item_source.list_items(filter)
heading = f"<h6><a href='/filecrawl.html?source={urllib.parse.quote_plus(item_source.name)}'>{item_source.name}</a></h6>"
return heading + render_item_tree(item_tree, item_source)

View File

@@ -1,15 +0,0 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
from flask import request
def querystring():
qs = request.query_string.decode()
return f"?{qs}" if len(qs) > 0 else ""

View File

@@ -6,16 +6,16 @@
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES # under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for # OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License. # the specific language governing permissions and limitations under the License.
# import BaseHTTPServer # import BaseHTTPServer
import json import datetime
import logging
import os import os
import urllib.parse import logging
from threading import Lock, Thread from threading import Thread, Lock
import json
from flask import ( from flask import (
Flask, Flask,
make_response,
redirect, redirect,
render_template, render_template,
request, request,
@@ -23,38 +23,21 @@ from flask import (
url_for, url_for,
) )
from flask_api import status from flask_api import status
from werkzeug.utils import secure_filename from werkzeug import secure_filename
from cellxgene_gateway import env from cellxgene_gateway import env
from cellxgene_gateway.backend_cache import BackendCache from cellxgene_gateway.backend_cache import BackendCache
from cellxgene_gateway.cache_entry import CacheEntryStatus
from cellxgene_gateway.cellxgene_exception import CellxgeneException from cellxgene_gateway.cellxgene_exception import CellxgeneException
from cellxgene_gateway.dir_util import create_dir, is_subdir from cellxgene_gateway.dir_util import create_dir, recurse_dir, render_entries, is_subdir
from cellxgene_gateway.extra_scripts import get_extra_scripts from cellxgene_gateway.extra_scripts import get_extra_scripts
from cellxgene_gateway.filecrawl import render_item_source from cellxgene_gateway.path_util import get_dataset, get_file_path
from cellxgene_gateway.process_exception import ProcessException from cellxgene_gateway.process_exception import ProcessException
from cellxgene_gateway.prune_process_cache import PruneProcessCache from cellxgene_gateway.prune_process_cache import PruneProcessCache
from cellxgene_gateway.util import current_time_stamp from cellxgene_gateway.util import current_time_stamp
from cellxgene_gateway.cache_key import CacheKey
app = Flask(__name__) app = Flask(__name__)
item_sources = []
default_item_source = None
def _force_https(app):
def wrapper(environ, start_response):
environ["wsgi.url_scheme"] = env.external_protocol
return app(environ, start_response)
return wrapper
app.wsgi_app = _force_https(app.wsgi_app)
cache = BackendCache() cache = BackendCache()
location = f"{env.external_protocol}://{env.external_host}" location = f"{env.gateway_protocol}://{env.gateway_host}"
@app.errorhandler(CellxgeneException) @app.errorhandler(CellxgeneException)
@@ -90,8 +73,7 @@ def handle_invalid_process(error):
http_status=error.http_status, http_status=error.http_status,
stdout=error.stdout, stdout=error.stdout,
stderr=error.stderr, stderr=error.stderr,
dataset=error.key.h5ad_item.descriptor, dataset=error.dataset,
annotation_file=error.key.annotation_descriptor,
), ),
error.http_status, error.http_status,
) )
@@ -108,94 +90,105 @@ def favicon():
@app.route("/") @app.route("/")
def index(): def index():
users = [
name
for name in os.listdir(env.cellxgene_data)
if os.path.isdir(os.path.join(env.cellxgene_data, name))
]
return render_template( return render_template(
"index.html", "index.html",
ip=env.ip, ip=env.ip,
cellxgene_data=env.cellxgene_data, cellxgene_data=env.cellxgene_data,
extra_scripts=get_extra_scripts(), extra_scripts=get_extra_scripts(),
users=users,
enable_upload=env.enable_upload,
) )
def make_user():
dir_name = request.form["directory"]
create_dir(env.cellxgene_data, dir_name)
return redirect(location, code=302)
def make_subdir():
parent_path = os.path.join(env.cellxgene_data, request.form["usernames"])
dir_name = request.form["directory"]
create_dir(parent_path, dir_name)
return redirect(location, code=302)
def upload_file():
upload_dir = request.form["path"]
full_upload_path = os.path.join(env.cellxgene_data, upload_dir)
if is_subdir(full_upload_path, env.cellxgene_data) and os.path.isdir(full_upload_path):
if request.method == "POST":
if "file" in request.files:
f = request.files["file"]
if f and f.filename.endswith(".h5ad"):
f.save(
os.path.join(full_upload_path, secure_filename(f.filename))
)
return redirect("/filecrawl.html", code=302)
else:
raise CellxgeneException(
"Uploaded file must be in anndata (.h5ad) format.",
status.HTTP_400_BAD_REQUEST,
)
else:
raise CellxgeneException(
"A file must be chosen to upload.",
status.HTTP_400_BAD_REQUEST,
)
else:
raise CellxgeneException(
"Invalid directory.", status.HTTP_400_BAD_REQUEST
)
return redirect(env.location, code=302)
if env.enable_upload:
app.add_url_rule('/make_user', 'make_user', make_user, methods=["POST"])
app.add_url_rule('/make_subdir', 'make_subdir', make_subdir, methods=["POST"])
app.add_url_rule('/upload_file', 'upload_file', upload_file, methods=["POST"])
@app.route("/filecrawl.html") @app.route("/filecrawl.html")
@app.route("/filecrawl/<path:path>") def filecrawl():
def filecrawl(path=None):
source_name = request.args.get("source")
sources = (
filter(
lambda x: x.name == urllib.parse.unquote_plus(source_name),
item_sources,
)
if source_name
else item_sources
)
# loop all data sources --
rendered_sources = [
render_item_source(item_source, path) for item_source in sources
] # will we need to make this async in the page???
rendered_html = "\n".join(rendered_sources)
resp = make_response( entries = recurse_dir(env.cellxgene_data)
render_template( rendered_html = render_entries(entries)
"filecrawl.html", return render_template(
extra_scripts=get_extra_scripts(), "filecrawl.html",
rendered_html=rendered_html, extra_scripts=get_extra_scripts(),
path=path, rendered_html=rendered_html,
)
) )
resp.headers["Cache-Control"] = "no-cache, no-store, must-revalidate"
resp.headers["Pragma"] = "no-cache"
resp.headers["Expires"] = "0"
resp.headers["Cache-Control"] = "public, max-age=0"
return resp
entry_lock = Lock() entry_lock = Lock()
def matching_source(source_name):
if source_name is None:
source_name = default_item_source.name
matching = [i for i in item_sources if i.name == source_name]
if len(matching) != 1:
raise Exception(f"Could not find matching item source {source_name}")
source = matching[0]
return source
@app.route(
"/source/<path:source_name>/view/<path:path>",
methods=["GET", "PUT", "POST"],
)
@app.route("/view/<path:path>", methods=["GET", "PUT", "POST"]) @app.route("/view/<path:path>", methods=["GET", "PUT", "POST"])
def do_view(path, source_name=None): def do_view(path):
source = matching_source(source_name) dataset = get_dataset(path)
match = cache.check_path(source, path) file_path = get_file_path(dataset)
with entry_lock:
if match is None: match = cache.check_entry(dataset)
lookup = source.lookup(path) if match is None:
if lookup is None: uascripts = get_extra_scripts()
raise CellxgeneException( match = cache.create_entry(dataset, file_path, uascripts)
f"Could not find item for path {path} in source {source.name}",
404,
)
key = CacheKey.for_lookup(source, lookup)
print(
f"view path={path}, source_name={source_name}, dataset={key.file_path}, annotation_file= {key.annotation_file_path}, key={key.descriptor}, source={key.source_name}"
)
with entry_lock:
match = cache.check_entry(key)
if match is None:
uascripts = get_extra_scripts()
match = cache.create_entry(key, uascripts)
match.timestamp = current_time_stamp() match.timestamp = current_time_stamp()
if ( if match.status == "loaded":
match.status == CacheEntryStatus.loaded
or match.status == CacheEntryStatus.loading
):
return match.serve_content(path) return match.serve_content(path)
elif match.status == CacheEntryStatus.error: elif match.status == "loading":
launch_time = datetime.datetime.fromtimestamp(match.launchtime)
return render_template(
"loading.html", launchtime=launch_time, all_output=match.all_output
)
elif match.status == "error":
raise ProcessException.from_cache_entry(match) raise ProcessException.from_cache_entry(match)
@@ -203,92 +196,43 @@ def do_view(path, source_name=None):
def do_GET_status(): def do_GET_status():
return render_template("cache_status.html", entry_list=cache.entry_list) return render_template("cache_status.html", entry_list=cache.entry_list)
@app.route("/cache_status.json", methods=["GET"]) @app.route("/cache_status.json", methods=["GET"])
def do_GET_status_json(): def do_GET_status_json():
return json.dumps( return json.dumps({'launchtime':app.launchtime,
{ 'entry_list':[{
"launchtime": app.launchtime, 'dataset': entry.dataset,
"entry_list": [ 'launchtime': entry.launchtime,
{ 'last_access': entry.timestamp,
"dataset": entry.key.dataset, 'status': entry.status
"annotation_file": entry.key.annotation_file, } for entry in cache.entry_list]})
"launchtime": entry.launchtime,
"last_access": entry.timestamp,
"status": entry.status,
}
for entry in cache.entry_list
],
}
)
@app.route("/relaunch/<path:path>", methods=["GET"]) @app.route("/relaunch/<path:path>", methods=["GET"])
def do_relaunch(path): def do_relaunch(path):
source_name = request.args.get("source") or default_item_source.name dataset = get_dataset(path)
source = matching_source(source_name) match = cache.check_entry(dataset)
key = CacheKey.for_lookup(source, source.lookup(path))
match = cache.check_entry(key)
if not match is None: if not match is None:
match.terminate() match.terminate()
qs = request.query_string.decode() return redirect(url_for("do_view", path=path), code=302)
return redirect(
url_for("do_view", path=path) + (f"?{qs}" if len(qs) > 0 else ""),
code=302,
)
@app.route("/terminate/<path:path>", methods=["GET"]) @app.route("/terminate/<path:path>", methods=["GET"])
def do_terminate(path): def do_terminate(path):
source_name = request.args.get("source_name") or default_item_source.name dataset = get_dataset(path)
source = matching_source(source_name) match = cache.check_entry(dataset)
key = CacheKey.for_lookup(source, source.lookup(path))
match = cache.check_entry(key)
if not match is None: if not match is None:
match.terminate() match.terminate()
return redirect(url_for("do_GET_status"), code=302) return redirect(url_for("do_GET_status"), code=302)
def launch(): def main():
logging.basicConfig(level=logging.INFO, format='%(asctime)s:%(name)s:%(levelname)s:%(message)s')
env.validate() env.validate()
if not item_sources or not len(item_sources):
raise Exception("No data sources specified for Cellxgene Gateway")
global default_item_source
if default_item_source is None:
default_item_source = item_sources[0]
pruner = PruneProcessCache(cache) pruner = PruneProcessCache(cache)
background_thread = Thread(target=pruner) background_thread = Thread(target=pruner)
background_thread.start() background_thread.start()
app.launchtime = current_time_stamp() app.launchtime = current_time_stamp()
app.run(host="0.0.0.0", port=env.gateway_port, debug=False) app.run(host="0.0.0.0", port=5005, debug=False)
def main():
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s:%(name)s:%(levelname)s:%(message)s",
)
cellxgene_data = os.environ.get("CELLXGENE_DATA", None)
cellxgene_bucket = os.environ.get("CELLXGENE_BUCKET", None)
if cellxgene_bucket is not None:
from cellxgene_gateway.items.s3.s3item_source import S3ItemSource
item_sources.append(S3ItemSource(cellxgene_bucket, name="s3"))
default_item_source = "s3"
if cellxgene_data is not None:
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
item_sources.append(FileItemSource(cellxgene_data, name="local"))
default_item_source = "local"
if len(item_sources) == 0:
raise Exception("Please specify CELLXGENE_DATA or CELLXGENE_BUCKET")
launch()
if __name__ == "__main__": if __name__ == "__main__":

View File

@@ -1,27 +0,0 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
from cellxgene_gateway.items.item import Item
class FileItem(Item):
"""e.g. FileItem(subpath = subpath, name = filename, type = ItemType.h5ad)
The Item superclass expects a 'name' and 'type'.
"""
def __init__(self, subpath: str, *args, **kwargs):
super().__init__(*args, **kwargs)
self.subpath = subpath
@property
def descriptor(self) -> str:
return os.path.join(self.subpath, self.name).strip("/")

View File

@@ -1,185 +0,0 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
from typing import List
from cellxgene_gateway import dir_util
from cellxgene_gateway.items.file.fileitem import FileItem
from cellxgene_gateway.items.item import ItemTree, ItemType
from cellxgene_gateway.items.item_source import ItemSource, LookupResult
class FileItemSource(ItemSource):
def __init__(
self,
base_path,
name=None,
h5ad_suffix=dir_util.h5ad_suffix,
annotation_dir_suffix=dir_util.annotations_suffix,
annotation_file_suffix=".csv",
):
self._name = name
self.base_path = base_path
self.h5ad_suffix = h5ad_suffix
self.annotation_dir_suffix = annotation_dir_suffix
self.annotation_file_suffix = annotation_file_suffix
@property
def name(self):
return self._name or f"Files:{self.base_path}"
def is_h5ad_file(self, path: str) -> bool:
return path.endswith(self.h5ad_suffix) and os.path.isfile(path)
def convert_annotation_path_to_h5ad(self, path):
return path[: -len(self.annotation_dir_suffix)] + self.h5ad_suffix
def convert_h5ad_path_to_annotation(self, path):
return path[: -len(self.h5ad_suffix)] + self.annotation_dir_suffix
def get_local_path(self, item: FileItem) -> str:
return os.path.join(self.base_path, item.descriptor)
def get_annotations_subpath(self, item) -> str:
return self.convert_h5ad_path_to_annotation(item.descriptor)
def list_items(self, filter: str = None) -> ItemTree:
item_tree = self.scan_directory()
"""def get_items(dir):
if dir.branches:
return [*dir.items, *[item for subdir in dir.branches for item in get_items(subdir)]]
else:
return dir.items
return get_items(self.item_tree)"""
return item_tree
def scan_directory(self, subpath="") -> dict:
base_path = os.path.join(self.base_path, subpath)
if not os.path.exists(base_path):
raise Exception(
f"Path for local files '{base_path}' does not exist."
)
filepath_map = dict(
(filepath, os.path.join(base_path, filepath))
for filepath in sorted(os.listdir(base_path))
)
def is_annotation_dir(dir):
return (
dir.endswith(self.annotation_dir_suffix)
and self.convert_annotation_path_to_h5ad(dir) in h5ad_paths
)
h5ad_paths = [
filepath
for filepath, full_path in filepath_map.items()
if self.is_h5ad_file(full_path)
]
subdirs = [
filepath
for filepath, full_path in filepath_map.items()
if os.path.isdir(full_path) and not is_annotation_dir(filepath)
]
items = [
self.make_fileitem_from_path(filename, subpath)
for filename in h5ad_paths
]
branches = None
if len(subdirs) > 0:
branches = [
self.scan_directory(os.path.join(subpath, subdir))
for subdir in subdirs
]
return ItemTree(subpath, items, branches)
def create_annotation(self, item: FileItem, name: str) -> FileItem:
annotation = self.make_fileitem_from_path(
name, self.get_annotations_subpath(item), is_annotation=True
)
item.annotations = (item.annotations or []).append(annotation)
return annotation
def update(self, item: FileItem) -> None:
pass
def full_path(self, p):
return os.path.join(self.base_path, p)
def lookup_item(self, descriptor):
full_path = self.full_path(descriptor)
if self.is_h5ad_file(full_path):
return self.shallowitem_from_descriptor(descriptor)
def lookup(self, indescriptor: str) -> LookupResult:
descriptor = indescriptor.strip("/")
if descriptor.endswith(self.annotation_file_suffix):
annotation_item = self.shallowitem_from_descriptor(
descriptor, True
)
h5ad_descriptor = self.convert_annotation_path_to_h5ad(
annotation_item.subpath
)
item = self.lookup_item(h5ad_descriptor)
if item is not None:
return LookupResult(item, annotation_item)
else:
item = self.lookup_item(descriptor)
if item is not None:
return LookupResult(item)
def shallowitem_from_descriptor(self, descriptor, is_annotation=False):
filename = os.path.basename(descriptor)
subpath = os.path.dirname(descriptor)
return self.make_fileitem_from_path(
filename,
subpath,
is_annotation,
True,
)
def make_fileitem_from_path(
self, filename, subpath, is_annotation=False, is_shallow=False
) -> FileItem:
item = FileItem(
subpath=subpath,
name=filename,
type=ItemType.annotation if is_annotation else ItemType.h5ad,
)
if not is_annotation and not is_shallow:
annotations = self.make_annotations_for_fileitem(item)
item.annotations = annotations
return item
def make_annotations_for_fileitem(self, item: FileItem) -> List[FileItem]:
annotations_subpath = self.get_annotations_subpath(item)
annotations_fullpath = self.full_path(annotations_subpath)
if os.path.isdir(annotations_fullpath):
return [
self.make_fileitem_from_path(
annotation, annotations_subpath, True
)
for annotation in sorted(os.listdir(annotations_fullpath))
if annotation.endswith(self.annotation_file_suffix)
and os.path.isfile(
os.path.join(annotations_fullpath, annotation)
)
]
else:
return None

View File

@@ -1,43 +0,0 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
from abc import ABC, abstractmethod
from enum import Enum
from typing import List
class ItemType(Enum):
annotation = "annotation"
h5ad = "h5ad"
class Item(ABC):
def __init__(
self, name: str, type: ItemType, annotations: List["Item"] = None
):
self.name = name
self.type = type
self.annotations = annotations
@property
@abstractmethod
def descriptor(self):
raise Exception('"descriptor" not implemented')
class ItemTree:
def __init__(
self,
descriptor: str,
items: List[Item] = None,
branches: List["ItemTree"] = None,
):
self.descriptor = descriptor
self.items = items
self.branches = branches

View File

@@ -1,50 +0,0 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
from abc import ABC, abstractmethod
from typing import List
from cellxgene_gateway.items.item import Item
class LookupResult:
def __init__(self, h5ad_item: Item, annotation_item: Item = None):
self.h5ad_item = h5ad_item
self.annotation_item = annotation_item
class ItemSource(ABC):
@abstractmethod
def list_items(self, filter: str = None) -> List[Item]:
raise Exception('"list_items" unimplemented')
@abstractmethod
def get_local_path(self, item: Item) -> str:
raise Exception('"local_path" unimplemented')
@abstractmethod
def get_annotations_subpath(self, item) -> str:
raise Exception('"annotations_path" unimplemented')
@abstractmethod
def create_annotation(self, item: Item, name: str) -> Item:
raise Exception('"annotation" unimplemented')
@abstractmethod
def update(self, item: Item) -> None:
raise Exception('"update" unimplemented')
@abstractmethod
def lookup(self, descriptor: str) -> LookupResult:
raise Exception('"lookup" unimplemented')
@property
@abstractmethod
def name(self):
pass

View File

@@ -1,27 +0,0 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
import os
from cellxgene_gateway.items.item import Item
class S3Item(Item):
"""e.g. FileItem(subpath = subpath, name = filename, type = ItemType.h5ad)
The Item superclass expects a 'name' and 'type'.
"""
def __init__(self, s3key: str, *args, **kwargs):
super().__init__(*args, **kwargs)
self.s3key = s3key
@property
def descriptor(self) -> str:
return self.s3key

View File

@@ -1,171 +0,0 @@
# Copyright 2019 Novartis Institutes for BioMedical Research Inc. Licensed
# under the Apache License, Version 2.0 (the "License"); you may not use
# this file except in compliance with the License. You may obtain a copy
# of the License at http://www.apache.org/licenses/LICENSE-2.0. Unless
# required by applicable law or agreed to in writing, software distributed
# under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License.
from typing import List
from os.path import join, dirname, basename
from cellxgene_gateway import dir_util
import s3fs
from cellxgene_gateway.items.s3.s3item import S3Item
from cellxgene_gateway.items.item import ItemTree, ItemType
from cellxgene_gateway.items.item_source import ItemSource, LookupResult
class S3ItemSource(ItemSource):
def __init__(
self,
bucket,
name=None,
h5ad_suffix=dir_util.h5ad_suffix,
annotation_dir_suffix=dir_util.annotations_suffix,
annotation_file_suffix=".csv",
):
self._name = name
self.s3 = s3fs.S3FileSystem()
self.bucket = bucket
self.h5ad_suffix = h5ad_suffix
self.annotation_dir_suffix = annotation_dir_suffix
self.annotation_file_suffix = annotation_file_suffix
def url(self, path):
return "s3://" + join(self.bucket, path)
@property
def name(self):
return self._name or f"Items:{self.url('')}"
def is_h5ad_url(self, s3url: str) -> bool:
return s3url.endswith(self.h5ad_suffix) and self.s3.exists(s3url)
def convert_annotation_key_to_h5ad(self, s3key):
return s3key[: -len(self.annotation_dir_suffix)] + self.h5ad_suffix
def convert_h5ad_key_to_annotation(self, s3key):
return s3key[: -len(self.h5ad_suffix)] + self.annotation_dir_suffix
def get_local_path(self, item: S3Item) -> str:
return self.url(item.descriptor)
def get_annotations_subpath(self, item) -> str:
return self.convert_h5ad_key_to_annotation(item.descriptor)
def list_items(self, filter: str = None) -> ItemTree:
item_tree = self.scan_directory()
return item_tree
def scan_directory(self, subpath="") -> dict:
url = self.url(subpath)
if not self.s3.exists(url):
raise Exception(f"S3 url '{url}' does not exist.")
s3key_map = dict(
(filepath[len(self.bucket) :], "s3://" + filepath)
for filepath in sorted(self.s3.ls(url))
)
def is_annotation_dir(dir_s3key):
return (
dir_s3key.endswith(self.annotation_dir_suffix)
and self.convert_annotation_key_to_h5ad(dir_s3key)
in h5ad_paths
)
h5ad_paths = [
filepath
for filepath, item_url in s3key_map.items()
if self.is_h5ad_url(item_url)
]
subdirs = [
filepath
for filepath, item_url in s3key_map.items()
if self.s3.isdir(item_url) and not is_annotation_dir(filepath)
]
items = [
self.make_s3item_from_key(filename, join(subpath, filename))
for filename in h5ad_paths
]
branches = None
if len(subdirs) > 0:
branches = [
self.scan_directory(join(subpath, subdir))
for subdir in subdirs
]
return ItemTree(subpath, items, branches)
def create_annotation(self, item: S3Item, name: str) -> S3Item:
annotation = self.make_s3item_from_key(
name, self.get_annotations_subpath(item), is_annotation=True
)
item.annotations = (item.annotations or []).append(annotation)
return annotation
def update(self, item: S3Item) -> None:
pass
def lookup_item(self, descriptor):
full_path = self.url(descriptor)
if self.is_h5ad_url(full_path):
return self.shallowitem_from_descriptor(descriptor)
def lookup(self, indescriptor: str) -> LookupResult:
descriptor = indescriptor.strip("/")
if descriptor.endswith(self.annotation_file_suffix):
annotation_item = self.shallowitem_from_descriptor(
descriptor, True
)
if not self.s3.exists(self.url(annotation_item.s3key)):
with self.s3.open(self.url(annotation_item.s3key), "w") as f:
f.write("")
h5ad_descriptor = self.convert_annotation_key_to_h5ad(
dirname(annotation_item.s3key)
)
item = self.shallowitem_from_descriptor(h5ad_descriptor)
return LookupResult(item, annotation_item)
else:
item = self.lookup_item(descriptor)
if item is not None:
return LookupResult(item)
def shallowitem_from_descriptor(self, descriptor, is_annotation=False):
return self.make_s3item_from_key(
basename(descriptor), descriptor, is_annotation, True
)
def make_s3item_from_key(
self, name, s3key, is_annotation=False, is_shallow=False
) -> S3Item:
item = S3Item(
s3key=s3key,
name=name,
type=ItemType.annotation if is_annotation else ItemType.h5ad,
)
if not is_annotation and not is_shallow:
annotations = self.make_annotations_for_fileitem(item)
item.annotations = annotations
return item
def make_annotations_for_fileitem(self, item: S3Item) -> List[S3Item]:
annotations_subpath = self.get_annotations_subpath(item)
annotations_fullpath = self.url(annotations_subpath)
if self.s3.isdir(annotations_fullpath):
return [
self.make_s3item_from_key(
annotation, join(annotations_subpath, annotation), True
)
for annotation in sorted(self.s3.ls(annotations_fullpath))
if annotation.endswith(self.annotation_file_suffix)
and self.s3.isfile(join(annotations_fullpath, annotation))
]
else:
return None

View File

@@ -11,20 +11,31 @@ import os
from flask_api import status from flask_api import status
from cellxgene_gateway.cache_key import CacheKey from cellxgene_gateway import env
from cellxgene_gateway.cellxgene_exception import CellxgeneException from cellxgene_gateway.cellxgene_exception import CellxgeneException
from cellxgene_gateway.dir_util import make_h5ad
def validate_exists(file_path): def get_dataset(path):
if path == "/" or path == "":
raise CellxgeneException(
"No matching dataset found.", status.HTTP_404_NOT_FOUND
)
trimmed = path[:-1] if path[-1] == "/" else path
try:
get_file_path(trimmed)
return trimmed
except CellxgeneException:
split = os.path.split(trimmed)
return get_dataset(split[0])
def validate_path(file_path):
if not os.path.exists(file_path): if not os.path.exists(file_path):
raise CellxgeneException( raise CellxgeneException(
"File does not exist: " + file_path, status.HTTP_400_BAD_REQUEST "File does not exist: " + file_path, status.HTTP_400_BAD_REQUEST
) )
def validate_is_file(file_path):
validate_exists(file_path)
if not os.path.isfile(file_path): if not os.path.isfile(file_path):
raise CellxgeneException( raise CellxgeneException(
"Path is not file: " + file_path, status.HTTP_400_BAD_REQUEST "Path is not file: " + file_path, status.HTTP_400_BAD_REQUEST
@@ -32,10 +43,7 @@ def validate_is_file(file_path):
return return
def validate_is_dir(file_path): def get_file_path(dataset):
validate_exists(file_path) file_path = os.path.join(env.cellxgene_data, dataset)
if not os.path.isdir(file_path): validate_path(file_path)
raise CellxgeneException( return file_path
"Path is not dir: " + file_path, status.HTTP_400_BAD_REQUEST
)
return

View File

@@ -9,13 +9,13 @@
class ProcessException(Exception): class ProcessException(Exception):
def __init__(self, message, stdout, stderr, http_status, key): def __init__(self, message, stdout, stderr, http_status, dataset):
Exception.__init__(self) Exception.__init__(self)
self.message = message self.message = message
self.stdout = stdout self.stdout = stdout
self.stderr = stderr self.stderr = stderr
self.http_status = http_status self.http_status = http_status
self.key = key self.dataset = dataset
@classmethod @classmethod
def from_cache_entry(cls, cache_entry): def from_cache_entry(cls, cache_entry):
@@ -24,5 +24,5 @@ class ProcessException(Exception):
cache_entry.all_output, cache_entry.all_output,
cache_entry.stderr, cache_entry.stderr,
cache_entry.http_status, cache_entry.http_status,
cache_entry.key, cache_entry.dataset,
) )

View File

@@ -7,17 +7,16 @@
# OR CONDITIONS OF ANY KIND, either express or implied. See the License for # OR CONDITIONS OF ANY KIND, either express or implied. See the License for
# the specific language governing permissions and limitations under the License. # the specific language governing permissions and limitations under the License.
import logging
import time import time
import logging
from cellxgene_gateway.env import ttl
from cellxgene_gateway.util import current_time_stamp from cellxgene_gateway.util import current_time_stamp
from cellxgene_gateway.env import ttl
class PruneProcessCache: class PruneProcessCache:
def __init__(self, cache): def __init__(self, cache):
self.cache = cache self.cache = cache
self.expire_seconds = 3600 if ttl is None else int(ttl) self.expire_seconds = (3600 if ttl is None else int(ttl))
def __call__(self): def __call__(self):
while True: while True:
@@ -27,24 +26,16 @@ class PruneProcessCache:
def prune(self): def prune(self):
timestamp = current_time_stamp() timestamp = current_time_stamp()
cutoff = timestamp - self.expire_seconds cutoff = timestamp - self.expire_seconds
processes_to_delete = [ def prunable(p):
p for p in self.cache.entry_list if p.timestamp < cutoff return p.timestamp < cutoff and p.pid != None
] processes_to_delete = [p for p in self.cache.entry_list if prunable(p)]
processes_to_keep = [ processes_to_keep = [p for p in self.cache.entry_list if not prunable(p)]
p for p in self.cache.entry_list if not p.timestamp < cutoff
]
logger = logging.getLogger("cellxgene_gateway") logger = logging.getLogger("cellxgene_gateway")
logger.debug( logger.debug(f"Cutoff {cutoff} = timestamp {timestamp} - expire seconds {self.expire_seconds} , keeping {processes_to_keep}")
f"Cutoff {cutoff} = timestamp {timestamp} - expire seconds {self.expire_seconds} , keeping {processes_to_keep}"
)
for process in processes_to_delete: for process in processes_to_delete:
try: try:
logger.info( logger.info(f"pruning process {process.pid} ({process.dataset})")
f"pruning process {process.pid} ({process.key.dataset})"
)
self.cache.prune(process) self.cache.prune(process)
except Exception: except Exception:
logger.exception( logger.exception("failed to prune process {process.pid} ({process.dataset})")
"failed to prune process {process.pid} ({process.dataset})"
)

View File

@@ -1,20 +0,0 @@
// neandertal javascript
// TODO: rewrite this --
const new_annotation_callback = (() =>{
const suffix = `.csv`;
return (e) => {
e.preventDefault();
const el = $(e.target);
const href = el.attr('href');
const base = prompt(`Name your annotations collection\nnote: the suffix "${suffix}" will be appended`);
if (base !== null && base.length > 0) {
if (/^[0-9a-zA-Z_]+$/.test(base)) {
window.location = `${href}/${base}${suffix}`;
} else {
alert("Error: name must match ^[0-9a-zA-Z_]+$\nthat is, only numbers, letters and underscore are allowed")
}
}
return false;
}
})()

View File

@@ -12,13 +12,6 @@ import subprocess
from flask_api import status from flask_api import status
from cellxgene_gateway.cache_entry import CacheEntryStatus
from cellxgene_gateway.dir_util import make_annotations
from cellxgene_gateway.env import (
enable_annotations,
enable_backed_mode,
cellxgene_args,
)
from cellxgene_gateway.process_exception import ProcessException from cellxgene_gateway.process_exception import ProcessException
@@ -26,28 +19,13 @@ class SubprocessBackend:
def __init__(self): def __init__(self):
pass pass
def create_cmd( def create_cmd(self, cellxgene_loc, file_path, port, scripts):
self, cellxgene_loc, file_path, port, scripts, annotation_file_path
):
if enable_annotations and not annotation_file_path is None:
if annotation_file_path == "":
extra_args = (
f" --annotations-dir {make_annotations(file_path)}"
)
else:
extra_args = f" --annotations-file {annotation_file_path}"
else:
extra_args = " --disable-annotations"
if enable_backed_mode:
extra_args += " --backed"
if not cellxgene_args is None:
extra_args += f" {cellxgene_args}"
cmd = ( cmd = (
f"yes | {cellxgene_loc} launch {file_path}" f"yes | {cellxgene_loc} launch {file_path}"
+ f" --port {port}" + " --port "
+ str(port)
+ " --host 127.0.0.1" + " --host 127.0.0.1"
+ extra_args
) )
for s in scripts: for s in scripts:
@@ -56,12 +34,9 @@ class SubprocessBackend:
return cmd return cmd
def launch(self, cellxgene_loc, scripts, cache_entry): def launch(self, cellxgene_loc, scripts, cache_entry):
cmd = self.create_cmd( cmd = self.create_cmd(
cellxgene_loc, cellxgene_loc, cache_entry.file_path, cache_entry.port, scripts
cache_entry.key.file_path,
cache_entry.port,
scripts,
cache_entry.key.annotation_file_path,
) )
logging.getLogger("cellxgene_gateway").info(f"launching {cmd}") logging.getLogger("cellxgene_gateway").info(f"launching {cmd}")
process = subprocess.Popen( process = subprocess.Popen(
@@ -84,7 +59,7 @@ class SubprocessBackend:
message = "Cellxgene failed to launch dataset." message = "Cellxgene failed to launch dataset."
http_status = status.HTTP_500_INTERNAL_SERVER_ERROR http_status = status.HTTP_500_INTERNAL_SERVER_ERROR
cache_entry.status = CacheEntryStatus.error cache_entry.status = "error"
cache_entry.set_error(message, stderr, http_status) cache_entry.set_error(message, stderr, http_status)
raise ProcessException.from_cache_entry(cache_entry) raise ProcessException.from_cache_entry(cache_entry)

View File

@@ -10,68 +10,57 @@
--> -->
<html> <html>
<head> <head>
<title>Cellxgene Gateway - FILE CRAWLER</title> <title>Cellxgene Gateway - FILE CRAWLER</title>
<script src="https://ajax.googleapis.com/ajax/libs/jquery/3.3.1/jquery.min.js"></script> <script src="https://ajax.googleapis.com/ajax/libs/jquery/3.3.1/jquery.min.js"></script>
<link rel="icon" type="image/png" href="{{ url_for('static', filename='nibr.ico') }}"> <link rel="icon" type="image/png" href="{{ url_for('static', filename='nibr.ico') }}">
{% for script in extra_scripts %} {% for script in extra_scripts %}
<script src="{{ script }}"></script> <script src="{{ script }}"></script>
{% endfor %} {% endfor %}
<link rel="stylesheet" href="https://stackpath.bootstrapcdn.com/bootstrap/4.1.3/css/bootstrap.min.css" <link rel="stylesheet" href="https://stackpath.bootstrapcdn.com/bootstrap/4.1.3/css/bootstrap.min.css" integrity="sha384-MCw98/SFnGE8fJT3GXwEOngsV7Zt27NXFoaoApmYm81iuXoPkFOJwJ8ERdknLPMO" crossorigin="anonymous">
integrity="sha384-MCw98/SFnGE8fJT3GXwEOngsV7Zt27NXFoaoApmYm81iuXoPkFOJwJ8ERdknLPMO" crossorigin="anonymous">
</head> </head>
<body> <body>
<header class="navbar navbar-expand navbar-dark flex-column flex-md-row bd-navbar"> <header class="navbar navbar-expand navbar-dark flex-column flex-md-row bd-navbar">
<h3>Cellxgene Gateway - Cache Status</h3> <h3>Cellxgene Gateway - Cache Status</h3>
</header> </header>
<br> <br>
<table class="table"> <table class="table">
<thead> <thead>
<tr> <tr>
<th>PID</th> <th>PID</th>
<th>dataset</th> <th>dataset</th>
<th>annotation_file</th> <th>port</th>
<th>source</th> <th>launchtime</th>
<th>port</th> <th>last access</th>
<th>launchtime</th> <th>status</th>
<th>last access</th> <th>message</th>
<th>status</th> <th>http_status</th>
<th>message</th> <th>actions</th>
<th>http_status</th> </tr>
<th>actions</th> </thead>
</tr> <tbody>
</thead> {% for entry in entry_list %}
<tbody> <tr>
{% for entry in entry_list %} <td>{{ entry.pid }}</td>
<tr> <td><a href="{{ url_for('do_view', path=entry.dataset) }}">{{ entry.dataset }}</a></td>
<td>{{ entry.pid }}</td> <td>{{ entry.port }}</td>
<td><a <td class="timestamp">{{ entry.launchtime }}</td>
href="{{ url_for('do_view', path=entry.key.descriptor, source_name=entry.key.source_name) }}">{{ entry.key.h5ad_item.descriptor }}</a> <td class="timestamp">{{ entry.timestamp }}</td>
</td> <td>{{ entry.status }}</td>
<td>{{ entry.key.annotation_descriptor }}</td> <td>{{ entry.message }}</td>
<td>{{ entry.source_name }}</td> <td>{{ entry.http_status }}</td>
<td>{{ entry.port }}</td> <td>
<td class="timestamp">{{ entry.launchtime }}</td> {% if entry.status == 'loaded' %}
<td class="timestamp">{{ entry.timestamp }}</td> <a href="{{ url_for('do_terminate', path=entry.dataset) }}"> terminate </a>
<td>{{ entry.status.name }}</td> {% endif %}
<td>{{ entry.message }}</td> </td>
<td>{{ entry.http_status }}</td> </tr>
<td> {% endfor %}
{% if entry.status.name == 'loaded' %} </tbody>
<a
href="{{ url_for('do_terminate', path=entry.key.descriptor, source_name=entry.key.source_name) }}">
terminate </a>
{% endif %}
</td>
</tr>
{% endfor %}
</tbody>
</table> </table>
<script> <script>
$(() => { $(() => {
$(".timestamp").each(function () { $(".timestamp").each(function(){
const el = $(this); const el = $(this);
const ts = el.text(); const ts = el.text();
const dt = new Date(parseInt(ts * 1000)); const dt = new Date(parseInt(ts * 1000));
@@ -80,5 +69,4 @@
}) })
</script> </script>
</body> </body>
</html> </html>

View File

@@ -16,36 +16,19 @@
<link rel="icon" type="image/png" href="{{ url_for('static', filename='nibr.ico') }}"> <link rel="icon" type="image/png" href="{{ url_for('static', filename='nibr.ico') }}">
{% for script in extra_scripts %} {% for script in extra_scripts %}
<script src="{{ script }}"></script> <script src="{{ script }}"></script>
{% endfor %} {% endfor %}
<script src="{{ url_for('static', filename='js/annotation.js') }}"></script>
<link rel="stylesheet" href="https://stackpath.bootstrapcdn.com/bootstrap/4.1.3/css/bootstrap.min.css" integrity="sha384-MCw98/SFnGE8fJT3GXwEOngsV7Zt27NXFoaoApmYm81iuXoPkFOJwJ8ERdknLPMO" crossorigin="anonymous"> <link rel="stylesheet" href="https://stackpath.bootstrapcdn.com/bootstrap/4.1.3/css/bootstrap.min.css" integrity="sha384-MCw98/SFnGE8fJT3GXwEOngsV7Zt27NXFoaoApmYm81iuXoPkFOJwJ8ERdknLPMO" crossorigin="anonymous">
</head> </head>
<body> <body>
<header class="navbar navbar-expand navbar-dark flex-column flex-md-row bd-navbar"> <header class="navbar navbar-expand navbar-dark flex-column flex-md-row bd-navbar">
{% if path %} <h3>Cellxgene Gateway - FILE CRAWLER</h3>
<h3>Cellxgene Gateway - {{ path }}</h3>
{% else %}
<h3>Cellxgene Gateway - FILE CRAWLER</h3>
{% endif %}
</header> </header>
<br>
<h4>Please click on a dataset to view it in Cellxgene Server.</h4> <h4>Please click on a dataset to view it in Cellxgene Server.</h4>
<br>
{{ rendered_html|safe }} {{ rendered_html|safe }}
<p>
Navigation:
<ul>
{% if path %}
<li><a href="/filecrawl.html">top level</a></li>
{% else %}
{% endif %}
<li><a href="/">homepage</a></li>
</ul>
</p>
<script>
$(() => {
$("a.new").click(new_annotation_callback);
})
</script>
</body> </body>
</html> </html>

View File

@@ -44,6 +44,47 @@
<u>Cache Status: view status of launched cellxgene servers.</u></a> <u>Cache Status: view status of launched cellxgene servers.</u></a>
</div> </div>
{% if enable_upload %}
<br>
<h1 style="padding-left:35px">
How To Upload Data via HTTP:
</h1>
<ol style="padding-left:85px;">
<li>
Create a folder for your Username:
</li>
<br>
<form action="{{ url_for('make_user') }}" method="post">
Username <input type="text" name="directory">
<input type="submit" value="Create">
</form>
<li>
Create a subdirectory under the selected Folder:
</li>
<br>
<form action="{{ url_for('make_subdir') }}" method="post">
<select name="usernames" id="usernames">
{% for user in users %}
<option value="{{ user }}">{{ user }}</option>
{% endfor %}
</select>
<br>
Subdirectory Name <input type="text" name="directory">
<input type="submit" value="Create">
</form>
<li>Choose a folder to copy your data to, then upload your data file (must be in .h5ad format).</li>
<br>
<form action="{{ url_for('upload_file') }}" method="post" enctype="multipart/form-data">
Type in the name of the directory and subdirectory you wish to upload to, i.e. "USER/cells". <input type="text" name="path">
<br>
File: <input type="file" name="file"><br>
<input style="position:relative; top:10px;" type="submit" value="Upload">
</form>
<br>
<li>Take a look at your data using the file crawler link above</li>
</ol>
{% endif %}
<br> <br>
<h1 style="padding-left:35px"> <h1 style="padding-left:35px">

View File

@@ -44,13 +44,9 @@
</a> </a>
</div> </div>
<script> <script>
var count = 0;
window.setInterval(function(){ window.setInterval(function(){
var dots = document.getElementById('dots'); var dots = document.getElementById('dots');
dots.textContent = dots.textContent + '.'; dots.textContent = dots.textContent + '.';
if (count++ > 5) {
window.location.reload();
}
}, 1000); }, 1000);
</script> </script>
</body> </body>

View File

@@ -1,4 +1,4 @@
name: cellxgene-gateway name: cellxgene-dev
channels: channels:
- conda-forge - conda-forge
dependencies: dependencies:
@@ -6,8 +6,6 @@ dependencies:
- requests - requests
- flask - flask
- psutil - psutil
- black
- pip
- pip: - pip:
- flask-api - flask-api
- cellxgene>=0.15 - cellxgene

View File

@@ -1,4 +1,4 @@
cellxgene>=0.15 cellxgene
flask flask
flask_api flask_api
psutil psutil

View File

@@ -1,5 +1,8 @@
export CELLXGENE_LOCATION=$(pwd)/.cellxgene-gateway/bin/cellxgene export CELLXGENE_LOCATION=$(pwd)/.cellxgene-gateway/bin/cellxgene
export CELLXGENE_DATA=../cellxgene_data export CELLXGENE_DATA=../cellxgene_data
export DEPLOYMENT_ENV=dev
export GATEWAY_HOST=localhost:5005
export GATEWAY_PROTOCOL=http
export GATEWAY_IP=127.0.0.1 export GATEWAY_IP=127.0.0.1
#Once these are set, you run like a normal Flask app #Once these are set, you run like a normal Flask app

View File

@@ -1,2 +0,0 @@
[metadata]
description-file = README.md

View File

@@ -1,69 +1,40 @@
import codecs
import os import os
import sys from setuptools import setup
from setuptools import find_packages, setup
if sys.version_info < (3, 6):
sys.exit("Sorry, Python < 3.6 is not supported")
def read(rel_path):
here = os.path.abspath(os.path.dirname(__file__))
with codecs.open(os.path.join(here, rel_path), "r") as fp:
return fp.read()
def get_version(rel_path):
for line in read(rel_path).splitlines():
if line.startswith("__version__"):
delim = '"' if '"' in line else "'"
return line.split(delim)[1]
else:
raise RuntimeError("Unable to find version string.")
def parse_requirements(): def parse_requirements():
reqs = [] reqs = []
with open("requirements.txt", "r") as f: with open("requirements.txt", "r") as f:
for line in f.readlines(): for l in f.readlines():
reqs.append(line.strip("\n")) reqs.append(l.strip("\n"))
return reqs return reqs
with open("README.md", "r") as fh:
long_description = fh.read()
install_reqs = parse_requirements() install_reqs = parse_requirements()
setup( setup(
# mandatory # mandatory
name="cellxgene-gateway", name="cellxgene-gateway",
# mandatory # mandatory
version=get_version("cellxgene_gateway/__init__.py"), version="0.1",
# mandatory # mandatory
author="Niket Patel, Yohann Potier, Alok Saldanha", author="Niket Patel, Yohann Potier, Alok Saldanha",
author_email="alok.saldanha@novartis.com", author_email="alok.saldanha@novartis.com",
description=("Cellxgene Gateway"), description=("Cellxgene Gateway"),
long_description=long_description,
long_description_content_type="text/markdown",
license="MIT", license="MIT",
keywords="visualization, genomics", keywords="visualization, genomics",
url="http://github.com/Novartis/cellxgene-gateway", url="http://github.com/Novartis/cellxgene-gateway",
packages=find_packages(), packages=["cellxgene_gateway"],
package_data={ package_data={
"cellxgene_gateway": [ "cellxgene_gateway": [
"static/css/homepagestyle.css", "static/css/homepagestyle.css",
"static/js/annotation.js",
"static/nibr.ico", "static/nibr.ico",
"templates/*.html", "templates/*.html"
] ]},
}, data_files=[('', ['Readme.md', 'LICENSE.txt'])],
data_files=[("", ["README.md", "LICENSE"])],
install_requires=install_reqs, install_requires=install_reqs,
entry_points={ entry_points={
"console_scripts": ["cellxgene-gateway=cellxgene_gateway.gateway:main"] "console_scripts": ["cellxgene-gateway=cellxgene_gateway.gateway:main"]
}, },
classifiers=["Topic :: Scientific/Engineering :: Visualization"], classifiers=["Topic :: Scientific/Engineering :: Visualization"],
python_requires=">=3.6",
) )

View File

@@ -1,35 +0,0 @@
import unittest
from cellxgene_gateway.cache_entry import CacheEntry, CacheEntryStatus
from cellxgene_gateway.cache_key import CacheKey
from cellxgene_gateway.items.item import ItemType
from cellxgene_gateway.items.file.fileitem import FileItem
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
key = CacheKey(
FileItem("/czi/", "pbmc3k.h5ad", ItemType.h5ad),
FileItemSource("/tmp", "local"),
)
class TestRenderEntry(unittest.TestCase):
def test_GIVEN_key_and_port_THEN_returns_loading_CacheEntry(self):
entry = CacheEntry.for_key("some-key", 1)
self.assertEqual(entry.status, CacheEntryStatus.loading)
def test_GIVEN_absolute_static_url_THEN_include_path(self):
actual = CacheEntry.for_key(key, 8000).rewrite_text_content(
"src:url(/static/assets/"
)
expected = "src:url(http://localhost:5005/source/local/view/czi/pbmc3k.h5ad/static/assets/"
self.assertEqual(actual, expected)
def test_GIVEN_absolute_src_THEN_include_path(self):
actual = CacheEntry.for_key(key, 8000).rewrite_text_content(
'<link rel="shortcut icon" href="/static/assets/favicon.ico">'
)
expected = '<link rel="shortcut icon" href="http://localhost:5005/source/local/view/czi/pbmc3k.h5ad/static/assets/favicon.ico">'
self.assertEqual(actual, expected)
if __name__ == "__main__":
unittest.main()

38
tests/test_dir_util.py Normal file
View File

@@ -0,0 +1,38 @@
import unittest
from unittest.mock import MagicMock, patch
from cellxgene_gateway.dir_util import render_entry
class TestRenderEntry(unittest.TestCase):
def test_GIVEN_path_both_slash_THEN_view_has_single_slash(self):
entry = {
"path": "/somepath/",
"name": "entry",
"type": "file",
}
rendered = render_entry(entry)
self.assertIn('view/somepath', rendered)
def test_GIVEN_path_starts_slash_THEN_view_has_single_slash(self):
entry = {
"path": "/somepath",
"name": "entry",
"type": "file",
}
rendered = render_entry(entry)
self.assertIn('view/somepath', rendered)
def test_GIVEN_path_ends_slash_THEN_view_has_single_slash(self):
entry = {
"path": "somepath/",
"name": "entry",
"type": "file",
}
rendered = render_entry(entry)
self.assertIn('view/somepath', rendered)
def test_GIVEN_path_no_slash_THEN_view_has_single_slash(self):
entry = {
"path": "somepath",
"name": "entry",
"type": "file",
}
rendered = render_entry(entry)
self.assertIn('view/somepath', rendered)

View File

@@ -1,26 +1,23 @@
import unittest import unittest
from unittest.mock import MagicMock, patch from unittest.mock import MagicMock, patch
from cellxgene_gateway.extra_scripts import get_extra_scripts from cellxgene_gateway.extra_scripts import get_extra_scripts
class TestExtraScripts(unittest.TestCase): class TestExtraScripts(unittest.TestCase):
@patch("cellxgene_gateway.env.extra_scripts", new='["abc","def"]') @patch('cellxgene_gateway.env.extra_scripts', new='["abc","def"]')
def test_GIVEN_two_scripts_THEN_returns_two_strings(self): def test_GIVEN_two_scripts_THEN_returns_two_strings(self):
self.assertEqual(get_extra_scripts(), ["abc", "def"]) self.assertEqual(get_extra_scripts(), ['abc', 'def'])
@patch("cellxgene_gateway.env.extra_scripts", new='["abc", "def"]') @patch('cellxgene_gateway.env.extra_scripts', new='["abc", "def"]')
def test_GIVEN_two_scripts_space_THEN_returns_two_strings(self): def test_GIVEN_two_scripts_space_THEN_returns_two_strings(self):
self.assertEqual(get_extra_scripts(), ["abc", "def"]) self.assertEqual(get_extra_scripts(), ['abc', 'def'])
@patch("cellxgene_gateway.env.extra_scripts", new=None) @patch('cellxgene_gateway.env.extra_scripts', new=None)
def test_GIVEN_none_THEN_returns_empty_array(self): def test_GIVEN_none_THEN_returns_empty_array(self):
self.assertEqual(get_extra_scripts(), []) self.assertEqual(get_extra_scripts(), [])
@patch("cellxgene_gateway.env.extra_scripts", new="[]") @patch('cellxgene_gateway.env.extra_scripts', new='[]')
def test_GIVEN_empty_string_THEN_returns_empty_array(self): def test_GIVEN_empty_string_THEN_returns_empty_array(self):
self.assertEqual(get_extra_scripts(), []) self.assertEqual(get_extra_scripts(), [])
if __name__ == '__main__':
if __name__ == "__main__":
unittest.main() unittest.main()

View File

@@ -1,33 +0,0 @@
import unittest
from unittest.mock import MagicMock, patch
from cellxgene_gateway.filecrawl import render_item
from cellxgene_gateway.items.item import ItemType
from cellxgene_gateway.items.file.fileitem import FileItem
from cellxgene_gateway.items.file.fileitem_source import FileItemSource
source = FileItemSource("/tmp")
class TestRenderEntry(unittest.TestCase):
def test_GIVEN_path_both_slash_THEN_view_has_single_slash(self):
entry = FileItem(
subpath="/somepath/", name="entry", type=ItemType.h5ad
)
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry/'", rendered)
def test_GIVEN_path_starts_slash_THEN_view_has_single_slash(self):
entry = FileItem(subpath="/somepath", name="entry", type=ItemType.h5ad)
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry/'", rendered)
def test_GIVEN_path_ends_slash_THEN_view_has_single_slash(self):
entry = FileItem(subpath="somepath/", name="entry", type=ItemType.h5ad)
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry/'", rendered)
def test_GIVEN_path_no_slash_THEN_view_has_single_slash(self):
entry = FileItem(subpath="somepath", name="entry", type=ItemType.h5ad)
rendered = render_item(entry, source)
self.assertIn("view/somepath/entry/'", rendered)

View File

@@ -1,15 +1,13 @@
import unittest import unittest
from unittest.mock import MagicMock, patch from unittest.mock import MagicMock, patch
from cellxgene_gateway.backend_cache import BackendCache
from cellxgene_gateway.cache_entry import CacheEntry from cellxgene_gateway.cache_entry import CacheEntry
from cellxgene_gateway.backend_cache import BackendCache
class TestPruneProcessCache(unittest.TestCase): class TestPruneProcessCache(unittest.TestCase):
@patch("cellxgene_gateway.util.current_time_stamp", new=lambda: 0) @patch('cellxgene_gateway.util.current_time_stamp', new=lambda:0)
@patch("cellxgene_gateway.env.ttl", new="10") @patch('cellxgene_gateway.env.ttl', new='10')
@patch("cellxgene_gateway.cache_entry.CacheEntry") @patch('cellxgene_gateway.cache_entry.CacheEntry')
@patch("cellxgene_gateway.cache_entry.CacheEntry") @patch('cellxgene_gateway.cache_entry.CacheEntry')
def test_GIVEN_one_old_one_new_THEN_prune_old(self, old, new): def test_GIVEN_one_old_one_new_THEN_prune_old(self, old, new):
from cellxgene_gateway.prune_process_cache import PruneProcessCache from cellxgene_gateway.prune_process_cache import PruneProcessCache
@@ -24,6 +22,5 @@ class TestPruneProcessCache(unittest.TestCase):
self.assertEqual(len(cache.entry_list), 1) self.assertEqual(len(cache.entry_list), 1)
self.assertEqual(cache.entry_list[0], new) self.assertEqual(cache.entry_list[0], new)
if __name__ == '__main__':
if __name__ == "__main__":
unittest.main() unittest.main()