Restv2 feature branch merge to master (#284)

Move to new REST v0.2 communication between front and back-end.   This is a first cut implementation which is functional, but will need follow-up enhancements for performance, error checking, etc.    Protocol spec is in docs directory.

* Add filtering via indexing

* Using new filter specs

Indexing working

* Added filtering by annotation value

* factor out common methods

* Documentation

* create enum for axis (obs/var)

* Better description for filter's return

* Add boolean to enumerated types

* Augmented enum for scanpy axis

* Create schema for annotations

Based on datatype within scanpy/anndata
+ tests

* remove obsolete schema parse script

* Update rest api to remove old routes and add schema route

* Separate development requirements

* Warning for unsupported datatypes

* include -r requirements.txt in dev

* Merged downcast warnings

* Fixed bug where names were NaNs

Needed to include the index too when creating the series

* Add config endpoint

* Generate app features from CLI selections

* Move features to driver

* Add tests for schema

* Clearer version wording

* python3 version of super

* version from engine to package level

* move features to driver

* Revise layout function to match the new spec

* GET for layout/obs

* PUT Layout (#211)

* PUT Layout

* Csweaver/annotations (#212)


* Update scanpy engine to support the rest v0.2 annotation requests

* GET endpoint for obs annotations + tests

* Documentation

* Test annotations in scanpy engine

* Description for annotation-keys param

* annotation->annotations

* clarified return for annotations

* Use URL query list for annotations fields

* parse_filter parses v0.2 GET filters (#215)

* parse_filter parses v0.2 GET filters

* Don't allow index filters from query params

* Better variable conversion

* Parse filter improvements

- uses default dict
- renamed filter -> query_filter

* Cleanup Tasks (#216)

* Add test_api back into travis build

* Do custom JSON encoding the correct way

* Run cellxgene server in test setup

* Cleanup new tests too

* Option to bind to all interfaces (#225)

app.run("0.0.0.0") instead of app.run("127.0.0.1") binds to all interfaces.

Note: There are comments on the internet that says that the flask server is not up to the task of production serving.  I don't think that such scalability concerns apply here, but I was able to get cellxgene working with twistd relatively easily, and we could switch to that if there are scalability concerns.

Test plan: browsed to <ip>:5005/api/v0.2/config on a different host.

* Add filtering via indexing

* Using new filter specs

Indexing working

* Added filtering by annotation value

* factor out common methods

* Documentation

* create enum for axis (obs/var)

* Better description for filter's return

* Add boolean to enumerated types

* Augmented enum for scanpy axis

* Create schema for annotations

Based on datatype within scanpy/anndata
+ tests

* remove obsolete schema parse script

* Update rest api to remove old routes and add schema route

* Separate development requirements

* Warning for unsupported datatypes

* include -r requirements.txt in dev

* Merged downcast warnings

* Fixed bug where names were NaNs

Needed to include the index too when creating the series

* Add config endpoint

* Generate app features from CLI selections

* Move features to driver

* Add tests for schema

* Clearer version wording

* python3 version of super

* version from engine to package level

* move features to driver

* Revise layout function to match the new spec

* GET for layout/obs

* PUT Layout (#211)

* PUT Layout

* Csweaver/annotations (#212)


* Update scanpy engine to support the rest v0.2 annotation requests

* GET endpoint for obs annotations + tests

* Documentation

* Test annotations in scanpy engine

* Description for annotation-keys param

* annotation->annotations

* clarified return for annotations

* Use URL query list for annotations fields

* parse_filter parses v0.2 GET filters (#215)

* parse_filter parses v0.2 GET filters

* Don't allow index filters from query params

* Better variable conversion

* Parse filter improvements

- uses default dict
- renamed filter -> query_filter

* Cleanup Tasks (#216)

* Add test_api back into travis build

* Do custom JSON encoding the correct way

* Run cellxgene server in test setup

* Cleanup new tests too

* Option to bind to all interfaces (#225)

app.run("0.0.0.0") instead of app.run("127.0.0.1") binds to all interfaces.

Note: There are comments on the internet that says that the flask server is not up to the task of production serving.  I don't think that such scalability concerns apply here, but I was able to get cellxgene working with twistd relatively easily, and we could switch to that if there are scalability concerns.

Test plan: browsed to <ip>:5005/api/v0.2/config on a different host.

* Fix merge errors

- import warnings was improperly deleted
- scanpy engine tests were totally wrong

* Fix merge error with driver

* PUT /annotations (#235)

* Add query param for annotation name

* fix descriptions, eliminate else clause

* first cut at initial data load on rest 0.2 api

* Annotation var (#248)

* Fix bug strings are always objects in pandas

* Add axis to annotation method

* Add /annotation/var to REST api

* Csweaver/expressiondata (#242)

* Refactor expression method for REST v2

* Add message to QueryStringError

* Fix range filters

* Add GET route for /data

* /data PUT route

* rename expression to data_frame

* clarification of error

* Improve accept type handling

* support all schema types for 0.2 REST API

* remove REST 0.1 code; connect var annotations loading

* config reducer; use config to set data set title; remove obsolete templating code for data set title

* REST 0.2 expression conversion support

* partial port of expression to REST 0.2

*  diffexp (#273)

* Add diffexp method to scanpy

and test

* Minor tweaks to diffexp

Get a minimal working version to unblock FE development

* Fixing things git deleted

* cleanup print statements

* Add index test

* additional, partial REST 0.2 bring up of diffexp

* Ignore unstructured annotations for data (#275)

This is a temp hack, need to figure out how to include data.uns if there is only one gene

* diffexp REST 0.2 port finish

* ignore unstructured annotaitons on all routes except layout

* correctly use varDataCache; maintain state during world rebuild

* correct varDataCache use

* temporarily disable all memoization

* refinements to expression data caching

* clear cell sets upon regraph/reset

* update version of REST to 0.2

* Travis build fixes

- comment out cache import
- fix duplicate test name

* Remove dependency from travis

* clarify semantics of config variables

* move generic action helpers into util
This commit is contained in:
Bruce Martin
2018-10-01 14:58:46 -07:00
committed by GitHub
parent f0d9d873be
commit eeec842ad0
31 changed files with 2139 additions and 1373 deletions
+27
View File
@@ -0,0 +1,27 @@
from enum import Enum
DEFAULT_TOP_N = 10
class AugmentedEnum(Enum):
def __hash__(self):
return self.value.__hash__()
def __eq__(self, other):
if isinstance(other, type(self)) or isinstance(other, str):
return self.value == other
return False
def __str__(self) -> str:
return self.value
class Axis(AugmentedEnum):
OBS = "obs"
VAR = "var"
class DiffExpMode(AugmentedEnum):
TOP_N = "topN"
VAR_FILTER = "varFilter"
+60 -43
View File
@@ -1,5 +1,16 @@
import json
from collections import defaultdict
from numpy import float32, int32
from server.app.util.constants import Axis
class QueryStringError(Exception):
pass
def __init__(self, key, message):
self.key = key
self.message = message
def _convert_variable(datatype, variable):
@@ -9,62 +20,68 @@ def _convert_variable(datatype, variable):
:param datatype: type to convert to
:param variable (string or None): value of variable
:return: converted variable
:raises: ValueError
:raises: AssertionError
"""
try:
if variable is None:
return variable
if datatype == "int":
variable = int(variable)
elif datatype == "float":
variable = float(variable)
assert datatype in ["boolean", "categorical", "float32", "int32", "string"]
if variable is None:
return variable
except ValueError:
raise
if datatype == "int32":
variable = int32(variable)
elif datatype == "float32":
variable = float32(variable)
elif datatype == "boolean":
variable = json.loads(variable)
assert isinstance(variable, bool)
return variable
def parse_filter(filter, schema):
def parse_filter(query_filter, schema):
"""
The filter comes in as arguments from a GET/POST request
For categorical metadata keys filter based on key=value
For continuous metadata keys filter by key=min,max
Either value can be replaced by a * To have only a minimum value key=min, To have only a maximum value key=*,max
The filter comes in as arguments from a GET request
For categorical metadata keys filter based on axis:key=value
For continuous metadata keys filter by axis:key=min,max
Either value can be replaced by a * To have only a minimum
value axis:key=min,* To have only a maximum value axis:key=*,max
They combine via AND so a cell's metadata would have to match every filter
The results is a matrix with the cells the pass the filter and at this point all the genes
:param filter: flask's request.args
:param query_filter: flask's request.args
:param schema: dictionary schema
:raises QueryStringError
:return:
"""
query = {}
for key in filter:
value = filter.getlist(key)
if key not in schema:
raise QueryStringError("Error: key {} not in metadata schema".format(key))
query[key] = {
"variable_type": schema[key]["variabletype"],
"value_type": schema[key]["type"]
}
if query[key]["variable_type"] == "categorical":
query[key]["query"] = [_convert_variable(query[key]["value_type"], v) for v in value]
elif query[key]["variable_type"] == "continuous":
value = value[0]
query = defaultdict(lambda: defaultdict(list))
for key in query_filter:
axis, annotation = key.split(":", 1)
try:
Axis(axis)
except ValueError:
raise QueryStringError(key, f"Error: key {key} not in metadata schema")
ann_filter = {"name": annotation}
for ann in schema[axis]:
if ann["name"] == annotation:
dtype = ann["type"]
break
else:
raise QueryStringError(key, f"Error: {annotation} not a valid annotation name")
if dtype in ["string", "categorical", "boolean"]:
ann_filter["values"] = [_convert_variable(dtype, i) for i in query_filter.getlist(key)]
else:
value = query_filter.get(key)
try:
min, max = value.split(",")
min_, max_ = value.split(",")
except ValueError:
raise QueryStringError("Error: min,max format required for range for key {}, got {}".format(key, value))
if min == "*":
min = None
if max == "*":
max = None
raise QueryStringError(key, f"Error: min,max format required for range for {annotation}, got {value}")
if min_ == "*":
min_ = None
if max_ == "*":
max_ = None
try:
query[key]["query"] = {
"min": _convert_variable(query[key]["value_type"], min),
"max": _convert_variable(query[key]["value_type"], max)
}
ann_filter["min"] = _convert_variable(dtype, min_)
ann_filter["max"] = _convert_variable(dtype, max_)
except ValueError:
raise QueryStringError(
"Error: expected type {} for key {}, got {}".format(query[key]["type"], key, value)
)
raise QueryStringError(key, f"Error: expected type {query[key]['type']} for key {key}, got {value}")
query[axis]["annotation_value"].append(ann_filter)
return query
+65
View File
@@ -0,0 +1,65 @@
from flask_restful_swagger_2 import Schema
class AnnotationModel(Schema):
type = "object"
description = "Filter by annotation key: value"
properties = {
"name": {
"type": "string"
},
# TODO update to OpenAPI v3.0 when a library is available that supports it
# Unfortunately 2.0 doesn't have a way to have a schema that accepts multiple types
# Overloading the type key with a list seems to work ok and makes it to the page
"values": {
"type": "array",
"items": {
"type": ["float32", "string", "int32", "bool"]
}
},
"min": {
"type": ["int32", "float32"],
},
"max": {
"type": ["int32", "float32"],
}
}
required = ["name"]
class IndexModel(Schema):
type = "object"
description = "Filter by index of observation/variable ex. [0, 5, 15]"
properties = {
"index": {
"type": "array",
"items": {
"format": "int32",
"type": "integer"
}
}
}
class AxisModel(Schema):
type = "object"
description = "Axis of data -- obs or var"
properties = {
"index": IndexModel,
"annotation_value": AnnotationModel.array()
}
class FilterModel(Schema):
type = "object"
description = "Complex filter"
properties = {
"filter": {
"type": "object",
"properties": {
"obs": AxisModel,
"var": AxisModel
}
}
}
-7
View File
@@ -1,7 +0,0 @@
import json
def parse_schema(filename):
with open(filename) as fh:
schema = json.load(fh)
return schema
-28
View File
@@ -1,7 +1,6 @@
import json
from numpy import float32, integer
from flask import make_response, jsonify, Response
class Float32JSONEncoder(json.JSONEncoder):
@@ -11,30 +10,3 @@ class Float32JSONEncoder(json.JSONEncoder):
elif isinstance(obj, integer):
return int(obj)
return json.JSONEncoder.default(self, obj)
def make_payload(data, errormessage="", errorcode=200):
"""
Creates JSON respons for requests
:param data: json data
:param errormessage: error message
:param errorcode: http error code
:return: flask json repsonse
"""
error = False
if errormessage:
error = True
# Questionable
data = json.loads(json.dumps(data, cls=Float32JSONEncoder))
return make_response(jsonify({
"data": data,
"status": {
"error": error,
"errormessage": errormessage,
}
}), errorcode)
def make_streaming_response(data_generator, errorcode=200, content_type="application/json"):
# TODO headers
return Response(data_generator, status=errorcode, content_type=content_type)