Restv2 feature branch merge to master (#284)

Move to new REST v0.2 communication between front and back-end.   This is a first cut implementation which is functional, but will need follow-up enhancements for performance, error checking, etc.    Protocol spec is in docs directory.

* Add filtering via indexing

* Using new filter specs

Indexing working

* Added filtering by annotation value

* factor out common methods

* Documentation

* create enum for axis (obs/var)

* Better description for filter's return

* Add boolean to enumerated types

* Augmented enum for scanpy axis

* Create schema for annotations

Based on datatype within scanpy/anndata
+ tests

* remove obsolete schema parse script

* Update rest api to remove old routes and add schema route

* Separate development requirements

* Warning for unsupported datatypes

* include -r requirements.txt in dev

* Merged downcast warnings

* Fixed bug where names were NaNs

Needed to include the index too when creating the series

* Add config endpoint

* Generate app features from CLI selections

* Move features to driver

* Add tests for schema

* Clearer version wording

* python3 version of super

* version from engine to package level

* move features to driver

* Revise layout function to match the new spec

* GET for layout/obs

* PUT Layout (#211)

* PUT Layout

* Csweaver/annotations (#212)


* Update scanpy engine to support the rest v0.2 annotation requests

* GET endpoint for obs annotations + tests

* Documentation

* Test annotations in scanpy engine

* Description for annotation-keys param

* annotation->annotations

* clarified return for annotations

* Use URL query list for annotations fields

* parse_filter parses v0.2 GET filters (#215)

* parse_filter parses v0.2 GET filters

* Don't allow index filters from query params

* Better variable conversion

* Parse filter improvements

- uses default dict
- renamed filter -> query_filter

* Cleanup Tasks (#216)

* Add test_api back into travis build

* Do custom JSON encoding the correct way

* Run cellxgene server in test setup

* Cleanup new tests too

* Option to bind to all interfaces (#225)

app.run("0.0.0.0") instead of app.run("127.0.0.1") binds to all interfaces.

Note: There are comments on the internet that says that the flask server is not up to the task of production serving.  I don't think that such scalability concerns apply here, but I was able to get cellxgene working with twistd relatively easily, and we could switch to that if there are scalability concerns.

Test plan: browsed to <ip>:5005/api/v0.2/config on a different host.

* Add filtering via indexing

* Using new filter specs

Indexing working

* Added filtering by annotation value

* factor out common methods

* Documentation

* create enum for axis (obs/var)

* Better description for filter's return

* Add boolean to enumerated types

* Augmented enum for scanpy axis

* Create schema for annotations

Based on datatype within scanpy/anndata
+ tests

* remove obsolete schema parse script

* Update rest api to remove old routes and add schema route

* Separate development requirements

* Warning for unsupported datatypes

* include -r requirements.txt in dev

* Merged downcast warnings

* Fixed bug where names were NaNs

Needed to include the index too when creating the series

* Add config endpoint

* Generate app features from CLI selections

* Move features to driver

* Add tests for schema

* Clearer version wording

* python3 version of super

* version from engine to package level

* move features to driver

* Revise layout function to match the new spec

* GET for layout/obs

* PUT Layout (#211)

* PUT Layout

* Csweaver/annotations (#212)


* Update scanpy engine to support the rest v0.2 annotation requests

* GET endpoint for obs annotations + tests

* Documentation

* Test annotations in scanpy engine

* Description for annotation-keys param

* annotation->annotations

* clarified return for annotations

* Use URL query list for annotations fields

* parse_filter parses v0.2 GET filters (#215)

* parse_filter parses v0.2 GET filters

* Don't allow index filters from query params

* Better variable conversion

* Parse filter improvements

- uses default dict
- renamed filter -> query_filter

* Cleanup Tasks (#216)

* Add test_api back into travis build

* Do custom JSON encoding the correct way

* Run cellxgene server in test setup

* Cleanup new tests too

* Option to bind to all interfaces (#225)

app.run("0.0.0.0") instead of app.run("127.0.0.1") binds to all interfaces.

Note: There are comments on the internet that says that the flask server is not up to the task of production serving.  I don't think that such scalability concerns apply here, but I was able to get cellxgene working with twistd relatively easily, and we could switch to that if there are scalability concerns.

Test plan: browsed to <ip>:5005/api/v0.2/config on a different host.

* Fix merge errors

- import warnings was improperly deleted
- scanpy engine tests were totally wrong

* Fix merge error with driver

* PUT /annotations (#235)

* Add query param for annotation name

* fix descriptions, eliminate else clause

* first cut at initial data load on rest 0.2 api

* Annotation var (#248)

* Fix bug strings are always objects in pandas

* Add axis to annotation method

* Add /annotation/var to REST api

* Csweaver/expressiondata (#242)

* Refactor expression method for REST v2

* Add message to QueryStringError

* Fix range filters

* Add GET route for /data

* /data PUT route

* rename expression to data_frame

* clarification of error

* Improve accept type handling

* support all schema types for 0.2 REST API

* remove REST 0.1 code; connect var annotations loading

* config reducer; use config to set data set title; remove obsolete templating code for data set title

* REST 0.2 expression conversion support

* partial port of expression to REST 0.2

*  diffexp (#273)

* Add diffexp method to scanpy

and test

* Minor tweaks to diffexp

Get a minimal working version to unblock FE development

* Fixing things git deleted

* cleanup print statements

* Add index test

* additional, partial REST 0.2 bring up of diffexp

* Ignore unstructured annotations for data (#275)

This is a temp hack, need to figure out how to include data.uns if there is only one gene

* diffexp REST 0.2 port finish

* ignore unstructured annotaitons on all routes except layout

* correctly use varDataCache; maintain state during world rebuild

* correct varDataCache use

* temporarily disable all memoization

* refinements to expression data caching

* clear cell sets upon regraph/reset

* update version of REST to 0.2

* Travis build fixes

- comment out cache import
- fix duplicate test name

* Remove dependency from travis

* clarify semantics of config variables

* move generic action helpers into util
This commit is contained in:
Bruce Martin
2018-10-01 14:58:46 -07:00
committed by GitHub
parent f0d9d873be
commit eeec842ad0
31 changed files with 2139 additions and 1373 deletions
+28
View File
@@ -0,0 +1,28 @@
/*
Catch unexpected errors and make sure we don't lose them!
*/
export function catchErrorsWrap(fn) {
return (dispatch, getState) => {
fn(dispatch, getState).catch(error => {
console.error(error);
dispatch({ type: "UNEXPECTED ERROR", error });
});
};
}
/*
Bootstrap application with the initial data loading.
* /config - application configuration
* /schema - schema of dataframe
* /annotations/obs - all metadata annotation
*/
export const doJsonRequest = async url => {
const res = await fetch(url, {
method: "get",
headers: new Headers({
"Content-Type": "application/json",
"Accept-Encoding": "gzip, deflate, br"
})
});
return res.json();
};
+20 -4
View File
@@ -4,13 +4,12 @@ import _ from "lodash";
/*
Very simple key/value cache for use by World & Universe.
* constructor(lowWatermark, cachekey):
* constructor(lowWatermark, minTTL):
- lowWatermark defines the number of cache elements below which
flushing will not occur.
- minTTL defines minimum time in MS that cache entries will live.
- minTTL defines minimum time in milliseconds that cache entries will live.
A value of -1 disables automatic flushing (flush() can still
be called by external user).
- cachekey is a key that will be assigned to any value to track age
* set() - add a key/val pair.
* get() - get a value or undefined if not present.
* flush(minAgeMs) - flush cache entries in excess of lowWatermark if those
@@ -69,4 +68,21 @@ function flush(kvcache, minAgeMs = 0) {
return kvcache;
}
export { create, get, set, flush };
/*
use to create a cache that is a transformation of another cache.
*/
function map(srcKvCache, cb, createOptions) {
const keysInSrcKvCache = _(srcKvCache)
.keys()
.filter(k => k !== cachePrivateKey)
.value();
const newKvCache = create(createOptions.lowWatermark, createOptions.minTTL);
_.forEach(keysInSrcKvCache, key => {
const val = cb(get(srcKvCache, key));
newKvCache[key] = val;
val[cachePrivateKey] = Date.now();
});
return newKvCache;
}
export { create, get, set, flush, map };
+159 -185
View File
@@ -3,6 +3,44 @@
import _ from "lodash";
import * as kvCache from "./keyvalcache";
/*
Private helper function - create and return a template Universe
*/
function templateUniverse() {
/* default universe template */
/* varDataCache config - see kvCache for semantics */
const VarDataCacheLowWatermark = 32; // cache element count
const VarDataCacheTTLMs = 1000; // min cache time in MS
return {
api: null,
finalized: false, // XXX: may not be needed
nObs: 0,
nVar: 0,
schema: {},
/*
Annotations
*/
obsAnnotations: [] /* all obs annotations, by obs index */,
varAnnotations: [] /* all var annotations, by var index */,
obsNameToIndexMap: {} /* reverse map 'name' to index */,
varNameToIndexMap: {} /* reverse map 'name' to index */,
obsLayout: { X: [], Y: [] } /* xy layout */,
/*
Cache of var data (expression), by var annotation name. Data can be
accesses as a POJO, but if you want caching semantics, use the kvCache
API (eg., kvCache.get(), kvCache.set(), ...), which will maintain the
LRU semantics.
*/
varDataCache: kvCache.create(VarDataCacheLowWatermark, VarDataCacheTTLMs)
};
}
/*
This module implements functions that support storage of "Universe",
aka all of the var/obs data and annotations.
@@ -12,130 +50,9 @@ build an internal POJO for use by the rendering components.
*/
/*
Cherry pick from /api/v0.1 response format to make somethign similar
to the v0.2 schema, which we use for internal interfaces.
generate any client-side transformations or summarization that
is independent of REST API response formats.
*/
function RESTv01ResponseToSchema(response) {
/*
Annotation schemas in V02 (our target) look like:
annotations: {
obs: [
{ name: "name", type: "string" },
{ name: "num_reads", type: "int32" },
{
name: "clusters",
type: "categorical",
categories=[ 99, 1, "unknown cluster" ]
},
{ name: "QScore", type: "float32" }
],
var: [
{ "name": "name", "type": "string" },
{ "name": "gene", "type": "string" }
]
}
In V01, our source, it looks like:
"schema": {
"CellName": {
"displayname": "Name",
"include": true,
"type": "string",
"variabletype": "categorical"
},
"Cluster_2d": {
"displayname": "Cluster2d",
"include": true,
"type": "string",
"variabletype": "categorical"
},
"ERCC_reads": {
"displayname": "ERCC Reads",
"include": true,
"type": "int",
"variabletype": "continuous"
},
...
}
Mapping between the two assumes:
- V01 only has schema for observations
- CellName is mapped to 'name'
- type conversion: float->float32, int->int32, string->string
*/
return {
annotations: {
obs: _.map(response.data.schema, (val, key) => {
const name = key === "CellName" ? "name" : key;
let { type } = val;
if (type === "int") {
type = "int32";
}
if (type === "float") {
type = "float32";
}
return {
name,
type
};
}),
var: [{ name: "name", type: "string" }]
}
};
}
function RESTv01ResponseToVarAnnotations(response) {
/*
v0.1 initialize response contains 'genes' - names of all genes
in order.
*/
return _.map(response.data.genes, (g, i) => ({ __varIndex__: i, name: g }));
}
function RESTv01ResponseToObsAnnotations(response) {
/*
v0.1 format for metadata:
metadata: [ { key: val, key: val, ... }, ... ]
Target format is essentially the same, except the CellName key becomes name.
*/
return _.map(response.data.metadata, (c, i) => ({
__obsIndex__: i,
name: c.CellName,
...c
}));
}
function RESTv01ResponseToLayout(obsAnnotations, response) {
/*
v0.1 format for the graph is:
[ [ 'cellname', x, y ], [ 'cellname', x, y, ], ... ]
NOTE XXX: this code does not assume any particular array ordering in the V0.1
response. But for Universe initial load, the layout will be in the same
order as annotations, so this extra work isn't really necessary.
*/
const obsAnnotationsByName = _.keyBy(obsAnnotations, "name");
const { graph } = response.data;
const layout = {
X: new Float32Array(graph.length),
Y: new Float32Array(graph.length)
};
for (let i = 0; i < graph.length; i += 1) {
const [name, x, y] = graph[i];
const anno = obsAnnotationsByName[name];
const idx = anno.__obsIndex__;
layout.X[idx] = x;
layout.Y[idx] = y;
}
return layout;
}
function finalize(universe) {
/* A bit of sanity checking! */
const { nObs, nVar } = universe;
@@ -147,7 +64,14 @@ function finalize(universe) {
) {
throw new Error("Universe dimensionality mismatch - failed to load");
}
// TODO: add more sanity checks, such as:
// - all annotations in the schema
// - layout has supported number of dimensions
// - ...
/*
Create all derived (convenience) data structures.
*/
universe.obsNameToIndexMap = _.transform(
universe.obsAnnotations,
(acc, value, idx) => {
@@ -166,85 +90,135 @@ function finalize(universe) {
return universe;
}
function templateUniverse() {
/* default universe template */
const VarDataCacheLowWatermark = 32;
const VarDataCacheTTLMs = 1000;
function RESTv02AnnotationsResponseToInternal(response) {
/*
Source per the spec:
{
names: [
'tissue_type', 'sex', 'num_reads', 'clusters'
],
data: [
[ 0, 'lung', 'F', 39844, 99 ],
[ 1, 'heart', 'M', 83, 1 ],
[ 49, 'spleen', null, 2, "unknown cluster" ],
// [ obsOrVarIndex, value, value, value, value ],
// ...
]
}
return {
api: "0.1",
finalized: true, // XXX: may not be needed
nObs: 0,
nVar: 0,
schema: {},
/*
Annotations
*/
obsAnnotations: [] /* all obs annotations, by obs index */,
varAnnotations: [] /* all var annotations, by var index */,
obsNameToIndexMap: {} /* reverse map 'name' to index */,
varNameToIndexMap: {} /* reverse map 'name' to index */,
obsLayout: { X: [], Y: [] } /* xy layout */,
varDataCache: kvCache.create(
VarDataCacheLowWatermark,
VarDataCacheTTLMs
) /* cache of var data (expression) */
};
Internal (target) format:
[
{ __index__: 0, tissue_type: "lung", sex: "F", ... },
...
]
*/
const { names, data } = response;
const keys = ["__index__", ...names];
return _(data)
.map(obs => _.zipObject(keys, obs))
.sortBy("__index__")
.value();
}
export function createUniverseFromRESTv01Response(initResponse, cellsResponse) {
function RESTv02LayoutResponseToInternal(response) {
/*
build & return universe from a REST 0.1 /init and /cells response
*/
Source per the spec:
{
layout: {
ndims: 2,
coordinates: [
[ 0, 0.284483, 0.983744 ],
[ 1, 0.038844, 0.739444 ],
// [ obsOrVarIndex, X_coord, Y_coord ],
// ...
]
}
}
Target (internal) format:
{
X: Float32Array(numObs),
Y: Float32Array(numObs)
}
In the same order as obsAnnotations
*/
const { ndims, coordinates } = response.layout;
if (ndims !== 2) {
throw new Error("Unsupported layout dimensionality");
}
const layout = {
X: new Float32Array(coordinates.length),
Y: new Float32Array(coordinates.length)
};
for (let i = 0; i < coordinates.length; i += 1) {
const [idx, x, y] = coordinates[i];
layout.X[idx] = x;
layout.Y[idx] = y;
}
return layout;
}
export function createUniverseFromRestV02Response(
configResponse,
schemaResponse,
annotationsObsResponse,
annotationsVarResponse,
layoutObsResponse
) {
/*
build & return universe from a REST 0.2 /config, /schema and /annotations/obs response
*/
const { schema } = schemaResponse;
const universe = templateUniverse();
/* extract information from init OTA response */
universe.schema = RESTv01ResponseToSchema(initResponse);
universe.varAnnotations = RESTv01ResponseToVarAnnotations(initResponse);
universe.nVar = universe.varAnnotations.length;
/* constants */
universe.api = "0.2";
/* extract information fron cells REST json response */
/*
NOTE: this code *assumes* that cell order in data.metadata and data.graph
are the same. TODO: error checking.
*/
universe.obsAnnotations = RESTv01ResponseToObsAnnotations(cellsResponse);
universe.nObs = universe.obsAnnotations.length;
universe.obsLayout = RESTv01ResponseToLayout(
universe.obsAnnotations,
cellsResponse
/* schema related */
universe.schema = schema;
universe.nObs = schema.dataframe.nObs;
universe.nVar = schema.dataframe.nVar;
/* annotations */
universe.obsAnnotations = RESTv02AnnotationsResponseToInternal(
annotationsObsResponse
);
universe.varAnnotations = RESTv02AnnotationsResponseToInternal(
annotationsVarResponse
);
/* layout */
universe.obsLayout = RESTv02LayoutResponseToInternal(layoutObsResponse);
return finalize(universe);
}
export function convertExpressionRESTv01ToObject(universe, response) {
export function convertExpressionRESTv02ToObject(universe, response) {
/*
v0.1 ota looks like:
{
genes: [ "name1", "name2", ... ],
cells: [
{ cellname: 'cell1', e: [ 3, 4, n, x, y, ... ] },
...
]
}
/data/obs response looks like:
{
var: [ varIndices fetched ],
obs: [
[ obsIndex, evalue, ... ],
...
]
}
convert expression to a simple Float32Array, and return
[ [geneName, array], [geneName, array], ... ]
*/
convert expression toa simple Float32Array, and return
{ geneName: array, geneName: array, ... }
NOTE: geneName, not varIndex
*/
const vars = response.var;
const { obs } = response;
const result = {};
const { genes, cells } = response.data;
for (let idx = 0; idx < genes.length; idx += 1) {
const gene = genes[idx];
// XXX TODO: could this use _.unzip and have less code?
for (let varIdx = 0; varIdx < vars.length; varIdx += 1) {
const gene = universe.varAnnotations[vars[varIdx]].name;
const data = new Float32Array(universe.nObs);
for (let c = 0; c < cells.length; c += 1) {
const obsIndex = universe.obsNameToIndexMap[cells[c].cellname];
data[obsIndex] = cells[c].e[idx];
for (let obsIdx = 0; obsIdx < obs.length; obsIdx += 1) {
data[obsIdx] = obs[obsIdx][varIdx + 1];
}
result[gene] = data;
}
+67 -52
View File
@@ -28,7 +28,7 @@ obs/cell.
NOTE: world.obsAnnotation should be identical to the old state.cells value,
EXCEPT that
* __cellIndex__ renamed to __obsIndex__
* __cellIndex__ renamed to __index__
* __x__ and __y__ are now in world.obsLayout
* __color__ and __colorRBG__ should be moved to controls reducer
@@ -46,47 +46,50 @@ obs/cell.
*/
/*
Summary information for each annotation, keyed by annotation name.
Value will be an object, containing either 'range' or 'options' object,
depending on the annotation schema type (categorical or continuous).
/* varDataCache config - see kvCache for semantics */
const VarDataCacheLowWatermark = 32; // cache element count
const VarDataCacheTTLMs = 1000; // min cache time in MS
Summarize for BOTH obs and var annotations. Result format:
{
obs: {
annotation_name: { ... },
...
},
var: {
annotation_name: { ... },
...
}
}
Example:
{
"Splice_sites_Annotated": {
"range": {
"min": 26,
"max": 1075869
}
},
"Selection": {
"options": {
"Astrocytes(HEPACAM)": 714,
"Endothelial(BSC)": 123,
"Oligodendrocytes(GC)": 294,
"Neurons(Thy1)": 685,
"Microglia(CD45)": 1108,
"Unpanned": 665
}
}
}
*/
function summarizeAnnotations(schema, obsAnnotations) {
/*
Build and return obs/var summary using any annotation in the schema
Summary information for each annotation, keyed by annotation name.
Value will be an object, containing either 'range' or 'options' object,
depending on the annotation schema type (categorical or continuous).
Summarize for BOTH obs and var annotations. Result format:
{
obs: {
annotation_name: { ... },
...
},
var: {
annotation_name: { ... },
...
}
}
Example:
{
"Splice_sites_Annotated": {
"range": {
"min": 26,
"max": 1075869
}
},
"Selection": {
"options": {
"Astrocytes(HEPACAM)": 714,
"Endothelial(BSC)": 123,
"Oligodendrocytes(GC)": 294,
"Neurons(Thy1)": 685,
"Microglia(CD45)": 1108,
"Unpanned": 665
}
}
}
*/
const obsSummary = _(schema.annotations.obs)
.keyBy("name")
@@ -115,7 +118,8 @@ function summarizeAnnotations(schema, obsAnnotations) {
})
.value();
const varSummary = {}; // TODO XXX - not currently used, so skip it
// TODO XXX - not currently used, so skip it
const varSummary = {};
return {
obs: obsSummary,
@@ -124,9 +128,6 @@ function summarizeAnnotations(schema, obsAnnotations) {
}
function templateWorld() {
const VarDataCacheLowWatermark = 32;
const VarDataCacheTTLMs = 1000;
return {
// map from universe obsIndex to world offset.
// Undefined / null indicates identity mapping.
@@ -186,6 +187,13 @@ export function createWorldFromEntireUniverse(universe) {
/* derived data & summaries */
world.summary = summarizeAnnotations(world.schema, world.obsAnnotations);
/* build the varDataCache */
world.varDataCache = kvCache.map(
universe.varDataCache,
val => subsetVarData(world, universe, val),
{ lowWatermark: VarDataCacheLowWatermark, minTTL: VarDataCacheTTLMs }
);
return world;
}
@@ -227,13 +235,21 @@ export function createWorldFromCurrentSelection(universe, world, crossfilter) {
// build index to our world offset
newWorld.worldObsIndex.fill(-1); // default - aka unused
for (let i = 0; i < newWorld.nObs; i += 1) {
newWorld.worldObsIndex[newWorld.obsAnnotations[i].__obsIndex__] = i;
newWorld.worldObsIndex[newWorld.obsAnnotations[i].__index__] = i;
}
/* derived data & summaries */
newWorld.summary = summarizeAnnotations(
newWorld.schema,
newWorld.obsAnnotations
);
/* build the varDataCache */
newWorld.varDataCache = kvCache.map(
universe.varDataCache,
val => subsetVarData(newWorld, universe, val),
{ lowWatermark: VarDataCacheLowWatermark, minTTL: VarDataCacheTTLMs }
);
return newWorld;
}
@@ -243,20 +259,19 @@ export function createWorldFromCurrentSelection(universe, world, crossfilter) {
*/
function deduceDimensionType(attributes, fieldName) {
let dimensionType;
if (attributes.type === "string") {
const { type } = attributes;
if (type === "string" || type === "categorical" || type === "boolean") {
dimensionType = "enum";
} else if (attributes.type === "int32") {
} else if (type === "int32") {
dimensionType = Int32Array;
} else if (attributes.type === "float32") {
} else if (type === "float32") {
dimensionType = Float32Array;
} else {
/*
Currently not supporting boolean and categorical types.
*/
console.error(
`Warning - REST API returned unknown metadata schema (${
attributes.type
}) for field ${fieldName}.`
`Warning - REST API returned unknown metadata schema (${type}) for field ${fieldName}.`
);
// skip it - we don't know what to do with this type
}
@@ -286,11 +301,11 @@ export function createObsDimensionMap(crossfilter, world) {
*/
const worldIndex = worldObsIndex ? idx => worldObsIndex[idx] : idx => idx;
dimensionMap.x = crossfilter.dimension(
r => obsLayout.X[worldIndex(r.__obsIndex__)],
r => obsLayout.X[worldIndex(r.__index__)],
Float32Array
);
dimensionMap.y = crossfilter.dimension(
r => obsLayout.Y[worldIndex(r.__obsIndex__)],
r => obsLayout.Y[worldIndex(r.__index__)],
Float32Array
);
@@ -309,7 +324,7 @@ export function subsetVarData(world, universe, varData) {
const newVarData = new Float32Array(world.nObs);
for (let i = 0; i < world.nObs; i += 1) {
newVarData[i] = varData[world.obsAnnotations[i].__obsIndex__];
newVarData[i] = varData[world.obsAnnotations[i].__index__];
}
return newVarData;
}