mirror of
https://github.com/chanzuckerberg/cellxgene.git
synced 2026-09-29 15:58:11 +08:00
Restv2 feature branch merge to master (#284)
Move to new REST v0.2 communication between front and back-end. This is a first cut implementation which is functional, but will need follow-up enhancements for performance, error checking, etc. Protocol spec is in docs directory. * Add filtering via indexing * Using new filter specs Indexing working * Added filtering by annotation value * factor out common methods * Documentation * create enum for axis (obs/var) * Better description for filter's return * Add boolean to enumerated types * Augmented enum for scanpy axis * Create schema for annotations Based on datatype within scanpy/anndata + tests * remove obsolete schema parse script * Update rest api to remove old routes and add schema route * Separate development requirements * Warning for unsupported datatypes * include -r requirements.txt in dev * Merged downcast warnings * Fixed bug where names were NaNs Needed to include the index too when creating the series * Add config endpoint * Generate app features from CLI selections * Move features to driver * Add tests for schema * Clearer version wording * python3 version of super * version from engine to package level * move features to driver * Revise layout function to match the new spec * GET for layout/obs * PUT Layout (#211) * PUT Layout * Csweaver/annotations (#212) * Update scanpy engine to support the rest v0.2 annotation requests * GET endpoint for obs annotations + tests * Documentation * Test annotations in scanpy engine * Description for annotation-keys param * annotation->annotations * clarified return for annotations * Use URL query list for annotations fields * parse_filter parses v0.2 GET filters (#215) * parse_filter parses v0.2 GET filters * Don't allow index filters from query params * Better variable conversion * Parse filter improvements - uses default dict - renamed filter -> query_filter * Cleanup Tasks (#216) * Add test_api back into travis build * Do custom JSON encoding the correct way * Run cellxgene server in test setup * Cleanup new tests too * Option to bind to all interfaces (#225) app.run("0.0.0.0") instead of app.run("127.0.0.1") binds to all interfaces. Note: There are comments on the internet that says that the flask server is not up to the task of production serving. I don't think that such scalability concerns apply here, but I was able to get cellxgene working with twistd relatively easily, and we could switch to that if there are scalability concerns. Test plan: browsed to <ip>:5005/api/v0.2/config on a different host. * Add filtering via indexing * Using new filter specs Indexing working * Added filtering by annotation value * factor out common methods * Documentation * create enum for axis (obs/var) * Better description for filter's return * Add boolean to enumerated types * Augmented enum for scanpy axis * Create schema for annotations Based on datatype within scanpy/anndata + tests * remove obsolete schema parse script * Update rest api to remove old routes and add schema route * Separate development requirements * Warning for unsupported datatypes * include -r requirements.txt in dev * Merged downcast warnings * Fixed bug where names were NaNs Needed to include the index too when creating the series * Add config endpoint * Generate app features from CLI selections * Move features to driver * Add tests for schema * Clearer version wording * python3 version of super * version from engine to package level * move features to driver * Revise layout function to match the new spec * GET for layout/obs * PUT Layout (#211) * PUT Layout * Csweaver/annotations (#212) * Update scanpy engine to support the rest v0.2 annotation requests * GET endpoint for obs annotations + tests * Documentation * Test annotations in scanpy engine * Description for annotation-keys param * annotation->annotations * clarified return for annotations * Use URL query list for annotations fields * parse_filter parses v0.2 GET filters (#215) * parse_filter parses v0.2 GET filters * Don't allow index filters from query params * Better variable conversion * Parse filter improvements - uses default dict - renamed filter -> query_filter * Cleanup Tasks (#216) * Add test_api back into travis build * Do custom JSON encoding the correct way * Run cellxgene server in test setup * Cleanup new tests too * Option to bind to all interfaces (#225) app.run("0.0.0.0") instead of app.run("127.0.0.1") binds to all interfaces. Note: There are comments on the internet that says that the flask server is not up to the task of production serving. I don't think that such scalability concerns apply here, but I was able to get cellxgene working with twistd relatively easily, and we could switch to that if there are scalability concerns. Test plan: browsed to <ip>:5005/api/v0.2/config on a different host. * Fix merge errors - import warnings was improperly deleted - scanpy engine tests were totally wrong * Fix merge error with driver * PUT /annotations (#235) * Add query param for annotation name * fix descriptions, eliminate else clause * first cut at initial data load on rest 0.2 api * Annotation var (#248) * Fix bug strings are always objects in pandas * Add axis to annotation method * Add /annotation/var to REST api * Csweaver/expressiondata (#242) * Refactor expression method for REST v2 * Add message to QueryStringError * Fix range filters * Add GET route for /data * /data PUT route * rename expression to data_frame * clarification of error * Improve accept type handling * support all schema types for 0.2 REST API * remove REST 0.1 code; connect var annotations loading * config reducer; use config to set data set title; remove obsolete templating code for data set title * REST 0.2 expression conversion support * partial port of expression to REST 0.2 * diffexp (#273) * Add diffexp method to scanpy and test * Minor tweaks to diffexp Get a minimal working version to unblock FE development * Fixing things git deleted * cleanup print statements * Add index test * additional, partial REST 0.2 bring up of diffexp * Ignore unstructured annotations for data (#275) This is a temp hack, need to figure out how to include data.uns if there is only one gene * diffexp REST 0.2 port finish * ignore unstructured annotaitons on all routes except layout * correctly use varDataCache; maintain state during world rebuild * correct varDataCache use * temporarily disable all memoization * refinements to expression data caching * clear cell sets upon regraph/reset * update version of REST to 0.2 * Travis build fixes - comment out cache import - fix duplicate test name * Remove dependency from travis * clarify semantics of config variables * move generic action helpers into util
This commit is contained in:
@@ -1,69 +1,229 @@
|
||||
import json
|
||||
from os import path
|
||||
import pytest
|
||||
import time
|
||||
import unittest
|
||||
|
||||
import numpy as np
|
||||
from pandas import Series
|
||||
|
||||
from server.app.scanpy_engine.scanpy_engine import ScanpyEngine
|
||||
|
||||
|
||||
class UtilTest(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.data = ScanpyEngine("example-dataset/", schema="data_schema.json")
|
||||
self.data = ScanpyEngine("example-dataset/", layout_method="umap", diffexp_method="ttest")
|
||||
self.data._create_schema()
|
||||
|
||||
def test_init(self):
|
||||
self.assertEqual(self.data.cell_count, 2638)
|
||||
self.assertEqual(self.data.gene_count, 1838)
|
||||
epsilon = 0.000005
|
||||
self.assertTrue(self.data.data.X[0,0] - -0.17146951 < epsilon)
|
||||
self.assertTrue(self.data.data.X[0, 0] - -0.17146951 < epsilon)
|
||||
|
||||
def test_mandatory_annotations(self):
|
||||
self.assertIn("name", self.data.data.obs)
|
||||
self.assertEqual(list(self.data.data.obs.index), list(range(2638)))
|
||||
self.assertIn("name", self.data.data.var)
|
||||
self.assertEqual(list(self.data.data.var.index), list(range(1838)))
|
||||
|
||||
@pytest.mark.filterwarnings("ignore:Scanpy data matrix")
|
||||
def test_data_type(self):
|
||||
self.data.data.X = self.data.data.X.astype("float64")
|
||||
self.assertWarns(UserWarning, self.data._validatate_data_types())
|
||||
|
||||
def test_filter_idx(self):
|
||||
filter_ = {
|
||||
"filter": {
|
||||
"var": {
|
||||
"index": [1, 99, [200, 300]]
|
||||
},
|
||||
"obs": {
|
||||
"index": [1, 99, [1000, 2000]]
|
||||
}
|
||||
}
|
||||
}
|
||||
data = self.data.filter_dataframe(filter_["filter"])
|
||||
self.assertEqual(data.shape, (1002, 102))
|
||||
|
||||
def test_filter_annotation(self):
|
||||
filter_ = {
|
||||
"filter": {
|
||||
"obs": {
|
||||
"annotation_value": [
|
||||
{"name": "louvain", "values": ["NK cells", "CD8 T cells"]},
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
data = self.data.filter_dataframe(filter_["filter"])
|
||||
self.assertEqual(data.shape, (470, 1838))
|
||||
filter_ = {
|
||||
"filter": {
|
||||
"obs": {
|
||||
"annotation_value": [
|
||||
{"name": "n_counts", "min": 3000},
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
data = self.data.filter_dataframe(filter_["filter"])
|
||||
self.assertEqual(data.shape, (497, 1838))
|
||||
|
||||
def test_filter_annotation_no_uns(self):
|
||||
filter_ = {
|
||||
"filter": {
|
||||
"var": {
|
||||
"annotation_value": [
|
||||
{"name": "name", "values": ["RER1"]},
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
data = self.data.filter_dataframe(filter_["filter"], include_uns=False)
|
||||
self.assertEqual(data.shape[1], 1)
|
||||
|
||||
def test_filter_complex(self):
|
||||
filter_ = {
|
||||
"filter": {
|
||||
"var": {
|
||||
"index": [1, 99, [200, 300]]
|
||||
},
|
||||
"obs": {
|
||||
"annotation_value": [
|
||||
{"name": "louvain", "values": ["NK cells", "CD8 T cells"]},
|
||||
{"name": "n_counts", "min": 3000},
|
||||
],
|
||||
"index": [1, 99, [1000, 2000]]
|
||||
}
|
||||
}
|
||||
}
|
||||
data = self.data.filter_dataframe(filter_["filter"])
|
||||
self.assertEqual(data.shape, (15, 102))
|
||||
|
||||
def test_obs_and_var_names(self):
|
||||
self.assertEqual(np.sum(self.data.data.var["name"].isna()), 0)
|
||||
self.assertEqual(np.sum(self.data.data.obs["name"].isna()), 0)
|
||||
|
||||
def test_schema(self):
|
||||
self.assertEqual(self.data.schema, {'CellName': {'type': 'string', 'variabletype': 'categorical', 'displayname': 'Name', 'include': True}, 'n_genes': {'type': 'int', 'variabletype': 'continuous', 'displayname': 'Num Genes', 'include': True}, 'percent_mito': {'type': 'float', 'variabletype': 'continuous', 'displayname': 'Mitochondrial Percentage', 'include': True}, 'n_counts': {'type': 'float', 'variabletype': 'continuous', 'displayname': 'Num Counts', 'include': True}, 'louvain': {'type': 'string', 'variabletype': 'categorical', 'displayname': 'Louvain Cluster', 'include': True}})
|
||||
with open(path.join(path.dirname(__file__), "schema.json")) as fh:
|
||||
schema = json.load(fh)
|
||||
self.assertEqual(self.data.schema, schema)
|
||||
|
||||
def test_cells(self):
|
||||
cells = self.data.cells()
|
||||
self.assertIn("AAACATACAACCAC-1", cells)
|
||||
self.assertEqual(len(cells), 2638)
|
||||
def test_schema_produces_error(self):
|
||||
self.data.data.obs["time"] = Series(list([time.time() for i in range(self.data.cell_count)]),
|
||||
dtype="datetime64[ns]")
|
||||
with pytest.raises(TypeError):
|
||||
self.data._create_schema()
|
||||
|
||||
def test_genes(self):
|
||||
genes = self.data.genes()
|
||||
self.assertIn("SEPT4", genes)
|
||||
self.assertEqual(len(genes), 1838)
|
||||
def test_config(self):
|
||||
self.assertEqual(self.data.features["layout"]["obs"], {'available': True, 'interactiveLimit': 15000})
|
||||
|
||||
def test_filter_categorical(self):
|
||||
filter = {"louvain": {"variable_type": "categorical", "value_type": "string", "query": ["B cells"]}}
|
||||
filtered_data = self.data.filter_cells(filter)
|
||||
self.assertEqual(filtered_data.shape, (342, 1838))
|
||||
louvain_vals = filtered_data.obs['louvain'].tolist()
|
||||
self.assertIn("B cells", louvain_vals)
|
||||
self.assertNotIn("NK cells", louvain_vals)
|
||||
def test_layout(self):
|
||||
layout = self.data.layout(self.data.data)
|
||||
self.assertEqual(layout["ndims"], 2)
|
||||
self.assertEqual(len(layout["coordinates"]), 2638)
|
||||
self.assertEqual(layout["coordinates"][0][0], 0)
|
||||
for idx, val in enumerate(layout["coordinates"]):
|
||||
self.assertLessEqual(val[1], 1)
|
||||
self.assertLessEqual(val[2], 1)
|
||||
|
||||
def test_filter_continuous(self):
|
||||
# print(self.data.data.obs["n_genes"].tolist())
|
||||
filter = {"n_genes": {"variable_type": "continuous", "value_type": "int", "query": {"min": 300, "max": 400}}}
|
||||
filtered_data = self.data.filter_cells(filter)
|
||||
self.assertEqual(filtered_data.shape, (71, 1838))
|
||||
n_genes_vals = filtered_data.obs['n_genes'].tolist()
|
||||
for val in n_genes_vals:
|
||||
self.assertTrue(300 <= val <= 400)
|
||||
def test_annotations(self):
|
||||
annotations = self.data.annotation(self.data.data, "obs")
|
||||
self.assertEqual(annotations["names"], ["n_genes", "percent_mito", "n_counts", "louvain", "name"])
|
||||
self.assertEqual(len(annotations["data"]), 2638)
|
||||
annotations = self.data.annotation(self.data.data, "var")
|
||||
self.assertEqual(annotations["names"], ["n_cells", "name"])
|
||||
self.assertEqual(len(annotations["data"]), 1838)
|
||||
|
||||
def test_metadata(self):
|
||||
metadata = self.data.metadata(df=self.data.data)
|
||||
self.assertEqual(len(metadata), 2638)
|
||||
self.assertIn('louvain', metadata[0])
|
||||
def test_annotation_fields(self):
|
||||
annotations = self.data.annotation(self.data.data, "obs", ["n_genes", "n_counts"])
|
||||
self.assertEqual(annotations["names"], ["n_genes", "n_counts"])
|
||||
self.assertEqual(len(annotations["data"]), 2638)
|
||||
annotations = self.data.annotation(self.data.data, "var", ["name"])
|
||||
self.assertEqual(annotations["names"], ["name"])
|
||||
self.assertEqual(len(annotations["data"]), 1838)
|
||||
|
||||
@unittest.skip("Umap not producing the same graph on different systems, even with the same seed. Skipping for now")
|
||||
def test_create_graph(self):
|
||||
graph = self.data.create_graph(df=self.data.data)
|
||||
self.assertEqual(graph[0][1], 0.5545382653143183)
|
||||
self.assertEqual(graph[0][2], 0.6021833809031731)
|
||||
def test_filtered_annotation(self):
|
||||
filter_ = {
|
||||
"filter": {
|
||||
"obs": {
|
||||
"annotation_value": [
|
||||
{"name": "n_counts", "min": 3000},
|
||||
]
|
||||
},
|
||||
"var": {
|
||||
"annotation_value": [
|
||||
{"name": "name", "values": ["ATAD3C", "RER1"]},
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
data = self.data.filter_dataframe(filter_["filter"])
|
||||
annotations = self.data.annotation(data, "obs")
|
||||
self.assertEqual(annotations["names"], ["n_genes", "percent_mito", "n_counts", "louvain", "name"])
|
||||
self.assertEqual(len(annotations["data"]), 497)
|
||||
annotations = self.data.annotation(data, "var")
|
||||
self.assertEqual(annotations["names"], ["n_cells", "name"])
|
||||
self.assertEqual(len(annotations["data"]), 2)
|
||||
|
||||
def test_filtered_layout(self):
|
||||
filter_ = {
|
||||
"filter": {
|
||||
"obs": {
|
||||
"annotation_value": [
|
||||
{"name": "n_counts", "min": 3000},
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
data = self.data.filter_dataframe(filter_["filter"])
|
||||
layout = self.data.layout(data)
|
||||
self.assertEqual(len(layout["coordinates"]), 497)
|
||||
|
||||
def test_diffexp(self):
|
||||
diffexp = self.data.diffexp(["AAACATACAACCAC-1", "AACCGATGGTCATG-1"], ["CCGATAGACCTAAG-1", "GGTGGAGAAGTAGA-1"], 0.5, 7)
|
||||
self.assertEqual(diffexp["celllist1"]["topgenes"], ['EBNA1BP2', 'DIAPH1', 'SLC25A11', 'SNRNP27', 'COMMD8', 'COTL1', 'GTF3A'])
|
||||
f1 = {
|
||||
"filter": {
|
||||
"obs": {
|
||||
"index": [[0, 500]]
|
||||
}
|
||||
}
|
||||
}
|
||||
df1 = self.data.filter_dataframe(f1["filter"])
|
||||
f2 = {
|
||||
"filter": {
|
||||
"obs": {
|
||||
"index": [[500, 1000]]
|
||||
}
|
||||
}
|
||||
}
|
||||
df2 = self.data.filter_dataframe(f2["filter"])
|
||||
result = self.data.diffexp(df1, df2)
|
||||
self.assertEqual(len(result), 10)
|
||||
var_idx = [i[0] for i in result]
|
||||
self.assertEqual(var_idx, sorted(var_idx))
|
||||
result = self.data.diffexp(df1, df2, 20)
|
||||
self.assertEqual(len(result), 20)
|
||||
|
||||
def test_expression(self):
|
||||
expression = self.data.expression(cells=["AAACATACAACCAC-1"])
|
||||
data_exp = self.data.data[["AAACATACAACCAC-1"], :].X
|
||||
for idx in range(len(expression["cells"][0]["e"])):
|
||||
self.assertEqual(expression["cells"][0]["e"][idx], data_exp[idx])
|
||||
def test_data_frame(self):
|
||||
data_frame = self.data.data_frame(self.data.data)
|
||||
self.assertEqual(len(data_frame["var"]), 1838)
|
||||
self.assertEqual(len(data_frame["obs"]), 2638)
|
||||
|
||||
def test_filtered_data_frame(self):
|
||||
filter_ = {
|
||||
"filter": {
|
||||
"obs": {
|
||||
"annotation_value": [
|
||||
{"name": "n_counts", "min": 3000},
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
data = self.data.filter_dataframe(filter_["filter"])
|
||||
data_frame = self.data.data_frame(data)
|
||||
self.assertEqual(len(data_frame["var"]), 1838)
|
||||
self.assertEqual(len(data_frame["obs"]), 497)
|
||||
|
||||
if __name__ == '__main__':
|
||||
unittest.main()
|
||||
if __name__ == '__main__':
|
||||
unittest.main()
|
||||
|
||||
Reference in New Issue
Block a user