Compare commits

..
1 Commits
Author SHA1 Message Date
Charlotte Weaver 07cee2c710 0.0.3 bump 2018-11-15 11:37:27 -08:00
25 changed files with 119 additions and 310 deletions
+1 -1
View File
@@ -1,5 +1,5 @@
[bumpversion] [bumpversion]
current_version = 0.2.0 current_version = 0.0.3
[bumpversion:file:setup.py] [bumpversion:file:setup.py]
search = version="{current_version}" search = version="{current_version}"
-1
View File
@@ -1,4 +1,3 @@
recursive-include server/app/web/templates * recursive-include server/app/web/templates *
recursive-include server/app/web/static * recursive-include server/app/web/static *
include server/requirements.txt
+78 -195
View File
@@ -1,228 +1,111 @@
# cellxgene # cellxgene
> an interactive explorer for single-cell transcriptomics data ### An interactive, performant explorer for single cell transcriptomics data.
`cellxgene` is an interactive data explorer for single-cell transcriptomics datasets, such as those coming from the [Human Cell Atlas](https://humancellatlas.org). Leveraging modern web development techniques to enable fast visualizations of at least 1 million cells, we hope to enable biologists and computational researchers to explore their data, and to demonstrate general, scalable, and reusable patterns for scientific data visualization. <img align="right" width="350" height="218" src="./example-dataset/cellxgene-demo.gif" pad="50px">
cellxgene is an open-source experiment in how to bring powerful tools from modern web development to visualize and explore large single-cell transcriptomics datasets.
Started in the context of the Human Cell Atlas Consortium, cellxgene hopes to both enable scientists to explore their data and to equip developers with scalable, reusable patterns and frameworks for visualizing large scientific datasets.
<img src="https://github.com/chanzuckerberg/cellxgene/blob/master/docs/cellxgene-demo-1.gif" width="200" height="200" hspace="30"><img src="https://github.com/chanzuckerberg/cellxgene/blob/master/docs/cellxgene-demo-2.gif" width="200" height="200" hspace="30"><img src="https://github.com/chanzuckerberg/cellxgene/blob/master/docs/cellxgene-demo-3.gif" width="200" height="200" hspace="30"> ## Features
## getting started - **Visualization at scale:** built with [WebGL](https://www.khronos.org/webgl/), [React](https://reactjs.org/) & [Redux](https://redux.js.org/) to handle visualization of at least 1 million cells.
You'll need **python 3.6** and **Google Chrome**. The web UI is tested on OSX and Windows using Chrome, and the python CLI is tested on OSX and Ubuntu (via WSL/Windows). It should work on other platforms, but if you run into trouble let us know (see [help](#help-and-contact) below). - **Interactive exploration:** select, cross-filter, and compare subsets of your data with performant indexing and data handling.
To install run - **Flexible API:** the cellxgene client-server model is designed to support a range of existing analysis packages for backend computational tasks (eg scanpy), integrated with client-side visualization via a [REST API](https://restfulapi.net/).
``` ## Getting Started
pip install cellxgene
```
To start exploring a dataset call **Requirements**
```
cellxgene launch dataset.h5ad --open
```
If you want an example dataset download [this file](https://github.com/chanzuckerberg/cellxgene/raw/master/example-dataset/pbmc3k.h5ad) and then call
```
cellxgene launch pbmc3k.h5ad --open
```
You should see your web browser open with the following
<img width="450" src="https://github.com/chanzuckerberg/cellxgene/blob/master/docs/cellxgene-opening-screenshot.png" pad="50px">
**Note**: automatic opening of the browser with the `--open` flag only works on OS X, on other platforms you'll need to directly point to the provided link in your browser.
There are several options available, such as:
- `--layout` to specify the layout as `tsne` or `umap`
- `--title` to show a title on the explorer
- `--open` to automatically open the web browser after launching (OS X only)
To see all options call
```
cellxgene launch --help
```
There is an additional subcommand called `cellxgene prepare` that takes an existing dataset in one of several formats and applies minimal preprocessing and reformatting so that `launch` can use it (see [the next section](##data-formatting) for more info on `prepare`).
## data formatting
### assumptions
The `launch` command assumes that the data is stored in the `.h5ad` format from the [`anndata`](https://anndata.readthedocs.io/en/latest/index.html) library. It also assumes that certain computations have already been performed. Briefly, the `.h5ad` format wraps a two-dimensional `ndarray` and stores additional metadata as "annotations" for either observations (referred to as `obs` and `obsm`) or variables (`var` and `varm`). `cellxgene launch` makes the following assumptions about your data (we recommend loading and inspecting your data using `scanpy` to validate these assumptions)
- an `obs` field has a unique identifier for every cell (you can specify which field to use with the `--obs-names` option, by default it will use the value of `data.obs_names`)
- a `var` field has a unique identifier for every gene (you can specify which field to use with the `--var-names` option, by default it will use the value of `data.var_names`)
- an `obsm` field contains the two-dimensional coordinates for the layout that you want to render (e.g. `X_tsne` for the `tsne` layout or `X_umap` for the `umap` layout)
- any additional `obs` fields will be rendered as per-cell continuous or categorical metadata by the app (e.g. `louvain` cluster assignments)
### prepare
The `prepare` command is included to help you format your data. It uses `scanpy` under the hood. This is especially useful if you are starting with raw unanalyzed data and are unfamiliar with `scanpy`.
To prepare from an existing `.h5ad` file use
```
cellxgene prepare dataset.h5ad --output=dataset-processed.h5ad
```
This will load the input data, perform PCA and nearest neighbor calculations, compute `umap` and `tsne` layouts and `louvain` cluster assignments, and save the results in a new file called `dataset-processed.h5ad` that can be loaded using `cellxgene launch`. Data can be loaded from several formats, including `.h5ad` `.loom` and a `10-Genomics-formatted` `mtx` directory. Several options are available, including running one of the preprocessing `recipes` included with `scanpy`, which include steps like cell filtering and gene selection.
Depending on the options chosen, `prepare` can take a long time to run (a few minutes for datasets with 10-100k cells, up to an hour or more for datasets with >100k cells). If you want `prepare` to run faster we recommend using the `sparse` option and only computing the layout for `umap`, using a call like this
```
cellxgene prepare dataset.h5ad --output=dataset-processed.h5ad --layout=umap --sparse
```
To see all options call
```
cellxgene prepare --help
```
**Note**: `cellxgene prepare` will only perform `louvain` clustering if you have the `python-igraph` and `louvain` packages installed. To make sure they are installed alongside `cellxgene` use
```
pip install cellxgene[louvain]
```
## conda and virtual environments
If you use conda and want to create a conda environment for `cellxgene` you can use the following commands
```
conda create --yes -n cellxgene python=3.6
conda activate cellxgene
pip install cellxgene
```
Or you can create a virtual environment by using
```
ENV_NAME=cellxgene
python3 -m venv ${ENV_NAME}
source ${ENV_NAME}/bin/activate
pip install cellxgene
```
## FAQ
> Someone sent me a directory of `10X-Genomics` data with a `mtx` file and I've never used `scanpy`, can I use `cellxgene`?
Yep! This should only take a couple steps. We'll assume your data is in a folder called `data/` and you've successfully installed `cellxgene` with the `louvain` packages as described above. Just run
```
cellxgene prepare data/ --output=data-processed.h5ad --layout=umap
```
Depending on the size of the dataset, this may take some time. Once it's done, call
```
cellxgene launch data-processed.h5ad --layout=umap --open
```
And your web browser should open with an interactive view of your data.
> In my `prepare` command I received the following error `Warning: louvain module is not installed, no clusters will be calculated. To fix this please install cellxgene with the optional feature louvain enabled`
Louvain clustering requires additional dependencies that are somewhat complex, so we don't include them by default. For now, you need to specify that you want these packages by using
```
pip install cellxgene[louvain]
```
> I ran `prepare` and I'm getting results that look unexpected
You might want to try running one of the preprocessing recipes included with `scanpy` (read more about them [here](https://scanpy.readthedocs.io/en/latest/api/index.html#recipes)). You can specify this with the `--recipe` option, such as
```
cellxgene prepare data/ --output=data-processed.h5ad --recipe=zheng17
```
It should be easy to run `prepare` then call `cellxgene launch` a few times with different settings to explore different behaviors. We may explore adding other preprocessing options in the future.
> I have extra metadata that I want to add to my dataset
Currently this is not supported directly, but you should be able to do this manually using `scanpy`. For example, this [notebook](https://github.com/falexwolf/fun-analyses/blob/master/tabula_muris/tabula_muris.ipynb) shows adding the contents of a `csv` file with metadata to an `anndata` object. For now, you could do this manually on your data in the same way and then save out the result before loading into `cellxgene`.
> I tried to `pip install cellxgene` and got a weird error I don't understand
This may happen, especially as we work out bugs in our installation process! Please create a new [Github issue](https://github.com/chanzuckerberg/cellxgene/issues), explain what you did, and include all the error messages you saw. It'd also be super helpful if you call `pip freeze` and include the full output alongside your issue.
> How are you computing and sorting differential expression results?
Currently we use a [Welch's *t*-test](https://en.wikipedia.org/wiki/Welch%27s_t-test) implementation including the same variance overestimation correction as used in `scanpy`. We sort the `tscore` to identify the top N genes, and then filter to remove any that fall below a cutoff log fold change value, which can help remove spurious test results. The default threshold is `0.01` and can be changed using the option `--diffexp-lfc-cutoff`. We can explore adding support for other test types in the future.
> I'm following the developer instructions and get an error about "missing files and directories” when trying to build the client
This is likely because you do not have node and npm installed, we recommend using [nvm](https://github.com/creationix/nvm) if you're new to using these tools.
## developer guide
This project has made a few key design choices
- The front-end is built with [`regl`](https://github.com/regl-project/regl) (a webgl library), [`react`](https://reactjs.org/), [`redux`](https://redux.js.org/), [`d3`](https://github.com/d3/d3), and [`blueprint`](https://blueprintjs.com/docs/#core) to handle rendering large numbers of cells with lots of complex interactivity
- The app is designed with a client-server model that can support a range of existing analysis packages for backend computational tasks (currently built for [scanpy](https://github.com/theislab/scanpy))
- The client uses fast cross-filtering to handle selections and comparisons across subsets of data
Depending on your background and interests, you might want to contribute to the frontend, or backend, or both!
If you are interested in working on `cellxgene` development, we recommend cloning the project from Gitub. First you'll need the following installed on your machine
- OS: OSX, Windows, Linux -- the developers are currently testing on OSX and Windows (via WSL using Ubuntu). It should work on other platforms but if you are using something different and need help, please let us know.
- python 3.6 - python 3.6
- node and npm (we recommend using [nvm](https://github.com/creationix/nvm) if this is your first time with node) - python3 tkinter
- npm
- Google Chrome
Then clone the project **Clone project**
``` git clone https://github.com/chanzuckerberg/cellxgene.git
git clone https://github.com/chanzuckerberg/cellxgene.git
```
Build the client web assets by calling this from inside the `cellxgene` folder **Install client**
``` cd cellxgene
./bin/build-client ./bin/build-client
```
Install all requirements (we recommend doing this inside a virtual environment) **To use with virtual env for python**
(optional, but recommended)
``` ENV_NAME=cellxgene
pip install -e . python3 -m venv ${ENV_NAME}
``` source ${ENV_NAME}/bin/activate
You can start the app while developing either by calling `cellxgene` or by calling `python -m server`. We recommend using the `--debug` flag to see more output, which you can include when reporting bugs. **Install server**
If you have any questions about developing or contributing, come hang out with us by joining the [CZI Science Slack](https://cziscience.slack.com/messages/CCTA8DF1T) and posting in the `#cellxgene-dev` channel. pip install -e .
## development roadmap **Run (with demo data)**
`cellxgene` is still very much in development, and we've love to include the community as we plan new features to work on. We are thinking about working on the following features over the next 3-12 months. If you are interested in updates, want to give feedback, want to contribute, or have ideas about other features we should work on, please [contact us](#help-and-contact) cellxgene launch --title PBMC3K example-dataset/pbmc3k.h5ad
- **Visualizaling spatial metadata** Image-based transcriptomics methods also generate large cell by gene matrices, alongside rich metadata about spatial location; we would like to render this information in `cellxgene` **Help**
- **Visualizing trajectories** Trajectory analyses infer progression along some ordering or pseudotime; we would like `cellxgene ` to render the results of these analyses when they have been performed
- **Deploy to web** Many projects release public data browser websites alongside their publicatons; we would like to make it easy for anyone to deploy `cellxgene` to a custom URL with their own dataset that they own and operate
- **HCA Integration** The [Human Cell Atlas](https://humancellatlas.org) is generating a large corpus of single-cell expression data and will make it available through the Data Coordination Platform; we would like `cellxgene` to be one of several different portals for browsing these data
## contributing cellxgene --help
We warmly welcome contributions from the community! Please submit any bug reports and feature requests through [Github issues](https://github.com/chanzuckerberg/cellxgene/issues). Please submit any direct contributions by forking the repository, creating a branch, and submitting a Pull Request. It'd be great for PRs to include test cases and documentation updates where relevant, though we know the core test suite is itself still a work in progress. And all code contributions and dependencies must be compatible with the project's open-source license (MIT). If you have any questions about this stuff, just ask! _For help with the scanpy engine_
## inspiration and collaboration cellxgene scanpy --help
We've been heavily inspired by several other related single-cell visualization projects, including the [UCSC Cell Browswer](http://cells.ucsc.edu/), [Cytoscape](http://www.cytoscape.org/), [Xena](https://xena.ucsc.edu/), [ASAP](https://asap.epfl.ch/), [Gene Pattern](http://genepattern-notebook.org/), and many others. We hope to explore collaborations where useful as this community works together on improving interactive visualization for single-cell data. ## Using your own data
We were inspired by Mike Bostock and the [crossfilter](https://github.com/crossfilter) team for the design of our filtering implementation. ### Scanpy
We have been working closely with the [`scanpy`](https://github.com/theislab/scanpy) team to integrate with their awesome analysis tools. Special thanks to Alex Wolf, Fabian Theis, and the rest of the team for their help during development and for providing an example dataset. To prepare your data you will need to format your data into AnnData format using scanpy and calculate PCA and nearest neighbors and save in h5ad format.
We are eager to explore integrations with other computational backends such as [`Seurat`](https://github.com/satijalab/seurat) or [`Bioconductor`](https://github.com/Bioconductor) 1. [Load data into scanpy](https://scanpy.readthedocs.io/en/latest/api/index.html#reading)
## help and contact - Ensure that `obs`'s index is the cell names: `print(data.obs_names)` should show your cell indices. If it shows gene names, you may need to just call `data.transpose()`.
Have questions, suggestions, or comments? You can come hang out with us by joining the [CZI Science Slack](https://cziscience.slack.com/messages/CCTA8DF1T) and posting in the `#cellxgene-users` channel. As mentioned above, please submit any feature requests or bugs as [Github issues](https://github.com/chanzuckerberg/cellxgene/issues). We'd love to hear from you! 2. Calculate PCA
## reuse sc.pp.pca(data) ## sc is scanpy.api
This project was started with the sole goal of empowering the scientific community to explore and understand their data. As such, we encourage other scientific tool builders in academia or industry to adopt the patterns, tools, and code from this project, and reach out to us with ideas or questions. All code is freely available for reuse under the [MIT license](https://opensource.org/licenses/MIT). 3. Calculate nearest neighbors (depending on layout algorithm)
```
# For umap layout algorithm, you need to use the "umap" method for neighbors
sc.pp.neighbors(data, method="umap", metric="euclidean", use_rep="X_pca")
# For tsne layout algorithm, you can use either "umap" or "gauss"; we recommend "gauss"
sc.pp.neighbors(data, method="gauss", metric="euclidean", use_rep="X_pca")
```
4. Save file
```
# cellxgene requires file to be named data.h5ad
data.write("data.h5ad")
```
## Contributing
We warmly welcome contributions from the community. Please submit any bug reports and feature requests through github issues. Please submit any direct contributions via a branch + pull request.
## Inspiration and collaboration
We’ve been inspired by several other related efforts in this space, including the [UCSC Cell Browswer](http://cells.ucsc.edu/), [Cytoscape](http://www.cytoscape.org/), [Xena](https://xena.ucsc.edu/), [ASAP](https://asap.epfl.ch/), [Gene Pattern](http://genepattern-notebook.org/), & many others; we hope to explore collaborations where useful.
## Help/Contact
Have questions, suggestions, or comments? You can contact us by joining [CZI Science Slack](https://cziscience.slack.com/messages/CCTA8DF1T) and posting in the #cellxgene channel. Please submit any feature requests or bugs as an issue in github. We'd love to hear from you!
## Reuse
This project was started with the sole goal of empowering the scientific community to explore and understand their data. As such, we whole-heartedly encourage other scientific tool builders to adopt the patterns, tools, and code from this project, and reach out to us with ideas or questions using Github Issues or Pull Requests. All code is freely available for reuse under the [MIT license](https://opensource.org/licenses/MIT).
## Acknowledgements
cellxgene is inspired by many innovative projects. We would like to specifically thank:
- Alex Wolf for the demo dataset.
- Mike Bostock and the [crossfilter](https://github.com/crossfilter) team for API inspiration.
-2
View File
@@ -7,8 +7,6 @@ echo "removing node_modules"
rm -rf $CELLXGENE_DIR/client/node_modules rm -rf $CELLXGENE_DIR/client/node_modules
echo "removing client_build" echo "removing client_build"
rm -rf $CELLXGENE_DIR/client/build rm -rf $CELLXGENE_DIR/client/build
echo "removing dist"
rm -rf $CELLXGENE_DIR/dist
echo "removing egg-info" echo "removing egg-info"
rm -rf $CELLXGENE_DIR/cellxgene.egg-info rm -rf $CELLXGENE_DIR/cellxgene.egg-info
echo "removing static files" echo "removing static files"
+1 -2
View File
@@ -34,8 +34,7 @@ module.exports = {
"object-curly-newline": ["error", { consistent: true }], "object-curly-newline": ["error", { consistent: true }],
"react/prop-types": [0], "react/prop-types": [0],
"space-before-function-paren": "off", "space-before-function-paren": "off",
"function-paren-newline": "off", "function-paren-newline": "off"
"prefer-destructuring": ["error", { object: true, array: false }]
}, },
overrides: [ overrides: [
{ {
+1 -1
View File
@@ -1,6 +1,6 @@
{ {
"name": "cellxgene", "name": "cellxgene",
"version": "0.1.0", "version": "0.0.3",
"lockfileVersion": 1, "lockfileVersion": 1,
"requires": true, "requires": true,
"dependencies": { "dependencies": {
+1 -1
View File
@@ -1,6 +1,6 @@
{ {
"name": "cellxgene", "name": "cellxgene",
"version": "0.2.0", "version": "0.0.3",
"license": "MIT", "license": "MIT",
"description": "cellxgene is a web application for the interactive exploration of single cell sequence data.", "description": "cellxgene is a web application for the interactive exploration of single cell sequence data.",
"repository": "https://github.com/chanzuckerberg/cellxgene", "repository": "https://github.com/chanzuckerberg/cellxgene",
+3 -16
View File
@@ -1,13 +1,11 @@
// jshint esversion: 6 // jshint esversion: 6
import { connect } from "react-redux"; import { connect } from "react-redux";
import React from "react"; import React from "react";
import _ from "lodash";
@connect(state => ({ @connect(state => ({
categoricalAsBooleansMap: state.controls.categoricalAsBooleansMap, categoricalAsBooleansMap: state.controls.categoricalAsBooleansMap,
colorScale: state.controls.colorScale, colorScale: state.controls.colorScale,
colorAccessor: state.controls.colorAccessor, colorAccessor: state.controls.colorAccessor
schema: _.get(state.controls.world, "schema", null)
})) }))
class CategoryValue extends React.Component { class CategoryValue extends React.Component {
toggleOff() { toggleOff() {
@@ -36,8 +34,7 @@ class CategoryValue extends React.Component {
value, value,
colorAccessor, colorAccessor,
colorScale, colorScale,
i, i
schema
} = this.props; } = this.props;
if (!categoricalAsBooleansMap) return null; if (!categoricalAsBooleansMap) return null;
@@ -45,13 +42,6 @@ class CategoryValue extends React.Component {
const selected = categoricalAsBooleansMap[metadataField][value]; const selected = categoricalAsBooleansMap[metadataField][value];
/* this is the color scale, so add swatches below */ /* this is the color scale, so add swatches below */
const c = metadataField === colorAccessor; const c = metadataField === colorAccessor;
let categories = null;
if (c && schema) {
categories = _.filter(schema.annotations.obs, {
name: colorAccessor
})[0].categories;
}
return ( return (
<div <div
@@ -88,10 +78,7 @@ class CategoryValue extends React.Component {
marginLeft: 5, marginLeft: 5,
width: 11, width: 11,
height: 11, height: 11,
backgroundColor: backgroundColor: c ? colorScale(value) : "inherit"
c && categories
? colorScale(categories.indexOf(value))
: "inherit"
}} }}
/> />
</span> </span>
@@ -2,7 +2,7 @@
import React from "react"; import React from "react";
import { connect } from "react-redux"; import { connect } from "react-redux";
import * as d3 from "d3"; import * as d3 from "d3";
import { interpolateViridis, interpolateCool } from "d3-scale-chromatic"; import { interpolateViridis } from "d3-scale-chromatic";
// create continuous color legend // create continuous color legend
// http://bl.ocks.org/syntagmatic/e8ccca52559796be775553b467593a9f // http://bl.ocks.org/syntagmatic/e8ccca52559796be775553b467593a9f
@@ -121,12 +121,12 @@ class ContinuousLegend extends React.Component {
.remove(); .remove();
} }
if (colorAccessor && colorScale && colorScale.range) { if (colorAccessor && colorScale) {
/* fragile! continuous range is 0 to 1, not [#fa4b2c, ...], make this a flag? */ /* fragile! continuous range is 0 to 1, not [#fa4b2c, ...], make this a flag? */
if (colorScale.range()[0][0] !== "#") { if (colorScale.range()[0][0] !== "#") {
continuous( continuous(
"#continuous_legend", "#continuous_legend",
d3.scaleSequential(interpolateCool).domain(colorScale.domain()), d3.scaleSequential(interpolateViridis).domain(colorScale.domain()),
colorAccessor colorAccessor
); );
} }
+5 -17
View File
@@ -1,13 +1,7 @@
// jshint esversion: 6 // jshint esversion: 6
import _ from "lodash"; import _ from "lodash";
import * as d3 from "d3"; import * as d3 from "d3";
import { import { interpolateViridis } from "d3-scale-chromatic";
interpolateViridis,
interpolateSpectral,
interpolateRainbow,
interpolateBlues,
interpolateCool
} from "d3-scale-chromatic";
import * as globals from "../globals"; import * as globals from "../globals";
import parseRGB from "../util/parseRGB"; import parseRGB from "../util/parseRGB";
@@ -65,17 +59,11 @@ const updateCellColorsMiddleware = store => next => action => {
*/ */
if (action.type === "color by categorical metadata") { if (action.type === "color by categorical metadata") {
const categories = _.filter(s.controls.world.schema.annotations.obs, { colorScale = d3.scaleOrdinal().range(globals.ordinalColors);
name: action.colorAccessor
})[0].categories;
colorScale = d3
.scaleSequential(interpolateRainbow)
.domain([0, categories.length]);
for (let i = 0; i < obsAnnotations.length; i += 1) { for (let i = 0; i < obsAnnotations.length; i += 1) {
const obs = obsAnnotations[i]; const obs = obsAnnotations[i];
const c = colorScale(categories.indexOf(obs[action.colorAccessor])); const c = colorScale(obs[action.colorAccessor]);
colorsByName[i] = c; colorsByName[i] = c;
colorsByRGB[i] = parseRGB(c); colorsByRGB[i] = parseRGB(c);
} }
@@ -89,7 +77,7 @@ const updateCellColorsMiddleware = store => next => action => {
for (let i = 0; i < obsAnnotations.length; i += 1) { for (let i = 0; i < obsAnnotations.length; i += 1) {
const obs = obsAnnotations[i]; const obs = obsAnnotations[i];
const c = interpolateCool(colorScale(obs[action.colorAccessor])); const c = interpolateViridis(colorScale(obs[action.colorAccessor]));
colorsByName[i] = c; colorsByName[i] = c;
colorsByRGB[i] = parseRGB(c); colorsByRGB[i] = parseRGB(c);
} }
@@ -107,7 +95,7 @@ const updateCellColorsMiddleware = store => next => action => {
]); /* invert viridis... probably pass this scale through to others */ ]); /* invert viridis... probably pass this scale through to others */
for (let i = 0, len = expression.length; i < len; i += 1) { for (let i = 0, len = expression.length; i < len; i += 1) {
const c = interpolateCool(colorScale(expression[i])); const c = interpolateViridis(colorScale(expression[i]));
colorsByName[i] = c; colorsByName[i] = c;
colorsByRGB[i] = parseRGB(c); colorsByRGB[i] = parseRGB(c);
} }
+2 -2
View File
@@ -568,14 +568,14 @@ If differential expression is not supported by the server, must return an HTTP 5
**Response body:** **Response body:**
- For 200 Success, differential expression statistics returned as array of arrays, where each contains the following values: - For 200 Success, differential expression statistics returned as array of arrays sorted by varindex, where each contains the following values:
- **varIndex**: variable index for the computed results - **varIndex**: variable index for the computed results
- **logfoldchange**: log fold-change of the average expression between the two groups. Positive values indicate that the gene is more highly expressed in the first group, - **logfoldchange**: log fold-change of the average expression between the two groups. Positive values indicate that the gene is more highly expressed in the first group,
- **pVal**: unadjusted p-value, - **pVal**: unadjusted p-value,
- **pValAdj**: adjusted p-value - **pValAdj**: adjusted p-value
Values ordered as: Statistics are encoded as an array of arrays, with fields ordered as:
_varIndex_, _logfoldchange_, _pVal_, _pValAdj_ _varIndex_, _logfoldchange_, _pVal_, _pValAdj_
Binary file not shown.

Before

Width:  |  Height:  |  Size: 1.1 MiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 312 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 285 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 644 KiB

+2 -10
View File
@@ -24,20 +24,12 @@ Follow these steps to create a release.
2. Create a release branch, eg, `release-version` 2. Create a release branch, eg, `release-version`
3. In the release branch: 3. In the release branch:
- run `bumpversion --config-file .bumpversion.cfg [major | minor | patch]` - run `bumpversion --config-file .bumpversion.cfg [major | minor | patch]`
- clean up existing environment using `bin/clean`
- build the JS asserts using `bin/build-client` - build the JS asserts using `bin/build-client`
4. Commit and push the new branch 4. Commit and push the new branch
5. Create a PR for the release. 5. Create a PR for the release.
- [optional] As needed, conduct PR review. - [optional] As needed, conduct PR review.
6. Merge to master 6. Create Github release using the version number and release notes ([instructions](https://help.github.com/articles/creating-releases/)).
7. Create Github release using the version number and release notes ([instructions](https://help.github.com/articles/creating-releases/)). 7. Publish to pypi by performing the following steps
- Draft new release
- Type version name matching release version number from (1)
- Select `master` as release branch (ensure you merged the release PR)
- Type title `Release {version num}`
- [optional] check pre-release if this release is not ready for production
- Publish Release
8. Publish to pypi by performing the following steps
(assumes you have `setuptools` and `twine` installed and that you have (assumes you have `setuptools` and `twine` installed and that you have
registered for pypi and have write access to the cellxgene pypi package) registered for pypi and have write access to the cellxgene pypi package)
- build the distribution by calling - build the distribution by calling

Before

Width:  |  Height:  |  Size: 6.0 MiB

After

Width:  |  Height:  |  Size: 6.0 MiB

-1
View File
@@ -17,7 +17,6 @@ class CXGDriver(metaclass=ABCMeta):
self.layout_method = args["layout"] self.layout_method = args["layout"]
self.diffexp_method = args["diffexp"] self.diffexp_method = args["diffexp"]
self.max_category_items = args["max_category_items"] self.max_category_items = args["max_category_items"]
self.diffexp_lfc_cutoff = args["diffexp_lfc_cutoff"]
self.cluster = None self.cluster = None
@property @property
+2 -2
View File
@@ -346,7 +346,7 @@ class DataObsAPI(Resource):
def get(self): def get(self):
accept_type = request.args.get("accept-type", None) accept_type = request.args.get("accept-type", None)
# request.args is immutable # request.args is immutable
args = request.args.copy() args = dict(request.args)
args.pop("accept-type", None) args.pop("accept-type", None)
try: try:
filter_ = parse_filter(ImmutableMultiDict(args), current_app.data.schema['annotations']) filter_ = parse_filter(ImmutableMultiDict(args), current_app.data.schema['annotations'])
@@ -450,7 +450,7 @@ class DataVarAPI(Resource):
def get(self): def get(self):
accept_type = request.args.get("accept-type", None) accept_type = request.args.get("accept-type", None)
# request.args is immutable # request.args is immutable
args = request.args.copy() args = dict(request.args)
args.pop("accept-type", None) args.pop("accept-type", None)
try: try:
filter_ = parse_filter(ImmutableMultiDict(args), current_app.data.schema['annotations']) filter_ = parse_filter(ImmutableMultiDict(args), current_app.data.schema['annotations'])
+12 -42
View File
@@ -25,42 +25,25 @@ def _mean_var_n(X):
return mean, v, n return mean, v, n
def diffexp_ttest(adata, maskA, maskB, top_n=8, diffexp_lfc_cutoff=0.01): def diffexp_ttest(adata, maskA, maskB, top_n=8):
""" """
Return differential expression statistics for top N variables. Return differential expression statistics for top N variables, sorted by
t statistic. Implemented as a unequal variance t-test.
Algorithm:
- compute log fold change (log2(meanA/meanB))
- compute Welch's t-test statistic and pvalue (w/ Bonferroni correction)
- return top N abs(logfoldchange) where lfc > diffexp_lfc_cutoff
If there are not N which meet criteria, augment by removing the logfoldchange
threshold requirement.
Notes on alogrithm:
- Welch's ttest provides basic statistics test.
https://en.wikipedia.org/wiki/Welch%27s_t-test
- p-values adjusted with Bonferroni correction.
https://en.wikipedia.org/wiki/Bonferroni_correction
:param adata: anndata dataframe :param adata: anndata dataframe
:param maskA: observation selection mask for set 1 :param maskA: observation selection mask for set 1
:param maskB: observation selection mask for set 2 :param maskB: observation selection mask for set 2
:param top_n: number of variables to return stats for :param top_n: number of variables to return stats for
:param diffexp_lfc_cutoff: minimum
:return: for top N genes, [ varindex, logfoldchange, pval, pval_adj ] :return: for top N genes, [ varindex, logfoldchange, pval, pval_adj ]
""" """
if top_n > adata.n_obs: # mean, variance, N
top_n = adata.n_obs
# mean, variance, N - calculate for both selections
meanA, vA, nA = _mean_var_n(adata._X[maskA]) meanA, vA, nA = _mean_var_n(adata._X[maskA])
meanB, vB, nB = _mean_var_n(adata._X[maskB]) meanB, vB, nB = _mean_var_n(adata._X[maskB])
# variance / N # variance / N
vnA = vA / min(nA, nB) # overestimate variance, would normally be nA vnA = vA / nA
vnB = vB / min(nA, nB) # overestimate variance, would normally be nB vnB = vB / nB
sum_vn = vnA + vnB sum_vn = vnA + vnB
# degrees of freedom for Welch's t-test # degrees of freedom for Welch's t-test
@@ -76,31 +59,18 @@ def diffexp_ttest(adata, maskA, maskB, top_n=8, diffexp_lfc_cutoff=0.01):
# p-value # p-value
pvals = stats.t.sf(np.abs(tscores), dof) * 2 pvals = stats.t.sf(np.abs(tscores), dof) * 2
pvals_adj = pvals * adata._X.shape[1] pvals_adj = pvals * adata._X.shape[1]
pvals_adj[pvals_adj > 1] = 1 # cap adjusted p-value at 1
# logfoldchanges: log2(meanA / meanB) # logfoldchanges: log2(meanA / meanB)
logfoldchanges = np.log2(np.abs((meanA + 1e-9) / (meanB + 1e-9))) logfoldchanges = np.log2(np.abs((meanA + 1e-9) / (meanB + 1e-9)))
# find all with lfc > cutoff # top n sort
lfc_above_cutoff_idx = np.nonzero(np.abs(logfoldchanges) > diffexp_lfc_cutoff)[0]
stats_to_sort = np.abs(tscores) stats_to_sort = np.abs(tscores)
partition = np.argpartition(stats_to_sort, -top_n)[-top_n:]
rel_sort_order = np.argsort(stats_to_sort[partition])[::-1]
vars_indices = np.arange(adata.n_vars, dtype=int)
sort_order = vars_indices[partition][rel_sort_order]
# derive sort order # top n slice
if lfc_above_cutoff_idx.shape[0] > top_n:
# partition top N
rel_t_partition = np.argpartition(stats_to_sort[lfc_above_cutoff_idx], -top_n)[-top_n:]
t_partition = lfc_above_cutoff_idx[rel_t_partition]
# sort the top N partition
rel_sort_order = np.argsort(stats_to_sort[t_partition])[::-1]
sort_order = t_partition[rel_sort_order]
else:
# partition and sort top N, ignoring lfc cutoff
partition = np.argpartition(stats_to_sort, -top_n)[-top_n:]
rel_sort_order = np.argsort(stats_to_sort[partition])[::-1]
indices = np.indices(stats_to_sort.shape)[0]
sort_order = indices[partition][rel_sort_order]
# top n slice based upon sort order
logfoldchanges_top_n = logfoldchanges[sort_order] logfoldchanges_top_n = logfoldchanges[sort_order]
pvals_top_n = pvals[sort_order] pvals_top_n = pvals[sort_order]
pvals_adj_top_n = pvals_adj[sort_order] pvals_adj_top_n = pvals_adj[sort_order]
+2 -2
View File
@@ -325,8 +325,8 @@ class ScanpyEngine(CXGDriver):
raise FilterError(f"Error parsing filter: {e}") from e raise FilterError(f"Error parsing filter: {e}") from e
if top_n is None: if top_n is None:
top_n = DEFAULT_TOP_N top_n = DEFAULT_TOP_N
result = diffexp_ttest(self.data, obs_mask_A, obs_mask_B, top_n, self.diffexp_lfc_cutoff) result = diffexp_ttest(self.data, obs_mask_A, obs_mask_B, top_n)
return result return sorted(result, key=lambda r: r[0])
def layout(self, filter, interactive_limit=None): def layout(self, filter, interactive_limit=None):
""" """
+1 -1
View File
@@ -5,7 +5,7 @@ from .prepare import prepare
@click.group(name="cellxgene", context_settings=dict(max_content_width=85)) @click.group(name="cellxgene", context_settings=dict(max_content_width=85))
@click.version_option(version="0.2.0", prog_name="cellxgene", message="[%(prog)s] Version %(version)s") @click.version_option(version="0.0.3", prog_name="cellxgene", message="[%(prog)s] Version %(version)s")
def cli(): def cli():
pass pass
+1 -9
View File
@@ -1,7 +1,6 @@
import sys import sys
import click import click
import logging import logging
from os import devnull
from os.path import splitext, basename from os.path import splitext, basename
import webbrowser import webbrowser
@@ -28,10 +27,8 @@ from server.app.util.errors import ScanpyFileError
help="Bind to all interfaces (this makes the server accessible beyond this computer).") help="Bind to all interfaces (this makes the server accessible beyond this computer).")
@click.option("--max-category-items", default=100, metavar="", show_default=True, @click.option("--max-category-items", default=100, metavar="", show_default=True,
help="Limits the number of categorical annotation items displayed.") help="Limits the number of categorical annotation items displayed.")
@click.option("--diffexp-lfc-cutoff", default=0.01, show_default=True,
help="Relative expression cutoff used when selecting top N differentially expressed genes")
def launch(data, layout, diffexp, title, verbose, debug, obs_names, var_names, def launch(data, layout, diffexp, title, verbose, debug, obs_names, var_names,
open_browser, port, listen_all, max_category_items, diffexp_lfc_cutoff): open_browser, port, listen_all, max_category_items):
"""Launch the cellxgene data viewer. """Launch the cellxgene data viewer.
This web app lets you explore single-cell expression data. This web app lets you explore single-cell expression data.
Data must be in a format that cellxgene expects, read the Data must be in a format that cellxgene expects, read the
@@ -95,7 +92,6 @@ def launch(data, layout, diffexp, title, verbose, debug, obs_names, var_names,
"layout": layout, "layout": layout,
"diffexp": diffexp, "diffexp": diffexp,
"max_category_items": max_category_items, "max_category_items": max_category_items,
"diffexp_lfc_cutoff": diffexp_lfc_cutoff,
"obs_names": obs_names, "obs_names": obs_names,
"var_names": var_names "var_names": var_names
} }
@@ -113,8 +109,4 @@ def launch(data, layout, diffexp, title, verbose, debug, obs_names, var_names,
click.echo("[cellxgene] Type CTRL-C at any time to exit.") click.echo("[cellxgene] Type CTRL-C at any time to exit.")
if not verbose:
f = open(devnull, 'w')
sys.stdout = f
app.run(host=host, debug=debug, port=port, threaded=True) app.run(host=host, debug=debug, port=port, threaded=True)
+3 -1
View File
@@ -14,7 +14,7 @@ from server.app.scanpy_engine.scanpy_engine import ScanpyEngine
class UtilTest(unittest.TestCase): class UtilTest(unittest.TestCase):
def setUp(self): def setUp(self):
args = {'layout': 'umap', 'diffexp': 'ttest', 'max_category_items': 100, args = {'layout': 'umap', 'diffexp': 'ttest', 'max_category_items': 100,
'obs_names': None, 'var_names': None, 'diffexp_lfc_cutoff': 0.01} 'obs_names': None, 'var_names': None}
self.data = ScanpyEngine("example-dataset/pbmc3k.h5ad", args) self.data = ScanpyEngine("example-dataset/pbmc3k.h5ad", args)
self.data._create_schema() self.data._create_schema()
@@ -201,6 +201,8 @@ class UtilTest(unittest.TestCase):
} }
result = self.data.diffexp_topN(f1["filter"], f2["filter"]) result = self.data.diffexp_topN(f1["filter"], f2["filter"])
self.assertEqual(len(result), 10) self.assertEqual(len(result), 10)
var_idx = [i[0] for i in result]
self.assertEqual(var_idx, sorted(var_idx))
result = self.data.diffexp_topN(f1["filter"], f2["filter"], 20) result = self.data.diffexp_topN(f1["filter"], f2["filter"], 20)
self.assertEqual(len(result), 20) self.assertEqual(len(result), 20)
+1 -1
View File
@@ -8,7 +8,7 @@ with open("server/requirements.txt") as fh:
setup( setup(
name="cellxgene", name="cellxgene",
version="0.2.0", version="0.0.3",
packages=find_packages(), packages=find_packages(),
url="https://github.com/chanzuckerberg/cellxgene", url="https://github.com/chanzuckerberg/cellxgene",
license="MIT", license="MIT",