Merge pull request #157 from chanzuckerberg/csweaver/readme-data

Added instructions on how to load your own data in scanpy
This commit is contained in:
Charlotte Weaver
2018-08-07 10:19:31 -07:00
committed by GitHub
+89 -22
View File
@@ -21,41 +21,108 @@ Started in the context of the Human Cell Atlas Consortium, cellxgene hopes to bo
- OS: OSX, Windows, Linux
- python 3.6
- npm
- Google Chrome
- Google Chrome
**Clone project**
git clone https://github.com/chanzuckerberg/cellxgene.git
**Clone project**
git clone https://github.com/chanzuckerberg/cellxgene.git
**Install client**
**Install client**
cd cellxgene
./bin/build-client
./bin/build-client
**To use with virtual env for python**
(optional, but recommended)
ENV_NAME=cellxgene
python3 -m venv ${ENV_NAME}
source ${ENV_NAME}/bin/activate
**To use with virtual env for python**
(optional, but recommended)
**Install server**
pip install -e .
ENV_NAME=cellxgene
python3 -m venv ${ENV_NAME}
source ${ENV_NAME}/bin/activate
**Install server**
pip install -e .
**Run (with demo data)**
**Run (with demo data)**
cellxgene --title PBMC3K scanpy example-dataset/
*In google chrome, navigate to the viewer via the web address printed in your console.
*In google chrome, navigate to the viewer via the web address printed in your console.
E.g.,* `Running on http://0.0.0.0:5005/`
**Help**
cellxgene --help
_For help with the scanpy engine_
_For help with the scanpy engine_
cellxgene scanpy --help
## Using your own data
### Scanpy
To prepare your data you will need to format your data into AnnData format using scanpy and calculate PCA and nearest neighbors and save in h5ad format.
1. [Load data into scanpy](https://scanpy.readthedocs.io/en/latest/api/index.html#reading)
- Ensure that `obs`'s index is the cell names: `print(data.obs_names)` should show your cell indices. If it shows gene names, you may need to just call `data.transpose()`.
2. Calculate PCA
sc.pp.pca(data) ## sc is scanpy.api
3. Calculate nearest neighbors (depending on layout algorithm)
```
# For umap layout algorithm, you need to use the "umap" method for neighbors
sc.pp.neighbors(data, method="umap", metric="euclidean", use_rep="X_pca")
# For tsne layout algorithm, you can use either "umap" or "gauss"; we recommend "gauss"
sc.pp.neighbors(data, method="gauss", metric="euclidean", use_rep="X_pca")
```
4. Save file
```
# cellxgene requires file to be named data.h5ad
data.write("data.h5ad")
```
5. Create config file (optional)
If you do not have a config file, the schema (metadata names, types, and categorical/continuous) will be inferred from the observations in the data file. Config file is required to be named 'data_schema.json' and located in the same directory as data file.
- The config file is a JSON format file with information on the metadata associated with the cells. The key is the column name in obs. The value is an object
```
type: string, int, or float (what type the values are),
variabletype: categorical or continuous (categorical values are displayed as checkboxes, continuous values are displayed as a histogram)
displayname: (what the heading should be displayed as)
include: True/False (whether to display values on web interface)
```
```
Example
{
"CellName": {
"type": "string",
"variabletype": "categorical",
"displayname": "Name",
"include": true
},
"clusters": {
"type": "string",
"variabletype": "categorical",
"displayname": "Clusters",
"include": true
},
"num_genes": {
"type": "int",
"variabletype": "continuous",
"displayname": "Number Genes",
"include": true
}
}
```
## Contributing
We warmly welcome contributions from the community. Please submit any bug reports and feature requests through github issues. Please submit any direct contributions via a branch + pull request.