mirror of
https://github.com/chanzuckerberg/cellxgene.git
synced 2026-09-22 22:38:11 +08:00
Merge pull request #157 from chanzuckerberg/csweaver/readme-data
Added instructions on how to load your own data in scanpy
This commit is contained in:
@@ -21,41 +21,108 @@ Started in the context of the Human Cell Atlas Consortium, cellxgene hopes to bo
|
||||
- OS: OSX, Windows, Linux
|
||||
- python 3.6
|
||||
- npm
|
||||
- Google Chrome
|
||||
- Google Chrome
|
||||
|
||||
**Clone project**
|
||||
|
||||
git clone https://github.com/chanzuckerberg/cellxgene.git
|
||||
**Clone project**
|
||||
|
||||
git clone https://github.com/chanzuckerberg/cellxgene.git
|
||||
|
||||
**Install client**
|
||||
|
||||
**Install client**
|
||||
|
||||
cd cellxgene
|
||||
./bin/build-client
|
||||
./bin/build-client
|
||||
|
||||
**To use with virtual env for python**
|
||||
(optional, but recommended)
|
||||
|
||||
ENV_NAME=cellxgene
|
||||
python3 -m venv ${ENV_NAME}
|
||||
source ${ENV_NAME}/bin/activate
|
||||
**To use with virtual env for python**
|
||||
(optional, but recommended)
|
||||
|
||||
**Install server**
|
||||
|
||||
pip install -e .
|
||||
ENV_NAME=cellxgene
|
||||
python3 -m venv ${ENV_NAME}
|
||||
source ${ENV_NAME}/bin/activate
|
||||
|
||||
**Install server**
|
||||
|
||||
|
||||
pip install -e .
|
||||
|
||||
**Run (with demo data)**
|
||||
|
||||
**Run (with demo data)**
|
||||
|
||||
cellxgene --title PBMC3K scanpy example-dataset/
|
||||
*In google chrome, navigate to the viewer via the web address printed in your console.
|
||||
*In google chrome, navigate to the viewer via the web address printed in your console.
|
||||
E.g.,* `Running on http://0.0.0.0:5005/`
|
||||
|
||||
**Help**
|
||||
|
||||
|
||||
cellxgene --help
|
||||
_For help with the scanpy engine_
|
||||
|
||||
_For help with the scanpy engine_
|
||||
|
||||
cellxgene scanpy --help
|
||||
|
||||
## Using your own data
|
||||
|
||||
### Scanpy
|
||||
|
||||
To prepare your data you will need to format your data into AnnData format using scanpy and calculate PCA and nearest neighbors and save in h5ad format.
|
||||
|
||||
1. [Load data into scanpy](https://scanpy.readthedocs.io/en/latest/api/index.html#reading)
|
||||
|
||||
- Ensure that `obs`'s index is the cell names: `print(data.obs_names)` should show your cell indices. If it shows gene names, you may need to just call `data.transpose()`.
|
||||
|
||||
2. Calculate PCA
|
||||
|
||||
sc.pp.pca(data) ## sc is scanpy.api
|
||||
|
||||
3. Calculate nearest neighbors (depending on layout algorithm)
|
||||
|
||||
```
|
||||
# For umap layout algorithm, you need to use the "umap" method for neighbors
|
||||
sc.pp.neighbors(data, method="umap", metric="euclidean", use_rep="X_pca")
|
||||
|
||||
# For tsne layout algorithm, you can use either "umap" or "gauss"; we recommend "gauss"
|
||||
sc.pp.neighbors(data, method="gauss", metric="euclidean", use_rep="X_pca")
|
||||
```
|
||||
|
||||
4. Save file
|
||||
|
||||
```
|
||||
# cellxgene requires file to be named data.h5ad
|
||||
data.write("data.h5ad")
|
||||
```
|
||||
|
||||
5. Create config file (optional)
|
||||
|
||||
If you do not have a config file, the schema (metadata names, types, and categorical/continuous) will be inferred from the observations in the data file. Config file is required to be named 'data_schema.json' and located in the same directory as data file.
|
||||
- The config file is a JSON format file with information on the metadata associated with the cells. The key is the column name in obs. The value is an object
|
||||
```
|
||||
type: string, int, or float (what type the values are),
|
||||
variabletype: categorical or continuous (categorical values are displayed as checkboxes, continuous values are displayed as a histogram)
|
||||
displayname: (what the heading should be displayed as)
|
||||
include: True/False (whether to display values on web interface)
|
||||
```
|
||||
|
||||
```
|
||||
Example
|
||||
{
|
||||
"CellName": {
|
||||
"type": "string",
|
||||
"variabletype": "categorical",
|
||||
"displayname": "Name",
|
||||
"include": true
|
||||
},
|
||||
"clusters": {
|
||||
"type": "string",
|
||||
"variabletype": "categorical",
|
||||
"displayname": "Clusters",
|
||||
"include": true
|
||||
},
|
||||
"num_genes": {
|
||||
"type": "int",
|
||||
"variabletype": "continuous",
|
||||
"displayname": "Number Genes",
|
||||
"include": true
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Contributing
|
||||
We warmly welcome contributions from the community. Please submit any bug reports and feature requests through github issues. Please submit any direct contributions via a branch + pull request.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user