mirror of
https://github.com/chanzuckerberg/cellxgene.git
synced 2026-10-06 20:58:12 +08:00
Merge pull request #157 from chanzuckerberg/csweaver/readme-data
Added instructions on how to load your own data in scanpy
This commit is contained in:
@@ -41,7 +41,8 @@ Started in the context of the Human Cell Atlas Consortium, cellxgene hopes to bo
|
|||||||
|
|
||||||
**Install server**
|
**Install server**
|
||||||
|
|
||||||
pip install -e .
|
|
||||||
|
pip install -e .
|
||||||
|
|
||||||
**Run (with demo data)**
|
**Run (with demo data)**
|
||||||
|
|
||||||
@@ -56,6 +57,72 @@ _For help with the scanpy engine_
|
|||||||
|
|
||||||
cellxgene scanpy --help
|
cellxgene scanpy --help
|
||||||
|
|
||||||
|
## Using your own data
|
||||||
|
|
||||||
|
### Scanpy
|
||||||
|
|
||||||
|
To prepare your data you will need to format your data into AnnData format using scanpy and calculate PCA and nearest neighbors and save in h5ad format.
|
||||||
|
|
||||||
|
1. [Load data into scanpy](https://scanpy.readthedocs.io/en/latest/api/index.html#reading)
|
||||||
|
|
||||||
|
- Ensure that `obs`'s index is the cell names: `print(data.obs_names)` should show your cell indices. If it shows gene names, you may need to just call `data.transpose()`.
|
||||||
|
|
||||||
|
2. Calculate PCA
|
||||||
|
|
||||||
|
sc.pp.pca(data) ## sc is scanpy.api
|
||||||
|
|
||||||
|
3. Calculate nearest neighbors (depending on layout algorithm)
|
||||||
|
|
||||||
|
```
|
||||||
|
# For umap layout algorithm, you need to use the "umap" method for neighbors
|
||||||
|
sc.pp.neighbors(data, method="umap", metric="euclidean", use_rep="X_pca")
|
||||||
|
|
||||||
|
# For tsne layout algorithm, you can use either "umap" or "gauss"; we recommend "gauss"
|
||||||
|
sc.pp.neighbors(data, method="gauss", metric="euclidean", use_rep="X_pca")
|
||||||
|
```
|
||||||
|
|
||||||
|
4. Save file
|
||||||
|
|
||||||
|
```
|
||||||
|
# cellxgene requires file to be named data.h5ad
|
||||||
|
data.write("data.h5ad")
|
||||||
|
```
|
||||||
|
|
||||||
|
5. Create config file (optional)
|
||||||
|
|
||||||
|
If you do not have a config file, the schema (metadata names, types, and categorical/continuous) will be inferred from the observations in the data file. Config file is required to be named 'data_schema.json' and located in the same directory as data file.
|
||||||
|
- The config file is a JSON format file with information on the metadata associated with the cells. The key is the column name in obs. The value is an object
|
||||||
|
```
|
||||||
|
type: string, int, or float (what type the values are),
|
||||||
|
variabletype: categorical or continuous (categorical values are displayed as checkboxes, continuous values are displayed as a histogram)
|
||||||
|
displayname: (what the heading should be displayed as)
|
||||||
|
include: True/False (whether to display values on web interface)
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
Example
|
||||||
|
{
|
||||||
|
"CellName": {
|
||||||
|
"type": "string",
|
||||||
|
"variabletype": "categorical",
|
||||||
|
"displayname": "Name",
|
||||||
|
"include": true
|
||||||
|
},
|
||||||
|
"clusters": {
|
||||||
|
"type": "string",
|
||||||
|
"variabletype": "categorical",
|
||||||
|
"displayname": "Clusters",
|
||||||
|
"include": true
|
||||||
|
},
|
||||||
|
"num_genes": {
|
||||||
|
"type": "int",
|
||||||
|
"variabletype": "continuous",
|
||||||
|
"displayname": "Number Genes",
|
||||||
|
"include": true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
## Contributing
|
## Contributing
|
||||||
We warmly welcome contributions from the community. Please submit any bug reports and feature requests through github issues. Please submit any direct contributions via a branch + pull request.
|
We warmly welcome contributions from the community. Please submit any bug reports and feature requests through github issues. Please submit any direct contributions via a branch + pull request.
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user