mirror of
https://github.com/chhylp123/hifiasm.git
synced 2026-09-29 02:58:12 +08:00
update README
This commit is contained in:
+1
-1
@@ -3,7 +3,7 @@
|
||||
|
||||
#include <pthread.h>
|
||||
|
||||
#define HA_VERSION "0.3.0-dirty-r235"
|
||||
#define HA_VERSION "0.4.0"
|
||||
|
||||
#define VERBOSE 0
|
||||
|
||||
|
||||
@@ -21,53 +21,41 @@ outputs consist of:
|
||||
(*prefix*.r\_utg.gfa). This graph keeps all haplotype information, including
|
||||
somatic mutations and recurrent sequencing errors.
|
||||
2. Haplotype-resolved processed unitig graph without small bubbles
|
||||
(*prefix*.p\_utg.gfa). This is usually the preferred output for highly
|
||||
heterozygous genomes.
|
||||
3. Primary assembly [contig][unitig] graph (*prefix*.p\_ctg.gfa). This is the
|
||||
preferred output for inbred strains or human. For highly heterozygous
|
||||
genomes, this graph may represent multiple haplotypes. We plan to change
|
||||
this to represent one set of haplotypes.
|
||||
4. Alternate assembly contig graph (*prefix*.a\_ctg.gfa).
|
||||
5. Haplotype-aware error corrected reads in fasta format (*prefix*.ec.fa).
|
||||
6. All-to-all overlaps in the [PAF][paf] format (*prefix*.ovlp.paf).
|
||||
(*prefix*.p\_utg.gfa). Small bubbles might be caused by somatic mutations or noise in data,
|
||||
which are not the real haplotype information
|
||||
3. Primary assembly [contig][unitig] graph (*prefix*.p\_ctg.gfa). This graph collapses different
|
||||
haplotypes.
|
||||
4. Alternate assembly contig graph (*prefix*.a\_ctg.gfa). This graph consists of all assemblies that
|
||||
are discarded in primary contig graph.
|
||||
|
||||
For trio assembly, the input of hifiasm is the PacBio Hifi reads in fasta/fastq format, and the paternal/maternal trio indexes generated by `yak count` (see https://github.com/lh3/yak). The outputs consist of:
|
||||
1. Haplotype-resolved raw [unitig][unitig] graph in [GFA][gfa] format
|
||||
(*prefix*.r\_utg.gfa). This graph keeps all haplotype information.
|
||||
|
||||
2. Phased maternal unitig graph (*prefix*.m.r\_utg.gfa).
|
||||
This graph keeps the phased maternal assembly.
|
||||
2. Phased paternal/haplotype1 contig graph (*prefix*.hap1.p\_ctg.gfa). This graph keeps the phased
|
||||
paternal/haplotype1 assembly.
|
||||
|
||||
3. Phased maternal/haplotype2 contig graph (*prefix*.hap2.p\_ctg.gfa). This graph keeps the phased
|
||||
maternal/haplotype2 assembly.
|
||||
|
||||
3. Phased paternal unitig graph (*prefix*.p.r\_utg.gfa).
|
||||
This graph keeps the phased paternal assembly.
|
||||
|
||||
|
||||
In addition, hifiasm also outputs three binary files that save all overlap inforamtion
|
||||
(hifiasm.asm.ovlp, hifiasm.asm.ovlp.source, hifiasm.asm.ovlp.reverse in default). With these files, hifiasm can avoid the time-consuming all-to-all overlap calculation step, and do the assembly
|
||||
(*prefix*.ec.bin, *prefix*.ovlp.reverse.bin, *prefix*.ovlp.source.bin). With these files, hifiasm can avoid the time-consuming all-to-all overlap calculation step, and do the assembly
|
||||
directly and quickly. This might be helpful when you want to get an optimized
|
||||
assembly by multiple rounds of experiments with different parameters.
|
||||
|
||||
Hifiasm is a standalone and lightweight assembler, which does not need external
|
||||
libraries (except zlib). For large genomes, it can generate high-quality
|
||||
assembly in a few hours. Hifiasm has been tested on the following datasets:
|
||||
assembly in a few hours. Hifiasm has been tested on human, butterfly, rice and drosophila.
|
||||
In particular, hifiasm is able to assemble the 26.5Gb California redwood tree in a few days.
|
||||
The results are as follows:
|
||||
|
||||
|<sub>Dataset<sub>|<sub>GSize<sub>|<sub>Cov<sub>|<sub>Asm options<sub>|<sub>CPU time<sub>|<sub>Wall time<sub>|<sub>RAM<sub>|<sub>[unitig][unitig]/[contig][unitig] N50<sup>[1]</sup><sub>|
|
||||
|:---------------|-----:|-----:|:---------------------|-------:|--------:|----:|----------------:|
|
||||
|<sub>[Human NA12878]<sub>|<sub>3Gb<sub>|<sub>x28<sub>|<sub>-k 40 -t 42 -r 2<sub>|<sub>200h<sub>| <sub>5h32m<sub>|<sub>114G<sub>|<sub>93.5Kb/28.2Mb<sub>|
|
||||
|<sub>[Human HG002]<sub>|<sub>3Gb<sub>|<sub>x43<sub>|<sub>-k 40 -t 42 -r 2<sub>|<sub>405h10m<sub>|<sub>12h7m<sub>|<sub>146G<sub>|<sub>320kb/46.0Mb<sub>|
|
||||
|<sub>[Human CHM13]<sub>|<sub>3Gb<sub>|<sub>x27<sub>|<sub>-k 40 -t 42 -r 2<sub>|<sub>157h28m<sub>|<sub>5h10m<sub>|<sub>85.8G<sub>|<sub>NA<sup>[2]</sup>/41.4Mb<sub>|
|
||||
|<sub>[Butterfly]<sub>|<sub>358Mb<sub>|<sub>x35<sub>|<sub>-k 40 -t 42 -r 2 -z 20<sub>|<sub>17h6m<sub>|<sub>36m<sub>|<sub>16G<sub>|<sub>7.5Mb/NA<sup>[3]</sup><sub>|
|
||||
|<sub>[\[Redwood\]](https://downloads.pacbcloud.com/public/dataset/redwood2020/)<sub>|<sub>26.5Gb<sub>|<sub>x23<sub>|<sub>-k 40 -t 64 -r 2<sub>|<sub>7274h30m<sub>|<sub>141h30m<sub>|<sub>512G<sub>|<sub>1.7Mb/1.9Mb<sub>|
|
||||
|
||||
<sub>[1] unitig N50 is the N50 of assembly graph with haplotype information (i.e., bubbles), while the contig N50 is the N50 of haplotype collapsed assembly (i.e., without bubbles).
|
||||
[2] CHM13 is a homozygous sample, so that unitig N50 makes no sense.
|
||||
[3] Butterfly has high heterozygous rate, so that most chromosomes have been fully separated into two haplotypes. In this case, contig N50 makes no sense.<sub>
|
||||
|
||||
Note that different species need different assembly graphs. For homozygous genomes (i.e., Human CHM13), the primary assembly contig graph is the best choice.
|
||||
For species with high heterozygous rate (i.e., Butterfly), different haplotypes can be fully separated. It is important to remove small bubbles from the haplotype-resolved unitig graph. The
|
||||
reason is that some small bubbles are caused by somatic mutations or noise in data, which are not
|
||||
the real haplotype information. In this case, haplotype-resolved processed unitig graph
|
||||
without small bubbles should be better. For ordinary human genome (i.e., Human NA12878 and HG002), different haplotypes cannot be fully separated due to the low heterozygous rate. There are many small bubbles including haplotype information, which cannot be simply removed. Thus, it is necessary to use the haplotype-resolved raw unitig graph. **Hifiasm will generate a universal haplotype-resolved contig graph for all species in the near future.**
|
||||
<sub>[1] unitig N50 is the N50 of assembly graph with haplotype information (i.e., bubbles), while the contig N50 is the N50 of haplotype collapsed assembly (i.e., without bubbles).<sub>
|
||||
|
||||
## Usage
|
||||
|
||||
@@ -109,7 +97,7 @@ For trio assembly, first the trio indexes of paternal/maternal should be generat
|
||||
and then run hifiasm as follow:
|
||||
|
||||
```sh
|
||||
./hifiasm -o NA12878.asm -t 32 -P pat.yak -M mat.yak NA12878_1.fq.gz NA12878_2.fq.gz
|
||||
./hifiasm -o NA12878.asm -t 32 -1 pat.yak -2 mat.yak NA12878_1.fq.gz NA12878_2.fq.gz
|
||||
```
|
||||
|
||||
[unitig]: http://wgs-assembler.sourceforge.net/wiki/index.php/Celera_Assembler_Terminology
|
||||
@@ -128,9 +116,6 @@ have further questions, please raise an issue at the issue page.
|
||||
haplotype-resolved assembly graph, instead of the phased chromosome-level
|
||||
assembly (will support such output in future).
|
||||
|
||||
2. For different species, hifiasm outputs different assembly graphs, which are not easy to use.
|
||||
Hifiasm will generate a universal haplotype-resolved contig graph for all species in future.
|
||||
2. The running time and memory usage should be further reduced.
|
||||
|
||||
3. The running time and memory usage should be further reduced.
|
||||
|
||||
4. The N50 should be further improved.
|
||||
3. The N50 should be further improved.
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
.TH hifiasm 1 "22 Mar 2020" "hifiasm-0.3.0" "Bioinformatics tools"
|
||||
.TH hifiasm 1 "12 Apr 2020" "hifiasm-0.4.0" "Bioinformatics tools"
|
||||
|
||||
.SH NAME
|
||||
.PP
|
||||
@@ -307,16 +307,4 @@ the ease of visualization. Hifiasm keeps corrected reads and overlaps in three
|
||||
binary files such as it can regenerate assembly graphs from the binary files
|
||||
without redoing error correction.
|
||||
|
||||
.PP
|
||||
Note that different species need different assembly graphs. For homozygous genomes,
|
||||
the primary assembly contig graph is the best choice.
|
||||
For species with high heterozygous rate, different haplotypes can be fully separated.
|
||||
It is important to remove small bubbles from the haplotype-resolved unitig graph. The
|
||||
reason is that some small bubbles are caused by somatic mutations or noise in data,
|
||||
which are not the real haplotype information. In this case, haplotype-resolved processed
|
||||
unitig graph without small bubbles should be better.
|
||||
For ordinary human genome, different haplotypes cannot be fully separated due to the low
|
||||
heterozygous rate. There are many small bubbles including haplotype information,
|
||||
which cannot be simply removed. Thus, it is necessary to use the haplotype-resolved raw
|
||||
unitig graph.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user