mirror of
https://github.com/lh3/minimap2.git
synced 2026-09-24 19:38:12 +08:00
Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
a05edfa5ec | ||
|
|
8e81145817 | ||
|
|
e37f5ffe39 | ||
|
|
8a1d52bcbe | ||
|
|
f7271a7c24 | ||
|
|
70393eb46e | ||
|
|
9d049f0562 | ||
|
|
5180b70ff3 | ||
|
|
2392e54fe2 | ||
|
|
629c11728e | ||
|
|
7e33fde82b | ||
|
|
c4fe52fb07 | ||
|
|
ead1cfbaca | ||
|
|
83a535f148 | ||
|
|
f3af29a8aa | ||
|
|
cf7eaef367 | ||
|
|
2411887d8e | ||
|
|
1a8373bb84 | ||
|
|
161ae7ff73 | ||
|
|
8a6edab847 | ||
|
|
15118dd521 | ||
|
|
2546999639 | ||
|
|
52fafe0fed | ||
|
|
5f449c5cae | ||
|
|
b046052d82 | ||
|
|
581f2d7123 | ||
|
|
52dbd439bc | ||
|
|
260a68d232 | ||
|
|
177eef259d | ||
|
|
459ce04c84 | ||
|
|
e6cce019e4 | ||
|
|
7025b0b941 | ||
|
|
fe6a0bb337 | ||
|
|
3f7147864b | ||
|
|
c83589b9ea | ||
|
|
ce7a59f412 | ||
|
|
15471bd629 | ||
|
|
ca19463268 | ||
|
|
4f8d1bc360 | ||
|
|
1776c0c645 | ||
|
|
9febf532c1 | ||
|
|
cec23131e4 | ||
|
|
ef09ccf104 | ||
|
|
e74dfd1aa9 | ||
|
|
f31705bb4a | ||
|
|
41d7ccb191 | ||
|
|
34a41197d7 | ||
|
|
9626b3e716 | ||
|
|
379728726a | ||
|
|
4f91558160 | ||
|
|
ec3bc6efd7 | ||
|
|
8ec8866100 | ||
|
|
9e7247cff9 | ||
|
|
cd66777bfb | ||
|
|
5d7d25e92d | ||
|
|
2a3793bbd2 | ||
|
|
f97008a10e | ||
|
|
10502e2a78 | ||
|
|
4422c0c6f9 | ||
|
|
42a11e1d58 | ||
|
|
76df351fa8 | ||
|
|
d065d3bead | ||
|
|
ac146fe7bc | ||
|
|
6c96078ed0 | ||
|
|
bbb4f97e52 | ||
|
|
b7f4d8a0f4 | ||
|
|
e81927e7a1 | ||
|
|
f7dc5799c5 | ||
|
|
817cb81cb0 | ||
|
|
0f5608c4a4 | ||
|
|
e8823a3709 | ||
|
|
7edeec67b0 | ||
|
|
feb92d32ea | ||
|
|
cdbd96be0c | ||
|
|
ba52c79024 | ||
|
|
cd9ccfa069 | ||
|
|
86b716448c | ||
|
|
9ab95be1bb | ||
|
|
b51e859945 | ||
|
|
a9037dc16c | ||
|
|
9729fa99ad | ||
|
|
b6ff332de1 | ||
|
|
77abafaaf3 | ||
|
|
507d39af15 | ||
|
|
827ca4b461 | ||
|
|
d3dde2fdd4 | ||
|
|
7db2e8d21a | ||
|
|
0b41dd26a2 | ||
|
|
2b47846cd6 | ||
|
|
67dd906a80 | ||
|
|
1b0bb7b0ba | ||
|
|
1c4b7e8a48 | ||
|
|
ecbc399fa2 | ||
|
|
4dfd495cc2 | ||
|
|
194b457e79 | ||
|
|
75c8933511 | ||
|
|
1025993469 | ||
|
|
a3253d1a6b | ||
|
|
2da649d1d7 | ||
|
|
f995f55610 | ||
|
|
c9874e2dc5 | ||
|
|
28a37a017a | ||
|
|
ccb0f7b05d | ||
|
|
66db9da7d8 | ||
|
|
cd2b19035b | ||
|
|
9c0e2c67f8 | ||
|
|
da7109fd29 | ||
|
|
2b3403f094 | ||
|
|
3e16e4e39d | ||
|
|
f47e8a525e | ||
|
|
c172df7d2d | ||
|
|
9e6fdd376b | ||
|
|
29f67a1666 | ||
|
|
adde608a42 | ||
|
|
f10dff78dc | ||
|
|
d97bba9f27 | ||
|
|
50775362bb | ||
|
|
0a5e386359 | ||
|
|
cb56fb762a | ||
|
|
e2451e497a | ||
|
|
d2de282d21 | ||
|
|
48cb80ea94 | ||
|
|
6a4b9f9082 | ||
|
|
a7a01fe5bd | ||
|
|
9dceae59a0 | ||
|
|
20a3987082 | ||
|
|
eb3ed6993d | ||
|
|
7996f04008 | ||
|
|
d2e14705e7 | ||
|
|
24f50f38e8 | ||
|
|
04e015d803 | ||
|
|
040f74102c | ||
|
|
cdb7857841 | ||
|
|
3c0d05d272 | ||
|
|
47b646acbf | ||
|
|
a79cb3e991 | ||
|
|
367aed4271 | ||
|
|
081df6ac7d | ||
|
|
a3e7a575fb | ||
|
|
d90583b83c | ||
|
|
7fc03b0c32 | ||
|
|
20c104ce8d | ||
|
|
238b6bb3ea | ||
|
|
e026e18439 | ||
|
|
58c2251b18 | ||
|
|
03dc8d5d97 | ||
|
|
5cb61f8ee6 | ||
|
|
c16a1742a3 | ||
|
|
4bd5a018c2 | ||
|
|
05974c80f1 | ||
|
|
7bc87b4175 | ||
|
|
6762368cf0 | ||
|
|
c2aec88b84 | ||
|
|
97f67a2a0a | ||
|
|
189555503a | ||
|
|
69af86657e | ||
|
|
49c6d83a8e | ||
|
|
f64e426a5a | ||
|
|
2bb8cbbeef | ||
|
|
e80759c97a | ||
|
|
f4c844b143 | ||
|
|
be171aa2dc | ||
|
|
cdc730d573 | ||
|
|
6420acca6d | ||
|
|
371bc9513a | ||
|
|
169216bfff | ||
|
|
6b391e3373 | ||
|
|
55e39c2d30 | ||
|
|
90b7b83ec7 | ||
|
|
d431dc0181 | ||
|
|
ccf1680aaf | ||
|
|
ea84fc0a53 | ||
|
|
19208fb06b | ||
|
|
e02bebd96d | ||
|
|
32ab6ce15b | ||
|
|
1739a260fb | ||
|
|
aaf3233818 | ||
|
|
8b05880f73 | ||
|
|
eba237f39d | ||
|
|
a8e1e3cbb8 | ||
|
|
597212b9f3 | ||
|
|
30abcf3cf9 | ||
|
|
48e230f40d | ||
|
|
c404f49569 | ||
|
|
cf2bae6e9b | ||
|
|
5b2fdfff9c | ||
|
|
ea2b1c5b2a | ||
|
|
eef1cee9b7 | ||
|
|
2c52364527 | ||
|
|
128476efc9 | ||
|
|
1b3a6a0fe5 | ||
|
|
83a8ee7038 | ||
|
|
62bbadf668 | ||
|
|
91f548b497 | ||
|
|
cdaf46665a | ||
|
|
6596c63dcd | ||
|
|
59f23f7579 | ||
|
|
5e55e397e9 | ||
|
|
88c421e8de | ||
|
|
3db5bfe6e5 | ||
|
|
83dfdd5f50 | ||
|
|
8a2b1cd4c9 | ||
|
|
1ede8ca170 | ||
|
|
13981404e2 | ||
|
|
fd64dd26f6 | ||
|
|
24df95e4b8 | ||
|
|
a8ee48c2ce | ||
|
|
09e089c3dc | ||
|
|
e46cbb7d84 | ||
|
|
57ec73ec6c | ||
|
|
9e27575387 | ||
|
|
b4ad8d8bf0 | ||
|
|
e315b9fada | ||
|
|
42baf287a4 | ||
|
|
2ceba22a7a | ||
|
|
9ed56b4a25 | ||
|
|
ecb6c5c36c | ||
|
|
377c7099a8 | ||
|
|
51e2abfa60 | ||
|
|
7b0a49732e | ||
|
|
20268a6068 | ||
|
|
d04ac068fd | ||
|
|
5d5d392c02 | ||
|
|
170863e553 | ||
|
|
97f97306a4 | ||
|
|
1077b7ddc8 | ||
|
|
c57b59f02f | ||
|
|
34be359e25 | ||
|
|
8b12da8b0f | ||
|
|
c63a33904f | ||
|
|
70b0fede64 | ||
|
|
0b681e51e7 | ||
|
|
7d80d6de4a | ||
|
|
791e89ce0f | ||
|
|
98c48a1c45 | ||
|
|
63d397120a | ||
|
|
7998fe9906 | ||
|
|
3a119d606f | ||
|
|
a5eafb75f9 | ||
|
|
9a567e4b37 | ||
|
|
8e606bcc06 | ||
|
|
a1b7219b5d | ||
|
|
b0f39a1a61 | ||
|
|
5ab6538757 | ||
|
|
b32296e18f | ||
|
|
99ecdf7b5d | ||
|
|
ff9917a1c4 | ||
|
|
8c064a5f29 | ||
|
|
0e137670fc | ||
|
|
395c8d678a | ||
|
|
830da7fa27 | ||
|
|
a655cbef86 | ||
|
|
4b707aac92 | ||
|
|
951c0d1d35 | ||
|
|
3545e35a42 | ||
|
|
f3417da838 | ||
|
|
e5277dbf5c | ||
|
|
1a55227d5a | ||
|
|
5cfa621b2d | ||
|
|
a609a07f8c | ||
|
|
bcf92b3c46 | ||
|
|
097378ab90 | ||
|
|
10bbbe28c5 | ||
|
|
c92a6866f3 | ||
|
|
6908dc59a5 | ||
|
|
50dae10421 | ||
|
|
0517972d02 | ||
|
|
d46e68e6ad | ||
|
|
2584a4149a | ||
|
|
66674afd09 | ||
|
|
e9ca0c9dab | ||
|
|
7e6e8ca73f | ||
|
|
408e098859 | ||
|
|
57f37551f8 | ||
|
|
4c66b689c3 | ||
|
|
1bde2cf076 | ||
|
|
31fc0f218a | ||
|
|
3d3bcc29a8 | ||
|
|
99dcd75f64 | ||
|
|
154d2caf5b | ||
|
|
a3afeec0b2 | ||
|
|
3573784b4d | ||
|
|
872f300955 | ||
|
|
d7b61a039e | ||
|
|
248158a3e1 | ||
|
|
9f4309c376 | ||
|
|
463f9309f9 | ||
|
|
abe989e355 | ||
|
|
881b4ca3a2 | ||
|
|
10c6dd2551 | ||
|
|
7ec6721c44 | ||
|
|
e61812ee55 | ||
|
|
734ac379bb | ||
|
|
759f8e4ac9 | ||
|
|
aef7b0744c | ||
|
|
39f836eac8 | ||
|
|
cbeb86dad6 | ||
|
|
372c90ceb5 | ||
|
|
2e2e69107c | ||
|
|
ee4cd089f7 | ||
|
|
2d7ec75d50 | ||
|
|
7938ed4893 | ||
|
|
4740423afa | ||
|
|
1776311a9b | ||
|
|
c1a3e05cb0 | ||
|
|
ecb6703d4a | ||
|
|
0bf97a367b | ||
|
|
1c504d72e4 | ||
|
|
5ef9580b17 | ||
|
|
08bd2123b6 | ||
|
|
8766d286df | ||
|
|
623b5d9d48 | ||
|
|
18659118cd | ||
|
|
d1050f4eaf | ||
|
|
b81d45510e | ||
|
|
d135feb1a5 | ||
|
|
242ff4e91d | ||
|
|
7a0c1316ce | ||
|
|
77ebd479f4 | ||
|
|
e3f226a9d9 | ||
|
|
bdc615c1d4 | ||
|
|
ad1beaf255 | ||
|
|
de0480ac5b | ||
|
|
f2866533a8 | ||
|
|
0173850ef0 | ||
|
|
acea3594fb | ||
|
|
f78a247749 | ||
|
|
ccaf12e1a2 | ||
|
|
96b132c97d | ||
|
|
70428ca3a8 | ||
|
|
9aea79d621 | ||
|
|
1770988627 | ||
|
|
2bfdad34bb | ||
|
|
953766cedd | ||
|
|
dc61301d9f | ||
|
|
19e05a099d | ||
|
|
0238caa8b1 | ||
|
|
a22ebb9836 | ||
|
|
eeb314edd6 | ||
|
|
83c57a9d98 | ||
|
|
24a4808826 | ||
|
|
29ed675ee5 | ||
|
|
7dc7097208 | ||
|
|
f434653432 | ||
|
|
090361c25b | ||
|
|
e1f18690f6 | ||
|
|
54a42aafe5 | ||
|
|
8fc5f8dc90 | ||
|
|
a0d62519c1 | ||
|
|
b71c01b316 | ||
|
|
1372977a37 | ||
|
|
c0e0d5d84b | ||
|
|
b328795051 | ||
|
|
3b17d62ccd | ||
|
|
874b8c4795 | ||
|
|
7ef5490884 | ||
|
|
f66de7df59 | ||
|
|
fbbd4e0968 | ||
|
|
8c89ba005e | ||
|
|
50775a1e6f | ||
|
|
87cf650168 | ||
|
|
6e65c5e631 | ||
|
|
a58b05a61b | ||
|
|
42dab6319b | ||
|
|
8428809369 | ||
|
|
642af5591e | ||
|
|
2025a6279a | ||
|
|
3465b04724 | ||
|
|
66c5e71fa8 | ||
|
|
01560f1db0 | ||
|
|
39535565ee | ||
|
|
a8d476c6ad | ||
|
|
29b4a1786c | ||
|
|
dbf284b2d9 | ||
|
|
3df5015668 | ||
|
|
756379bf83 | ||
|
|
41fd8a966a | ||
|
|
86e4933b1a | ||
|
|
a633a744b6 | ||
|
|
ddc31f57ba | ||
|
|
35d3e064bf | ||
|
|
997ab9bb2e |
@@ -0,0 +1,21 @@
|
||||
name: CI
|
||||
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- master
|
||||
pull_request:
|
||||
|
||||
jobs:
|
||||
build:
|
||||
runs-on: ubuntu-latest
|
||||
strategy:
|
||||
matrix:
|
||||
compiler: [gcc, clang]
|
||||
|
||||
steps:
|
||||
- name: Checkout minimap2
|
||||
uses: actions/checkout@v2
|
||||
|
||||
- name: Compile with ${{ matrix.compiler }}
|
||||
run: make CC=${{ matrix.compiler }}
|
||||
@@ -0,0 +1,3 @@
|
||||
[submodule "lib/simde"]
|
||||
path = lib/simde
|
||||
url = https://github.com/nemequ/simde.git
|
||||
+5
-1
@@ -6,6 +6,10 @@ matrix:
|
||||
- language: c
|
||||
compiler: clang
|
||||
script: make
|
||||
- arch: arm64
|
||||
language: c
|
||||
compiler: gcc
|
||||
script: make arm_neon=1 aarch64=1
|
||||
- language: python
|
||||
python: "2.7"
|
||||
before_install: pip install cython
|
||||
@@ -15,6 +19,6 @@ matrix:
|
||||
before_install: pip install cython
|
||||
script: python setup.py build_ext
|
||||
- language: python
|
||||
python: "3.6"
|
||||
python: "3.9"
|
||||
before_install: pip install cython
|
||||
script: python setup.py build_ext
|
||||
|
||||
@@ -0,0 +1,46 @@
|
||||
#### 1. Alignment different with option `-a` or `-c`?
|
||||
|
||||
Without `-a`, `-c` or `--cs`, minimap2 only finds *approximate* mapping
|
||||
locations without detailed base alignment. In particular, the start and end
|
||||
positions of the alignment are impricise. With one of those options, minimap2
|
||||
will perform base alignment, which is generally more accurate but is much
|
||||
slower.
|
||||
|
||||
#### 2. How to map Illumina short reads to noisy long reads?
|
||||
|
||||
No good solutions. The better approach is to assemble short reads into contigs
|
||||
and then map noisy reads to contigs.
|
||||
|
||||
#### 3. The output SAM doesn't have a header.
|
||||
|
||||
By default, minimap2 indexes 4 billion reference bases (4Gb) in a batch and map
|
||||
all reads against each reference batch. Given a reference longer than 4Gb,
|
||||
minimap2 is unable to see all the sequences and thus can't produce a correct
|
||||
SAM header. In this case, minimap2 doesn't output any SAM header. There are two
|
||||
solutions to this issue. First, you may increase option `-I` to, for example,
|
||||
`-I8g` to index more reference bases in a batch. This is preferred if your
|
||||
machine has enough memory. Second, if your machines doesn't have enough memory
|
||||
to hold the reference index, you can use the `--split-prefix` option in a
|
||||
command line like:
|
||||
```sh
|
||||
minimap2 -ax map-ont --split-prefix=tmp ref.fa reads.fq
|
||||
```
|
||||
This second approach uses less memory, but it is slower and requires temporary
|
||||
disk space.
|
||||
|
||||
#### 4. The output SAM is malformatted.
|
||||
|
||||
This typically happens when you use nohup to wrap a minimap2 command line.
|
||||
Nohup is discouraged as it breaks piping. If you have to use nohup, please
|
||||
specify an output file with option `-o`.
|
||||
|
||||
#### 5. How to output one alignment per read?
|
||||
|
||||
You can use `--secondary=no` to suppress secondary alignments (aka multiple
|
||||
mappings), but you can't suppress supplementary alignment (aka split or
|
||||
chimeric alignment) this way. You can use samtools to filter out these
|
||||
alignments:
|
||||
```sh
|
||||
minimap2 -ax map-out ref.fa reads.fq | samtools view -F0x900
|
||||
```
|
||||
However, this is discouraged as supplementary alignment is informative.
|
||||
+2
-1
@@ -1,6 +1,7 @@
|
||||
The MIT License
|
||||
|
||||
Copyright (c) 2017 Broad Institute, Inc.
|
||||
Copyright (c) 2018- Dana-Farber Cancer Institute
|
||||
2017-2018 Broad Institute, Inc.
|
||||
|
||||
Permission is hereby granted, free of charge, to any person obtaining
|
||||
a copy of this software and associated documentation files (the
|
||||
|
||||
+1
-2
@@ -1,10 +1,9 @@
|
||||
include *.h
|
||||
include Makefile
|
||||
include ksw2_dispatch.c
|
||||
include getopt.c
|
||||
include main.c
|
||||
include README.md
|
||||
include python/mappy.c
|
||||
include sse2neon/emmintrin.h
|
||||
include python/cmappy.h
|
||||
include python/cmappy.pxd
|
||||
include python/mappy.pyx
|
||||
|
||||
@@ -1,21 +1,37 @@
|
||||
CFLAGS= -g -Wall -O2 -Wc++-compat
|
||||
CFLAGS= -g -Wall -O2 -Wc++-compat #-Wextra
|
||||
CPPFLAGS= -DHAVE_KALLOC
|
||||
INCLUDES=
|
||||
OBJS= kthread.o kalloc.o misc.o bseq.o sketch.o sdust.o options.o index.o chain.o align.o hit.o map.o format.o pe.o esterr.o ksw2_ll_sse.o
|
||||
OBJS= kthread.o kalloc.o misc.o bseq.o sketch.o sdust.o options.o index.o \
|
||||
lchain.o align.o hit.o seed.o map.o format.o pe.o esterr.o splitidx.o \
|
||||
ksw2_ll_sse.o
|
||||
PROG= minimap2
|
||||
PROG_EXTRA= sdust minimap2-lite
|
||||
LIBS= -lm -lz -lpthread
|
||||
|
||||
ifeq ($(arm_neon),)
|
||||
ifeq ($(sse2only),)
|
||||
ifeq ($(arm_neon),) # if arm_neon is not defined
|
||||
ifeq ($(sse2only),) # if sse2only is not defined
|
||||
OBJS+=ksw2_extz2_sse41.o ksw2_extd2_sse41.o ksw2_exts2_sse41.o ksw2_extz2_sse2.o ksw2_extd2_sse2.o ksw2_exts2_sse2.o ksw2_dispatch.o
|
||||
else
|
||||
else # if sse2only is defined
|
||||
OBJS+=ksw2_extz2_sse.o ksw2_extd2_sse.o ksw2_exts2_sse.o
|
||||
endif
|
||||
else
|
||||
OBJS+=ksw2_extz2_neon.o ksw2_extd2_neon.o ksw2_exts2_neon.o
|
||||
CFLAGS+=-D_FILE_OFFSET_BITS=64 -mfpu=neon -fsigned-char
|
||||
INCLUDES+=-I sse2neon
|
||||
else # if arm_neon is defined
|
||||
OBJS+=ksw2_extz2_neon.o ksw2_extd2_neon.o ksw2_exts2_neon.o
|
||||
INCLUDES+=-Isse2neon
|
||||
ifeq ($(aarch64),) #if aarch64 is not defined
|
||||
CFLAGS+=-D_FILE_OFFSET_BITS=64 -mfpu=neon -fsigned-char
|
||||
else #if aarch64 is defined
|
||||
CFLAGS+=-D_FILE_OFFSET_BITS=64 -fsigned-char
|
||||
endif
|
||||
endif
|
||||
|
||||
ifneq ($(asan),)
|
||||
CFLAGS+=-fsanitize=address
|
||||
LIBS+=-fsanitize=address
|
||||
endif
|
||||
|
||||
ifneq ($(tsan),)
|
||||
CFLAGS+=-fsanitize=thread
|
||||
LIBS+=-fsanitize=thread
|
||||
endif
|
||||
|
||||
.PHONY:all extra clean depend
|
||||
@@ -28,8 +44,8 @@ all:$(PROG)
|
||||
|
||||
extra:all $(PROG_EXTRA)
|
||||
|
||||
minimap2:main.o getopt.o libminimap2.a
|
||||
$(CC) $(CFLAGS) main.o getopt.o -o $@ -L. -lminimap2 $(LIBS)
|
||||
minimap2:main.o libminimap2.a
|
||||
$(CC) $(CFLAGS) main.o -o $@ -L. -lminimap2 $(LIBS)
|
||||
|
||||
minimap2-lite:example.o libminimap2.a
|
||||
$(CC) $(CFLAGS) $< -o $@ -L. -lminimap2 $(LIBS)
|
||||
@@ -37,31 +53,36 @@ minimap2-lite:example.o libminimap2.a
|
||||
libminimap2.a:$(OBJS)
|
||||
$(AR) -csru $@ $(OBJS)
|
||||
|
||||
sdust:sdust.c getopt.o kalloc.o kalloc.h kdq.h kvec.h kseq.h sdust.h
|
||||
$(CC) -D_SDUST_MAIN $(CFLAGS) $< getopt.o kalloc.o -o $@ -lz
|
||||
sdust:sdust.c kalloc.o kalloc.h kdq.h kvec.h kseq.h ketopt.h sdust.h
|
||||
$(CC) -D_SDUST_MAIN $(CFLAGS) $< kalloc.o -o $@ -lz
|
||||
|
||||
# SSE-specific targets on x86/x86_64
|
||||
|
||||
ifeq ($(arm_neon),) # if arm_neon is defined, compile this target with the default setting (i.e. no -msse2)
|
||||
ksw2_ll_sse.o:ksw2_ll_sse.c ksw2.h kalloc.h
|
||||
$(CC) -c $(CFLAGS) -msse2 $(CPPFLAGS) $(INCLUDES) $< -o $@
|
||||
endif
|
||||
|
||||
ksw2_extz2_sse41.o:ksw2_extz2_sse.c ksw2.h kalloc.h
|
||||
$(CC) -c -msse4 $(CFLAGS) $(CPPFLAGS) -DKSW_CPU_DISPATCH $(INCLUDES) $< -o $@
|
||||
$(CC) -c $(CFLAGS) -msse4.1 $(CPPFLAGS) -DKSW_CPU_DISPATCH $(INCLUDES) $< -o $@
|
||||
|
||||
ksw2_extz2_sse2.o:ksw2_extz2_sse.c ksw2.h kalloc.h
|
||||
$(CC) -c $(CFLAGS) $(CPPFLAGS) -DKSW_CPU_DISPATCH -DKSW_SSE2_ONLY $(INCLUDES) $< -o $@
|
||||
$(CC) -c $(CFLAGS) -msse2 -mno-sse4.1 $(CPPFLAGS) -DKSW_CPU_DISPATCH -DKSW_SSE2_ONLY $(INCLUDES) $< -o $@
|
||||
|
||||
ksw2_extd2_sse41.o:ksw2_extd2_sse.c ksw2.h kalloc.h
|
||||
$(CC) -c -msse4 $(CFLAGS) $(CPPFLAGS) -DKSW_CPU_DISPATCH $(INCLUDES) $< -o $@
|
||||
$(CC) -c $(CFLAGS) -msse4.1 $(CPPFLAGS) -DKSW_CPU_DISPATCH $(INCLUDES) $< -o $@
|
||||
|
||||
ksw2_extd2_sse2.o:ksw2_extd2_sse.c ksw2.h kalloc.h
|
||||
$(CC) -c $(CFLAGS) $(CPPFLAGS) -DKSW_CPU_DISPATCH -DKSW_SSE2_ONLY $(INCLUDES) $< -o $@
|
||||
$(CC) -c $(CFLAGS) -msse2 -mno-sse4.1 $(CPPFLAGS) -DKSW_CPU_DISPATCH -DKSW_SSE2_ONLY $(INCLUDES) $< -o $@
|
||||
|
||||
ksw2_exts2_sse41.o:ksw2_exts2_sse.c ksw2.h kalloc.h
|
||||
$(CC) -c -msse4 $(CFLAGS) $(CPPFLAGS) -DKSW_CPU_DISPATCH $(INCLUDES) $< -o $@
|
||||
$(CC) -c $(CFLAGS) -msse4.1 $(CPPFLAGS) -DKSW_CPU_DISPATCH $(INCLUDES) $< -o $@
|
||||
|
||||
ksw2_exts2_sse2.o:ksw2_exts2_sse.c ksw2.h kalloc.h
|
||||
$(CC) -c $(CFLAGS) $(CPPFLAGS) -DKSW_CPU_DISPATCH -DKSW_SSE2_ONLY $(INCLUDES) $< -o $@
|
||||
$(CC) -c $(CFLAGS) -msse2 -mno-sse4.1 $(CPPFLAGS) -DKSW_CPU_DISPATCH -DKSW_SSE2_ONLY $(INCLUDES) $< -o $@
|
||||
|
||||
ksw2_dispatch.o:ksw2_dispatch.c ksw2.h
|
||||
$(CC) -c $(CFLAGS) $(CPPFLAGS) -DKSW_CPU_DISPATCH $(INCLUDES) $< -o $@
|
||||
$(CC) -c $(CFLAGS) -msse4.1 $(CPPFLAGS) -DKSW_CPU_DISPATCH $(INCLUDES) $< -o $@
|
||||
|
||||
# NEON-specific targets on ARM
|
||||
|
||||
@@ -84,26 +105,28 @@ depend:
|
||||
|
||||
# DO NOT DELETE
|
||||
|
||||
align.o: minimap.h mmpriv.h bseq.h ksw2.h kalloc.h
|
||||
align.o: minimap.h mmpriv.h bseq.h kseq.h ksw2.h kalloc.h
|
||||
bseq.o: bseq.h kvec.h kalloc.h kseq.h
|
||||
chain.o: minimap.h mmpriv.h bseq.h kalloc.h
|
||||
esterr.o: mmpriv.h minimap.h bseq.h
|
||||
esterr.o: mmpriv.h minimap.h bseq.h kseq.h
|
||||
example.o: minimap.h kseq.h
|
||||
format.o: kalloc.h mmpriv.h minimap.h bseq.h
|
||||
getopt.o: getopt.h
|
||||
hit.o: mmpriv.h minimap.h bseq.h kalloc.h khash.h
|
||||
index.o: kthread.h bseq.h minimap.h mmpriv.h kvec.h kalloc.h khash.h
|
||||
format.o: kalloc.h mmpriv.h minimap.h bseq.h kseq.h
|
||||
hit.o: mmpriv.h minimap.h bseq.h kseq.h kalloc.h khash.h
|
||||
index.o: kthread.h bseq.h minimap.h mmpriv.h kseq.h kvec.h kalloc.h khash.h
|
||||
index.o: ksort.h
|
||||
kalloc.o: kalloc.h
|
||||
ksw2_extd2_sse.o: ksw2.h kalloc.h
|
||||
ksw2_exts2_sse.o: ksw2.h kalloc.h
|
||||
ksw2_extz2_sse.o: ksw2.h kalloc.h
|
||||
ksw2_ll_sse.o: ksw2.h kalloc.h
|
||||
kthread.o: kthread.h
|
||||
main.o: bseq.h minimap.h mmpriv.h getopt.h
|
||||
map.o: kthread.h kvec.h kalloc.h sdust.h mmpriv.h minimap.h bseq.h khash.h
|
||||
map.o: ksort.h
|
||||
misc.o: mmpriv.h minimap.h bseq.h ksort.h
|
||||
options.o: mmpriv.h minimap.h bseq.h
|
||||
pe.o: mmpriv.h minimap.h bseq.h kvec.h kalloc.h ksort.h
|
||||
lchain.o: mmpriv.h minimap.h bseq.h kseq.h kalloc.h krmq.h
|
||||
main.o: bseq.h minimap.h mmpriv.h kseq.h ketopt.h
|
||||
map.o: kthread.h kvec.h kalloc.h sdust.h mmpriv.h minimap.h bseq.h kseq.h
|
||||
map.o: khash.h ksort.h
|
||||
misc.o: mmpriv.h minimap.h bseq.h kseq.h ksort.h
|
||||
options.o: mmpriv.h minimap.h bseq.h kseq.h
|
||||
pe.o: mmpriv.h minimap.h bseq.h kseq.h kvec.h kalloc.h ksort.h
|
||||
sdust.o: kalloc.h kdq.h kvec.h sdust.h
|
||||
sketch.o: kvec.h kalloc.h mmpriv.h minimap.h bseq.h
|
||||
seed.o: mmpriv.h minimap.h bseq.h kseq.h kalloc.h ksort.h
|
||||
sketch.o: kvec.h kalloc.h mmpriv.h minimap.h bseq.h kseq.h
|
||||
splitidx.o: mmpriv.h minimap.h bseq.h kseq.h
|
||||
|
||||
@@ -0,0 +1,97 @@
|
||||
CFLAGS= -g -Wall -O2 -Wc++-compat #-Wextra
|
||||
CPPFLAGS= -DHAVE_KALLOC -DUSE_SIMDE -DSIMDE_ENABLE_NATIVE_ALIASES
|
||||
INCLUDES= -Ilib/simde
|
||||
OBJS= kthread.o kalloc.o misc.o bseq.o sketch.o sdust.o options.o index.o chain.o align.o hit.o map.o format.o pe.o esterr.o splitidx.o \
|
||||
ksw2_extz2_simde.o ksw2_extd2_simde.o ksw2_exts2_simde.o ksw2_ll_simde.o
|
||||
PROG= minimap2
|
||||
PROG_EXTRA= sdust minimap2-lite
|
||||
LIBS= -lm -lz -lpthread
|
||||
|
||||
|
||||
ifneq ($(arm_neon),) # if arm_neon is defined
|
||||
ifeq ($(aarch64),) #if aarch64 is not defined
|
||||
CFLAGS+=-D_FILE_OFFSET_BITS=64 -mfpu=neon -fsigned-char
|
||||
else #if aarch64 is defined
|
||||
CFLAGS+=-D_FILE_OFFSET_BITS=64 -fsigned-char
|
||||
endif
|
||||
endif
|
||||
|
||||
ifneq ($(asan),)
|
||||
CFLAGS+=-fsanitize=address
|
||||
LIBS+=-fsanitize=address
|
||||
endif
|
||||
|
||||
ifneq ($(tsan),)
|
||||
CFLAGS+=-fsanitize=thread
|
||||
LIBS+=-fsanitize=thread
|
||||
endif
|
||||
|
||||
.PHONY:all extra clean depend
|
||||
.SUFFIXES:.c .o
|
||||
|
||||
.c.o:
|
||||
$(CC) -c $(CFLAGS) $(CPPFLAGS) $(INCLUDES) $< -o $@
|
||||
|
||||
all:$(PROG)
|
||||
|
||||
extra:all $(PROG_EXTRA)
|
||||
|
||||
minimap2:main.o libminimap2.a
|
||||
$(CC) $(CFLAGS) main.o -o $@ -L. -lminimap2 $(LIBS)
|
||||
|
||||
minimap2-lite:example.o libminimap2.a
|
||||
$(CC) $(CFLAGS) $< -o $@ -L. -lminimap2 $(LIBS)
|
||||
|
||||
libminimap2.a:$(OBJS)
|
||||
$(AR) -csru $@ $(OBJS)
|
||||
|
||||
sdust:sdust.c kalloc.o kalloc.h kdq.h kvec.h kseq.h ketopt.h sdust.h
|
||||
$(CC) -D_SDUST_MAIN $(CFLAGS) $< kalloc.o -o $@ -lz
|
||||
|
||||
ksw2_ll_simde.o:ksw2_ll_sse.c ksw2.h kalloc.h
|
||||
$(CC) -c $(CFLAGS) -msse2 $(CPPFLAGS) $(INCLUDES) $< -o $@
|
||||
|
||||
ksw2_extz2_simde.o:ksw2_extz2_sse.c ksw2.h kalloc.h
|
||||
$(CC) -c $(CFLAGS) -msse4.1 $(CPPFLAGS) $(INCLUDES) $< -o $@
|
||||
|
||||
ksw2_extd2_simde.o:ksw2_extd2_sse.c ksw2.h kalloc.h
|
||||
$(CC) -c $(CFLAGS) -msse4.1 $(CPPFLAGS) $(INCLUDES) $< -o $@
|
||||
|
||||
ksw2_exts2_simde.o:ksw2_exts2_sse.c ksw2.h kalloc.h
|
||||
$(CC) -c $(CFLAGS) -msse4.1 $(CPPFLAGS) $(INCLUDES) $< -o $@
|
||||
|
||||
# other non-file targets
|
||||
|
||||
clean:
|
||||
rm -fr gmon.out *.o a.out $(PROG) $(PROG_EXTRA) *~ *.a *.dSYM build dist mappy*.so mappy.c python/mappy.c mappy.egg*
|
||||
|
||||
depend:
|
||||
(LC_ALL=C; export LC_ALL; makedepend -Y -- $(CFLAGS) $(CPPFLAGS) -- *.c)
|
||||
|
||||
# DO NOT DELETE
|
||||
|
||||
align.o: minimap.h mmpriv.h bseq.h kseq.h ksw2.h kalloc.h
|
||||
bseq.o: bseq.h kvec.h kalloc.h kseq.h
|
||||
chain.o: minimap.h mmpriv.h bseq.h kseq.h kalloc.h
|
||||
esterr.o: mmpriv.h minimap.h bseq.h kseq.h
|
||||
example.o: minimap.h kseq.h
|
||||
format.o: kalloc.h mmpriv.h minimap.h bseq.h kseq.h
|
||||
hit.o: mmpriv.h minimap.h bseq.h kseq.h kalloc.h khash.h
|
||||
index.o: kthread.h bseq.h minimap.h mmpriv.h kseq.h kvec.h kalloc.h khash.h
|
||||
index.o: ksort.h
|
||||
kalloc.o: kalloc.h
|
||||
ksw2_extd2_sse.o: ksw2.h kalloc.h
|
||||
ksw2_exts2_sse.o: ksw2.h kalloc.h
|
||||
ksw2_extz2_sse.o: ksw2.h kalloc.h
|
||||
ksw2_ll_sse.o: ksw2.h kalloc.h
|
||||
kthread.o: kthread.h
|
||||
main.o: bseq.h minimap.h mmpriv.h kseq.h ketopt.h
|
||||
map.o: kthread.h kvec.h kalloc.h sdust.h mmpriv.h minimap.h bseq.h kseq.h
|
||||
map.o: khash.h ksort.h
|
||||
misc.o: mmpriv.h minimap.h bseq.h kseq.h ksort.h
|
||||
options.o: mmpriv.h minimap.h bseq.h kseq.h
|
||||
pe.o: mmpriv.h minimap.h bseq.h kseq.h kvec.h kalloc.h ksort.h
|
||||
sdust.o: kalloc.h kdq.h kvec.h sdust.h
|
||||
self-chain.o: minimap.h kseq.h
|
||||
sketch.o: kvec.h kalloc.h mmpriv.h minimap.h bseq.h kseq.h
|
||||
splitidx.o: mmpriv.h minimap.h bseq.h kseq.h
|
||||
@@ -1,3 +1,462 @@
|
||||
Release 2.21-r1071 (6 July 2021)
|
||||
--------------------------------
|
||||
|
||||
This release fixed a regression in short-read mapping introduced in v2.19
|
||||
(#776). It also fixed invalid comparisons of uninitialized variables, though
|
||||
these are harmless (#752). Long-read alignment should be identical to v2.20.
|
||||
|
||||
(2.21: 6 July 2021)
|
||||
|
||||
|
||||
|
||||
Release 2.20-r1061 (27 May 2021)
|
||||
--------------------------------
|
||||
|
||||
This release fixed a bug in the Python module and improves the command-line
|
||||
compatibiliity with v2.18. In v2.19, if `-r` is specified with an `asm*` preset,
|
||||
users would get alignments more fragmented than v2.18. This could be an issue
|
||||
for existing pipelines specifying `-r`. This release resolves this issue.
|
||||
|
||||
(2.20: 27 May 2021, r1061)
|
||||
|
||||
|
||||
|
||||
Release 2.19-r1057 (26 May 2021)
|
||||
--------------------------------
|
||||
|
||||
This release includes a few important improvements backported from unimap:
|
||||
|
||||
* Improvement: more contiguous alignment through long INDELs. This is enabled
|
||||
by the minigraph chaining algorithm. All `asm*` presets now use the new
|
||||
algorithm. They can find INDELs up to 100kb and may be faster for
|
||||
chromosome-long contigs. The default mode and `map*` presets use this
|
||||
algorithm to replace the long-join heuristic.
|
||||
|
||||
* Improvement: better alignment in highly repetitive regions by rescuing
|
||||
high-occurrence seeds. If the distance between two adjacent seeds is too
|
||||
large, attempt to choose a fraction of high-occurrence seeds in-between.
|
||||
Minimap2 now produces fewer clippings and alignment break points in long
|
||||
satellite regions.
|
||||
|
||||
* Improvement: allow to specify an interval of k-mer occurrences with `-U`.
|
||||
For repeat-rich genomes, the automatic k-mer occurrence threshold determined
|
||||
by `-f` may be too large and makes alignment impractically slow. The new
|
||||
option protects against such cases. Enabled for `asm*` and `map-hifi`.
|
||||
|
||||
* New feature: added the `map-hifi` preset for maping PacBio High-Fidelity
|
||||
(HiFi) reads.
|
||||
|
||||
* Change to the default: apply `--cap-sw-mem=100m` for genomic alignment.
|
||||
|
||||
* Bugfix: minimap2 could not generate an index file with `-xsr` (#734).
|
||||
|
||||
This release represents the most signficant algorithmic change since v2.1 in
|
||||
2017. With features backported from unimap, minimap2 now has similar power to
|
||||
unimap for contig alignment. Unimap will remain an experimental project and is
|
||||
no longer recommended over minimap2. Sorry for reverting the recommendation in
|
||||
short time.
|
||||
|
||||
(2.19: 26 May 2021, r1057)
|
||||
|
||||
|
||||
|
||||
Release 2.18-r1015 (9 April 2021)
|
||||
---------------------------------
|
||||
|
||||
This release fixes multiple rare bugs in minimap2 and adds additional
|
||||
functionality to paftools.js.
|
||||
|
||||
Changes to minimap2:
|
||||
|
||||
* Bugfix: a rare segfault caused by an off-by-one error (#489)
|
||||
|
||||
* Bugfix: minimap2 segfaulted due to an uninitilized variable (#622 and #625).
|
||||
|
||||
* Bugfix: minimap2 parsed spaces as field separators in BED (#721). This led
|
||||
to issues when the BED name column contains spaces.
|
||||
|
||||
* Bugfix: minimap2 `--split-prefix` did not work with long reference names
|
||||
(#394).
|
||||
|
||||
* Bugfix: option `--junc-bonus` didn't work (#513)
|
||||
|
||||
* Bugfix: minimap2 didn't return 1 on I/O errors (#532)
|
||||
|
||||
* Bugfix: the `de:f` tag (sequence divergence) could be negative if there were
|
||||
ambiguous bases
|
||||
|
||||
* Bugfix: fixed two undefined behaviors caused by calling memcpy() on
|
||||
zero-length blocks (#443)
|
||||
|
||||
* Bugfix: there were duplicated SAM @SQ lines if option `--split-prefix` is in
|
||||
use (#400 and #527)
|
||||
|
||||
* Bugfix: option -K had to be smaller than 2 billion (#491). This was caused
|
||||
by a 32-bit integer overflow.
|
||||
|
||||
* Improvement: optionally compile against SIMDe (#597). Minimap2 should work
|
||||
with IBM POWER CPUs, though this has not been tested. To compile with SIMDe,
|
||||
please use `make -f Makefile.simde`.
|
||||
|
||||
* Improvement: more informative error message for I/O errors (#454) and for
|
||||
FASTQ parsing errors (#510)
|
||||
|
||||
* Improvement: abort given malformatted RG line (#541)
|
||||
|
||||
* Improvement: better formula to estimate the `dv:f` tag (approximate sequence
|
||||
divergence). See DOI:10.1101/2021.01.15.426881.
|
||||
|
||||
* New feature: added the `--mask-len` option to fine control the removal of
|
||||
redundant hits (#659). The default behavior is unchanged.
|
||||
|
||||
Changes to mappy:
|
||||
|
||||
* Bugfix: mappy caused segmentation fault if the reference index is not
|
||||
present (#413).
|
||||
|
||||
* Bugfix: fixed a memory leak via 238b6bb3
|
||||
|
||||
* Change: always require Cython to compile the mappy module (#723). Older
|
||||
mappy packages at PyPI bundled the C source code generated by Cython such
|
||||
that end users did not need to install Cython to compile mappy. However, as
|
||||
Python 3.9 is breaking backward compatibility, older mappy does not work
|
||||
with Python 3.9 anymore. We have to add this Cython dependency as a
|
||||
workaround.
|
||||
|
||||
Changes to paftools.js:
|
||||
|
||||
* Bugfix: the "part10-" line from asmgene was wrong (#581)
|
||||
|
||||
* Improvement: compatibility with GTF files from GenBank (#422)
|
||||
|
||||
* New feature: asmgene also checks missing multi-copy genes
|
||||
|
||||
* New feature: added the misjoin command to evaluate large-scale misjoins and
|
||||
megabase-long inversions.
|
||||
|
||||
Although given the many bug fixes and minor improvements, the core algorithm
|
||||
stays the same. This version of minimap2 produces nearly identical alignments
|
||||
to v2.17 except very rare corner cases.
|
||||
|
||||
Now unimap is recommended over minimap2 for aligning long contigs against a
|
||||
reference genome. It often takes less wall-clock time and is much more
|
||||
sensitive to long insertions and deletions.
|
||||
|
||||
(2.18: 9 April 2021, r1015)
|
||||
|
||||
|
||||
|
||||
Release 2.17-r941 (4 May 2019)
|
||||
------------------------------
|
||||
|
||||
Changes since the last release:
|
||||
|
||||
* Fixed flawed CIGARs like `5I6D7I` (#392).
|
||||
|
||||
* Bugfix: TLEN should be 0 when either end is unmapped (#373 and #365).
|
||||
|
||||
* Bugfix: mappy is unable to write index (#372).
|
||||
|
||||
* Added option `--junc-bed` to load known gene annotations in the BED12
|
||||
format. Minimap2 prefers annotated junctions over novel junctions (#197 and
|
||||
#348). GTF can be converted to BED12 with `paftools.js gff2bed`.
|
||||
|
||||
* Added option `--sam-hit-only` to suppress unmapped hits in SAM (#377).
|
||||
|
||||
* Added preset `splice:hq` for high-quality CCS or mRNA sequences. It applies
|
||||
better scoring and improves the sensitivity to small exons. This preset may
|
||||
introduce false small introns, but the overall accuracy should be higher.
|
||||
|
||||
This version produces nearly identical alignments to v2.16, except for CIGARs
|
||||
affected by the bug mentioned above.
|
||||
|
||||
(2.17: 5 May 2019, r941)
|
||||
|
||||
|
||||
|
||||
Release 2.16-r922 (28 February 2019)
|
||||
------------------------------------
|
||||
|
||||
This release is 50% faster for mapping ultra-long nanopore reads at comparable
|
||||
accuracy. For short-read mapping, long-read overlapping and ordinary long-read
|
||||
mapping, the performance and accuracy remain similar. This speedup is achieved
|
||||
with a new heuristic to limit the number of chaining iterations (#324). Users
|
||||
can disable the heuristic by increasing a new option `--max-chain-iter` to a
|
||||
huge number.
|
||||
|
||||
Other changes to minimap2:
|
||||
|
||||
* Implemented option `--paf-no-hit` to output unmapped query sequences in PAF.
|
||||
The strand and reference name columns are both `*` at an unmapped line. The
|
||||
hidden option is available in earlier minimap2 but had a different 2-column
|
||||
output format instead of PAF.
|
||||
|
||||
* Fixed a bug that leads to wrongly calculated `de` tags when ambiguous bases
|
||||
are involved (#309). This bug only affects v2.15.
|
||||
|
||||
* Fixed a bug when parsing command-line option `--splice` (#344). This bug was
|
||||
introduced in v2.13.
|
||||
|
||||
* Fixed two division-by-zero cases (#326). They don't affect final alignments
|
||||
because the results of the divisions are not used in both case.
|
||||
|
||||
* Added an option `-o` to output alignments to a specified file. It is still
|
||||
recommended to use UNIX pipes for on-the-fly conversion or compression.
|
||||
|
||||
* Output a new `rl` tag to give the length of query regions harboring
|
||||
repetitive seeds.
|
||||
|
||||
Changes to paftool.js:
|
||||
|
||||
* Added a new option to convert the MD tag to the long form of the cs tag.
|
||||
|
||||
Changes to mappy:
|
||||
|
||||
* Added the `mappy.Aligner.seq_names` method to return sequence names (#312).
|
||||
|
||||
For NA12878 ultra-long reads, this release changes the alignments of <0.1% of
|
||||
reads in comparison to v2.15. All these reads have highly fragmented alignments
|
||||
and are likely to be problematic anyway. For shorter or well aligned reads,
|
||||
this release should produce mostly identical alignments to v2.15.
|
||||
|
||||
(2.16: 28 February 2019, r922)
|
||||
|
||||
|
||||
|
||||
Release 2.15-r905 (10 January 2019)
|
||||
-----------------------------------
|
||||
|
||||
Changes to minimap2:
|
||||
|
||||
* Fixed a rare segmentation fault when option -H is in use (#307). This may
|
||||
happen when there are very long homopolymers towards the 5'-end of a read.
|
||||
|
||||
* Fixed wrong CIGARs when option --eqx is used (#266).
|
||||
|
||||
* Fixed a typo in the base encoding table (#264). This should have no
|
||||
practical effect.
|
||||
|
||||
* Fixed a typo in the example code (#265).
|
||||
|
||||
* Improved the C++ compatibility by removing "register" (#261). However,
|
||||
minimap2 still can't be compiled in the pedantic C++ mode (#306).
|
||||
|
||||
* Output a new "de" tag for gap-compressed sequence divergence.
|
||||
|
||||
Changes to paftools.js:
|
||||
|
||||
* Added "asmgene" to evaluate the completeness of an assembly by measuring the
|
||||
uniquely mapped single-copy genes. This command learns the idea of BUSCO.
|
||||
|
||||
* Added "vcfpair" to call a phased VCF from phased whole-genome assemblies. An
|
||||
earlier version of this script is used to produce the ground truth for the
|
||||
syndip benchmark [PMID:30013044].
|
||||
|
||||
This release produces identical alignment coordinates and CIGARs in comparison
|
||||
to v2.14. Users are advised to upgrade due to the several bug fixes.
|
||||
|
||||
(2.15: 10 Janurary 2019, r905)
|
||||
|
||||
|
||||
|
||||
Release 2.14-r883 (5 November 2018)
|
||||
-----------------------------------
|
||||
|
||||
Notable changes:
|
||||
|
||||
* Fixed two minor bugs caused by typos (#254 and #266).
|
||||
|
||||
* Fixed a bug that made minimap2 abort when --eqx was used together with --MD
|
||||
or --cs (#257).
|
||||
|
||||
* Added --cap-sw-mem to cap the size of DP matrices (#259). Base alignment may
|
||||
take a lot of memory in the splicing mode. This may lead to issues when we
|
||||
run minimap2 on a cluster with a hard memory limit. The new option avoids
|
||||
unlimited memory usage at the cost of missing a few long introns.
|
||||
|
||||
* Conforming to C99 and C11 when possible (#261).
|
||||
|
||||
* Warn about malformatted FASTA or FASTQ (#252 and #255).
|
||||
|
||||
This release occasionally produces base alignments different from v2.13. The
|
||||
overall alignment accuracy remain similar.
|
||||
|
||||
(2.14: 5 November 2018, r883)
|
||||
|
||||
|
||||
|
||||
Release 2.13-r850 (11 October 2018)
|
||||
-----------------------------------
|
||||
|
||||
Changes to minimap2:
|
||||
|
||||
* Fixed wrongly formatted SAM when -L is in use (#231 and #233).
|
||||
|
||||
* Fixed an integer overflow in rare cases.
|
||||
|
||||
* Added --hard-mask-level to fine control split alignments (#244).
|
||||
|
||||
* Made --MD work with spliced alignment (#139).
|
||||
|
||||
* Replaced musl's getopt with ketopt for portability.
|
||||
|
||||
* Log peak memory usage on exit.
|
||||
|
||||
This release should produce alignments identical to v2.12 and v2.11.
|
||||
|
||||
(2.13: 11 October 2018, r850)
|
||||
|
||||
|
||||
|
||||
Release 2.12-r827 (6 August 2018)
|
||||
---------------------------------
|
||||
|
||||
Changes to minimap2:
|
||||
|
||||
* Added option --split-prefix to write proper alignments (correct mapping
|
||||
quality and clustered query sequences) given a multi-part index (#141 and
|
||||
#189; mostly by @hasindu2008).
|
||||
|
||||
* Fixed a memory leak when option -y is in use.
|
||||
|
||||
Changes to mappy:
|
||||
|
||||
* Support the MD/cs tag (#183 and #203).
|
||||
|
||||
* Allow mappy to index a single sequence, to add extra flags and to change the
|
||||
scoring system.
|
||||
|
||||
Minimap2 should produce alignments identical to v2.11.
|
||||
|
||||
(2.12: 6 August 2018, r827)
|
||||
|
||||
|
||||
|
||||
Release 2.11-r797 (20 June 2018)
|
||||
--------------------------------
|
||||
|
||||
Changes to minimap2:
|
||||
|
||||
* Improved alignment accuracy in low-complexity regions for SV calling. Thank
|
||||
@armintoepfer for multiple offline examples.
|
||||
|
||||
* Added option --eqx to encode sequence match/mismatch with the =/X CIGAR
|
||||
operators (#156, #157 and #175).
|
||||
|
||||
* When compiled with VC++, minimap2 generated wrong alignments due to a
|
||||
comparison between a signed integer and an unsigned integer (#184). Also
|
||||
fixed warnings reported by "clang -Wextra".
|
||||
|
||||
* Fixed incorrect anchor filtering due to a missing 64- to 32-bit cast.
|
||||
|
||||
* Fixed incorrect mapping quality for inversions (#148).
|
||||
|
||||
* Fixed incorrect alignment involving ambiguous bases (#155).
|
||||
|
||||
* Fixed incorrect presets: option `-r 2000` is intended to be used with
|
||||
ava-ont, not ava-pb. The bug was introduced in 2.10.
|
||||
|
||||
* Fixed a bug when --for-only/--rev-only is used together with --sr or
|
||||
--heap-sort=yes (#166).
|
||||
|
||||
* Fixed option -Y that was not working in the previous releases.
|
||||
|
||||
* Added option --lj-min-ratio to fine control the alignment of long gaps
|
||||
found by the "long-join" heuristic (#128).
|
||||
|
||||
* Exposed `mm_idx_is_idx`, `mm_idx_load` and `mm_idx_dump` C APIs (#177).
|
||||
Also fixed a bug when indexing without reference names (this feature is not
|
||||
exposed to the command line).
|
||||
|
||||
Changes to mappy:
|
||||
|
||||
* Added `__version__` (#165).
|
||||
|
||||
* Exposed the maximum fragment length parameter to mappy (#174).
|
||||
|
||||
Changes to paftools:
|
||||
|
||||
* Don't crash when there is no "cg" tag (#153).
|
||||
|
||||
* Fixed wrong coverage report by "paftools.js call" (#145).
|
||||
|
||||
This version may produce slightly different base-level alignment. The overall
|
||||
alignment statistics should remain similar.
|
||||
|
||||
(2.11: 20 June 2018, r797)
|
||||
|
||||
|
||||
|
||||
Release 2.10-r761 (27 March 2018)
|
||||
---------------------------------
|
||||
|
||||
Changes to minimap2:
|
||||
|
||||
* Optionally output the MD tag for compatibility with existing tools (#63,
|
||||
#118 and #137).
|
||||
|
||||
* Use SSE compiler flags more precisely to prevent compiling errors on certain
|
||||
machines (#127).
|
||||
|
||||
* Added option --min-occ-floor to set a minimum occurrence threshold. Presets
|
||||
intended for assembly-to-reference alignment set this option to 100. This
|
||||
option alleviates issues with regions having high copy numbers (#107).
|
||||
|
||||
* Exit with non-zero code on file writing errors (e.g. disk full; #103 and
|
||||
#132).
|
||||
|
||||
* Added option -y to copy FASTA/FASTQ comments in query sequences to the
|
||||
output (#136).
|
||||
|
||||
* Added the asm20 preset for alignments between genomes at 5-10% sequence
|
||||
divergence.
|
||||
|
||||
* Changed the band-width in the ava-ont preset from 500 to 2000. Oxford
|
||||
Nanopore reads may contain long deletion sequencing errors that break
|
||||
chaining.
|
||||
|
||||
Changes to mappy, the Python binding:
|
||||
|
||||
* Fixed a typo in Align.seq() (#126).
|
||||
|
||||
Changes to paftools.js, the companion script:
|
||||
|
||||
* Command sam2paf now converts the MD tag to cs.
|
||||
|
||||
* Support VCF output for assembly-to-reference variant calling (#109).
|
||||
|
||||
This version should produce identical alignment for read overlapping, RNA-seq
|
||||
read mapping, and genomic read mapping. We have also added a cook book to show
|
||||
the variety uses of minimap2 on real datasets. Please see cookbook.md in the
|
||||
minimap2 source code directory.
|
||||
|
||||
(2.10: 27 March 2017, r761)
|
||||
|
||||
|
||||
|
||||
Release 2.9-r720 (23 February 2018)
|
||||
-----------------------------------
|
||||
|
||||
This release fixed multiple minor bugs.
|
||||
|
||||
* Fixed two bugs that lead to incorrect inversion alignment. Also improved the
|
||||
sensitivity to small inversions by using double Z-drop cutoff (#112).
|
||||
|
||||
* Fixed an issue that may cause the end of a query sequence unmapped (#104).
|
||||
|
||||
* Added a mappy API to retrieve sequences from the index (#126) and to reverse
|
||||
complement DNA sequences. Fixed a bug where the `best_n` parameter did not
|
||||
work (#117).
|
||||
|
||||
* Avoided segmentation fault given incorrect FASTQ input (#111).
|
||||
|
||||
* Combined all auxiliary javascripts to paftools.js. Fixed several bugs in
|
||||
these scripts at the same time.
|
||||
|
||||
(2.9: 24 February 2018, r720)
|
||||
|
||||
|
||||
|
||||
Release 2.8-r672 (1 February 2018)
|
||||
----------------------------------
|
||||
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
[](https://github.com/lh3/minimap2/releases)
|
||||
[](https://anaconda.org/bioconda/minimap2)
|
||||
[](https://pypi.python.org/pypi/mappy)
|
||||
[](https://travis-ci.org/lh3/minimap2)
|
||||
[](https://github.com/lh3/minimap2/actions)
|
||||
## <a name="started"></a>Getting Started
|
||||
```sh
|
||||
git clone https://github.com/lh3/minimap2
|
||||
@@ -9,20 +9,25 @@ cd minimap2 && make
|
||||
# long sequences against a reference genome
|
||||
./minimap2 -a test/MT-human.fa test/MT-orang.fa > test.sam
|
||||
# create an index first and then map
|
||||
./minimap2 -d MT-human.mmi test/MT-human.fa
|
||||
./minimap2 -a MT-human.mmi test/MT-orang.fa > test.sam
|
||||
./minimap2 -x map-ont -d MT-human-ont.mmi test/MT-human.fa
|
||||
./minimap2 -a MT-human-ont.mmi test/MT-orang.fa > test.sam
|
||||
# use presets (no test data)
|
||||
./minimap2 -ax map-pb ref.fa pacbio.fq.gz > aln.sam # PacBio genomic reads
|
||||
./minimap2 -ax map-pb ref.fa pacbio.fq.gz > aln.sam # PacBio CLR genomic reads
|
||||
./minimap2 -ax map-ont ref.fa ont.fq.gz > aln.sam # Oxford Nanopore genomic reads
|
||||
./minimap2 -ax map-hifi ref.fa pacbio-ccs.fq.gz > aln.sam # PacBio HiFi/CCS genomic reads (v2.19 or later)
|
||||
./minimap2 -ax asm20 ref.fa pacbio-ccs.fq.gz > aln.sam # PacBio HiFi/CCS genomic reads (v2.18 or earlier)
|
||||
./minimap2 -ax sr ref.fa read1.fa read2.fa > aln.sam # short genomic paired-end reads
|
||||
./minimap2 -ax splice ref.fa rna-reads.fa > aln.sam # spliced long reads
|
||||
./minimap2 -ax splice -k14 -uf ref.fa reads.fa > aln.sam # Nanopore Direct RNA-seq
|
||||
./minimap2 -ax splice ref.fa rna-reads.fa > aln.sam # spliced long reads (strand unknown)
|
||||
./minimap2 -ax splice -uf -k14 ref.fa reads.fa > aln.sam # noisy Nanopore Direct RNA-seq
|
||||
./minimap2 -ax splice:hq -uf ref.fa query.fa > aln.sam # Final PacBio Iso-seq or traditional cDNA
|
||||
./minimap2 -ax splice --junc-bed anno.bed12 ref.fa query.fa > aln.sam # prioritize on annotated junctions
|
||||
./minimap2 -cx asm5 asm1.fa asm2.fa > aln.paf # intra-species asm-to-asm alignment
|
||||
./minimap2 -x ava-pb reads.fa reads.fa > overlaps.paf # PacBio read overlap
|
||||
./minimap2 -x ava-ont reads.fa reads.fa > overlaps.paf # Nanopore read overlap
|
||||
# man page for detailed command line options
|
||||
man ./minimap2.1
|
||||
```
|
||||
|
||||
## Table of Contents
|
||||
|
||||
- [Getting Started](#started)
|
||||
@@ -38,7 +43,7 @@ man ./minimap2.1
|
||||
- [Advanced features](#advanced)
|
||||
- [Working with >65535 CIGAR operations](#long-cigar)
|
||||
- [The cs optional tag](#cs)
|
||||
- [Evaluation scripts](#eval)
|
||||
- [Working with the PAF format](#paftools)
|
||||
- [Algorithm overview](#algo)
|
||||
- [Getting help](#help)
|
||||
- [Citing minimap2](#cite)
|
||||
@@ -61,16 +66,16 @@ mainstream long-read mappers such as BLASR, BWA-MEM, NGMLR and GMAP. It is more
|
||||
accurate on simulated long reads and produces biologically meaningful alignment
|
||||
ready for downstream analyses. For >100bp Illumina short reads, minimap2 is
|
||||
three times as fast as BWA-MEM and Bowtie2, and as accurate on simulated data.
|
||||
Detailed evaluations are available from the [minimap2 preprint][preprint].
|
||||
Detailed evaluations are available from the [minimap2 paper][doi] or the
|
||||
[preprint][preprint].
|
||||
|
||||
### <a name="install"></a>Installation
|
||||
|
||||
Minimap2 is optimized for x86-64 CPUs. You can acquire precompiled binaries from
|
||||
the [release page][release] with:
|
||||
```sh
|
||||
curl -L https://github.com/lh3/minimap2/releases/download/v2.8/minimap2-2.8_x64-linux.tar.bz2 \
|
||||
| tar -jxvf -
|
||||
./minimap2-2.8_x64-linux/minimap2
|
||||
curl -L https://github.com/lh3/minimap2/releases/download/v2.21/minimap2-2.21_x64-linux.tar.bz2 | tar -jxvf -
|
||||
./minimap2-2.21_x64-linux/minimap2
|
||||
```
|
||||
If you want to compile from the source, you need to have a C compiler, GNU make
|
||||
and zlib development files installed. Then type `make` in the source code
|
||||
@@ -78,13 +83,20 @@ directory to compile. If you see compilation errors, try `make sse2only=1`
|
||||
to disable SSE4 code, which will make minimap2 slightly slower.
|
||||
|
||||
Minimap2 also works with ARM CPUs supporting the NEON instruction sets. To
|
||||
compile, use `make arm_neon=1`.
|
||||
compile for 32 bit ARM architectures (such as ARMv7), use `make arm_neon=1`. To
|
||||
compile for for 64 bit ARM architectures (such as ARMv8), use `make arm_neon=1
|
||||
aarch64=1`.
|
||||
|
||||
Minimap2 can use [SIMD Everywhere (SIMDe)][simde] library for porting
|
||||
implementation to the different SIMD instruction sets. To compile using SIMDe,
|
||||
use `make -f Makefile.simde`. To compile for ARM CPUs, use `Makefile.simde`
|
||||
with the ARM related command lines given above.
|
||||
|
||||
### <a name="general"></a>General usage
|
||||
|
||||
Without any options, minimap2 takes a reference database and a query sequence
|
||||
file as input and produce approximate mapping, without base-level alignment
|
||||
(i.e. no CIGAR), in the [PAF format][paf]:
|
||||
(i.e. coordinates are only approximate and no CIGAR in output), in the [PAF format][paf]:
|
||||
```sh
|
||||
minimap2 ref.fa query.fq > approx-mapping.paf
|
||||
```
|
||||
@@ -125,19 +137,19 @@ parameters at the same time. The default setting is the same as `map-ont`.
|
||||
#### <a name="map-long-genomic"></a>Map long noisy genomic reads
|
||||
|
||||
```sh
|
||||
minimap2 -ax map-pb ref.fa pacbio-reads.fq > aln.sam # for PacBio subreads
|
||||
minimap2 -ax map-pb ref.fa pacbio-reads.fq > aln.sam # for PacBio CLR reads
|
||||
minimap2 -ax map-ont ref.fa ont-reads.fq > aln.sam # for Oxford Nanopore reads
|
||||
```
|
||||
The difference between `map-pb` and `map-ont` is that `map-pb` uses
|
||||
homopolymer-compressed (HPC) minimizers as seeds, while `map-ont` uses ordinary
|
||||
minimizers as seeds. Emperical evaluation suggests HPC minimizers improve
|
||||
performance and sensitivity when aligning PacBio reads, but hurt when aligning
|
||||
performance and sensitivity when aligning PacBio CLR reads, but hurt when aligning
|
||||
Nanopore reads.
|
||||
|
||||
#### <a name="map-long-splice"></a>Map long mRNA/cDNA reads
|
||||
|
||||
```sh
|
||||
minimap2 -ax splice -uf ref.fa iso-seq.fq > aln.sam # PacBio Iso-seq/traditional cDNA
|
||||
minimap2 -ax splice:hq -uf ref.fa iso-seq.fq > aln.sam # PacBio Iso-seq/traditional cDNA
|
||||
minimap2 -ax splice ref.fa nanopore-cdna.fa > aln.sam # Nanopore 2D cDNA-seq
|
||||
minimap2 -ax splice -uf -k14 ref.fa direct-rna.fq > aln.sam # Nanopore Direct RNA-seq
|
||||
minimap2 -ax splice --splice-flank=no SIRV.fa SIRV-seq.fa # mapping against SIRV control
|
||||
@@ -176,10 +188,23 @@ This is because SIRV does not honor the evolutionarily conservative splicing
|
||||
signal. If you are studying SIRV, you may apply `--splice-flank=no` to let
|
||||
minimap2 only model GT..AG, ignoring the additional base.
|
||||
|
||||
Since v2.17, minimap2 can optionally take annotated genes as input and
|
||||
prioritize on annotated splice junctions. To use this feature, you can
|
||||
```sh
|
||||
paftools.js gff2bed anno.gff > anno.bed
|
||||
minimap2 -ax splice --junc-bed anno.bed ref.fa query.fa > aln.sam
|
||||
```
|
||||
Here, `anno.gff` is the gene annotation in the GTF or GFF3 format (`gff2bed`
|
||||
automatically tests the format). The output of `gff2bed` is in the 12-column
|
||||
BED format, or the BED12 format. With the `--junc-bed` option, minimap2 adds a
|
||||
bonus score (tuned by `--junc-bonus`) if an aligned junction matches a junction
|
||||
in the annotation. Option `--junc-bed` also takes 5-column BED, including the
|
||||
strand field. In this case, each line indicates an oriented junction.
|
||||
|
||||
#### <a name="long-overlap"></a>Find overlaps between long reads
|
||||
|
||||
```sh
|
||||
minimap2 -x ava-pb reads.fq reads.fq > ovlp.paf # PacBio read overlap
|
||||
minimap2 -x ava-pb reads.fq reads.fq > ovlp.paf # PacBio CLR read overlap
|
||||
minimap2 -x ava-ont reads.fq reads.fq > ovlp.paf # Oxford Nanopore read overlap
|
||||
```
|
||||
Similarly, `ava-pb` uses HPC minimizers while `ava-ont` uses ordinary
|
||||
@@ -226,10 +251,10 @@ To avoid this issue, you can add option `-L` at the minimap2 command line.
|
||||
This option moves a long CIGAR to the `CG` tag and leaves a fully clipped CIGAR
|
||||
at the SAM CIGAR column. Current tools that don't read CIGAR (e.g. merging and
|
||||
sorting) still work with such BAM records; tools that read CIGAR will
|
||||
effectively ignore these records. It has been decided that future tools will
|
||||
effectively ignore these records. It has been decided that future tools
|
||||
will seamlessly recognize long-cigar records generated by option `-L`.
|
||||
|
||||
**TD;DR**: if you work with ultra-long reads and use tools that only process
|
||||
**TL;DR**: if you work with ultra-long reads and use tools that only process
|
||||
BAM files, please add option `-L`.
|
||||
|
||||
#### <a name="cs"></a>The cs optional tag
|
||||
@@ -247,35 +272,24 @@ CGATCGATAAATAGAGTAG---GAATAGCA
|
||||
CGATCG---AATAGAGTAGGTCGAATtGCA
|
||||
```
|
||||
is represented as `:6-ata:10+gtc:4*at:3`, where `:[0-9]+` represents an
|
||||
identical block, `-ata` represents a deltion, `+gtc` an insertion and `*at`
|
||||
identical block, `-ata` represents a deletion, `+gtc` an insertion and `*at`
|
||||
indicates reference base `a` is substituted with a query base `t`. It is
|
||||
similar to the `MD` SAM tag but is standalone and easier to parse.
|
||||
|
||||
If `--cs=long` is used, the `cs` string also contains identical sequences in
|
||||
the alignment. The above example will become
|
||||
`=CGATCG-ata=AATAGAGTAG+gtc=GAAT*at=GCA`. The long form of `cs` encodes both
|
||||
reference and query sequences in one string.
|
||||
reference and query sequences in one string. The `cs` tag also encodes intron
|
||||
positions and splicing signals (see the [minimap2 manpage][manpage-cs] for
|
||||
details).
|
||||
|
||||
#### <a name="eval"></a>Evaluation scripts
|
||||
#### <a name="paftools"></a>Working with the PAF format
|
||||
|
||||
Minimap2 comes with several (java)scripts for evaluating the accuracy of
|
||||
minimap2. These scripts require the [k8][k8] javascript shell to run.
|
||||
Recent minimap2 binary release tar-balls contain a copy of k8 executable, a
|
||||
single file. Here are a few examples on how to use these scripts:
|
||||
|
||||
```sh
|
||||
# Generate reads from PBSIM alignment (truth encoded in read names)
|
||||
k8 misc/sim-pbsim.js ref.fa.fai pbsim-aln.maf > pbsim-reads.fq
|
||||
# Generate reads from mason2 alignment (not tested for simulated SVs)
|
||||
k8 misc/sim-mason2.js mason2-aln.sam > mason2-reads.fq
|
||||
# Evaluate mapping accuracy with ROC-like curve
|
||||
k8 misc/sim-eval.js my-aln.sam.gz > result.txt
|
||||
k8 misc/sim-eval.js my-aln.paf.gz > result.txt
|
||||
# Collect alignment statistics
|
||||
k8 misc/mapstat.js my-aln.sam > result.txt
|
||||
# Compare spliced junctions to existing gene annotations
|
||||
k8 misc/intron-eval.js anno.gtf my-spliced-aln.sam > result.txt
|
||||
```
|
||||
Minimap2 also comes with a (java)script [paftools.js](misc/paftools.js) that
|
||||
processes alignments in the PAF format. It calls variants from
|
||||
assembly-to-reference alignment, lifts over BED files based on alignment,
|
||||
converts between formats and provides utilities for various evaluations. For
|
||||
details, please see [misc/README.md](misc/README.md).
|
||||
|
||||
### <a name="algo"></a>Algorithm overview
|
||||
|
||||
@@ -324,15 +338,17 @@ highlighted in bold. The description may help to tune minimap2 parameters.
|
||||
### <a name="help"></a>Getting help
|
||||
|
||||
Manpage [minimap2.1][manpage] provides detailed description of minimap2
|
||||
command line options and optional tags. If you encounter bugs or have further
|
||||
questions or requests, you can raise an issue at the [issue page][issue].
|
||||
There is not a specific mailing list for the time being.
|
||||
command line options and optional tags. The [FAQ](FAQ.md) page answers several
|
||||
frequently asked questions. If you encounter bugs or have further questions or
|
||||
requests, you can raise an issue at the [issue page][issue]. There is not a
|
||||
specific mailing list for the time being.
|
||||
|
||||
### <a name="cite"></a>Citing minimap2
|
||||
|
||||
If you use minimap2 in your work, please consider to cite:
|
||||
If you use minimap2 in your work, please cite:
|
||||
|
||||
> Li, H. (2017). Minimap2: fast pairwise alignment for long nucleotide sequences. [arXiv:1708.01492][preprint]
|
||||
> Li, H. (2018). Minimap2: pairwise alignment for nucleotide sequences.
|
||||
> *Bioinformatics*, **34**:3094-3100. [doi:10.1093/bioinformatics/bty191][doi]
|
||||
|
||||
## <a name="dguide"></a>Developers' Guide
|
||||
|
||||
@@ -359,6 +375,12 @@ mappy` or [from BioConda][mappyconda] via `conda install -c bioconda mappy`.
|
||||
possible to add non-SIMD support, but it would make minimap2 slower by
|
||||
several times.
|
||||
|
||||
* Minimap2 does not work with a single query or database sequence ~2
|
||||
billion bases or longer (2,147,483,647 to be exact). The total length of all
|
||||
sequences can well exceed this threshold.
|
||||
|
||||
* Minimap2 often misses small exons.
|
||||
|
||||
|
||||
|
||||
[paf]: https://github.com/lh3/miniasm/blob/master/PAF.md
|
||||
@@ -375,3 +397,7 @@ mappy` or [from BioConda][mappyconda] via `conda install -c bioconda mappy`.
|
||||
[issue]: https://github.com/lh3/minimap2/issues
|
||||
[k8]: https://github.com/attractivechaos/k8
|
||||
[manpage]: https://lh3.github.io/minimap2/minimap2.html
|
||||
[manpage-cs]: https://lh3.github.io/minimap2/minimap2.html#10
|
||||
[doi]: https://doi.org/10.1093/bioinformatics/bty191
|
||||
[smide]: https://github.com/nemequ/simde
|
||||
[unimap]: https://github.com/lh3/unimap
|
||||
|
||||
@@ -6,18 +6,19 @@
|
||||
#include "mmpriv.h"
|
||||
#include "ksw2.h"
|
||||
|
||||
static void ksw_gen_simple_mat(int m, int8_t *mat, int8_t a, int8_t b)
|
||||
static void ksw_gen_simple_mat(int m, int8_t *mat, int8_t a, int8_t b, int8_t sc_ambi)
|
||||
{
|
||||
int i, j;
|
||||
a = a < 0? -a : a;
|
||||
b = b > 0? -b : b;
|
||||
sc_ambi = sc_ambi > 0? -sc_ambi : sc_ambi;
|
||||
for (i = 0; i < m - 1; ++i) {
|
||||
for (j = 0; j < m - 1; ++j)
|
||||
mat[i * m + j] = i == j? a : b;
|
||||
mat[i * m + m - 1] = 0;
|
||||
mat[i * m + m - 1] = sc_ambi;
|
||||
}
|
||||
for (j = 0; j < m; ++j)
|
||||
mat[(m - 1) * m + j] = 0;
|
||||
mat[(m - 1) * m + j] = sc_ambi;
|
||||
}
|
||||
|
||||
static inline void mm_seq_rev(uint32_t len, uint8_t *seq)
|
||||
@@ -28,56 +29,81 @@ static inline void mm_seq_rev(uint32_t len, uint8_t *seq)
|
||||
t = seq[i], seq[i] = seq[len - 1 - i], seq[len - 1 - i] = t;
|
||||
}
|
||||
|
||||
static inline int test_zdrop_aux(int32_t score, int i, int j, int32_t *max, int *max_i, int *max_j, int e, int zdrop)
|
||||
static inline void update_max_zdrop(int32_t score, int i, int j, int32_t *max, int *max_i, int *max_j, int e, int *max_zdrop, int pos[2][2])
|
||||
{
|
||||
if (score < *max) {
|
||||
int li = i - *max_i;
|
||||
int lj = j - *max_j;
|
||||
int diff = li > lj? li - lj : lj - li;
|
||||
if (*max - score > zdrop + diff * e)
|
||||
return 1;
|
||||
int z = *max - score - diff * e;
|
||||
if (z > *max_zdrop) {
|
||||
*max_zdrop = z;
|
||||
pos[0][0] = *max_i, pos[0][1] = i;
|
||||
pos[1][0] = *max_j, pos[1][1] = j;
|
||||
}
|
||||
} else *max = score, *max_i = i, *max_j = j;
|
||||
return 0;
|
||||
}
|
||||
|
||||
static int mm_check_zdrop(const uint8_t *qseq, const uint8_t *tseq, uint32_t n_cigar, uint32_t *cigar, const int8_t *mat, int8_t q, int8_t e, int zdrop)
|
||||
static int mm_test_zdrop(void *km, const mm_mapopt_t *opt, const uint8_t *qseq, const uint8_t *tseq, uint32_t n_cigar, uint32_t *cigar, const int8_t *mat)
|
||||
{
|
||||
uint32_t k;
|
||||
int32_t score = 0, max = 0, max_i = -1, max_j = -1, i = 0, j = 0;
|
||||
for (k = 0; k < n_cigar; ++k) {
|
||||
int32_t score = 0, max = INT32_MIN, max_i = -1, max_j = -1, i = 0, j = 0, max_zdrop = 0;
|
||||
int pos[2][2] = {{-1, -1}, {-1, -1}}, q_len, t_len;
|
||||
|
||||
// find the score and the region where score drops most along diagonal
|
||||
for (k = 0, score = 0; k < n_cigar; ++k) {
|
||||
uint32_t l, op = cigar[k]&0xf, len = cigar[k]>>4;
|
||||
if (op == 0) {
|
||||
if (op == MM_CIGAR_MATCH) {
|
||||
for (l = 0; l < len; ++l) {
|
||||
score += mat[tseq[i + l] * 5 + qseq[j + l]];
|
||||
if (test_zdrop_aux(score, i+l, j+l, &max, &max_i, &max_j, e, zdrop)) return 1;
|
||||
update_max_zdrop(score, i+l, j+l, &max, &max_i, &max_j, opt->e, &max_zdrop, pos);
|
||||
}
|
||||
i += len, j += len;
|
||||
} else if (op == 1) {
|
||||
score -= q + e * len, j += len;
|
||||
if (test_zdrop_aux(score, i, j, &max, &max_i, &max_j, e, zdrop)) return 1;
|
||||
} else if (op == 2 || op == 3) {
|
||||
score -= q + e * len, i += len;
|
||||
if (test_zdrop_aux(score, i, j, &max, &max_i, &max_j, e, zdrop)) return 1;
|
||||
} else if (op == MM_CIGAR_INS || op == MM_CIGAR_DEL || op == MM_CIGAR_N_SKIP) {
|
||||
score -= opt->q + opt->e * len;
|
||||
if (op == MM_CIGAR_INS) j += len;
|
||||
else i += len;
|
||||
update_max_zdrop(score, i, j, &max, &max_i, &max_j, opt->e, &max_zdrop, pos);
|
||||
}
|
||||
}
|
||||
return 0;
|
||||
|
||||
// test if there is an inversion in the most dropped region
|
||||
q_len = pos[1][1] - pos[1][0], t_len = pos[0][1] - pos[0][0];
|
||||
if (!(opt->flag&(MM_F_SPLICE|MM_F_SR|MM_F_FOR_ONLY|MM_F_REV_ONLY)) && max_zdrop > opt->zdrop_inv && q_len < opt->max_gap && t_len < opt->max_gap) {
|
||||
uint8_t *qseq2;
|
||||
void *qp;
|
||||
int q_off, t_off;
|
||||
qseq2 = (uint8_t*)kmalloc(km, q_len);
|
||||
for (i = 0; i < q_len; ++i) {
|
||||
int c = qseq[pos[1][1] - i - 1];
|
||||
qseq2[i] = c >= 4? 4 : 3 - c;
|
||||
}
|
||||
qp = ksw_ll_qinit(km, 2, q_len, qseq2, 5, mat);
|
||||
score = ksw_ll_i16(qp, t_len, tseq + pos[0][0], opt->q, opt->e, &q_off, &t_off);
|
||||
kfree(km, qseq2);
|
||||
kfree(km, qp);
|
||||
if (score >= opt->min_chain_score * opt->a && score >= opt->min_dp_max)
|
||||
return 2; // there is a potential inversion
|
||||
}
|
||||
return max_zdrop > opt->zdrop? 1 : 0;
|
||||
}
|
||||
|
||||
static void mm_fix_cigar(mm_reg1_t *r, const uint8_t *qseq, const uint8_t *tseq, int *qshift, int *tshift)
|
||||
{
|
||||
mm_extra_t *p = r->p;
|
||||
int32_t k, toff = 0, qoff = 0, to_shrink = 0;
|
||||
int32_t toff = 0, qoff = 0, to_shrink = 0;
|
||||
uint32_t k;
|
||||
*qshift = *tshift = 0;
|
||||
if (p->n_cigar <= 1) return;
|
||||
for (k = 0; k < p->n_cigar; ++k) { // indel left alignment
|
||||
uint32_t op = p->cigar[k]&0xf, len = p->cigar[k]>>4;
|
||||
if (len == 0) to_shrink = 1;
|
||||
if (op == 0) {
|
||||
if (op == MM_CIGAR_MATCH) {
|
||||
toff += len, qoff += len;
|
||||
} else if (op == 1 || op == 2) { // insertion or deletion
|
||||
} else if (op == MM_CIGAR_INS || op == MM_CIGAR_DEL) {
|
||||
if (k > 0 && k < p->n_cigar - 1 && (p->cigar[k-1]&0xf) == 0 && (p->cigar[k+1]&0xf) == 0) {
|
||||
int l, prev_len = p->cigar[k-1] >> 4;
|
||||
if (op == 1) {
|
||||
if (op == MM_CIGAR_INS) {
|
||||
for (l = 0; l < prev_len; ++l)
|
||||
if (qseq[qoff - 1 - l] != qseq[qoff + len - 1 - l])
|
||||
break;
|
||||
@@ -90,13 +116,32 @@ static void mm_fix_cigar(mm_reg1_t *r, const uint8_t *qseq, const uint8_t *tseq,
|
||||
p->cigar[k-1] -= l<<4, p->cigar[k+1] += l<<4, qoff -= l, toff -= l;
|
||||
if (l == prev_len) to_shrink = 1;
|
||||
}
|
||||
if (op == 1) qoff += len;
|
||||
if (op == MM_CIGAR_INS) qoff += len;
|
||||
else toff += len;
|
||||
} else if (op == 3) {
|
||||
} else if (op == MM_CIGAR_N_SKIP) {
|
||||
toff += len;
|
||||
}
|
||||
}
|
||||
assert(qoff == r->qe - r->qs && toff == r->re - r->rs);
|
||||
for (k = 0; k < p->n_cigar - 2; ++k) { // fix CIGAR like 5I6D7I
|
||||
if ((p->cigar[k]&0xf) > 0 && (p->cigar[k]&0xf) + (p->cigar[k+1]&0xf) == 3) {
|
||||
uint32_t l, s[3] = {0,0,0};
|
||||
for (l = k; l < p->n_cigar; ++l) { // count number of adjacent I and D
|
||||
uint32_t op = p->cigar[l]&0xf;
|
||||
if (op == MM_CIGAR_INS || op == MM_CIGAR_DEL || p->cigar[l]>>4 == 0)
|
||||
s[op] += p->cigar[l] >> 4;
|
||||
else break;
|
||||
}
|
||||
if (s[1] > 0 && s[2] > 0 && l - k > 2) { // turn to a single I and a single D
|
||||
p->cigar[k] = s[1]<<4|MM_CIGAR_INS;
|
||||
p->cigar[k+1] = s[2]<<4|MM_CIGAR_DEL;
|
||||
for (k += 2; k < l; ++k)
|
||||
p->cigar[k] &= 0xf;
|
||||
to_shrink = 1;
|
||||
}
|
||||
k = l;
|
||||
}
|
||||
}
|
||||
if (to_shrink) { // squeeze out zero-length operations
|
||||
int32_t l = 0;
|
||||
for (k = 0; k < p->n_cigar; ++k) // squeeze out zero-length operations
|
||||
@@ -109,9 +154,9 @@ static void mm_fix_cigar(mm_reg1_t *r, const uint8_t *qseq, const uint8_t *tseq,
|
||||
else p->cigar[k+1] += p->cigar[k]>>4<<4; // add length to the next CIGAR operator
|
||||
p->n_cigar = l;
|
||||
}
|
||||
if ((p->cigar[0]&0xf) == 1 || (p->cigar[0]&0xf) == 2) { // get rid of leading I or D
|
||||
if ((p->cigar[0]&0xf) == MM_CIGAR_INS || (p->cigar[0]&0xf) == MM_CIGAR_DEL) { // get rid of leading I or D
|
||||
int32_t l = p->cigar[0] >> 4;
|
||||
if ((p->cigar[0]&0xf) == 1) {
|
||||
if ((p->cigar[0]&0xf) == MM_CIGAR_INS) {
|
||||
if (r->rev) r->qe -= l;
|
||||
else r->qs += l;
|
||||
*qshift = l;
|
||||
@@ -121,10 +166,82 @@ static void mm_fix_cigar(mm_reg1_t *r, const uint8_t *qseq, const uint8_t *tseq,
|
||||
}
|
||||
}
|
||||
|
||||
static void mm_update_extra(mm_reg1_t *r, const uint8_t *qseq, const uint8_t *tseq, const int8_t *mat, int8_t q, int8_t e)
|
||||
static void mm_update_cigar_eqx(mm_reg1_t *r, const uint8_t *qseq, const uint8_t *tseq) // written by @armintoepfer
|
||||
{
|
||||
uint32_t k, l, toff = 0, qoff = 0;
|
||||
int32_t s = 0, max = 0, qshift, tshift;
|
||||
uint32_t n_EQX = 0;
|
||||
uint32_t k, l, m, cap, toff = 0, qoff = 0, n_M = 0;
|
||||
mm_extra_t *p;
|
||||
if (r->p == 0) return;
|
||||
for (k = 0; k < r->p->n_cigar; ++k) {
|
||||
uint32_t op = r->p->cigar[k]&0xf, len = r->p->cigar[k]>>4;
|
||||
if (op == MM_CIGAR_MATCH) {
|
||||
while (len > 0) {
|
||||
for (l = 0; l < len && qseq[qoff + l] == tseq[toff + l]; ++l) {} // run of "="; TODO: N<=>N is converted to "="
|
||||
if (l > 0) { ++n_EQX; len -= l; toff += l; qoff += l; }
|
||||
|
||||
for (l = 0; l < len && qseq[qoff + l] != tseq[toff + l]; ++l) {} // run of "X"
|
||||
if (l > 0) { ++n_EQX; len -= l; toff += l; qoff += l; }
|
||||
}
|
||||
++n_M;
|
||||
} else if (op == MM_CIGAR_INS) {
|
||||
qoff += len;
|
||||
} else if (op == MM_CIGAR_DEL) {
|
||||
toff += len;
|
||||
} else if (op == MM_CIGAR_N_SKIP) {
|
||||
toff += len;
|
||||
}
|
||||
}
|
||||
// update in-place if we can
|
||||
if (n_EQX == n_M) {
|
||||
for (k = 0; k < r->p->n_cigar; ++k) {
|
||||
uint32_t op = r->p->cigar[k]&0xf, len = r->p->cigar[k]>>4;
|
||||
if (op == MM_CIGAR_MATCH) r->p->cigar[k] = len << 4 | MM_CIGAR_EQ_MATCH;
|
||||
}
|
||||
return;
|
||||
}
|
||||
// allocate new storage
|
||||
cap = r->p->n_cigar + (n_EQX - n_M) + sizeof(mm_extra_t);
|
||||
kroundup32(cap);
|
||||
p = (mm_extra_t*)calloc(cap, 4);
|
||||
memcpy(p, r->p, sizeof(mm_extra_t));
|
||||
p->capacity = cap;
|
||||
// update cigar while copying
|
||||
toff = qoff = m = 0;
|
||||
for (k = 0; k < r->p->n_cigar; ++k) {
|
||||
uint32_t op = r->p->cigar[k]&0xf, len = r->p->cigar[k]>>4;
|
||||
if (op == MM_CIGAR_MATCH) {
|
||||
while (len > 0) {
|
||||
// match
|
||||
for (l = 0; l < len && qseq[qoff + l] == tseq[toff + l]; ++l) {}
|
||||
if (l > 0) p->cigar[m++] = l << 4 | MM_CIGAR_EQ_MATCH;
|
||||
len -= l;
|
||||
toff += l, qoff += l;
|
||||
// mismatch
|
||||
for (l = 0; l < len && qseq[qoff + l] != tseq[toff + l]; ++l) {}
|
||||
if (l > 0) p->cigar[m++] = l << 4 | MM_CIGAR_X_MISMATCH;
|
||||
len -= l;
|
||||
toff += l, qoff += l;
|
||||
}
|
||||
continue;
|
||||
} else if (op == MM_CIGAR_INS) {
|
||||
qoff += len;
|
||||
} else if (op == MM_CIGAR_DEL) {
|
||||
toff += len;
|
||||
} else if (op == MM_CIGAR_N_SKIP) {
|
||||
toff += len;
|
||||
}
|
||||
p->cigar[m++] = r->p->cigar[k];
|
||||
}
|
||||
p->n_cigar = m;
|
||||
free(r->p);
|
||||
r->p = p;
|
||||
}
|
||||
|
||||
static void mm_update_extra(mm_reg1_t *r, const uint8_t *qseq, const uint8_t *tseq, const int8_t *mat, int8_t q, int8_t e, int is_eqx, int log_gap)
|
||||
{
|
||||
uint32_t k, l;
|
||||
int32_t qshift, tshift, toff = 0, qoff = 0;
|
||||
double s = 0.0, max = 0.0;
|
||||
mm_extra_t *p = r->p;
|
||||
if (p == 0) return;
|
||||
mm_fix_cigar(r, qseq, tseq, &qshift, &tshift);
|
||||
@@ -132,7 +249,7 @@ static void mm_update_extra(mm_reg1_t *r, const uint8_t *qseq, const uint8_t *ts
|
||||
r->blen = r->mlen = 0;
|
||||
for (k = 0; k < p->n_cigar; ++k) {
|
||||
uint32_t op = p->cigar[k]&0xf, len = p->cigar[k]>>4;
|
||||
if (op == 0) { // match/mismatch
|
||||
if (op == MM_CIGAR_MATCH) {
|
||||
int n_ambi = 0, n_diff = 0;
|
||||
for (l = 0; l < len; ++l) {
|
||||
int cq = qseq[qoff + l], ct = tseq[toff + l];
|
||||
@@ -144,28 +261,31 @@ static void mm_update_extra(mm_reg1_t *r, const uint8_t *qseq, const uint8_t *ts
|
||||
}
|
||||
r->blen += len - n_ambi, r->mlen += len - (n_ambi + n_diff), p->n_ambi += n_ambi;
|
||||
toff += len, qoff += len;
|
||||
} else if (op == 1) { // insertion
|
||||
} else if (op == MM_CIGAR_INS) {
|
||||
int n_ambi = 0;
|
||||
for (l = 0; l < len; ++l)
|
||||
if (qseq[qoff + l] > 3) ++n_ambi;
|
||||
r->blen += len - n_ambi, p->n_ambi += n_ambi;
|
||||
s -= q + e * len;
|
||||
if (log_gap) s -= q + (double)e * mg_log2(1.0 + len);
|
||||
else s -= q + e;
|
||||
if (s < 0) s = 0;
|
||||
qoff += len;
|
||||
} else if (op == 2) { // deletion
|
||||
} else if (op == MM_CIGAR_DEL) {
|
||||
int n_ambi = 0;
|
||||
for (l = 0; l < len; ++l)
|
||||
if (tseq[toff + l] > 3) ++n_ambi;
|
||||
r->blen += len - n_ambi, p->n_ambi += n_ambi;
|
||||
s -= q + e * len;
|
||||
if (log_gap) s -= q + (double)e * mg_log2(1.0 + len);
|
||||
else s -= q + e;
|
||||
if (s < 0) s = 0;
|
||||
toff += len;
|
||||
} else if (op == 3) { // intron
|
||||
} else if (op == MM_CIGAR_N_SKIP) {
|
||||
toff += len;
|
||||
}
|
||||
}
|
||||
p->dp_max = max;
|
||||
p->dp_max = (int32_t)(max + .499);
|
||||
assert(qoff == r->qe - r->qs && toff == r->re - r->rs);
|
||||
if (is_eqx) mm_update_cigar_eqx(r, qseq, tseq); // NB: it has to be called here as changes to qseq and tseq are not returned
|
||||
}
|
||||
|
||||
static void mm_append_cigar(mm_reg1_t *r, uint32_t n_cigar, uint32_t *cigar) // TODO: this calls the libc realloc()
|
||||
@@ -173,12 +293,12 @@ static void mm_append_cigar(mm_reg1_t *r, uint32_t n_cigar, uint32_t *cigar) //
|
||||
mm_extra_t *p;
|
||||
if (n_cigar == 0) return;
|
||||
if (r->p == 0) {
|
||||
uint32_t capacity = n_cigar + sizeof(mm_extra_t);
|
||||
uint32_t capacity = n_cigar + sizeof(mm_extra_t)/4;
|
||||
kroundup32(capacity);
|
||||
r->p = (mm_extra_t*)calloc(capacity, 4);
|
||||
r->p->capacity = capacity;
|
||||
} else if (r->p->n_cigar + n_cigar + sizeof(mm_extra_t) > r->p->capacity) {
|
||||
r->p->capacity = r->p->n_cigar + n_cigar + sizeof(mm_extra_t);
|
||||
} else if (r->p->n_cigar + n_cigar + sizeof(mm_extra_t)/4 > r->p->capacity) {
|
||||
r->p->capacity = r->p->n_cigar + n_cigar + sizeof(mm_extra_t)/4;
|
||||
kroundup32(r->p->capacity);
|
||||
r->p = (mm_extra_t*)realloc(r->p, r->p->capacity * 4);
|
||||
}
|
||||
@@ -193,9 +313,8 @@ static void mm_append_cigar(mm_reg1_t *r, uint32_t n_cigar, uint32_t *cigar) //
|
||||
}
|
||||
}
|
||||
|
||||
static void mm_align_pair(void *km, const mm_mapopt_t *opt, int qlen, const uint8_t *qseq, int tlen, const uint8_t *tseq, const int8_t *mat, int w, int end_bonus, int flag, ksw_extz_t *ez)
|
||||
static void mm_align_pair(void *km, const mm_mapopt_t *opt, int qlen, const uint8_t *qseq, int tlen, const uint8_t *tseq, const uint8_t *junc, const int8_t *mat, int w, int end_bonus, int zdrop, int flag, ksw_extz_t *ez)
|
||||
{
|
||||
int zdrop = opt->zdrop;
|
||||
if (mm_dbg_flag & MM_DBG_PRINT_ALN_SEQ) {
|
||||
int i;
|
||||
fprintf(stderr, "===> q=(%d,%d), e=(%d,%d), bw=%d, flag=%d, zdrop=%d <===\n", opt->q, opt->q2, opt->e, opt->e2, w, flag, opt->zdrop);
|
||||
@@ -204,8 +323,11 @@ static void mm_align_pair(void *km, const mm_mapopt_t *opt, int qlen, const uint
|
||||
for (i = 0; i < qlen; ++i) fputc("ACGTN"[qseq[i]], stderr);
|
||||
fputc('\n', stderr);
|
||||
}
|
||||
if (opt->flag & MM_F_SPLICE)
|
||||
ksw_exts2_sse(km, qlen, qseq, tlen, tseq, 5, mat, opt->q, opt->e, opt->q2, opt->noncan, zdrop, flag, ez);
|
||||
if (opt->max_sw_mat > 0 && (int64_t)tlen * qlen > opt->max_sw_mat) {
|
||||
ksw_reset_extz(ez);
|
||||
ez->zdropped = 1;
|
||||
} else if (opt->flag & MM_F_SPLICE)
|
||||
ksw_exts2_sse(km, qlen, qseq, tlen, tseq, 5, mat, opt->q, opt->e, opt->q2, opt->noncan, zdrop, opt->junc_bonus, flag, junc, ez);
|
||||
else if (opt->q == opt->q2 && opt->e == opt->e2)
|
||||
ksw_extz2_sse(km, qlen, qseq, tlen, tseq, 5, mat, opt->q, opt->e, w, zdrop, end_bonus, flag, ez);
|
||||
else
|
||||
@@ -214,7 +336,7 @@ static void mm_align_pair(void *km, const mm_mapopt_t *opt, int qlen, const uint
|
||||
int i;
|
||||
fprintf(stderr, "score=%d, cigar=", ez->score);
|
||||
for (i = 0; i < ez->n_cigar; ++i)
|
||||
fprintf(stderr, "%d%c", ez->cigar[i]>>4, "MIDN"[ez->cigar[i]&0xf]);
|
||||
fprintf(stderr, "%d%c", ez->cigar[i]>>4, MM_CIGAR_STR[ez->cigar[i]&0xf]);
|
||||
fprintf(stderr, "\n");
|
||||
}
|
||||
}
|
||||
@@ -245,20 +367,30 @@ static inline void mm_adjust_minier(const mm_idx_t *mi, uint8_t *const qseq0[2],
|
||||
}
|
||||
}
|
||||
|
||||
static void mm_filter_bad_seeds(void *km, int as1, int cnt1, mm128_t *a, int min_gap, int diff_thres, int max_ext_len, int max_ext_cnt)
|
||||
static int *collect_long_gaps(void *km, int as1, int cnt1, mm128_t *a, int min_gap, int *n_)
|
||||
{
|
||||
int max_st, max_en, n, i, k, max, *K;
|
||||
int i, n, *K;
|
||||
*n_ = 0;
|
||||
for (i = 1, n = 0; i < cnt1; ++i) { // count the number of gaps longer than min_gap
|
||||
int gap = ((int32_t)a[as1 + i].y - a[as1 + i - 1].y) - ((int32_t)a[as1 + i].x - a[as1 + i - 1].x);
|
||||
if (gap < -min_gap || gap > min_gap) ++n;
|
||||
}
|
||||
if (n <= 1) return;
|
||||
if (n <= 1) return 0;
|
||||
K = (int*)kmalloc(km, n * sizeof(int));
|
||||
for (i = 1, n = 0; i < cnt1; ++i) { // store the positions of long gaps
|
||||
int gap = ((int32_t)a[as1 + i].y - a[as1 + i - 1].y) - ((int32_t)a[as1 + i].x - a[as1 + i - 1].x);
|
||||
if (gap < -min_gap || gap > min_gap)
|
||||
K[n++] = i;
|
||||
}
|
||||
*n_ = n;
|
||||
return K;
|
||||
}
|
||||
|
||||
static void mm_filter_bad_seeds(void *km, int as1, int cnt1, mm128_t *a, int min_gap, int diff_thres, int max_ext_len, int max_ext_cnt)
|
||||
{
|
||||
int max_st, max_en, n, i, k, max, *K;
|
||||
K = collect_long_gaps(km, as1, cnt1, a, min_gap, &n);
|
||||
if (K == 0) return;
|
||||
max = 0, max_st = max_en = -1;
|
||||
for (k = 0;; ++k) { // traverse long gaps
|
||||
int gap, l, n_ins = 0, n_del = 0, qs, rs, max_diff = 0, max_diff_l = -1;
|
||||
@@ -270,7 +402,7 @@ static void mm_filter_bad_seeds(void *km, int as1, int cnt1, mm128_t *a, int min
|
||||
if (k == n) break;
|
||||
}
|
||||
i = K[k];
|
||||
gap = ((int32_t)a[as1 + i].y - a[as1 + i - 1].y) - ((int32_t)a[as1 + i].x - a[as1 + i - 1].x);
|
||||
gap = ((int32_t)a[as1 + i].y - (int32_t)a[as1 + i - 1].y) - (int32_t)(a[as1 + i].x - a[as1 + i - 1].x);
|
||||
if (gap > 0) n_ins += gap;
|
||||
else n_del += -gap;
|
||||
qs = (int32_t)a[as1 + i - 1].y;
|
||||
@@ -278,7 +410,7 @@ static void mm_filter_bad_seeds(void *km, int as1, int cnt1, mm128_t *a, int min
|
||||
for (l = k + 1; l < n && l <= k + max_ext_cnt; ++l) {
|
||||
int j = K[l], diff;
|
||||
if ((int32_t)a[as1 + j].y - qs > max_ext_len || (int32_t)a[as1 + j].x - rs > max_ext_len) break;
|
||||
gap = ((int32_t)a[as1 + j].y - (int32_t)a[as1 + j - 1].y) - (a[as1 + j].x - a[as1 + j - 1].x);
|
||||
gap = ((int32_t)a[as1 + j].y - (int32_t)a[as1 + j - 1].y) - (int32_t)(a[as1 + j].x - a[as1 + j - 1].x);
|
||||
if (gap > 0) n_ins += gap;
|
||||
else n_del += -gap;
|
||||
diff = n_ins + n_del - abs(n_ins - n_del);
|
||||
@@ -291,37 +423,79 @@ static void mm_filter_bad_seeds(void *km, int as1, int cnt1, mm128_t *a, int min
|
||||
kfree(km, K);
|
||||
}
|
||||
|
||||
static void mm_fix_bad_ends(const mm_reg1_t *r, const mm128_t *a, int bw, int32_t *as, int32_t *cnt)
|
||||
static void mm_filter_bad_seeds_alt(void *km, int as1, int cnt1, mm128_t *a, int min_gap, int max_ext)
|
||||
{
|
||||
int32_t i, l;
|
||||
int n, k, *K;
|
||||
K = collect_long_gaps(km, as1, cnt1, a, min_gap, &n);
|
||||
if (K == 0) return;
|
||||
for (k = 0; k < n;) {
|
||||
int i = K[k], l;
|
||||
int gap1 = ((int32_t)a[as1 + i].y - (int32_t)a[as1 + i - 1].y) - ((int32_t)a[as1 + i].x - (int32_t)a[as1 + i - 1].x);
|
||||
int re1 = (int32_t)a[as1 + i].x;
|
||||
int qe1 = (int32_t)a[as1 + i].y;
|
||||
gap1 = gap1 > 0? gap1 : -gap1;
|
||||
for (l = k + 1; l < n; ++l) {
|
||||
int j = K[l], gap2, q_span_pre, rs2, qs2, m;
|
||||
if ((int32_t)a[as1 + j].y - qe1 > max_ext || (int32_t)a[as1 + j].x - re1 > max_ext) break;
|
||||
gap2 = ((int32_t)a[as1 + j].y - (int32_t)a[as1 + j - 1].y) - (int32_t)(a[as1 + j].x - a[as1 + j - 1].x);
|
||||
q_span_pre = a[as1 + j - 1].y >> 32 & 0xff;
|
||||
rs2 = (int32_t)a[as1 + j - 1].x + q_span_pre;
|
||||
qs2 = (int32_t)a[as1 + j - 1].y + q_span_pre;
|
||||
m = rs2 - re1 < qs2 - qe1? rs2 - re1 : qs2 - qe1;
|
||||
gap2 = gap2 > 0? gap2 : -gap2;
|
||||
if (m > gap1 + gap2) break;
|
||||
re1 = (int32_t)a[as1 + j].x;
|
||||
qe1 = (int32_t)a[as1 + j].y;
|
||||
gap1 = gap2;
|
||||
}
|
||||
if (l > k + 1) {
|
||||
int j, end = K[l - 1];
|
||||
for (j = K[k]; j < end; ++j)
|
||||
a[as1 + j].y |= MM_SEED_IGNORE;
|
||||
a[as1 + end].y |= MM_SEED_LONG_JOIN;
|
||||
}
|
||||
k = l;
|
||||
}
|
||||
kfree(km, K);
|
||||
}
|
||||
|
||||
static void mm_fix_bad_ends(const mm_reg1_t *r, const mm128_t *a, int bw, int min_match, int32_t *as, int32_t *cnt)
|
||||
{
|
||||
int32_t i, l, m;
|
||||
*as = r->as, *cnt = r->cnt;
|
||||
if (r->cnt < 3) return;
|
||||
l = a[r->as].y >> 32 & 0xff;
|
||||
m = l = a[r->as].y >> 32 & 0xff;
|
||||
for (i = r->as + 1; i < r->as + r->cnt - 1; ++i) {
|
||||
int32_t lq, lr, min, max;
|
||||
int32_t q_span = a[i].y >> 32 & 0xff;
|
||||
if (a[i].y & MM_SEED_LONG_JOIN) break;
|
||||
lr = (int32_t)a[i].x - (int32_t)a[i-1].x;
|
||||
lq = (int32_t)a[i].y - (int32_t)a[i-1].y;
|
||||
min = lr < lq? lr : lq;
|
||||
max = lr > lq? lr : lq;
|
||||
if (max - min > l >> 1) *as = i;
|
||||
l += min;
|
||||
if (l >= bw << 1) break;
|
||||
m += min < q_span? min : q_span;
|
||||
if (l >= bw << 1 || (m >= min_match && m >= bw) || m >= r->mlen >> 1) break;
|
||||
}
|
||||
*cnt = r->as + r->cnt - *as;
|
||||
l = a[r->as + r->cnt - 1].y >> 32 & 0xff;
|
||||
m = l = a[r->as + r->cnt - 1].y >> 32 & 0xff;
|
||||
for (i = r->as + r->cnt - 2; i > *as; --i) {
|
||||
int32_t lq, lr, min, max;
|
||||
int32_t q_span = a[i+1].y >> 32 & 0xff;
|
||||
if (a[i+1].y & MM_SEED_LONG_JOIN) break;
|
||||
lr = (int32_t)a[i+1].x - (int32_t)a[i].x;
|
||||
lq = (int32_t)a[i+1].y - (int32_t)a[i].y;
|
||||
min = lr < lq? lr : lq;
|
||||
max = lr > lq? lr : lq;
|
||||
if (max - min > l >> 1) *cnt = i + 1 - *as;
|
||||
l += min;
|
||||
if (l >= bw) break;
|
||||
m += min < q_span? min : q_span;
|
||||
if (l >= bw << 1 || (m >= min_match && m >= bw) || m >= r->mlen >> 1) break;
|
||||
}
|
||||
}
|
||||
|
||||
static void mm_max_stretch(const mm_mapopt_t *opt, const mm_reg1_t *r, const mm128_t *a, int32_t *as, int32_t *cnt)
|
||||
static void mm_max_stretch(const mm_reg1_t *r, const mm128_t *a, int32_t *as, int32_t *cnt)
|
||||
{
|
||||
int32_t i, score, max_score, len, max_i, max_len;
|
||||
|
||||
@@ -359,11 +533,16 @@ static int mm_seed_ext_score(void *km, const mm_mapopt_t *opt, const mm_idx_t *m
|
||||
qe = (uint32_t)a->y + 1, qs = qe - q_span;
|
||||
rs = rs - ext_len > 0? rs - ext_len : 0;
|
||||
qs = qs - ext_len > 0? qs - ext_len : 0;
|
||||
re = re + ext_len < mi->seq[rid].len? re + ext_len : mi->seq[rid].len;
|
||||
re = re + ext_len < (int32_t)mi->seq[rid].len? re + ext_len : mi->seq[rid].len;
|
||||
qe = qe + ext_len < qlen? qe + ext_len : qlen;
|
||||
tseq = (uint8_t*)kmalloc(km, re - rs);
|
||||
mm_idx_getseq(mi, rid, rs, re, tseq);
|
||||
qseq = qseq0[a->x>>63] + qs;
|
||||
if (opt->flag & MM_F_QSTRAND) {
|
||||
qseq = qseq0[0] + qs;
|
||||
mm_idx_getseq2(mi, a->x>>63, rid, rs, re, tseq);
|
||||
} else {
|
||||
qseq = qseq0[a->x>>63] + qs;
|
||||
mm_idx_getseq(mi, rid, rs, re, tseq);
|
||||
}
|
||||
qp = ksw_ll_qinit(km, 2, qe - qs, qseq, 5, mat);
|
||||
score = ksw_ll_i16(qp, re - rs, tseq, opt->q, opt->e, &q_off, &t_off);
|
||||
kfree(km, tseq);
|
||||
@@ -395,8 +574,8 @@ static void mm_align1(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int
|
||||
{
|
||||
int is_sr = !!(opt->flag & MM_F_SR), is_splice = !!(opt->flag & MM_F_SPLICE);
|
||||
int32_t rid = a[r->as].x<<1>>33, rev = a[r->as].x>>63, as1, cnt1;
|
||||
uint8_t *tseq, *qseq;
|
||||
int32_t i, l, bw, dropped = 0, extra_flag = 0, rs0, re0, qs0, qe0;
|
||||
uint8_t *tseq, *qseq, *junc;
|
||||
int32_t i, l, bw, bw_long, dropped = 0, extra_flag = 0, rs0, re0, qs0, qe0;
|
||||
int32_t rs, re, qs, qe;
|
||||
int32_t rs1, qs1, re1, qe1;
|
||||
int8_t mat[25];
|
||||
@@ -405,22 +584,26 @@ static void mm_align1(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int
|
||||
|
||||
r2->cnt = 0;
|
||||
if (r->cnt == 0) return;
|
||||
ksw_gen_simple_mat(5, mat, opt->a, opt->b);
|
||||
ksw_gen_simple_mat(5, mat, opt->a, opt->b, opt->sc_ambi);
|
||||
bw = (int)(opt->bw * 1.5 + 1.);
|
||||
bw_long = (int)(opt->bw_long * 1.5 + 1.);
|
||||
if (bw_long < bw) bw_long = bw;
|
||||
|
||||
if (is_sr && !(mi->flag & MM_I_HPC)) {
|
||||
mm_max_stretch(opt, r, a, &as1, &cnt1);
|
||||
mm_max_stretch(r, a, &as1, &cnt1);
|
||||
rs = (int32_t)a[as1].x + 1 - (int32_t)(a[as1].y>>32&0xff);
|
||||
qs = (int32_t)a[as1].y + 1 - (int32_t)(a[as1].y>>32&0xff);
|
||||
re = (int32_t)a[as1+cnt1-1].x + 1;
|
||||
qe = (int32_t)a[as1+cnt1-1].y + 1;
|
||||
} else {
|
||||
if (is_splice) {
|
||||
mm_fix_bad_ends_splice(km, opt, mi, r, mat, qlen, qseq0, a, &as1, &cnt1);
|
||||
} else {
|
||||
mm_fix_bad_ends(r, a, opt->bw, &as1, &cnt1);
|
||||
}
|
||||
if (!(opt->flag & MM_F_NO_END_FLT)) {
|
||||
if (is_splice)
|
||||
mm_fix_bad_ends_splice(km, opt, mi, r, mat, qlen, qseq0, a, &as1, &cnt1);
|
||||
else
|
||||
mm_fix_bad_ends(r, a, opt->bw, opt->min_chain_score * 2, &as1, &cnt1);
|
||||
} else as1 = r->as, cnt1 = r->cnt;
|
||||
mm_filter_bad_seeds(km, as1, cnt1, a, 10, 40, opt->max_gap>>1, 10);
|
||||
mm_filter_bad_seeds_alt(km, as1, cnt1, a, 30, opt->max_gap>>1);
|
||||
mm_adjust_minier(mi, qseq0, &a[as1], &rs, &qs);
|
||||
mm_adjust_minier(mi, qseq0, &a[as1 + cnt1 - 1], &re, &qe);
|
||||
}
|
||||
@@ -444,7 +627,7 @@ static void mm_align1(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int
|
||||
rs0 = rs - l > 0? rs - l : 0;
|
||||
l = qlen - qe;
|
||||
l += l * opt->a + opt->end_bonus > opt->q? (l * opt->a + opt->end_bonus - opt->q) / opt->e : 0;
|
||||
re0 = re + l < mi->seq[rid].len? re + l : mi->seq[rid].len;
|
||||
re0 = re + l < (int32_t)mi->seq[rid].len? re + l : mi->seq[rid].len;
|
||||
} else {
|
||||
// compute rs0 and qs0
|
||||
rs0 = (int32_t)a[r->as].x + 1 - (int32_t)(a[r->as].y>>32&0xff);
|
||||
@@ -459,6 +642,7 @@ static void mm_align1(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int
|
||||
if (++l > opt->min_cnt) {
|
||||
l = rs0 - x > qs0 - y? rs0 - x : qs0 - y;
|
||||
rs1 = rs0 - l, qs1 = qs0 - l;
|
||||
if (rs1 < 0) rs1 = 0; // not strictly necessary; better have this guard for explicit
|
||||
break;
|
||||
}
|
||||
}
|
||||
@@ -472,6 +656,7 @@ static void mm_align1(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int
|
||||
l = l < rs? l : rs;
|
||||
rs1 = rs1 > rs - l? rs1 : rs - l;
|
||||
rs0 = rs0 < rs1? rs0 : rs1;
|
||||
rs0 = rs0 < rs? rs0 : rs;
|
||||
} else rs0 = rs, qs0 = qs;
|
||||
// compute re0 and qe0
|
||||
re0 = (int32_t)a[r->as + r->cnt - 1].x + 1;
|
||||
@@ -488,13 +673,13 @@ static void mm_align1(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int
|
||||
}
|
||||
}
|
||||
}
|
||||
if (qe < qlen && re < mi->seq[rid].len) {
|
||||
if (qe < qlen && re < (int32_t)mi->seq[rid].len) {
|
||||
l = qlen - qe < opt->max_gap? qlen - qe : opt->max_gap;
|
||||
qe1 = qe1 < qe + l? qe1 : qe + l;
|
||||
qe0 = qe0 > qe1? qe0 : qe1; // at least include qe0
|
||||
l += l * opt->a > opt->q? (l * opt->a - opt->q) / opt->e : 0;
|
||||
l = l < opt->max_gap? l : opt->max_gap;
|
||||
l = l < mi->seq[rid].len - re? l : mi->seq[rid].len - re;
|
||||
l = l < (int32_t)mi->seq[rid].len - re? l : mi->seq[rid].len - re;
|
||||
re1 = re1 < re + l? re1 : re + l;
|
||||
re0 = re0 > re1? re0 : re1;
|
||||
} else re0 = re, qe0 = qe;
|
||||
@@ -510,13 +695,21 @@ static void mm_align1(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int
|
||||
|
||||
assert(re0 > rs0);
|
||||
tseq = (uint8_t*)kmalloc(km, re0 - rs0);
|
||||
junc = (uint8_t*)kmalloc(km, re0 - rs0);
|
||||
|
||||
if (qs > 0 && rs > 0) { // left extension
|
||||
qseq = &qseq0[rev][qs0];
|
||||
mm_idx_getseq(mi, rid, rs0, rs, tseq);
|
||||
if (qs > 0 && rs > 0) { // left extension; probably the condition can be changed to "qs > qs0 && rs > rs0"
|
||||
if (opt->flag & MM_F_QSTRAND) {
|
||||
qseq = &qseq0[0][qs0];
|
||||
mm_idx_getseq2(mi, rev, rid, rs0, rs, tseq);
|
||||
} else {
|
||||
qseq = &qseq0[rev][qs0];
|
||||
mm_idx_getseq(mi, rid, rs0, rs, tseq);
|
||||
}
|
||||
mm_idx_bed_junc(mi, rid, rs0, rs, junc);
|
||||
mm_seq_rev(qs - qs0, qseq);
|
||||
mm_seq_rev(rs - rs0, tseq);
|
||||
mm_align_pair(km, opt, qs - qs0, qseq, rs - rs0, tseq, mat, bw, opt->end_bonus, extra_flag|KSW_EZ_EXTZ_ONLY|KSW_EZ_RIGHT|KSW_EZ_REV_CIGAR, ez);
|
||||
mm_seq_rev(rs - rs0, junc);
|
||||
mm_align_pair(km, opt, qs - qs0, qseq, rs - rs0, tseq, junc, mat, bw, opt->end_bonus, r->split_inv? opt->zdrop_inv : opt->zdrop, extra_flag|KSW_EZ_EXTZ_ONLY|KSW_EZ_RIGHT|KSW_EZ_REV_CIGAR, ez);
|
||||
if (ez->n_cigar > 0) {
|
||||
mm_append_cigar(r, ez->n_cigar, ez->cigar);
|
||||
r->p->dp_score += ez->max;
|
||||
@@ -536,11 +729,18 @@ static void mm_align1(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int
|
||||
} else mm_adjust_minier(mi, qseq0, &a[as1 + i], &re, &qe);
|
||||
re1 = re, qe1 = qe;
|
||||
if (i == cnt1 - 1 || (a[as1+i].y&MM_SEED_LONG_JOIN) || (qe - qs >= opt->min_ksw_len && re - rs >= opt->min_ksw_len)) {
|
||||
int j, bw1 = bw;
|
||||
int j, bw1 = bw_long, zdrop_code;
|
||||
if (a[as1+i].y & MM_SEED_LONG_JOIN)
|
||||
bw1 = qe - qs > re - rs? qe - qs : re - rs;
|
||||
qseq = &qseq0[rev][qs];
|
||||
mm_idx_getseq(mi, rid, rs, re, tseq);
|
||||
// perform alignment
|
||||
if (opt->flag & MM_F_QSTRAND) {
|
||||
qseq = &qseq0[0][qs];
|
||||
mm_idx_getseq2(mi, rev, rid, rs, re, tseq);
|
||||
} else {
|
||||
qseq = &qseq0[rev][qs];
|
||||
mm_idx_getseq(mi, rid, rs, re, tseq);
|
||||
}
|
||||
mm_idx_bed_junc(mi, rid, rs, re, junc);
|
||||
if (is_sr) { // perform ungapped alignment
|
||||
assert(qe - qs == re - rs);
|
||||
ksw_reset_extz(ez);
|
||||
@@ -548,15 +748,24 @@ static void mm_align1(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int
|
||||
if (qseq[j] >= 4 || tseq[j] >= 4) ez->score += opt->e2;
|
||||
else ez->score += qseq[j] == tseq[j]? opt->a : -opt->b;
|
||||
}
|
||||
ez->cigar = ksw_push_cigar(km, &ez->n_cigar, &ez->m_cigar, ez->cigar, 0, qe - qs);
|
||||
ez->cigar = ksw_push_cigar(km, &ez->n_cigar, &ez->m_cigar, ez->cigar, MM_CIGAR_MATCH, qe - qs);
|
||||
} else { // perform normal gapped alignment
|
||||
mm_align_pair(km, opt, qe - qs, qseq, re - rs, tseq, mat, bw1, -1, extra_flag|KSW_EZ_APPROX_MAX, ez); // first pass: with approximate Z-drop
|
||||
mm_align_pair(km, opt, qe - qs, qseq, re - rs, tseq, junc, mat, bw1, -1, opt->zdrop, extra_flag|KSW_EZ_APPROX_MAX, ez); // first pass: with approximate Z-drop
|
||||
}
|
||||
if (mm_check_zdrop(qseq, tseq, ez->n_cigar, ez->cigar, mat, opt->q, opt->e, opt->zdrop))
|
||||
mm_align_pair(km, opt, qe - qs, qseq, re - rs, tseq, mat, bw1, -1, extra_flag, ez); // second pass: lift approximate
|
||||
// test Z-drop and inversion Z-drop
|
||||
if ((zdrop_code = mm_test_zdrop(km, opt, qseq, tseq, ez->n_cigar, ez->cigar, mat)) != 0)
|
||||
mm_align_pair(km, opt, qe - qs, qseq, re - rs, tseq, junc, mat, bw1, -1, zdrop_code == 2? opt->zdrop_inv : opt->zdrop, extra_flag, ez); // second pass: lift approximate
|
||||
// update CIGAR
|
||||
if (ez->n_cigar > 0)
|
||||
mm_append_cigar(r, ez->n_cigar, ez->cigar);
|
||||
if (ez->zdropped) { // truncated by Z-drop; TODO: sometimes Z-drop kicks in because the next seed placement is wrong. This can be fixed in principle.
|
||||
if (!r->p) {
|
||||
assert(ez->n_cigar == 0);
|
||||
uint32_t capacity = sizeof(mm_extra_t)/4;
|
||||
kroundup32(capacity);
|
||||
r->p = (mm_extra_t*)calloc(capacity, 4);
|
||||
r->p->capacity = capacity;
|
||||
}
|
||||
for (j = i - 1; j >= 0; --j)
|
||||
if ((int32_t)a[as1 + j].x <= rs + ez->max_t)
|
||||
break;
|
||||
@@ -565,8 +774,10 @@ static void mm_align1(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int
|
||||
r->p->dp_score += ez->max;
|
||||
re1 = rs + (ez->max_t + 1);
|
||||
qe1 = qs + (ez->max_q + 1);
|
||||
if (cnt1 - (j + 1) >= opt->min_cnt)
|
||||
mm_split_reg(r, r2, as1 + j + 1 - r->as, qlen, a);
|
||||
if (cnt1 - (j + 1) >= opt->min_cnt) {
|
||||
mm_split_reg(r, r2, as1 + j + 1 - r->as, qlen, a, !!(opt->flag&MM_F_QSTRAND));
|
||||
if (zdrop_code == 2) r2->split_inv = 1;
|
||||
}
|
||||
break;
|
||||
} else r->p->dp_score += ez->score;
|
||||
rs = re, qs = qe;
|
||||
@@ -574,9 +785,15 @@ static void mm_align1(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int
|
||||
}
|
||||
|
||||
if (!dropped && qe < qe0 && re < re0) { // right extension
|
||||
qseq = &qseq0[rev][qe];
|
||||
mm_idx_getseq(mi, rid, re, re0, tseq);
|
||||
mm_align_pair(km, opt, qe0 - qe, qseq, re0 - re, tseq, mat, bw, opt->end_bonus, extra_flag|KSW_EZ_EXTZ_ONLY, ez);
|
||||
if (opt->flag & MM_F_QSTRAND) {
|
||||
qseq = &qseq0[0][qe];
|
||||
mm_idx_getseq2(mi, rev, rid, re, re0, tseq);
|
||||
} else {
|
||||
qseq = &qseq0[rev][qe];
|
||||
mm_idx_getseq(mi, rid, re, re0, tseq);
|
||||
}
|
||||
mm_idx_bed_junc(mi, rid, re, re0, junc);
|
||||
mm_align_pair(km, opt, qe0 - qe, qseq, re0 - re, tseq, junc, mat, bw, opt->end_bonus, opt->zdrop, extra_flag|KSW_EZ_EXTZ_ONLY, ez);
|
||||
if (ez->n_cigar > 0) {
|
||||
mm_append_cigar(r, ez->n_cigar, ez->cigar);
|
||||
r->p->dp_score += ez->max;
|
||||
@@ -587,22 +804,29 @@ static void mm_align1(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int
|
||||
assert(qe1 <= qlen);
|
||||
|
||||
r->rs = rs1, r->re = re1;
|
||||
if (rev) r->qs = qlen - qe1, r->qe = qlen - qs1;
|
||||
else r->qs = qs1, r->qe = qe1;
|
||||
if (!rev || (opt->flag & MM_F_QSTRAND)) r->qs = qs1, r->qe = qe1;
|
||||
else r->qs = qlen - qe1, r->qe = qlen - qs1;
|
||||
|
||||
assert(re1 - rs1 <= re0 - rs0);
|
||||
if (r->p) {
|
||||
mm_idx_getseq(mi, rid, rs1, re1, tseq);
|
||||
mm_update_extra(r, &qseq0[r->rev][qs1], tseq, mat, opt->q, opt->e);
|
||||
if (opt->flag & MM_F_QSTRAND) {
|
||||
mm_idx_getseq2(mi, r->rev, rid, rs1, re1, tseq);
|
||||
qseq = &qseq0[0][qs1];
|
||||
} else {
|
||||
mm_idx_getseq(mi, rid, rs1, re1, tseq);
|
||||
qseq = &qseq0[r->rev][qs1];
|
||||
}
|
||||
mm_update_extra(r, qseq, tseq, mat, opt->q, opt->e, opt->flag & MM_F_EQX, !(opt->flag & MM_F_SR));
|
||||
if (rev && r->p->trans_strand)
|
||||
r->p->trans_strand ^= 3; // flip to the read strand
|
||||
}
|
||||
|
||||
kfree(km, tseq);
|
||||
kfree(km, junc);
|
||||
}
|
||||
|
||||
static int mm_align1_inv(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int qlen, uint8_t *qseq0[2], const mm_reg1_t *r1, const mm_reg1_t *r2, mm_reg1_t *r_inv, ksw_extz_t *ez)
|
||||
{
|
||||
{ // NB: this doesn't work with the qstrand mode
|
||||
int tl, ql, score, ret = 0, q_off, t_off;
|
||||
uint8_t *tseq, *qseq;
|
||||
int8_t mat[25];
|
||||
@@ -613,15 +837,15 @@ static int mm_align1_inv(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, i
|
||||
if (r1->id != r1->parent && r1->parent != MM_PARENT_TMP_PRI) return 0;
|
||||
if (r2->id != r2->parent && r2->parent != MM_PARENT_TMP_PRI) return 0;
|
||||
if (r1->rid != r2->rid || r1->rev != r2->rev) return 0;
|
||||
ql = r2->qs - r1->qe;
|
||||
ql = r1->rev? r1->qs - r2->qe : r2->qs - r1->qe;
|
||||
tl = r2->rs - r1->re;
|
||||
if (ql < opt->min_chain_score || ql > opt->max_gap) return 0;
|
||||
if (tl < opt->min_chain_score || tl > opt->max_gap) return 0;
|
||||
|
||||
ksw_gen_simple_mat(5, mat, opt->a, opt->b);
|
||||
ksw_gen_simple_mat(5, mat, opt->a, opt->b, opt->sc_ambi);
|
||||
tseq = (uint8_t*)kmalloc(km, tl);
|
||||
mm_idx_getseq(mi, r1->rid, r1->re, r2->rs, tseq);
|
||||
qseq = &qseq0[!r1->rev][qlen - r2->qs];
|
||||
qseq = r1->rev? &qseq0[0][r2->qe] : &qseq0[1][qlen - r2->qs];
|
||||
|
||||
mm_seq_rev(ql, qseq);
|
||||
mm_seq_rev(tl, tseq);
|
||||
@@ -632,7 +856,7 @@ static int mm_align1_inv(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, i
|
||||
mm_seq_rev(tl, tseq);
|
||||
if (score < opt->min_dp_max) goto end_align1_inv;
|
||||
q_off = ql - (q_off + 1), t_off = tl - (t_off + 1);
|
||||
mm_align_pair(km, opt, ql - q_off, qseq + q_off, tl - t_off, tseq + t_off, mat, (int)(opt->bw * 1.5), -1, KSW_EZ_EXTZ_ONLY, ez);
|
||||
mm_align_pair(km, opt, ql - q_off, qseq + q_off, tl - t_off, tseq + t_off, 0, mat, (int)(opt->bw * 1.5), -1, opt->zdrop, KSW_EZ_EXTZ_ONLY, ez);
|
||||
if (ez->n_cigar == 0) goto end_align1_inv; // should never be here
|
||||
mm_append_cigar(r_inv, ez->n_cigar, ez->cigar);
|
||||
r_inv->p->dp_score = ez->max;
|
||||
@@ -642,9 +866,16 @@ static int mm_align1_inv(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, i
|
||||
r_inv->rev = !r1->rev;
|
||||
r_inv->rid = r1->rid;
|
||||
r_inv->div = -1.0f;
|
||||
r_inv->qs = r1->qe + q_off, r_inv->qe = r_inv->qs + ez->max_q + 1;
|
||||
r_inv->rs = r1->re + t_off, r_inv->re = r_inv->rs + ez->max_t + 1;
|
||||
mm_update_extra(r_inv, &qseq[q_off], &tseq[t_off], mat, opt->q, opt->e);
|
||||
if (r_inv->rev == 0) {
|
||||
r_inv->qs = r2->qe + q_off;
|
||||
r_inv->qe = r_inv->qs + ez->max_q + 1;
|
||||
} else {
|
||||
r_inv->qe = r2->qs - q_off;
|
||||
r_inv->qs = r_inv->qe - (ez->max_q + 1);
|
||||
}
|
||||
r_inv->rs = r1->re + t_off;
|
||||
r_inv->re = r_inv->rs + ez->max_t + 1;
|
||||
mm_update_extra(r_inv, &qseq[q_off], &tseq[t_off], mat, opt->q, opt->e, opt->flag & MM_F_EQX, !(opt->flag & MM_F_SR));
|
||||
ret = 1;
|
||||
end_align1_inv:
|
||||
kfree(km, tseq);
|
||||
@@ -661,6 +892,71 @@ static inline mm_reg1_t *mm_insert_reg(const mm_reg1_t *r, int i, int *n_regs, m
|
||||
return regs;
|
||||
}
|
||||
|
||||
static inline void mm_count_gaps(const mm_reg1_t *r, int32_t *n_gap_, int32_t *n_gapo_)
|
||||
{
|
||||
uint32_t i;
|
||||
int32_t n_gapo = 0, n_gap = 0;
|
||||
*n_gap_ = *n_gapo_ = -1;
|
||||
if (r->p == 0) return;
|
||||
for (i = 0; i < r->p->n_cigar; ++i) {
|
||||
int32_t op = r->p->cigar[i] & 0xf, len = r->p->cigar[i] >> 4;
|
||||
if (op == MM_CIGAR_INS || op == MM_CIGAR_DEL)
|
||||
++n_gapo, n_gap += len;
|
||||
}
|
||||
*n_gap_ = n_gap, *n_gapo_ = n_gapo;
|
||||
}
|
||||
|
||||
double mm_event_identity(const mm_reg1_t *r)
|
||||
{
|
||||
int32_t n_gap, n_gapo;
|
||||
if (r->p == 0) return -1.0f;
|
||||
mm_count_gaps(r, &n_gap, &n_gapo);
|
||||
return (double)r->mlen / (r->blen + r->p->n_ambi - n_gap + n_gapo);
|
||||
}
|
||||
|
||||
static int32_t mm_recal_max_dp(const mm_reg1_t *r, double b2, int32_t match_sc)
|
||||
{
|
||||
uint32_t i;
|
||||
int32_t n_gap = 0, n_gapo = 0, n_mis;
|
||||
double gap_cost = 0.0;
|
||||
if (r->p == 0) return -1;
|
||||
for (i = 0; i < r->p->n_cigar; ++i) {
|
||||
int32_t op = r->p->cigar[i] & 0xf, len = r->p->cigar[i] >> 4;
|
||||
if (op == MM_CIGAR_INS || op == MM_CIGAR_DEL) {
|
||||
gap_cost += b2 + (double)mg_log2(1.0 + len);
|
||||
++n_gapo, n_gap += len;
|
||||
}
|
||||
}
|
||||
n_mis = r->blen + r->p->n_ambi - r->mlen - n_gap;
|
||||
return (int32_t)(match_sc * (r->mlen - b2 * n_mis - gap_cost) + .499);
|
||||
}
|
||||
|
||||
void mm_update_dp_max(int qlen, int n_regs, mm_reg1_t *regs, float frac, int a, int b)
|
||||
{
|
||||
int32_t max = -1, max2 = -1, i, max_i = -1;
|
||||
double div, b2;
|
||||
if (n_regs < 2) return;
|
||||
for (i = 0; i < n_regs; ++i) {
|
||||
mm_reg1_t *r = ®s[i];
|
||||
if (r->p == 0) continue;
|
||||
if (r->p->dp_max > max) max2 = max, max = r->p->dp_max, max_i = i;
|
||||
else if (r->p->dp_max > max2) max2 = r->p->dp_max;
|
||||
}
|
||||
if (max_i < 0 || max < 0 || max2 < 0) return;
|
||||
if (regs[max_i].qe - regs[max_i].qs < (double)qlen * frac) return;
|
||||
if (max2 < (double)max * frac) return;
|
||||
div = 1. - mm_event_identity(®s[max_i]);
|
||||
if (div < 0.02) div = 0.02;
|
||||
b2 = 0.5 / div; // max value: 25
|
||||
if (b2 * a < b) b2 = (double)a / b;
|
||||
for (i = 0; i < n_regs; ++i) {
|
||||
mm_reg1_t *r = ®s[i];
|
||||
if (r->p == 0) continue;
|
||||
r->p->dp_max = mm_recal_max_dp(r, b2, a);
|
||||
if (r->p->dp_max < 0) r->p->dp_max = 0;
|
||||
}
|
||||
}
|
||||
|
||||
mm_reg1_t *mm_align_skeleton(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int qlen, const char *qstr, int *n_regs_, mm_reg1_t *regs, mm128_t *a)
|
||||
{
|
||||
extern unsigned char seq_nt4_table[256];
|
||||
@@ -704,7 +1000,7 @@ mm_reg1_t *mm_align_skeleton(void *km, const mm_mapopt_t *opt, const mm_idx_t *m
|
||||
regs[i].p->trans_strand = opt->flag&MM_F_SPLICE_FOR? 1 : 2;
|
||||
}
|
||||
if (r2.cnt > 0) regs = mm_insert_reg(&r2, i, &n_regs, regs);
|
||||
if (!(opt->flag&(MM_F_SPLICE|MM_F_SR)) && !(opt->flag&(MM_F_FOR_ONLY|MM_F_REV_ONLY)) && i > 0) { // don't try inversion alignment for -xsplice or -xsr, or --for-only/rev-only
|
||||
if (i > 0 && regs[i].split_inv && !(opt->flag & MM_F_NO_INV)) {
|
||||
if (mm_align1_inv(km, opt, mi, qlen, qseq0, ®s[i-1], ®s[i], &r2, &ez)) {
|
||||
regs = mm_insert_reg(&r2, i, &n_regs, regs);
|
||||
++i; // skip the inserted INV alignment
|
||||
@@ -714,7 +1010,11 @@ mm_reg1_t *mm_align_skeleton(void *km, const mm_mapopt_t *opt, const mm_idx_t *m
|
||||
*n_regs_ = n_regs;
|
||||
kfree(km, qseq0[0]);
|
||||
kfree(km, ez.cigar);
|
||||
mm_filter_regs(km, opt, n_regs_, regs);
|
||||
mm_hit_sort_by_dp(km, n_regs_, regs);
|
||||
mm_filter_regs(opt, qlen, n_regs_, regs);
|
||||
if (!(opt->flag&MM_F_SR) && !opt->split_prefix && qlen >= opt->rank_min_len) {
|
||||
mm_update_dp_max(qlen, *n_regs_, regs, opt->rank_frac, opt->a, opt->b);
|
||||
mm_filter_regs(opt, qlen, n_regs_, regs);
|
||||
}
|
||||
mm_hit_sort(km, n_regs_, regs, opt->alt_drop);
|
||||
return regs;
|
||||
}
|
||||
|
||||
@@ -15,7 +15,7 @@ unsigned char seq_comp_table[256] = {
|
||||
48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63,
|
||||
64, 'T', 'V', 'G', 'H', 'E', 'F', 'C', 'D', 'I', 'J', 'M', 'L', 'K', 'N', 'O',
|
||||
'P', 'Q', 'Y', 'S', 'A', 'A', 'B', 'W', 'X', 'R', 'Z', 91, 92, 93, 94, 95,
|
||||
64, 't', 'v', 'g', 'h', 'e', 'f', 'c', 'd', 'i', 'j', 'm', 'l', 'k', 'n', 'o',
|
||||
96, 't', 'v', 'g', 'h', 'e', 'f', 'c', 'd', 'i', 'j', 'm', 'l', 'k', 'n', 'o',
|
||||
'p', 'q', 'y', 's', 'a', 'a', 'b', 'w', 'x', 'r', 'z', 123, 124, 125, 126, 127,
|
||||
128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143,
|
||||
144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159,
|
||||
@@ -39,7 +39,7 @@ mm_bseq_file_t *mm_bseq_open(const char *fn)
|
||||
{
|
||||
mm_bseq_file_t *fp;
|
||||
gzFile f;
|
||||
f = fn && strcmp(fn, "-")? gzopen(fn, "r") : gzdopen(fileno(stdin), "r");
|
||||
f = fn && strcmp(fn, "-")? gzopen(fn, "r") : gzdopen(0, "r");
|
||||
if (f == 0) return 0;
|
||||
fp = (mm_bseq_file_t*)calloc(1, sizeof(mm_bseq_file_t));
|
||||
fp->fp = f;
|
||||
@@ -54,21 +54,33 @@ void mm_bseq_close(mm_bseq_file_t *fp)
|
||||
free(fp);
|
||||
}
|
||||
|
||||
static inline void kseq2bseq(kseq_t *ks, mm_bseq1_t *s, int with_qual)
|
||||
static inline char *kstrdup(const kstring_t *s)
|
||||
{
|
||||
char *t;
|
||||
t = (char*)malloc(s->l + 1);
|
||||
memcpy(t, s->s, s->l + 1);
|
||||
return t;
|
||||
}
|
||||
|
||||
static inline void kseq2bseq(kseq_t *ks, mm_bseq1_t *s, int with_qual, int with_comment)
|
||||
{
|
||||
int i;
|
||||
s->name = strdup(ks->name.s);
|
||||
s->seq = strdup(ks->seq.s);
|
||||
for (i = 0; i < ks->seq.l; ++i) // convert U to T
|
||||
if (ks->name.l == 0)
|
||||
fprintf(stderr, "[WARNING]\033[1;31m empty sequence name in the input.\033[0m\n");
|
||||
s->name = kstrdup(&ks->name);
|
||||
s->seq = kstrdup(&ks->seq);
|
||||
for (i = 0; i < (int)ks->seq.l; ++i) // convert U to T
|
||||
if (s->seq[i] == 'u' || s->seq[i] == 'U')
|
||||
--s->seq[i];
|
||||
s->qual = with_qual && ks->qual.l? strdup(ks->qual.s) : 0;
|
||||
s->qual = with_qual && ks->qual.l? kstrdup(&ks->qual) : 0;
|
||||
s->comment = with_comment && ks->comment.l? kstrdup(&ks->comment) : 0;
|
||||
s->l_seq = ks->seq.l;
|
||||
}
|
||||
|
||||
mm_bseq1_t *mm_bseq_read2(mm_bseq_file_t *fp, int chunk_size, int with_qual, int frag_mode, int *n_)
|
||||
mm_bseq1_t *mm_bseq_read3(mm_bseq_file_t *fp, int64_t chunk_size, int with_qual, int with_comment, int frag_mode, int *n_)
|
||||
{
|
||||
int64_t size = 0;
|
||||
int ret;
|
||||
kvec_t(mm_bseq1_t) a = {0,0,0};
|
||||
kseq_t *ks = fp->ks;
|
||||
*n_ = 0;
|
||||
@@ -78,17 +90,17 @@ mm_bseq1_t *mm_bseq_read2(mm_bseq_file_t *fp, int chunk_size, int with_qual, int
|
||||
size = fp->s.l_seq;
|
||||
memset(&fp->s, 0, sizeof(mm_bseq1_t));
|
||||
}
|
||||
while (kseq_read(ks) >= 0) {
|
||||
while ((ret = kseq_read(ks)) >= 0) {
|
||||
mm_bseq1_t *s;
|
||||
assert(ks->seq.l <= INT32_MAX);
|
||||
if (a.m == 0) kv_resize(mm_bseq1_t, 0, a, 256);
|
||||
kv_pushp(mm_bseq1_t, 0, a, &s);
|
||||
kseq2bseq(ks, s, with_qual);
|
||||
kseq2bseq(ks, s, with_qual, with_comment);
|
||||
size += s->l_seq;
|
||||
if (size >= chunk_size) {
|
||||
if (frag_mode && a.a[a.n-1].l_seq < CHECK_PAIR_THRES) {
|
||||
while (kseq_read(ks) >= 0) {
|
||||
kseq2bseq(ks, &fp->s, with_qual);
|
||||
while ((ret = kseq_read(ks)) >= 0) {
|
||||
kseq2bseq(ks, &fp->s, with_qual, with_comment);
|
||||
if (mm_qname_same(fp->s.name, a.a[a.n-1].name)) {
|
||||
kv_push(mm_bseq1_t, 0, a, fp->s);
|
||||
memset(&fp->s, 0, sizeof(mm_bseq1_t));
|
||||
@@ -98,16 +110,25 @@ mm_bseq1_t *mm_bseq_read2(mm_bseq_file_t *fp, int chunk_size, int with_qual, int
|
||||
break;
|
||||
}
|
||||
}
|
||||
if (ret < -1) {
|
||||
if (a.n) fprintf(stderr, "[WARNING]\033[1;31m failed to parse the FASTA/FASTQ record next to '%s'. Continue anyway.\033[0m\n", a.a[a.n-1].name);
|
||||
else fprintf(stderr, "[WARNING]\033[1;31m failed to parse the first FASTA/FASTQ record. Continue anyway.\033[0m\n");
|
||||
}
|
||||
*n_ = a.n;
|
||||
return a.a;
|
||||
}
|
||||
|
||||
mm_bseq1_t *mm_bseq_read(mm_bseq_file_t *fp, int chunk_size, int with_qual, int *n_)
|
||||
mm_bseq1_t *mm_bseq_read2(mm_bseq_file_t *fp, int64_t chunk_size, int with_qual, int frag_mode, int *n_)
|
||||
{
|
||||
return mm_bseq_read3(fp, chunk_size, with_qual, 0, frag_mode, n_);
|
||||
}
|
||||
|
||||
mm_bseq1_t *mm_bseq_read(mm_bseq_file_t *fp, int64_t chunk_size, int with_qual, int *n_)
|
||||
{
|
||||
return mm_bseq_read2(fp, chunk_size, with_qual, 0, n_);
|
||||
}
|
||||
|
||||
mm_bseq1_t *mm_bseq_read_frag(int n_fp, mm_bseq_file_t **fp, int chunk_size, int with_qual, int *n_)
|
||||
mm_bseq1_t *mm_bseq_read_frag2(int n_fp, mm_bseq_file_t **fp, int64_t chunk_size, int with_qual, int with_comment, int *n_)
|
||||
{
|
||||
int i;
|
||||
int64_t size = 0;
|
||||
@@ -128,7 +149,7 @@ mm_bseq1_t *mm_bseq_read_frag(int n_fp, mm_bseq_file_t **fp, int chunk_size, int
|
||||
for (i = 0; i < n_fp; ++i) {
|
||||
mm_bseq1_t *s;
|
||||
kv_pushp(mm_bseq1_t, 0, a, &s);
|
||||
kseq2bseq(fp[i]->ks, s, with_qual);
|
||||
kseq2bseq(fp[i]->ks, s, with_qual, with_comment);
|
||||
size += s->l_seq;
|
||||
}
|
||||
if (size >= chunk_size) break;
|
||||
@@ -137,6 +158,11 @@ mm_bseq1_t *mm_bseq_read_frag(int n_fp, mm_bseq_file_t **fp, int chunk_size, int
|
||||
return a.a;
|
||||
}
|
||||
|
||||
mm_bseq1_t *mm_bseq_read_frag(int n_fp, mm_bseq_file_t **fp, int64_t chunk_size, int with_qual, int *n_)
|
||||
{
|
||||
return mm_bseq_read_frag2(n_fp, fp, chunk_size, with_qual, 0, n_);
|
||||
}
|
||||
|
||||
int mm_bseq_eof(mm_bseq_file_t *fp)
|
||||
{
|
||||
return (ks_eof(fp->ks->f) && fp->s.seq == 0);
|
||||
|
||||
@@ -13,14 +13,16 @@ typedef struct mm_bseq_file_s mm_bseq_file_t;
|
||||
|
||||
typedef struct {
|
||||
int l_seq, rid;
|
||||
char *name, *seq, *qual;
|
||||
char *name, *seq, *qual, *comment;
|
||||
} mm_bseq1_t;
|
||||
|
||||
mm_bseq_file_t *mm_bseq_open(const char *fn);
|
||||
void mm_bseq_close(mm_bseq_file_t *fp);
|
||||
mm_bseq1_t *mm_bseq_read2(mm_bseq_file_t *fp, int chunk_size, int with_qual, int frag_mode, int *n_);
|
||||
mm_bseq1_t *mm_bseq_read(mm_bseq_file_t *fp, int chunk_size, int with_qual, int *n_);
|
||||
mm_bseq1_t *mm_bseq_read_frag(int n_fp, mm_bseq_file_t **fp, int chunk_size, int with_qual, int *n_);
|
||||
mm_bseq1_t *mm_bseq_read3(mm_bseq_file_t *fp, int64_t chunk_size, int with_qual, int with_comment, int frag_mode, int *n_);
|
||||
mm_bseq1_t *mm_bseq_read2(mm_bseq_file_t *fp, int64_t chunk_size, int with_qual, int frag_mode, int *n_);
|
||||
mm_bseq1_t *mm_bseq_read(mm_bseq_file_t *fp, int64_t chunk_size, int with_qual, int *n_);
|
||||
mm_bseq1_t *mm_bseq_read_frag2(int n_fp, mm_bseq_file_t **fp, int64_t chunk_size, int with_qual, int with_comment, int *n_);
|
||||
mm_bseq1_t *mm_bseq_read_frag(int n_fp, mm_bseq_file_t **fp, int64_t chunk_size, int with_qual, int *n_);
|
||||
int mm_bseq_eof(mm_bseq_file_t *fp);
|
||||
|
||||
extern unsigned char seq_nt4_table[256];
|
||||
|
||||
@@ -1,157 +0,0 @@
|
||||
#include <stdint.h>
|
||||
#include <string.h>
|
||||
#include <stdio.h>
|
||||
#include "minimap.h"
|
||||
#include "mmpriv.h"
|
||||
#include "kalloc.h"
|
||||
|
||||
static const char LogTable256[256] = {
|
||||
#define LT(n) n, n, n, n, n, n, n, n, n, n, n, n, n, n, n, n
|
||||
-1, 0, 1, 1, 2, 2, 2, 2, 3, 3, 3, 3, 3, 3, 3, 3,
|
||||
LT(4), LT(5), LT(5), LT(6), LT(6), LT(6), LT(6),
|
||||
LT(7), LT(7), LT(7), LT(7), LT(7), LT(7), LT(7), LT(7)
|
||||
};
|
||||
|
||||
static inline int ilog2_32(uint32_t v)
|
||||
{
|
||||
register uint32_t t, tt;
|
||||
if ((tt = v>>16)) return (t = tt>>8) ? 24 + LogTable256[t] : 16 + LogTable256[tt];
|
||||
return (t = v>>8) ? 8 + LogTable256[t] : LogTable256[v];
|
||||
}
|
||||
|
||||
mm128_t *mm_chain_dp(int max_dist_x, int max_dist_y, int bw, int max_skip, int min_cnt, int min_sc, int is_cdna, int n_segs, int64_t n, mm128_t *a, int *n_u_, uint64_t **_u, void *km)
|
||||
{ // TODO: make sure this works when n has more than 32 bits
|
||||
int32_t k, *f, *p, *t, *v, n_u, n_v;
|
||||
int64_t i, j, st = 0;
|
||||
uint64_t *u, *u2, sum_qspan = 0;
|
||||
float avg_qspan;
|
||||
mm128_t *b, *w;
|
||||
|
||||
if (_u) *_u = 0, *n_u_ = 0;
|
||||
f = (int32_t*)kmalloc(km, n * 4);
|
||||
p = (int32_t*)kmalloc(km, n * 4);
|
||||
t = (int32_t*)kmalloc(km, n * 4);
|
||||
v = (int32_t*)kmalloc(km, n * 4);
|
||||
memset(t, 0, n * 4);
|
||||
|
||||
for (i = 0; i < n; ++i) sum_qspan += a[i].y>>32&0xff;
|
||||
avg_qspan = (float)sum_qspan / n;
|
||||
|
||||
// fill the score and backtrack arrays
|
||||
for (i = 0; i < n; ++i) {
|
||||
uint64_t ri = a[i].x;
|
||||
int64_t max_j = -1;
|
||||
int32_t qi = (int32_t)a[i].y, q_span = a[i].y>>32&0xff; // NB: only 8 bits of span is used!!!
|
||||
int32_t max_f = q_span, n_skip = 0, min_d;
|
||||
int32_t sidi = (a[i].y & MM_SEED_SEG_MASK) >> MM_SEED_SEG_SHIFT;
|
||||
while (st < i && ri - a[st].x > max_dist_x) ++st;
|
||||
for (j = i - 1; j >= st; --j) {
|
||||
int64_t dr = ri - a[j].x;
|
||||
int32_t dq = qi - (int32_t)a[j].y, dd, sc, log_dd;
|
||||
int32_t sidj = (a[j].y & MM_SEED_SEG_MASK) >> MM_SEED_SEG_SHIFT;
|
||||
if ((sidi == sidj && dr == 0) || dq <= 0) continue; // don't skip if an anchor is used by multiple segments; see below
|
||||
if ((sidi == sidj && dq > max_dist_y) || dq > max_dist_x) continue;
|
||||
dd = dr > dq? dr - dq : dq - dr;
|
||||
if (sidi == sidj && dd > bw) continue;
|
||||
if (n_segs > 1 && !is_cdna && sidi == sidj && dr > max_dist_y) continue;
|
||||
min_d = dq < dr? dq : dr;
|
||||
sc = min_d > q_span? q_span : dq < dr? dq : dr;
|
||||
log_dd = dd? ilog2_32(dd) : 0;
|
||||
if (is_cdna || sidi != sidj) {
|
||||
int c_log, c_lin;
|
||||
c_lin = (int)(dd * .01 * avg_qspan);
|
||||
c_log = log_dd;
|
||||
if (sidi != sidj && dr == 0) ++sc; // possibly due to overlapping paired ends; give a minor bonus
|
||||
else if (dr > dq || sidi != sidj) sc -= c_lin < c_log? c_lin : c_log;
|
||||
else sc -= c_lin + (c_log>>1);
|
||||
} else sc -= (int)(dd * .01 * avg_qspan) + (log_dd>>1);
|
||||
sc += f[j];
|
||||
if (sc > max_f) {
|
||||
max_f = sc, max_j = j;
|
||||
if (n_skip > 0) --n_skip;
|
||||
} else if (t[j] == i) {
|
||||
if (++n_skip > max_skip)
|
||||
break;
|
||||
}
|
||||
if (p[j] >= 0) t[p[j]] = i;
|
||||
}
|
||||
f[i] = max_f, p[i] = max_j;
|
||||
v[i] = max_j >= 0 && v[max_j] > max_f? v[max_j] : max_f; // v[] keeps the peak score up to i; f[] is the score ending at i, not always the peak
|
||||
}
|
||||
|
||||
// find the ending positions of chains
|
||||
memset(t, 0, n * 4);
|
||||
for (i = 0; i < n; ++i)
|
||||
if (p[i] >= 0) t[p[i]] = 1;
|
||||
for (i = n_u = 0; i < n; ++i)
|
||||
if (t[i] == 0 && v[i] >= min_sc)
|
||||
++n_u;
|
||||
if (n_u == 0) {
|
||||
kfree(km, a); kfree(km, f); kfree(km, p); kfree(km, t); kfree(km, v);
|
||||
return 0;
|
||||
}
|
||||
u = (uint64_t*)kmalloc(km, n_u * 8);
|
||||
for (i = n_u = 0; i < n; ++i) {
|
||||
if (t[i] == 0 && v[i] >= min_sc) {
|
||||
j = i;
|
||||
while (j >= 0 && f[j] < v[j]) j = p[j]; // find the peak that maximizes f[]
|
||||
if (j < 0) j = i; // TODO: this should really be assert(j>=0)
|
||||
u[n_u++] = (uint64_t)f[j] << 32 | j;
|
||||
}
|
||||
}
|
||||
radix_sort_64(u, u + n_u);
|
||||
for (i = 0; i < n_u>>1; ++i) { // reverse, s.t. the highest scoring chain is the first
|
||||
uint64_t t = u[i];
|
||||
u[i] = u[n_u - i - 1], u[n_u - i - 1] = t;
|
||||
}
|
||||
|
||||
// backtrack
|
||||
memset(t, 0, n * 4);
|
||||
for (i = n_v = k = 0; i < n_u; ++i) { // starting from the highest score
|
||||
int32_t n_v0 = n_v, k0 = k;
|
||||
j = (int32_t)u[i];
|
||||
do {
|
||||
v[n_v++] = j;
|
||||
t[j] = 1;
|
||||
j = p[j];
|
||||
} while (j >= 0 && t[j] == 0);
|
||||
if (j < 0) {
|
||||
if (n_v - n_v0 >= min_cnt) u[k++] = u[i]>>32<<32 | (n_v - n_v0);
|
||||
} else if ((int32_t)(u[i]>>32) - f[j] >= min_sc) {
|
||||
if (n_v - n_v0 >= min_cnt) u[k++] = ((u[i]>>32) - f[j]) << 32 | (n_v - n_v0);
|
||||
}
|
||||
if (k0 == k) n_v = n_v0; // no new chain added, reset
|
||||
}
|
||||
*n_u_ = n_u = k, *_u = u; // NB: note that u[] may not be sorted by score here
|
||||
|
||||
// free temporary arrays
|
||||
kfree(km, f); kfree(km, p); kfree(km, t);
|
||||
|
||||
// write the result to b[]
|
||||
b = (mm128_t*)kmalloc(km, n_v * sizeof(mm128_t));
|
||||
for (i = 0, k = 0; i < n_u; ++i) {
|
||||
int32_t k0 = k, ni = (int32_t)u[i];
|
||||
for (j = 0; j < ni; ++j)
|
||||
b[k] = a[v[k0 + (ni - j - 1)]], ++k;
|
||||
}
|
||||
kfree(km, v);
|
||||
|
||||
// sort u[] and a[] by a[].x, such that adjacent chains may be joined (required by mm_join_long)
|
||||
w = (mm128_t*)kmalloc(km, n_u * sizeof(mm128_t));
|
||||
for (i = k = 0; i < n_u; ++i) {
|
||||
w[i].x = b[k].x, w[i].y = (uint64_t)k<<32|i;
|
||||
k += (int32_t)u[i];
|
||||
}
|
||||
radix_sort_128x(w, w + n_u);
|
||||
u2 = (uint64_t*)kmalloc(km, n_u * 8);
|
||||
for (i = k = 0; i < n_u; ++i) {
|
||||
int32_t j = (int32_t)w[i].y, n = (int32_t)u[j];
|
||||
u2[i] = u[j];
|
||||
memcpy(&a[k], &b[w[i].y>>32], n * sizeof(mm128_t));
|
||||
k += n;
|
||||
}
|
||||
memcpy(u, u2, n_u * 8);
|
||||
memcpy(b, a, k * sizeof(mm128_t)); // write _a_ to _b_ and deallocate _a_ because _a_ is oversized, sometimes a lot
|
||||
kfree(km, a); kfree(km, w); kfree(km, u2);
|
||||
return b;
|
||||
}
|
||||
@@ -0,0 +1,30 @@
|
||||
## Contributor Code of Conduct
|
||||
|
||||
As contributors and maintainers of this project, we pledge to respect all
|
||||
people who contribute through reporting issues, posting feature requests,
|
||||
updating documentation, submitting pull requests or patches, and other
|
||||
activities.
|
||||
|
||||
We are committed to making participation in this project a harassment-free
|
||||
experience for everyone, regardless of level of experience, gender, gender
|
||||
identity and expression, sexual orientation, disability, personal appearance,
|
||||
body size, race, age, or religion.
|
||||
|
||||
Examples of unacceptable behavior by participants include the use of sexual
|
||||
language or imagery, derogatory comments or personal attacks, trolling, public
|
||||
or private harassment, insults, or other unprofessional conduct.
|
||||
|
||||
Project maintainers have the right and responsibility to remove, edit, or
|
||||
reject comments, commits, code, wiki edits, issues, and other contributions
|
||||
that are not aligned to this Code of Conduct. Project maintainers or
|
||||
contributors who do not follow the Code of Conduct may be removed from the
|
||||
project team.
|
||||
|
||||
Instances of abusive, harassing, or otherwise unacceptable behavior may be
|
||||
reported by opening an issue or contacting the maintainer via email.
|
||||
|
||||
This Code of Conduct is adapted from the [Contributor Covenant][cc], [version
|
||||
1.0.0][v1].
|
||||
|
||||
[cc]: http://contributor-covenant.org/
|
||||
[v1]: http://contributor-covenant.org/version/1/0/0/
|
||||
+243
@@ -0,0 +1,243 @@
|
||||
## Table of Contents
|
||||
|
||||
- [Introduction & Installation](#intro)
|
||||
- [Mapping Genomic Reads](#map-reads)
|
||||
* [Mapping long reads](#map-pb)
|
||||
* [Mapping Illumina paired-end reads](#map-sr)
|
||||
* [Evaluating mapping accuracy with simulated reads (for developers)](#mapeval)
|
||||
- [Mapping Long RNA-seq Reads](#map-rna)
|
||||
* [Mapping Nanopore 2D cDNA reads](#map-ont-cdna-2d)
|
||||
* [Mapping Nanopore direct-RNA reads](#map-direct-rna)
|
||||
* [Mapping PacBio Iso-seq reads](#map-iso-seq)
|
||||
- [Full-Genome Alignment](#genome-aln)
|
||||
* [Intra-species assembly alignment](#asm-to-ref)
|
||||
* [Cross-species full-genome alignment](#x-species)
|
||||
* [Eyeballing alignment](#view-aln)
|
||||
* [Calling variants from assembly-to-reference alignment](#asm-var)
|
||||
* [Constructing self-homology map](#hom-map)
|
||||
* [Lift Over (for developers)](#liftover)
|
||||
- [Read Overlap](#read-overlap)
|
||||
* [Long-read overlap](#long-read-overlap)
|
||||
* [Evaluating overlap sensitivity (for developers)](#ov-eval)
|
||||
|
||||
## <a name="intro"></a>Introduction & Installation
|
||||
|
||||
This cookbook walks you through a variety of applications of minimap2 and its
|
||||
companion script `paftools.js`. All data here are freely available from the
|
||||
minimap2 release page at version tag [v2.10][v2.10]. Some examples only work
|
||||
with v2.10 or later.
|
||||
|
||||
To acquire the data used in this cookbook and to install minimap2 and paftools,
|
||||
please follow the command lines below:
|
||||
```sh
|
||||
# install minimap2 executables
|
||||
curl -L https://github.com/lh3/minimap2/releases/download/v2.21/minimap2-2.21_x64-linux.tar.bz2 | tar jxf -
|
||||
cp minimap2-2.21_x64-linux/{minimap2,k8,paftools.js} . # copy executables
|
||||
export PATH="$PATH:"`pwd` # put the current directory on PATH
|
||||
# download example datasets
|
||||
curl -L https://github.com/lh3/minimap2/releases/download/v2.10/cookbook-data.tgz | tar zxf -
|
||||
```
|
||||
|
||||
## <a name="map-reads"></a>Mapping Genomic Reads
|
||||
|
||||
### <a name="map-pb"></a>Mapping long reads
|
||||
```sh
|
||||
minimap2 -ax map-pb -t4 ecoli_ref.fa ecoli_p6_25x_canu.fa > mapped.sam
|
||||
```
|
||||
Alternatively, you can create a minimap2 index first and then map:
|
||||
```sh
|
||||
minimap2 -x map-pb -d ecoli-pb.mmi ecoli_ref.fa # create an index
|
||||
minimap2 -ax map-pb ecoli-pb.mmi ecoli_p6_25x_canu.fa > mapped.sam
|
||||
```
|
||||
This will save you a couple of minutes when you map against the human genome.
|
||||
**HOWEVER**, key algorithm parameters such as the k-mer length and window
|
||||
size can't be changed after indexing. Minimap2 will give you a warning if
|
||||
parameters used in a pre-built index doesn't match parameters on the command
|
||||
line. **Please always make sure you are using an intended pre-built index.**
|
||||
|
||||
### <a name="map-sr"></a>Mapping Illumina paired-end reads:
|
||||
```sh
|
||||
minimap2 -ax sr -t4 ecoli_ref.fa ecoli_mason_1.fq ecoli_mason_2.fq > mapped-sr.sam
|
||||
```
|
||||
|
||||
### <a name="mapeval"></a>Evaluating mapping accuracy with simulated reads (for developers)
|
||||
```sh
|
||||
minimap2 -ax sr ecoli_ref.fa ecoli_mason_1.fq ecoli_mason_2.fq | paftools.js mapeval -
|
||||
```
|
||||
The output is:
|
||||
```
|
||||
Q 60 19712 0 0.000000000 19712
|
||||
Q 0 282 219 0.010953286 19994
|
||||
U 6
|
||||
```
|
||||
where a `U`-line gives the number of unmapped reads (for SAM input only); a
|
||||
`Q`-line gives:
|
||||
|
||||
1. Mapping quality (mapQ) threshold
|
||||
2. Number of mapped reads between this threshold and the previous mapQ threshold.
|
||||
3. Number of wrong mappings in the same mapQ interval
|
||||
4. Accumulative mapping error rate
|
||||
5. Accumulative number of mappings
|
||||
|
||||
For `paftools.js mapeval` to work, you need to encode the true read positions
|
||||
in read names in the right format. For [pbsim2][pbsim] and [mason2][mason2], we
|
||||
provide scripts to generate the right format. Simulated reads in this cookbook
|
||||
were created with the following command lines:
|
||||
```sh
|
||||
# in the pbsim2 source code directory:
|
||||
src/pbsim --depth 1 --length-min 5000 --length-mean 20000 --accuracy-mean 0.95 --hmm_model data/R94.model ../ecoli_ref.fa
|
||||
paftools.js pbsim2fq ../ecoli_ref.fa.fai sd_0001.maf > ../ecoli_pbsim.fa
|
||||
|
||||
# mason2 simulation
|
||||
mason_simulator --illumina-prob-mismatch-scale 2.5 -ir ecoli_ref.fa -n 10000 -o tmp-l.fq -or tmp-r.fq -oa tmp.sam
|
||||
paftools.js mason2fq tmp.sam | seqtk seq -1 > ecoli_mason_1.fq
|
||||
paftools.js mason2fq tmp.sam | seqtk seq -2 > ecoli_mason_2.fq
|
||||
```
|
||||
|
||||
|
||||
|
||||
## <a name="map-rna"></a>Mapping Long RNA-seq Reads
|
||||
|
||||
### <a name="map-ont-cdna-2d"></a>Mapping Nanopore 2D cDNA reads
|
||||
```sh
|
||||
minimap2 -ax splice SIRV_E2.fa SIRV_ont-cdna.fa > aln.sam
|
||||
```
|
||||
You can compare the alignment to the true annotations with:
|
||||
```sh
|
||||
paftools.js junceval SIRV_E2C.gtf aln.sam
|
||||
```
|
||||
It gives the percentage of introns found in the annotation. For SIRV data, it
|
||||
is possible to achieve higher junction accuracy with
|
||||
```sh
|
||||
minimap2 -ax splice --splice-flank=no SIRV_E2.fa SIRV_ont-cdna.fa | paftools.js junceval SIRV_E2C.gtf
|
||||
```
|
||||
This is because minimap2 models one additional evolutionarily conserved base
|
||||
around a canonical junction, but SIRV doesn't honor this signal. Option
|
||||
`--splice-flank=no` asks minimap2 no to model this additional base.
|
||||
|
||||
In the output a tag `ts:A:+` indicates that the read strand is the same as the
|
||||
transcript strand; `ts:A:-` indicates the read strand is opposite to the
|
||||
transcript strand. This tag is inferred from the GT-AG signal and is thus only
|
||||
available to spliced reads.
|
||||
|
||||
### <a name="map-direct-rna"></a>Mapping Nanopore direct-RNA reads
|
||||
```sh
|
||||
minimap2 -ax splice -k14 -uf SIRV_E2.fa SIRV_ont-drna.fa > aln.sam
|
||||
```
|
||||
Direct-RNA reads are noisier, so we use a shorter k-mer for improved
|
||||
sensitivity. Here, option `-uf` forces minimap2 to map reads to the forward
|
||||
transcript strand only because direct-RNA reads are stranded. Again, applying
|
||||
`--splice-flank=no` helps junction accuracy for SIRV data.
|
||||
|
||||
### <a name="map-iso-seq"></a>Mapping PacBio Iso-seq reads
|
||||
```sh
|
||||
minimap2 -ax splice -uf -C5 SIRV_E2.fa SIRV_iso-seq.fq > aln.sam
|
||||
```
|
||||
Option `-C5` reduces the penalty on non-canonical splicing sites. It helps
|
||||
to align such sites correctly for data with low error rate such as Iso-seq
|
||||
reads and traditional cDNAs. On this example, minimap2 makes one junction
|
||||
error. Applying `--splice-flank=no` fixes this alignment error.
|
||||
|
||||
Note that the command line above is optimized for the final Iso-seq reads.
|
||||
PacBio's Iso-seq pipeline produces intermediate sequences at varying quality.
|
||||
For example, some intermediate reads are not stranded. For these reads, option
|
||||
`-uf` will lead to more errors. Please revise the minimap2 command line
|
||||
accordingly.
|
||||
|
||||
|
||||
|
||||
## <a name="genome-aln"></a>Full-Genome Alignment
|
||||
|
||||
### <a name="asm-to-ref"></a>Intra-species assembly alignment
|
||||
```sh
|
||||
# option "--cs" is recommended as paftools.js may need it
|
||||
minimap2 -cx asm5 --cs ecoli_ref.fa ecoli_canu.fa > ecoli_canu.paf
|
||||
```
|
||||
Here `ecoli_canu.fa` is the Canu assembly of `ecoli_p6_25x_canu.fa`. This
|
||||
command line outputs alignments in the [PAF format][paf]. Use `-a` instead of
|
||||
`-c` to get output in the SAM format.
|
||||
|
||||
### <a name="x-species"></a>Cross-species full-genome alignment
|
||||
```sh
|
||||
minimap2 -cx asm20 --cs ecoli_ref.fa ecoli_O104:H4.fa > ecoli_O104:H4.paf
|
||||
sort -k6,6 -k8,8n ecoli_O104:H4.paf | paftools.js call -f ecoli_ref.fa -L10000 -l1000 - > out.vcf
|
||||
```
|
||||
Minimap2 has three presets for full-genome alignment: "asm5" for sequence
|
||||
divergence below 1%, "asm10" for divergence around a couple of percent and
|
||||
"asm20" for divergence not more than 10%. In theory, with the right setting,
|
||||
minimap2 should work for sequence pairs with sequence divergence up to ~15%,
|
||||
but this has not been carefully evaluated.
|
||||
|
||||
### <a name="view-aln"></a>Eyeballing alignment
|
||||
```sh
|
||||
# option "--cs" required; minimap2-r741 or higher required for the "asm20" preset
|
||||
minimap2 -cx asm20 --cs ecoli_ref.fa ecoli_O104:H4.fa | paftools.js view - | less -S
|
||||
```
|
||||
This prints the alignment in a BLAST-like format.
|
||||
|
||||
### <a name="asm-var"></a>Calling variants from assembly-to-reference alignment
|
||||
```sh
|
||||
# don't forget the "--cs" option; otherwise it doesn't work
|
||||
minimap2 -cx asm5 --cs ecoli_ref.fa ecoli_canu.fa \
|
||||
| sort -k6,6 -k8,8n \
|
||||
| paftools.js call -f ecoli_ref.fa - > out.vcf
|
||||
```
|
||||
Without option `-f`, `paftools.js call` outputs in a custom format. In this
|
||||
format, lines starting with `R` give the regions covered by one contig only.
|
||||
This information is not available in the VCF output.
|
||||
|
||||
### <a name="hom-map"></a>Constructing self-homology map
|
||||
```sh
|
||||
minimap2 -DP -k19 -w19 -m200 ecoli_ref.fa ecoli_ref.fa > out.paf
|
||||
```
|
||||
Option `-D` asks minimap2 to ignore anchors from perfect self match and `-P`
|
||||
outputs all chains. For large nomes, we don't recommend to perform base-level
|
||||
alignment (with `-c`, `-a` or `--cs`) when `-P` is applied. This is because
|
||||
base-alignment is slow and occasionally gives wrong alignments close to the
|
||||
diagonal of a dotter plot. For E. coli, though, base-alignment is still fast.
|
||||
|
||||
### <a name="liftover"></a>Lift over (for developers)
|
||||
```sh
|
||||
minimap2 -cx asm5 --cs ecoli_ref.fa ecoli_canu.fa > ecoli_canu.paf
|
||||
echo -e 'tig00000001\t200000\t300000' | paftools.js liftover ecoli_canu.paf -
|
||||
```
|
||||
This lifts over a region on query sequences to one or multiple regions on
|
||||
reference sequences. Note that this paftools.js command may not be efficient
|
||||
enough to lift millions of regions.
|
||||
|
||||
|
||||
|
||||
## <a name="read-overlap"></a>Read Overlap
|
||||
|
||||
### <a name="long-read-overlap"></a>Long read overlap
|
||||
```sh
|
||||
# For pacbio reads:
|
||||
minimap2 -x ava-pb ecoli_p6_25x_canu.fa ecoli_p6_25x_canu.fa > overlap.paf
|
||||
# For Nanopore reads (ava-ont also works with PacBio but not as good):
|
||||
minimap2 -x ava-ont -r 10000 ecoli_p6_25x_canu.fa ecoli_p6_25x_canu.fa > overlap.paf
|
||||
# If you have miniasm installed:
|
||||
miniasm -f ecoli_p6_25x_canu.fa overlap.paf > asm.gfa
|
||||
```
|
||||
Here we explicitly applied `-r 10000`. We are considering to set this as the
|
||||
default for the `ava-ont` mode as this seems to improve the contiguity for
|
||||
nanopore read assembly (Loman, personal communication).
|
||||
|
||||
*Minimap2 doesn't work well with short-read overlap.*
|
||||
|
||||
### <a name="ov-eval"></a>Evaluating overlap sensitivity (for developers)
|
||||
|
||||
```sh
|
||||
# read to reference mapping
|
||||
minimap2 -cx map-pb ecoli_ref.fa ecoli_p6_25x_canu.fa > to-ref.paf
|
||||
# evaluate overlap sensitivity
|
||||
sort -k6,6 -k8,8n to-ref.paf | paftools.js ov-eval - overlap.paf
|
||||
```
|
||||
You can see that for PacBio reads, minimap2 achieves higher overlap sensitivity
|
||||
with `-x ava-pb` (99% vs 93% with `-x ava-ont`).
|
||||
|
||||
|
||||
|
||||
[pbsim]: https://github.com/yukiteruono/pbsim2
|
||||
[mason2]: https://github.com/seqan/seqan/tree/master/apps/mason2
|
||||
[paf]: https://github.com/lh3/miniasm/blob/master/PAF.md
|
||||
[v2.10]: https://github.com/lh3/minimap2/releases/tag/v2.10
|
||||
@@ -59,6 +59,6 @@ void mm_est_err(const mm_idx_t *mi, int qlen, int n_regs, mm_reg1_t *regs, const
|
||||
n_tot = en - st + 1;
|
||||
if (r->qs > avg_k && r->rs > avg_k) ++n_tot;
|
||||
if (qlen - r->qs > avg_k && l_ref - r->re > avg_k) ++n_tot;
|
||||
r->div = logf((float)n_tot / n_match) / avg_k;
|
||||
r->div = n_match >= n_tot? 0.0f : (float)(1.0 - pow((double)n_match / n_tot, 1.0 / avg_k));
|
||||
}
|
||||
}
|
||||
|
||||
@@ -35,6 +35,8 @@ int main(int argc, char *argv[])
|
||||
while ((mi = mm_idx_reader_read(r, n_threads)) != 0) { // traverse each part of the index
|
||||
mm_mapopt_update(&mopt, mi); // this sets the maximum minimizer occurrence; TODO: set a better default in mm_mapopt_init()!
|
||||
mm_tbuf_t *tbuf = mm_tbuf_init(); // thread buffer; for multi-threading, allocate one tbuf for each thread
|
||||
gzrewind(f);
|
||||
kseq_rewind(ks);
|
||||
while (kseq_read(ks) >= 0) { // each kseq_read() call reads one query sequence
|
||||
mm_reg1_t *reg;
|
||||
int j, i, n_reg;
|
||||
@@ -45,7 +47,7 @@ int main(int argc, char *argv[])
|
||||
printf("%s\t%d\t%d\t%d\t%c\t", ks->name.s, ks->seq.l, r->qs, r->qe, "+-"[r->rev]);
|
||||
printf("%s\t%d\t%d\t%d\t%d\t%d\t%d\tcg:Z:", mi->seq[r->rid].name, mi->seq[r->rid].len, r->rs, r->re, r->mlen, r->blen, r->mapq);
|
||||
for (i = 0; i < r->p->n_cigar; ++i) // IMPORTANT: this gives the CIGAR in the aligned regions. NO soft/hard clippings!
|
||||
printf("%d%c", r->p->cigar[i]>>4, "MIDSHN"[r->p->cigar[i]&0xf]);
|
||||
printf("%d%c", r->p->cigar[i]>>4, MM_CIGAR_STR[r->p->cigar[i]&0xf]);
|
||||
putchar('\n');
|
||||
free(r->p);
|
||||
}
|
||||
|
||||
@@ -79,11 +79,11 @@ static char *mm_escape(char *s)
|
||||
return s;
|
||||
}
|
||||
|
||||
static void sam_write_rg_line(kstring_t *str, const char *s)
|
||||
static int sam_write_rg_line(kstring_t *str, const char *s)
|
||||
{
|
||||
char *p, *q, *r, *rg_line = 0;
|
||||
memset(mm_rg_id, 0, 256);
|
||||
if (s == 0) return;
|
||||
if (s == 0) return 0;
|
||||
if (strstr(s, "@RG") != s) {
|
||||
if (mm_verbose >= 1) fprintf(stderr, "[ERROR] the read group line is not started with @RG\n");
|
||||
goto err_set_rg;
|
||||
@@ -92,7 +92,8 @@ static void sam_write_rg_line(kstring_t *str, const char *s)
|
||||
if (mm_verbose >= 1) fprintf(stderr, "[ERROR] the read group line contained literal <tab> characters -- replace with escaped tabs: \\t\n");
|
||||
goto err_set_rg;
|
||||
}
|
||||
rg_line = strdup(s);
|
||||
rg_line = (char*)malloc(strlen(s) + 1);
|
||||
strcpy(rg_line, s);
|
||||
mm_escape(rg_line);
|
||||
if ((p = strstr(rg_line, "\tID:")) == 0) {
|
||||
if (mm_verbose >= 1) fprintf(stderr, "[ERROR] no ID within the read group line\n");
|
||||
@@ -107,20 +108,23 @@ static void sam_write_rg_line(kstring_t *str, const char *s)
|
||||
for (q = p, r = mm_rg_id; *q && *q != '\t' && *q != '\n'; ++q)
|
||||
*r++ = *q;
|
||||
mm_sprintf_lite(str, "%s\n", rg_line);
|
||||
return 0;
|
||||
|
||||
err_set_rg:
|
||||
free(rg_line);
|
||||
return -1;
|
||||
}
|
||||
|
||||
void mm_write_sam_hdr(const mm_idx_t *idx, const char *rg, const char *ver, int argc, char *argv[])
|
||||
int mm_write_sam_hdr(const mm_idx_t *idx, const char *rg, const char *ver, int argc, char *argv[])
|
||||
{
|
||||
kstring_t str = {0,0,0};
|
||||
int ret = 0;
|
||||
if (idx) {
|
||||
uint32_t i;
|
||||
for (i = 0; i < idx->n_seq; ++i)
|
||||
printf("@SQ\tSN:%s\tLN:%d\n", idx->seq[i].name, idx->seq[i].len);
|
||||
mm_sprintf_lite(&str, "@SQ\tSN:%s\tLN:%d\n", idx->seq[i].name, idx->seq[i].len);
|
||||
}
|
||||
if (rg) sam_write_rg_line(&str, rg);
|
||||
if (rg) ret = sam_write_rg_line(&str, rg);
|
||||
mm_sprintf_lite(&str, "@PG\tID:minimap2\tPN:minimap2");
|
||||
if (ver) mm_sprintf_lite(&str, "\tVN:%s", ver);
|
||||
if (argc > 1) {
|
||||
@@ -129,36 +133,19 @@ void mm_write_sam_hdr(const mm_idx_t *idx, const char *rg, const char *ver, int
|
||||
for (i = 1; i < argc; ++i)
|
||||
mm_sprintf_lite(&str, " %s", argv[i]);
|
||||
}
|
||||
mm_sprintf_lite(&str, "\n");
|
||||
fputs(str.s, stdout);
|
||||
mm_err_puts(str.s);
|
||||
free(str.s);
|
||||
return ret;
|
||||
}
|
||||
|
||||
static void write_cs(void *km, kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, const mm_reg1_t *r, int no_iden)
|
||||
static void write_cs_core(kstring_t *s, const uint8_t *tseq, const uint8_t *qseq, const mm_reg1_t *r, char *tmp, int no_iden, int write_tag)
|
||||
{
|
||||
extern unsigned char seq_nt4_table[256];
|
||||
int i, q_off, t_off;
|
||||
uint8_t *qseq, *tseq;
|
||||
char *tmp;
|
||||
if (r->p == 0) return;
|
||||
mm_sprintf_lite(s, "\tcs:Z:");
|
||||
qseq = (uint8_t*)kmalloc(km, r->qe - r->qs);
|
||||
tseq = (uint8_t*)kmalloc(km, r->re - r->rs);
|
||||
tmp = (char*)kmalloc(km, r->re - r->rs > r->qe - r->qs? r->re - r->rs + 1 : r->qe - r->qs + 1);
|
||||
mm_idx_getseq(mi, r->rid, r->rs, r->re, tseq);
|
||||
if (!r->rev) {
|
||||
for (i = r->qs; i < r->qe; ++i)
|
||||
qseq[i - r->qs] = seq_nt4_table[(uint8_t)t->seq[i]];
|
||||
} else {
|
||||
for (i = r->qs; i < r->qe; ++i) {
|
||||
uint8_t c = seq_nt4_table[(uint8_t)t->seq[i]];
|
||||
qseq[r->qe - i - 1] = c >= 4? 4 : 3 - c;
|
||||
}
|
||||
}
|
||||
for (i = q_off = t_off = 0; i < r->p->n_cigar; ++i) {
|
||||
if (write_tag) mm_sprintf_lite(s, "\tcs:Z:");
|
||||
for (i = q_off = t_off = 0; i < (int)r->p->n_cigar; ++i) {
|
||||
int j, op = r->p->cigar[i]&0xf, len = r->p->cigar[i]>>4;
|
||||
assert(op >= 0 && op <= 3);
|
||||
if (op == 0) {
|
||||
assert((op >= MM_CIGAR_MATCH && op <= MM_CIGAR_N_SKIP) || op == MM_CIGAR_EQ_MATCH || op == MM_CIGAR_X_MISMATCH);
|
||||
if (op == MM_CIGAR_MATCH || op == MM_CIGAR_EQ_MATCH || op == MM_CIGAR_X_MISMATCH) {
|
||||
int l_tmp = 0;
|
||||
for (j = 0; j < len; ++j) {
|
||||
if (qseq[q_off + j] != tseq[t_off + j]) {
|
||||
@@ -179,17 +166,17 @@ static void write_cs(void *km, kstring_t *s, const mm_idx_t *mi, const mm_bseq1_
|
||||
} else mm_sprintf_lite(s, ":%d", l_tmp);
|
||||
}
|
||||
q_off += len, t_off += len;
|
||||
} else if (op == 1) {
|
||||
} else if (op == MM_CIGAR_INS) {
|
||||
for (j = 0, tmp[len] = 0; j < len; ++j)
|
||||
tmp[j] = "acgtn"[qseq[q_off + j]];
|
||||
mm_sprintf_lite(s, "+%s", tmp);
|
||||
q_off += len;
|
||||
} else if (op == 2) {
|
||||
} else if (op == MM_CIGAR_DEL) {
|
||||
for (j = 0, tmp[len] = 0; j < len; ++j)
|
||||
tmp[j] = "acgtn"[tseq[t_off + j]];
|
||||
mm_sprintf_lite(s, "-%s", tmp);
|
||||
t_off += len;
|
||||
} else {
|
||||
} else { // intron
|
||||
assert(len >= 2);
|
||||
mm_sprintf_lite(s, "~%c%c%d%c%c", "acgtn"[tseq[t_off]], "acgtn"[tseq[t_off+1]],
|
||||
len, "acgtn"[tseq[t_off+len-2]], "acgtn"[tseq[t_off+len-1]]);
|
||||
@@ -197,9 +184,93 @@ static void write_cs(void *km, kstring_t *s, const mm_idx_t *mi, const mm_bseq1_
|
||||
}
|
||||
}
|
||||
assert(t_off == r->re - r->rs && q_off == r->qe - r->qs);
|
||||
}
|
||||
|
||||
static void write_MD_core(kstring_t *s, const uint8_t *tseq, const uint8_t *qseq, const mm_reg1_t *r, char *tmp, int write_tag)
|
||||
{
|
||||
int i, q_off, t_off, l_MD = 0;
|
||||
if (write_tag) mm_sprintf_lite(s, "\tMD:Z:");
|
||||
for (i = q_off = t_off = 0; i < (int)r->p->n_cigar; ++i) {
|
||||
int j, op = r->p->cigar[i]&0xf, len = r->p->cigar[i]>>4;
|
||||
assert((op >= MM_CIGAR_MATCH && op <= MM_CIGAR_N_SKIP) || op == MM_CIGAR_EQ_MATCH || op == MM_CIGAR_X_MISMATCH);
|
||||
if (op == MM_CIGAR_MATCH || op == MM_CIGAR_EQ_MATCH || op == MM_CIGAR_X_MISMATCH) {
|
||||
for (j = 0; j < len; ++j) {
|
||||
if (qseq[q_off + j] != tseq[t_off + j]) {
|
||||
mm_sprintf_lite(s, "%d%c", l_MD, "ACGTN"[tseq[t_off + j]]);
|
||||
l_MD = 0;
|
||||
} else ++l_MD;
|
||||
}
|
||||
q_off += len, t_off += len;
|
||||
} else if (op == MM_CIGAR_INS) {
|
||||
q_off += len;
|
||||
} else if (op == MM_CIGAR_DEL) {
|
||||
for (j = 0, tmp[len] = 0; j < len; ++j)
|
||||
tmp[j] = "ACGTN"[tseq[t_off + j]];
|
||||
mm_sprintf_lite(s, "%d^%s", l_MD, tmp);
|
||||
l_MD = 0;
|
||||
t_off += len;
|
||||
} else if (op == MM_CIGAR_N_SKIP) {
|
||||
t_off += len;
|
||||
}
|
||||
}
|
||||
if (l_MD > 0) mm_sprintf_lite(s, "%d", l_MD);
|
||||
assert(t_off == r->re - r->rs && q_off == r->qe - r->qs);
|
||||
}
|
||||
|
||||
static void write_cs_or_MD(void *km, kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, const mm_reg1_t *r, int no_iden, int is_MD, int write_tag, int is_qstrand)
|
||||
{
|
||||
extern unsigned char seq_nt4_table[256];
|
||||
int i;
|
||||
uint8_t *qseq, *tseq;
|
||||
char *tmp;
|
||||
if (r->p == 0) return;
|
||||
qseq = (uint8_t*)kmalloc(km, r->qe - r->qs);
|
||||
tseq = (uint8_t*)kmalloc(km, r->re - r->rs);
|
||||
tmp = (char*)kmalloc(km, r->re - r->rs > r->qe - r->qs? r->re - r->rs + 1 : r->qe - r->qs + 1);
|
||||
if (is_qstrand) {
|
||||
mm_idx_getseq2(mi, r->rev, r->rid, r->rs, r->re, tseq);
|
||||
for (i = r->qs; i < r->qe; ++i)
|
||||
qseq[i - r->qs] = seq_nt4_table[(uint8_t)t->seq[i]];
|
||||
} else {
|
||||
mm_idx_getseq(mi, r->rid, r->rs, r->re, tseq);
|
||||
if (!r->rev) {
|
||||
for (i = r->qs; i < r->qe; ++i)
|
||||
qseq[i - r->qs] = seq_nt4_table[(uint8_t)t->seq[i]];
|
||||
} else {
|
||||
for (i = r->qs; i < r->qe; ++i) {
|
||||
uint8_t c = seq_nt4_table[(uint8_t)t->seq[i]];
|
||||
qseq[r->qe - i - 1] = c >= 4? 4 : 3 - c;
|
||||
}
|
||||
}
|
||||
}
|
||||
if (is_MD) write_MD_core(s, tseq, qseq, r, tmp, write_tag);
|
||||
else write_cs_core(s, tseq, qseq, r, tmp, no_iden, write_tag);
|
||||
kfree(km, qseq); kfree(km, tseq); kfree(km, tmp);
|
||||
}
|
||||
|
||||
int mm_gen_cs_or_MD(void *km, char **buf, int *max_len, const mm_idx_t *mi, const mm_reg1_t *r, const char *seq, int is_MD, int no_iden, int is_qstrand)
|
||||
{
|
||||
mm_bseq1_t t;
|
||||
kstring_t str;
|
||||
str.s = *buf, str.l = 0, str.m = *max_len;
|
||||
t.l_seq = strlen(seq);
|
||||
t.seq = (char*)seq;
|
||||
write_cs_or_MD(km, &str, mi, &t, r, no_iden, is_MD, 0, is_qstrand);
|
||||
*max_len = str.m;
|
||||
*buf = str.s;
|
||||
return str.l;
|
||||
}
|
||||
|
||||
int mm_gen_cs(void *km, char **buf, int *max_len, const mm_idx_t *mi, const mm_reg1_t *r, const char *seq, int no_iden)
|
||||
{
|
||||
return mm_gen_cs_or_MD(km, buf, max_len, mi, r, seq, 0, no_iden, 0);
|
||||
}
|
||||
|
||||
int mm_gen_MD(void *km, char **buf, int *max_len, const mm_idx_t *mi, const mm_reg1_t *r, const char *seq)
|
||||
{
|
||||
return mm_gen_cs_or_MD(km, buf, max_len, mi, r, seq, 1, 0, 0);
|
||||
}
|
||||
|
||||
static inline void write_tags(kstring_t *s, const mm_reg1_t *r)
|
||||
{
|
||||
int type;
|
||||
@@ -212,33 +283,57 @@ static inline void write_tags(kstring_t *s, const mm_reg1_t *r)
|
||||
}
|
||||
mm_sprintf_lite(s, "\ttp:A:%c\tcm:i:%d\ts1:i:%d", type, r->cnt, r->score);
|
||||
if (r->parent == r->id) mm_sprintf_lite(s, "\ts2:i:%d", r->subsc);
|
||||
if (r->div >= 0.0f && r->div <= 1.0f) {
|
||||
char buf[8];
|
||||
if (r->p) {
|
||||
char buf[16];
|
||||
double div;
|
||||
div = 1.0 - mm_event_identity(r);
|
||||
if (div == 0.0) buf[0] = '0', buf[1] = 0;
|
||||
else snprintf(buf, 16, "%.4f", 1.0 - mm_event_identity(r));
|
||||
mm_sprintf_lite(s, "\tde:f:%s", buf);
|
||||
} else if (r->div >= 0.0f && r->div <= 1.0f) {
|
||||
char buf[16];
|
||||
if (r->div == 0.0f) buf[0] = '0', buf[1] = 0;
|
||||
else sprintf(buf, "%.4f", r->div);
|
||||
else snprintf(buf, 16, "%.4f", r->div);
|
||||
mm_sprintf_lite(s, "\tdv:f:%s", buf);
|
||||
}
|
||||
if (r->split) mm_sprintf_lite(s, "\tzd:i:%d", r->split);
|
||||
}
|
||||
|
||||
void mm_write_paf(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, const mm_reg1_t *r, void *km, int opt_flag)
|
||||
void mm_write_paf3(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, const mm_reg1_t *r, void *km, int64_t opt_flag, int rep_len)
|
||||
{
|
||||
s->l = 0;
|
||||
if (r == 0) {
|
||||
mm_sprintf_lite(s, "%s\t%d\t0\t0\t*\t*\t0\t0\t0\t0\t0\t0", t->name, t->l_seq);
|
||||
if (rep_len >= 0) mm_sprintf_lite(s, "\trl:i:%d", rep_len);
|
||||
return;
|
||||
}
|
||||
mm_sprintf_lite(s, "%s\t%d\t%d\t%d\t%c\t", t->name, t->l_seq, r->qs, r->qe, "+-"[r->rev]);
|
||||
if (mi->seq[r->rid].name) mm_sprintf_lite(s, "%s", mi->seq[r->rid].name);
|
||||
else mm_sprintf_lite(s, "%d", r->rid);
|
||||
mm_sprintf_lite(s, "\t%d\t%d\t%d", mi->seq[r->rid].len, r->rs, r->re);
|
||||
mm_sprintf_lite(s, "\t%d", mi->seq[r->rid].len);
|
||||
if ((opt_flag & MM_F_QSTRAND) && r->rev)
|
||||
mm_sprintf_lite(s, "\t%d\t%d", mi->seq[r->rid].len - r->re, mi->seq[r->rid].len - r->rs);
|
||||
else
|
||||
mm_sprintf_lite(s, "\t%d\t%d", r->rs, r->re);
|
||||
mm_sprintf_lite(s, "\t%d\t%d", r->mlen, r->blen);
|
||||
mm_sprintf_lite(s, "\t%d", r->mapq);
|
||||
write_tags(s, r);
|
||||
if (rep_len >= 0) mm_sprintf_lite(s, "\trl:i:%d", rep_len);
|
||||
if (r->p && (opt_flag & MM_F_OUT_CG)) {
|
||||
uint32_t k;
|
||||
mm_sprintf_lite(s, "\tcg:Z:");
|
||||
for (k = 0; k < r->p->n_cigar; ++k)
|
||||
mm_sprintf_lite(s, "%d%c", r->p->cigar[k]>>4, "MIDN"[r->p->cigar[k]&0xf]);
|
||||
mm_sprintf_lite(s, "%d%c", r->p->cigar[k]>>4, MM_CIGAR_STR[r->p->cigar[k]&0xf]);
|
||||
}
|
||||
if (r->p && (opt_flag & MM_F_OUT_CS))
|
||||
write_cs(km, s, mi, t, r, !(opt_flag&MM_F_OUT_CS_LONG));
|
||||
if (r->p && (opt_flag & (MM_F_OUT_CS|MM_F_OUT_MD)))
|
||||
write_cs_or_MD(km, s, mi, t, r, !(opt_flag&MM_F_OUT_CS_LONG), opt_flag&MM_F_OUT_MD, 1, !!(opt_flag&MM_F_QSTRAND));
|
||||
if ((opt_flag & MM_F_COPY_COMMENT) && t->comment)
|
||||
mm_sprintf_lite(s, "\t%s", t->comment);
|
||||
}
|
||||
|
||||
void mm_write_paf(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, const mm_reg1_t *r, void *km, int64_t opt_flag)
|
||||
{
|
||||
mm_write_paf3(s, mi, t, r, km, opt_flag, -1);
|
||||
}
|
||||
|
||||
static void sam_write_sq(kstring_t *s, char *seq, int l, int rev, int comp)
|
||||
@@ -265,7 +360,7 @@ static inline const mm_reg1_t *get_sam_pri(int n_regs, const mm_reg1_t *regs)
|
||||
return NULL;
|
||||
}
|
||||
|
||||
static void write_sam_cigar(kstring_t *s, int sam_flag, int in_tag, int qlen, const mm_reg1_t *r, int opt_flag)
|
||||
static void write_sam_cigar(kstring_t *s, int sam_flag, int in_tag, int qlen, const mm_reg1_t *r, int64_t opt_flag)
|
||||
{
|
||||
if (r->p == 0) {
|
||||
mm_sprintf_lite(s, "*");
|
||||
@@ -282,19 +377,20 @@ static void write_sam_cigar(kstring_t *s, int sam_flag, int in_tag, int qlen, co
|
||||
if (clip_len[1]) mm_sprintf_lite(s, ",%u", clip_len[1]<<4|clip_char);
|
||||
} else {
|
||||
int clip_char = (sam_flag&0x800) && !(opt_flag&MM_F_SOFTCLIP)? 'H' : 'S';
|
||||
assert(clip_len[0] < qlen && clip_len[1] < qlen);
|
||||
if (clip_len[0]) mm_sprintf_lite(s, "%d%c", clip_len[0], clip_char);
|
||||
for (k = 0; k < r->p->n_cigar; ++k)
|
||||
mm_sprintf_lite(s, "%d%c", r->p->cigar[k]>>4, "MIDN"[r->p->cigar[k]&0xf]);
|
||||
mm_sprintf_lite(s, "%d%c", r->p->cigar[k]>>4, MM_CIGAR_STR[r->p->cigar[k]&0xf]);
|
||||
if (clip_len[1]) mm_sprintf_lite(s, "%d%c", clip_len[1], clip_char);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
void mm_write_sam2(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, int seg_idx, int reg_idx, int n_seg, const int *n_regss, const mm_reg1_t *const* regss, void *km, int opt_flag)
|
||||
void mm_write_sam3(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, int seg_idx, int reg_idx, int n_seg, const int *n_regss, const mm_reg1_t *const* regss, void *km, int64_t opt_flag, int rep_len)
|
||||
{
|
||||
const int max_bam_cigar_op = 65535;
|
||||
int flag, n_regs = n_regss[seg_idx], cigar_in_tag = 0;
|
||||
int this_rid = -1, this_pos = -1, this_rev = 0;
|
||||
int this_rid = -1, this_pos = -1;
|
||||
const mm_reg1_t *regs = regss[seg_idx], *r_prev = NULL, *r_next;
|
||||
const mm_reg1_t *r = n_regs > 0 && reg_idx < n_regs && reg_idx >= 0? ®s[reg_idx] : NULL;
|
||||
|
||||
@@ -343,7 +439,7 @@ void mm_write_sam2(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, int se
|
||||
mm_sprintf_lite(s, "\t%s\t%d\t0\t*", mi->seq[this_rid].name, this_pos+1);
|
||||
} else mm_sprintf_lite(s, "\t*\t0\t0\t*");
|
||||
} else {
|
||||
this_rid = r->rid, this_pos = r->rs, this_rev = r->rev;
|
||||
this_rid = r->rid, this_pos = r->rs;
|
||||
mm_sprintf_lite(s, "\t%s\t%d\t%d\t", mi->seq[r->rid].name, r->rs+1, r->mapq);
|
||||
if ((opt_flag & MM_F_LONG_CIGAR) && r->p && r->p->n_cigar > max_bam_cigar_op - 2) {
|
||||
int n_cigar = r->p->n_cigar;
|
||||
@@ -353,9 +449,11 @@ void mm_write_sam2(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, int se
|
||||
cigar_in_tag = 1;
|
||||
}
|
||||
if (cigar_in_tag) {
|
||||
if (flag & 0x100) mm_sprintf_lite(s, "0S"); // secondary alignment
|
||||
else if (flag & 0x800) mm_sprintf_lite(s, "%dS", r->re - r->rs); // supplementary alignment
|
||||
else mm_sprintf_lite(s, "%dS", t->l_seq);
|
||||
int slen;
|
||||
if ((flag & 0x900) == 0 || (opt_flag & MM_F_SOFTCLIP)) slen = t->l_seq;
|
||||
else if (flag & 0x100) slen = 0;
|
||||
else slen = r->qe - r->qs;
|
||||
mm_sprintf_lite(s, "%dS%dN", slen, r->re - r->rs);
|
||||
} else write_sam_cigar(s, flag, 0, t->l_seq, r, opt_flag);
|
||||
}
|
||||
|
||||
@@ -364,17 +462,17 @@ void mm_write_sam2(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, int se
|
||||
int tlen = 0;
|
||||
if (this_rid >= 0 && r_next) {
|
||||
if (this_rid == r_next->rid) {
|
||||
int this_pos5 = r && r->rev? r->re - 1 : this_pos;
|
||||
int next_pos5 = r_next->rev? r_next->re - 1 : r_next->rs;
|
||||
tlen = next_pos5 - this_pos5;
|
||||
if (r) {
|
||||
int this_pos5 = r->rev? r->re - 1 : this_pos;
|
||||
int next_pos5 = r_next->rev? r_next->re - 1 : r_next->rs;
|
||||
tlen = next_pos5 - this_pos5;
|
||||
}
|
||||
mm_sprintf_lite(s, "\t=\t");
|
||||
} else mm_sprintf_lite(s, "\t%s\t", mi->seq[r_next->rid].name);
|
||||
mm_sprintf_lite(s, "%d\t", r_next->rs + 1);
|
||||
} else if (r_next) { // && this_rid < 0
|
||||
mm_sprintf_lite(s, "\t%s\t%d\t", mi->seq[r_next->rid].name, r_next->rs + 1);
|
||||
} else if (this_rid >= 0) { // && r_next == NULL
|
||||
int this_pos5 = this_rev? r->re - 1 : this_pos; // this_rev is only true when r != NULL
|
||||
tlen = this_pos - this_pos5; // next_pos5 will be this_pos
|
||||
mm_sprintf_lite(s, "\t=\t%d\t", this_pos + 1); // next segment will take r's coordinate
|
||||
} else mm_sprintf_lite(s, "\t*\t0\t"); // neither has coordinates
|
||||
if (tlen > 0) ++tlen;
|
||||
@@ -434,15 +532,24 @@ void mm_write_sam2(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, int se
|
||||
}
|
||||
}
|
||||
}
|
||||
if (r->p && (opt_flag & MM_F_OUT_CS))
|
||||
write_cs(km, s, mi, t, r, !(opt_flag&MM_F_OUT_CS_LONG));
|
||||
if (r->p && (opt_flag & (MM_F_OUT_CS|MM_F_OUT_MD)))
|
||||
write_cs_or_MD(km, s, mi, t, r, !(opt_flag&MM_F_OUT_CS_LONG), opt_flag&MM_F_OUT_MD, 1, 0);
|
||||
if (cigar_in_tag)
|
||||
write_sam_cigar(s, flag, 1, t->l_seq, r, opt_flag);
|
||||
}
|
||||
if (rep_len >= 0) mm_sprintf_lite(s, "\trl:i:%d", rep_len);
|
||||
|
||||
if ((opt_flag & MM_F_COPY_COMMENT) && t->comment)
|
||||
mm_sprintf_lite(s, "\t%s", t->comment);
|
||||
|
||||
s->s[s->l] = 0; // we always have room for an extra byte (see str_enlarge)
|
||||
}
|
||||
|
||||
void mm_write_sam2(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, int seg_idx, int reg_idx, int n_seg, const int *n_regss, const mm_reg1_t *const* regss, void *km, int64_t opt_flag)
|
||||
{
|
||||
mm_write_sam3(s, mi, t, seg_idx, reg_idx, n_seg, n_regss, regss, km, opt_flag, -1);
|
||||
}
|
||||
|
||||
void mm_write_sam(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, const mm_reg1_t *r, int n_regs, const mm_reg1_t *regs)
|
||||
{
|
||||
int i;
|
||||
|
||||
@@ -1,216 +0,0 @@
|
||||
#include <stddef.h>
|
||||
#include <stdio.h>
|
||||
#include <string.h>
|
||||
#include "getopt.h"
|
||||
|
||||
char *optarg;
|
||||
int optind=1, opterr=1, optopt, __optpos, optreset=0;
|
||||
|
||||
#define optpos __optpos
|
||||
|
||||
static void __getopt_msg(const char *a, const char *b, const char *c, size_t l)
|
||||
{
|
||||
FILE *f = stderr;
|
||||
#if !defined(WIN32) && !defined(_WIN32)
|
||||
flockfile(f);
|
||||
#endif
|
||||
fputs(a, f);
|
||||
fwrite(b, strlen(b), 1, f);
|
||||
fwrite(c, 1, l, f);
|
||||
fputc('\n', f);
|
||||
#if !defined(WIN32) && !defined(_WIN32)
|
||||
funlockfile(f);
|
||||
#endif
|
||||
}
|
||||
|
||||
int getopt(int argc, char * const argv[], const char *optstring)
|
||||
{
|
||||
int i, c, d;
|
||||
int k, l;
|
||||
char *optchar;
|
||||
|
||||
if (!optind || optreset) {
|
||||
optreset = 0;
|
||||
__optpos = 0;
|
||||
optind = 1;
|
||||
}
|
||||
|
||||
if (optind >= argc || !argv[optind])
|
||||
return -1;
|
||||
|
||||
if (argv[optind][0] != '-') {
|
||||
if (optstring[0] == '-') {
|
||||
optarg = argv[optind++];
|
||||
return 1;
|
||||
}
|
||||
return -1;
|
||||
}
|
||||
|
||||
if (!argv[optind][1])
|
||||
return -1;
|
||||
|
||||
if (argv[optind][1] == '-' && !argv[optind][2])
|
||||
return optind++, -1;
|
||||
|
||||
if (!optpos) optpos++;
|
||||
c = argv[optind][optpos], k = 1;
|
||||
optchar = argv[optind]+optpos;
|
||||
optopt = c;
|
||||
optpos += k;
|
||||
|
||||
if (!argv[optind][optpos]) {
|
||||
optind++;
|
||||
optpos = 0;
|
||||
}
|
||||
|
||||
if (optstring[0] == '-' || optstring[0] == '+')
|
||||
optstring++;
|
||||
|
||||
i = 0;
|
||||
d = 0;
|
||||
do {
|
||||
d = optstring[i], l = 1;
|
||||
if (l>0) i+=l; else i++;
|
||||
} while (l && d != c);
|
||||
|
||||
if (d != c) {
|
||||
if (optstring[0] != ':' && opterr)
|
||||
__getopt_msg(argv[0], ": unrecognized option: ", optchar, k);
|
||||
return '?';
|
||||
}
|
||||
if (optstring[i] == ':') {
|
||||
if (optstring[i+1] == ':') optarg = 0;
|
||||
else if (optind >= argc) {
|
||||
if (optstring[0] == ':') return ':';
|
||||
if (opterr) __getopt_msg(argv[0],
|
||||
": option requires an argument: ",
|
||||
optchar, k);
|
||||
return '?';
|
||||
}
|
||||
if (optstring[i+1] != ':' || optpos) {
|
||||
optarg = argv[optind++] + optpos;
|
||||
optpos = 0;
|
||||
}
|
||||
}
|
||||
return c;
|
||||
}
|
||||
|
||||
static void permute(char *const *argv, int dest, int src)
|
||||
{
|
||||
char **av = (char **)argv;
|
||||
char *tmp = av[src];
|
||||
int i;
|
||||
for (i=src; i>dest; i--)
|
||||
av[i] = av[i-1];
|
||||
av[dest] = tmp;
|
||||
}
|
||||
|
||||
static int __getopt_long_core(int argc, char *const *argv, const char *optstring, const struct option *longopts, int *idx, int longonly)
|
||||
{
|
||||
optarg = 0;
|
||||
if (longopts && argv[optind][0] == '-' &&
|
||||
((longonly && argv[optind][1] && argv[optind][1] != '-') ||
|
||||
(argv[optind][1] == '-' && argv[optind][2])))
|
||||
{
|
||||
int colon = optstring[optstring[0]=='+'||optstring[0]=='-']==':';
|
||||
int i, cnt, match = -1;
|
||||
char *opt;
|
||||
for (cnt=i=0; longopts[i].name; i++) {
|
||||
const char *name = longopts[i].name;
|
||||
opt = argv[optind]+1;
|
||||
if (*opt == '-') opt++;
|
||||
for (; *name && *name == *opt; name++, opt++);
|
||||
if (*opt && *opt != '=') continue;
|
||||
match = i;
|
||||
if (!*name) {
|
||||
cnt = 1;
|
||||
break;
|
||||
}
|
||||
cnt++;
|
||||
}
|
||||
if (cnt==1) {
|
||||
i = match;
|
||||
optind++;
|
||||
optopt = longopts[i].val;
|
||||
if (*opt == '=') {
|
||||
if (!longopts[i].has_arg) {
|
||||
if (colon || !opterr)
|
||||
return '?';
|
||||
__getopt_msg(argv[0],
|
||||
": option does not take an argument: ",
|
||||
longopts[i].name,
|
||||
strlen(longopts[i].name));
|
||||
return '?';
|
||||
}
|
||||
optarg = opt+1;
|
||||
} else if (longopts[i].has_arg == required_argument) {
|
||||
if (!(optarg = argv[optind])) {
|
||||
if (colon) return ':';
|
||||
if (!opterr) return '?';
|
||||
__getopt_msg(argv[0],
|
||||
": option requires an argument: ",
|
||||
longopts[i].name,
|
||||
strlen(longopts[i].name));
|
||||
return '?';
|
||||
}
|
||||
optind++;
|
||||
}
|
||||
if (idx) *idx = i;
|
||||
if (longopts[i].flag) {
|
||||
*longopts[i].flag = longopts[i].val;
|
||||
return 0;
|
||||
}
|
||||
return longopts[i].val;
|
||||
}
|
||||
if (argv[optind][1] == '-') {
|
||||
if (!colon && opterr)
|
||||
__getopt_msg(argv[0], cnt ?
|
||||
": option is ambiguous: " :
|
||||
": unrecognized option: ",
|
||||
argv[optind]+2,
|
||||
strlen(argv[optind]+2));
|
||||
optind++;
|
||||
return '?';
|
||||
}
|
||||
}
|
||||
return getopt(argc, argv, optstring);
|
||||
}
|
||||
|
||||
static int __getopt_long(int argc, char *const *argv, const char *optstring, const struct option *longopts, int *idx, int longonly)
|
||||
{
|
||||
int ret, skipped, resumed;
|
||||
if (!optind || optreset) {
|
||||
optreset = 0;
|
||||
__optpos = 0;
|
||||
optind = 1;
|
||||
}
|
||||
if (optind >= argc || !argv[optind]) return -1;
|
||||
skipped = optind;
|
||||
if (optstring[0] != '+' && optstring[0] != '-') {
|
||||
int i;
|
||||
for (i=optind; ; i++) {
|
||||
if (i >= argc || !argv[i]) return -1;
|
||||
if (argv[i][0] == '-' && argv[i][1]) break;
|
||||
}
|
||||
optind = i;
|
||||
}
|
||||
resumed = optind;
|
||||
ret = __getopt_long_core(argc, argv, optstring, longopts, idx, longonly);
|
||||
if (resumed > skipped) {
|
||||
int i, cnt = optind-resumed;
|
||||
for (i=0; i<cnt; i++)
|
||||
permute(argv, skipped, optind-1);
|
||||
optind = skipped + cnt;
|
||||
}
|
||||
return ret;
|
||||
}
|
||||
|
||||
int getopt_long(int argc, char *const *argv, const char *optstring, const struct option *longopts, int *idx)
|
||||
{
|
||||
return __getopt_long(argc, argv, optstring, longopts, idx, 0);
|
||||
}
|
||||
|
||||
int getopt_long_only(int argc, char *const *argv, const char *optstring, const struct option *longopts, int *idx)
|
||||
{
|
||||
return __getopt_long(argc, argv, optstring, longopts, idx, 1);
|
||||
}
|
||||
@@ -1,53 +0,0 @@
|
||||
/*
|
||||
Copyright 2005-2014 Rich Felker, et al.
|
||||
|
||||
Permission is hereby granted, free of charge, to any person obtaining
|
||||
a copy of this software and associated documentation files (the
|
||||
"Software"), to deal in the Software without restriction, including
|
||||
without limitation the rights to use, copy, modify, merge, publish,
|
||||
distribute, sublicense, and/or sell copies of the Software, and to
|
||||
permit persons to whom the Software is furnished to do so, subject to
|
||||
the following conditions:
|
||||
|
||||
The above copyright notice and this permission notice shall be
|
||||
included in all copies or substantial portions of the Software.
|
||||
|
||||
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,
|
||||
EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF
|
||||
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.
|
||||
IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY
|
||||
CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT,
|
||||
TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE
|
||||
SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
|
||||
*/
|
||||
|
||||
#ifndef _GETOPT_H
|
||||
#define _GETOPT_H
|
||||
|
||||
#ifdef __cplusplus
|
||||
extern "C" {
|
||||
#endif
|
||||
|
||||
int getopt(int, char * const [], const char *);
|
||||
extern char *optarg;
|
||||
extern int optind, opterr, optopt, optreset;
|
||||
|
||||
struct option {
|
||||
const char *name;
|
||||
int has_arg;
|
||||
int *flag;
|
||||
int val;
|
||||
};
|
||||
|
||||
int getopt_long(int, char *const *, const char *, const struct option *, int *);
|
||||
int getopt_long_only(int, char *const *, const char *, const struct option *, int *);
|
||||
|
||||
#define no_argument 0
|
||||
#define required_argument 1
|
||||
#define optional_argument 2
|
||||
|
||||
#ifdef __cplusplus
|
||||
}
|
||||
#endif
|
||||
|
||||
#endif
|
||||
@@ -20,14 +20,14 @@ static inline void mm_cal_fuzzy_len(mm_reg1_t *r, const mm128_t *a)
|
||||
}
|
||||
}
|
||||
|
||||
static inline void mm_reg_set_coor(mm_reg1_t *r, int32_t qlen, const mm128_t *a)
|
||||
static inline void mm_reg_set_coor(mm_reg1_t *r, int32_t qlen, const mm128_t *a, int is_qstrand)
|
||||
{ // NB: r->as and r->cnt MUST BE set correctly for this function to work
|
||||
int32_t k = r->as, q_span = (int32_t)(a[k].y>>32&0xff);
|
||||
r->rev = a[k].x>>63;
|
||||
r->rid = a[k].x<<1>>33;
|
||||
r->rs = (int32_t)a[k].x + 1 > q_span? (int32_t)a[k].x + 1 - q_span : 0; // NB: target span may be shorter, so this test is necessary
|
||||
r->re = (int32_t)a[k + r->cnt - 1].x + 1;
|
||||
if (!r->rev) {
|
||||
if (!r->rev || is_qstrand) {
|
||||
r->qs = (int32_t)a[k].y + 1 - q_span;
|
||||
r->qe = (int32_t)a[k + r->cnt - 1].y + 1;
|
||||
} else {
|
||||
@@ -49,7 +49,7 @@ static inline uint64_t hash64(uint64_t key)
|
||||
return key;
|
||||
}
|
||||
|
||||
mm_reg1_t *mm_gen_regs(void *km, uint32_t hash, int qlen, int n_u, uint64_t *u, mm128_t *a) // convert chains to hits
|
||||
mm_reg1_t *mm_gen_regs(void *km, uint32_t hash, int qlen, int n_u, uint64_t *u, mm128_t *a, int is_qstrand) // convert chains to hits
|
||||
{
|
||||
mm128_t *z, tmp;
|
||||
mm_reg1_t *r;
|
||||
@@ -81,31 +81,48 @@ mm_reg1_t *mm_gen_regs(void *km, uint32_t hash, int qlen, int n_u, uint64_t *u,
|
||||
ri->cnt = (int32_t)z[i].y;
|
||||
ri->as = z[i].y >> 32;
|
||||
ri->div = -1.0f;
|
||||
mm_reg_set_coor(ri, qlen, a);
|
||||
mm_reg_set_coor(ri, qlen, a, is_qstrand);
|
||||
}
|
||||
kfree(km, z);
|
||||
return r;
|
||||
}
|
||||
|
||||
void mm_split_reg(mm_reg1_t *r, mm_reg1_t *r2, int n, int qlen, mm128_t *a)
|
||||
void mm_mark_alt(const mm_idx_t *mi, int n, mm_reg1_t *r)
|
||||
{
|
||||
int i;
|
||||
if (mi->n_alt == 0) return;
|
||||
for (i = 0; i < n; ++i)
|
||||
if (mi->seq[r[i].rid].is_alt)
|
||||
r[i].is_alt = 1;
|
||||
}
|
||||
|
||||
static inline int mm_alt_score(int score, float alt_diff_frac)
|
||||
{
|
||||
if (score < 0) return score;
|
||||
score = (int)(score * (1.0 - alt_diff_frac) + .499);
|
||||
return score > 0? score : 1;
|
||||
}
|
||||
|
||||
void mm_split_reg(mm_reg1_t *r, mm_reg1_t *r2, int n, int qlen, mm128_t *a, int is_qstrand)
|
||||
{
|
||||
if (n <= 0 || n >= r->cnt) return;
|
||||
*r2 = *r;
|
||||
r2->id = -1;
|
||||
r2->sam_pri = 0;
|
||||
r2->p = 0;
|
||||
r2->split_inv = 0;
|
||||
r2->cnt = r->cnt - n;
|
||||
r2->score = (int32_t)(r->score * ((float)r2->cnt / r->cnt) + .499);
|
||||
r2->as = r->as + n;
|
||||
if (r->parent == r->id) r2->parent = MM_PARENT_TMP_PRI;
|
||||
mm_reg_set_coor(r2, qlen, a);
|
||||
mm_reg_set_coor(r2, qlen, a, is_qstrand);
|
||||
r->cnt -= r2->cnt;
|
||||
r->score -= r2->score;
|
||||
mm_reg_set_coor(r, qlen, a);
|
||||
mm_reg_set_coor(r, qlen, a, is_qstrand);
|
||||
r->split |= 1, r2->split |= 2;
|
||||
}
|
||||
|
||||
void mm_set_parent(void *km, float mask_level, int n, mm_reg1_t *r, int sub_diff) // and compute mm_reg1_t::subsc
|
||||
void mm_set_parent(void *km, float mask_level, int mask_len, int n, mm_reg1_t *r, int sub_diff, int hard_mask_level, float alt_diff_frac) // and compute mm_reg1_t::subsc
|
||||
{
|
||||
int i, j, k, *w;
|
||||
uint64_t *cov;
|
||||
@@ -117,6 +134,7 @@ void mm_set_parent(void *km, float mask_level, int n, mm_reg1_t *r, int sub_diff
|
||||
for (i = 1, k = 1; i < n; ++i) {
|
||||
mm_reg1_t *ri = &r[i];
|
||||
int si = ri->qs, ei = ri->qe, n_cov = 0, uncov_len = 0;
|
||||
if (hard_mask_level) goto skip_uncov;
|
||||
for (j = 0; j < k; ++j) { // traverse existing primary hits to find overlapping hits
|
||||
mm_reg1_t *rp = &r[w[j]];
|
||||
int sj = rp->qs, ej = rp->qe;
|
||||
@@ -131,25 +149,29 @@ void mm_set_parent(void *km, float mask_level, int n, mm_reg1_t *r, int sub_diff
|
||||
int j, x = si;
|
||||
radix_sort_64(cov, cov + n_cov);
|
||||
for (j = 0; j < n_cov; ++j) {
|
||||
if (cov[j]>>32 > x) uncov_len += (cov[j]>>32) - x;
|
||||
if ((int)(cov[j]>>32) > x) uncov_len += (cov[j]>>32) - x;
|
||||
x = (int32_t)cov[j] > x? (int32_t)cov[j] : x;
|
||||
}
|
||||
if (ei > x) uncov_len += ei - x;
|
||||
}
|
||||
skip_uncov:
|
||||
for (j = 0; j < k; ++j) { // traverse existing primary hits again
|
||||
mm_reg1_t *rp = &r[w[j]];
|
||||
int sj = rp->qs, ej = rp->qe, min, max, ol;
|
||||
if (ej <= si || sj >= ei) continue; // no overlap
|
||||
min = ej - sj < ei - si? ej - sj : ei - si;
|
||||
max = ej - sj > ei - si? ej - sj : ei - si;
|
||||
ol = si < sj? (ei < sj? 0 : ei < ej? ei - sj : ej - sj) : (ej < si? 0 : ej < ei? ej - si : ei - si); // overlap length
|
||||
if ((float)ol / min - (float)uncov_len / max > mask_level) {
|
||||
int cnt_sub = 0;
|
||||
ol = si < sj? (ei < sj? 0 : ei < ej? ei - sj : ej - sj) : (ej < si? 0 : ej < ei? ej - si : ei - si); // overlap length; TODO: this can be simplified
|
||||
if ((float)ol / min - (float)uncov_len / max > mask_level && uncov_len <= mask_len) { // then this is a secondary hit
|
||||
int cnt_sub = 0, sci = ri->score;
|
||||
ri->parent = rp->parent;
|
||||
rp->subsc = rp->subsc > ri->score? rp->subsc : ri->score;
|
||||
if (!rp->is_alt && ri->is_alt) sci = mm_alt_score(sci, alt_diff_frac);
|
||||
rp->subsc = rp->subsc > sci? rp->subsc : sci;
|
||||
if (ri->cnt >= rp->cnt) cnt_sub = 1;
|
||||
if (rp->p && ri->p && (rp->rid != ri->rid || rp->rs != ri->rs || rp->re != ri->re || ol != min)) { // the last condition excludes identical hits after DP
|
||||
rp->p->dp_max2 = rp->p->dp_max2 > ri->p->dp_max? rp->p->dp_max2 : ri->p->dp_max;
|
||||
sci = ri->p->dp_max;
|
||||
if (!rp->is_alt && ri->is_alt) sci = mm_alt_score(sci, alt_diff_frac);
|
||||
rp->p->dp_max2 = rp->p->dp_max2 > sci? rp->p->dp_max2 : sci;
|
||||
if (rp->p->dp_max - ri->p->dp_max <= sub_diff) cnt_sub = 1;
|
||||
}
|
||||
if (cnt_sub) ++rp->n_sub;
|
||||
@@ -163,9 +185,9 @@ set_parent_test:
|
||||
kfree(km, w);
|
||||
}
|
||||
|
||||
void mm_hit_sort_by_dp(void *km, int *n_regs, mm_reg1_t *r)
|
||||
void mm_hit_sort(void *km, int *n_regs, mm_reg1_t *r, float alt_diff_frac)
|
||||
{
|
||||
int32_t i, n_aux, n = *n_regs;
|
||||
int32_t i, n_aux, n = *n_regs, has_cigar = 0, no_cigar = 0;
|
||||
mm128_t *aux;
|
||||
mm_reg1_t *t;
|
||||
|
||||
@@ -174,14 +196,18 @@ void mm_hit_sort_by_dp(void *km, int *n_regs, mm_reg1_t *r)
|
||||
t = (mm_reg1_t*)kmalloc(km, n * sizeof(mm_reg1_t));
|
||||
for (i = n_aux = 0; i < n; ++i) {
|
||||
if (r[i].inv || r[i].cnt > 0) { // squeeze out elements with cnt==0 (soft deleted)
|
||||
assert(r[i].p);
|
||||
aux[n_aux].x = (uint64_t)r[i].p->dp_max << 32 | r[i].hash;
|
||||
int score;
|
||||
if (r[i].p) score = r[i].p->dp_max, has_cigar = 1;
|
||||
else score = r[i].score, no_cigar = 1;
|
||||
if (r[i].is_alt) score = mm_alt_score(score, alt_diff_frac);
|
||||
aux[n_aux].x = (uint64_t)score << 32 | r[i].hash;
|
||||
aux[n_aux++].y = i;
|
||||
} else if (r[i].p) {
|
||||
free(r[i].p);
|
||||
r[i].p = 0;
|
||||
}
|
||||
}
|
||||
assert(has_cigar + no_cigar == 1);
|
||||
radix_sort_128x(aux, aux + n_aux);
|
||||
for (i = n_aux - 1; i >= 0; --i)
|
||||
t[n_aux - 1 - i] = r[aux[i].y];
|
||||
@@ -245,16 +271,17 @@ void mm_select_sub(void *km, float pri_ratio, int min_diff, int best_n, int *n_,
|
||||
}
|
||||
}
|
||||
|
||||
void mm_filter_regs(void *km, const mm_mapopt_t *opt, int *n_regs, mm_reg1_t *regs)
|
||||
void mm_filter_regs(const mm_mapopt_t *opt, int qlen, int *n_regs, mm_reg1_t *regs)
|
||||
{ // NB: after this call, mm_reg1_t::parent can be -1 if its parent filtered out
|
||||
int i, k;
|
||||
for (i = k = 0; i < *n_regs; ++i) {
|
||||
mm_reg1_t *r = ®s[i];
|
||||
int flt = 0;
|
||||
if (!r->inv && !r->seg_split && r->cnt < opt->min_cnt) flt = 1;
|
||||
if (r->p) {
|
||||
if (r->p) { // these filters are only applied when base-alignment is available
|
||||
if (r->mlen < opt->min_chain_score) flt = 1;
|
||||
else if (r->p->dp_max < opt->min_dp_max) flt = 1;
|
||||
else if (r->qs > qlen * opt->max_clip_ratio && qlen - r->qe > qlen * opt->max_clip_ratio) flt = 1;
|
||||
if (flt) free(r->p);
|
||||
}
|
||||
if (!flt) {
|
||||
@@ -285,63 +312,6 @@ int mm_squeeze_a(void *km, int n_regs, mm_reg1_t *regs, mm128_t *a)
|
||||
return as;
|
||||
}
|
||||
|
||||
void mm_join_long(void *km, const mm_mapopt_t *opt, int qlen, int *n_regs_, mm_reg1_t *regs, mm128_t *a)
|
||||
{
|
||||
int i, n_aux, n_regs = *n_regs_, n_drop = 0;
|
||||
uint64_t *aux;
|
||||
|
||||
if (n_regs < 2) return; // nothing to join
|
||||
mm_squeeze_a(km, n_regs, regs, a);
|
||||
|
||||
aux = (uint64_t*)kmalloc(km, n_regs * 8);
|
||||
for (i = n_aux = 0; i < n_regs; ++i)
|
||||
if (regs[i].parent == i || regs[i].parent < 0)
|
||||
aux[n_aux++] = (uint64_t)regs[i].as << 32 | i;
|
||||
radix_sort_64(aux, aux + n_aux);
|
||||
|
||||
for (i = n_aux - 1; i >= 1; --i) {
|
||||
mm_reg1_t *r0 = ®s[(int32_t)aux[i-1]], *r1 = ®s[(int32_t)aux[i]];
|
||||
mm128_t *a0e, *a1s;
|
||||
int max_gap, min_gap, sc_thres;
|
||||
|
||||
// test
|
||||
if (r0->as + r0->cnt != r1->as) continue; // not adjacent in a[]
|
||||
if (r0->rid != r1->rid || r0->rev != r1->rev) continue; // make sure on the same target and strand
|
||||
a0e = &a[r0->as + r0->cnt - 1];
|
||||
a1s = &a[r1->as];
|
||||
if (a1s->x <= a0e->x || (int32_t)a1s->y <= (int32_t)a0e->y) continue; // keep colinearity
|
||||
max_gap = min_gap = (int32_t)a1s->y - (int32_t)a0e->y;
|
||||
max_gap = max_gap > a1s->x - a0e->x? max_gap : a1s->x - a0e->x;
|
||||
min_gap = min_gap < a1s->x - a0e->x? min_gap : a1s->x - a0e->x;
|
||||
if (max_gap > opt->max_join_long || min_gap > opt->max_join_short) continue;
|
||||
sc_thres = (int)((float)opt->min_join_flank_sc / opt->max_join_long * max_gap + .499);
|
||||
if (r0->score < sc_thres || r1->score < sc_thres) continue; // require good flanking chains
|
||||
if (r0->re - r0->rs < max_gap>>1 || r0->qe - r0->qs < max_gap>>1) continue; // require enough flanking length
|
||||
if (r1->re - r1->rs < max_gap>>1 || r1->qe - r1->qs < max_gap>>1) continue;
|
||||
|
||||
// all conditions satisfied; join
|
||||
a[r1->as].y |= MM_SEED_LONG_JOIN;
|
||||
r0->cnt += r1->cnt, r0->score += r1->score;
|
||||
mm_reg_set_coor(r0, qlen, a);
|
||||
r1->cnt = 0;
|
||||
r1->parent = r0->id;
|
||||
++n_drop;
|
||||
}
|
||||
kfree(km, aux);
|
||||
|
||||
if (n_drop > 0) { // then fix the hits hierarchy
|
||||
for (i = 0; i < n_regs; ++i) { // adjust the mm_reg1_t::parent
|
||||
mm_reg1_t *r = ®s[i];
|
||||
if (r->parent >= 0 && r->id != r->parent) { // fix for secondary hits only
|
||||
if (regs[r->parent].parent >= 0 && regs[r->parent].parent != r->parent)
|
||||
r->parent = regs[r->parent].parent;
|
||||
}
|
||||
}
|
||||
mm_filter_regs(km, opt, n_regs_, regs);
|
||||
mm_sync_regs(km, *n_regs_, regs);
|
||||
}
|
||||
}
|
||||
|
||||
mm_seg_t *mm_seg_gen(void *km, uint32_t hash, int n_segs, const int *qlens, int n_regs0, const mm_reg1_t *regs0, int *n_regs, mm_reg1_t **regs, const mm128_t *a)
|
||||
{
|
||||
int s, i, j, acc_qlen[MM_MAX_SEG+1], qlen_sum = 0;
|
||||
@@ -388,7 +358,7 @@ mm_seg_t *mm_seg_gen(void *km, uint32_t hash, int n_segs, const int *qlens, int
|
||||
}
|
||||
}
|
||||
for (s = 0; s < n_segs; ++s) {
|
||||
regs[s] = mm_gen_regs(km, hash, qlens[s], seg[s].n_u, seg[s].u, seg[s].a);
|
||||
regs[s] = mm_gen_regs(km, hash, qlens[s], seg[s].n_u, seg[s].u, seg[s].a, 0);
|
||||
n_regs[s] = seg[s].n_u;
|
||||
for (i = 0; i < n_regs[s]; ++i) {
|
||||
regs[s][i].seg_split = 1;
|
||||
@@ -406,12 +376,39 @@ void mm_seg_free(void *km, int n_segs, mm_seg_t *segs)
|
||||
kfree(km, segs);
|
||||
}
|
||||
|
||||
void mm_set_mapq(int n_regs, mm_reg1_t *regs, int min_chain_sc, int match_sc, int rep_len, int is_sr)
|
||||
static void mm_set_inv_mapq(void *km, int n_regs, mm_reg1_t *regs)
|
||||
{
|
||||
int i, n_aux;
|
||||
mm128_t *aux;
|
||||
if (n_regs < 3) return;
|
||||
for (i = 0; i < n_regs; ++i)
|
||||
if (regs[i].inv) break;
|
||||
if (i == n_regs) return; // no inversion hits
|
||||
|
||||
aux = (mm128_t*)kmalloc(km, n_regs * 16);
|
||||
for (i = n_aux = 0; i < n_regs; ++i)
|
||||
if (regs[i].parent == i || regs[i].parent < 0)
|
||||
aux[n_aux].y = i, aux[n_aux++].x = (uint64_t)regs[i].rid << 32 | regs[i].rs;
|
||||
radix_sort_128x(aux, aux + n_aux);
|
||||
|
||||
for (i = 1; i < n_aux - 1; ++i) {
|
||||
mm_reg1_t *inv = ®s[aux[i].y];
|
||||
if (inv->inv) {
|
||||
mm_reg1_t *l = ®s[aux[i-1].y];
|
||||
mm_reg1_t *r = ®s[aux[i+1].y];
|
||||
inv->mapq = l->mapq < r->mapq? l->mapq : r->mapq;
|
||||
}
|
||||
}
|
||||
kfree(km, aux);
|
||||
}
|
||||
|
||||
void mm_set_mapq(void *km, int n_regs, mm_reg1_t *regs, int min_chain_sc, int match_sc, int rep_len, int is_sr)
|
||||
{
|
||||
static const float q_coef = 40.0f;
|
||||
int64_t sum_sc = 0;
|
||||
float uniq_ratio;
|
||||
int i;
|
||||
if (n_regs == 0) return;
|
||||
for (i = 0; i < n_regs; ++i)
|
||||
if (regs[i].parent == regs[i].id)
|
||||
sum_sc += regs[i].score;
|
||||
@@ -449,4 +446,5 @@ void mm_set_mapq(int n_regs, mm_reg1_t *regs, int min_chain_sc, int match_sc, in
|
||||
if (r->p && r->p->dp_max > r->p->dp_max2 && r->mapq == 0) r->mapq = 1;
|
||||
} else r->mapq = 0;
|
||||
}
|
||||
mm_set_inv_mapq(km, n_regs, regs);
|
||||
}
|
||||
|
||||
@@ -20,6 +20,8 @@
|
||||
KHASH_INIT(idx, uint64_t, uint64_t, 1, idx_hash, idx_eq)
|
||||
typedef khash_t(idx) idxhash_t;
|
||||
|
||||
KHASH_MAP_INIT_STR(str, uint32_t)
|
||||
|
||||
#define kroundup64(x) (--(x), (x)|=(x)>>1, (x)|=(x)>>2, (x)|=(x)>>4, (x)|=(x)>>8, (x)|=(x)>>16, (x)|=(x)>>32, ++(x))
|
||||
|
||||
typedef struct mm_idx_bucket_s {
|
||||
@@ -29,14 +31,15 @@ typedef struct mm_idx_bucket_s {
|
||||
void *h; // hash table indexing _p_ and minimizers appearing once
|
||||
} mm_idx_bucket_t;
|
||||
|
||||
void mm_idxopt_init(mm_idxopt_t *opt)
|
||||
{
|
||||
memset(opt, 0, sizeof(mm_idxopt_t));
|
||||
opt->k = 15, opt->w = 10, opt->flag = 0;
|
||||
opt->bucket_bits = 14;
|
||||
opt->mini_batch_size = 50000000;
|
||||
opt->batch_size = 4000000000ULL;
|
||||
}
|
||||
typedef struct {
|
||||
int32_t st, en, max; // max is not used for now
|
||||
int32_t score:30, strand:2;
|
||||
} mm_idx_intv1_t;
|
||||
|
||||
typedef struct mm_idx_intv_s {
|
||||
int32_t n, m;
|
||||
mm_idx_intv1_t *a;
|
||||
} mm_idx_intv_t;
|
||||
|
||||
mm_idx_t *mm_idx_init(int w, int k, int b, int flag)
|
||||
{
|
||||
@@ -52,12 +55,20 @@ mm_idx_t *mm_idx_init(int w, int k, int b, int flag)
|
||||
|
||||
void mm_idx_destroy(mm_idx_t *mi)
|
||||
{
|
||||
int i;
|
||||
uint32_t i;
|
||||
if (mi == 0) return;
|
||||
for (i = 0; i < 1<<mi->b; ++i) {
|
||||
free(mi->B[i].p);
|
||||
free(mi->B[i].a.a);
|
||||
kh_destroy(idx, (idxhash_t*)mi->B[i].h);
|
||||
if (mi->h) kh_destroy(str, (khash_t(str)*)mi->h);
|
||||
if (mi->B) {
|
||||
for (i = 0; i < 1U<<mi->b; ++i) {
|
||||
free(mi->B[i].p);
|
||||
free(mi->B[i].a.a);
|
||||
kh_destroy(idx, (idxhash_t*)mi->B[i].h);
|
||||
}
|
||||
}
|
||||
if (mi->I) {
|
||||
for (i = 0; i < mi->n_seq; ++i)
|
||||
free(mi->I[i].a);
|
||||
free(mi->I);
|
||||
}
|
||||
if (!mi->km) {
|
||||
for (i = 0; i < mi->n_seq; ++i)
|
||||
@@ -88,14 +99,15 @@ const uint64_t *mm_idx_get(const mm_idx_t *mi, uint64_t minier, int *n)
|
||||
|
||||
void mm_idx_stat(const mm_idx_t *mi)
|
||||
{
|
||||
int i, n = 0, n1 = 0;
|
||||
int n = 0, n1 = 0;
|
||||
uint32_t i;
|
||||
uint64_t sum = 0, len = 0;
|
||||
fprintf(stderr, "[M::%s] kmer size: %d; skip: %d; is_hpc: %d; #seq: %d\n", __func__, mi->k, mi->w, mi->flag&MM_I_HPC, mi->n_seq);
|
||||
for (i = 0; i < mi->n_seq; ++i)
|
||||
len += mi->seq[i].len;
|
||||
for (i = 0; i < 1<<mi->b; ++i)
|
||||
for (i = 0; i < 1U<<mi->b; ++i)
|
||||
if (mi->B[i].h) n += kh_size((idxhash_t*)mi->B[i].h);
|
||||
for (i = 0; i < 1<<mi->b; ++i) {
|
||||
for (i = 0; i < 1U<<mi->b; ++i) {
|
||||
idxhash_t *h = (idxhash_t*)mi->B[i].h;
|
||||
khint_t k;
|
||||
if (h == 0) continue;
|
||||
@@ -105,8 +117,36 @@ void mm_idx_stat(const mm_idx_t *mi)
|
||||
if (kh_key(h, k)&1) ++n1;
|
||||
}
|
||||
}
|
||||
fprintf(stderr, "[M::%s::%.3f*%.2f] distinct minimizers: %d (%.2f%% are singletons); average occurrences: %.3lf; average spacing: %.3lf\n",
|
||||
__func__, realtime() - mm_realtime0, cputime() / (realtime() - mm_realtime0), n, 100.0*n1/n, (double)sum / n, (double)len / sum);
|
||||
fprintf(stderr, "[M::%s::%.3f*%.2f] distinct minimizers: %d (%.2f%% are singletons); average occurrences: %.3lf; average spacing: %.3lf; total length: %ld\n",
|
||||
__func__, realtime() - mm_realtime0, cputime() / (realtime() - mm_realtime0), n, 100.0*n1/n, (double)sum / n, (double)len / sum, (long)len);
|
||||
}
|
||||
|
||||
int mm_idx_index_name(mm_idx_t *mi)
|
||||
{
|
||||
khash_t(str) *h;
|
||||
uint32_t i;
|
||||
int has_dup = 0, absent;
|
||||
if (mi->h) return 0;
|
||||
h = kh_init(str);
|
||||
for (i = 0; i < mi->n_seq; ++i) {
|
||||
khint_t k;
|
||||
k = kh_put(str, h, mi->seq[i].name, &absent);
|
||||
if (absent) kh_val(h, k) = i;
|
||||
else has_dup = 1;
|
||||
}
|
||||
mi->h = h;
|
||||
if (has_dup && mm_verbose >= 2)
|
||||
fprintf(stderr, "[WARNING] some database sequences have identical sequence names\n");
|
||||
return has_dup;
|
||||
}
|
||||
|
||||
int mm_idx_name2id(const mm_idx_t *mi, const char *name)
|
||||
{
|
||||
khash_t(str) *h = (khash_t(str)*)mi->h;
|
||||
khint_t k;
|
||||
if (h == 0) return -2;
|
||||
k = kh_get(str, h, name);
|
||||
return k == kh_end(h)? -1 : kh_val(h, k);
|
||||
}
|
||||
|
||||
int mm_idx_getseq(const mm_idx_t *mi, uint32_t rid, uint32_t st, uint32_t en, uint8_t *seq)
|
||||
@@ -121,6 +161,28 @@ int mm_idx_getseq(const mm_idx_t *mi, uint32_t rid, uint32_t st, uint32_t en, ui
|
||||
return en - st;
|
||||
}
|
||||
|
||||
int mm_idx_getseq_rev(const mm_idx_t *mi, uint32_t rid, uint32_t st, uint32_t en, uint8_t *seq)
|
||||
{
|
||||
uint64_t i, st1, en1;
|
||||
const mm_idx_seq_t *s;
|
||||
if (rid >= mi->n_seq || st >= mi->seq[rid].len) return -1;
|
||||
s = &mi->seq[rid];
|
||||
if (en > s->len) en = s->len;
|
||||
st1 = s->offset + (s->len - en);
|
||||
en1 = s->offset + (s->len - st);
|
||||
for (i = st1; i < en1; ++i) {
|
||||
uint8_t c = mm_seq4_get(mi->S, i);
|
||||
seq[en1 - i - 1] = c < 4? 3 - c : c;
|
||||
}
|
||||
return en - st;
|
||||
}
|
||||
|
||||
int mm_idx_getseq2(const mm_idx_t *mi, int is_rev, uint32_t rid, uint32_t st, uint32_t en, uint8_t *seq)
|
||||
{
|
||||
if (is_rev) return mm_idx_getseq_rev(mi, rid, st, en, seq);
|
||||
else return mm_idx_getseq(mi, rid, st, en, seq);
|
||||
}
|
||||
|
||||
int32_t mm_idx_cal_max_occ(const mm_idx_t *mi, float f)
|
||||
{
|
||||
int i;
|
||||
@@ -150,7 +212,8 @@ int32_t mm_idx_cal_max_occ(const mm_idx_t *mi, float f)
|
||||
|
||||
static void worker_post(void *g, long i, int tid)
|
||||
{
|
||||
int j, start_a, start_p, n, n_keys;
|
||||
int n, n_keys;
|
||||
size_t j, start_a, start_p;
|
||||
idxhash_t *h;
|
||||
mm_idx_t *mi = (mm_idx_t*)g;
|
||||
mm_idx_bucket_t *b = &mi->B[i];
|
||||
@@ -178,7 +241,7 @@ static void worker_post(void *g, long i, int tid)
|
||||
int absent;
|
||||
mm128_t *p = &b->a.a[j-1];
|
||||
itr = kh_put(idx, h, p->x>>8>>mi->b<<1, &absent);
|
||||
assert(absent && j - start_a == n);
|
||||
assert(absent && j == start_a + n);
|
||||
if (n == 1) {
|
||||
kh_key(h, itr) |= 1;
|
||||
kh_val(h, itr) = p->y;
|
||||
@@ -194,7 +257,7 @@ static void worker_post(void *g, long i, int tid)
|
||||
} else ++n;
|
||||
}
|
||||
b->h = h;
|
||||
assert(b->n == start_p);
|
||||
assert(b->n == (int32_t)start_p);
|
||||
|
||||
// deallocate and clear b->a
|
||||
kfree(0, b->a.a);
|
||||
@@ -275,6 +338,7 @@ static void *worker_pipeline(void *shared, int step, void *in)
|
||||
} else seq->name = 0;
|
||||
seq->len = s->seq[i].l_seq;
|
||||
seq->offset = p->sum_len;
|
||||
seq->is_alt = 0;
|
||||
// copy the sequence
|
||||
if (!(p->mi->flag & MM_I_NO_SEQ)) {
|
||||
for (j = 0; j < seq->len; ++j) { // TODO: this is not the fastest way, but let's first see if speed matters here
|
||||
@@ -314,7 +378,7 @@ mm_idx_t *mm_idx_gen(mm_bseq_file_t *fp, int w, int k, int b, int flag, int mini
|
||||
pipeline_t pl;
|
||||
if (fp == 0 || mm_bseq_eof(fp)) return 0;
|
||||
memset(&pl, 0, sizeof(pipeline_t));
|
||||
pl.mini_batch_size = mini_batch_size < batch_size? mini_batch_size : batch_size;
|
||||
pl.mini_batch_size = (uint64_t)mini_batch_size < batch_size? mini_batch_size : batch_size;
|
||||
pl.batch_size = batch_size;
|
||||
pl.fp = fp;
|
||||
pl.mi = mm_idx_init(w, k, b, flag);
|
||||
@@ -346,7 +410,9 @@ mm_idx_t *mm_idx_str(int w, int k, int is_hpc, int bucket_bits, int n, const cha
|
||||
uint64_t sum_len = 0;
|
||||
mm128_v a = {0,0,0};
|
||||
mm_idx_t *mi;
|
||||
khash_t(str) *h;
|
||||
int i, flag = 0;
|
||||
|
||||
if (n <= 0) return 0;
|
||||
for (i = 0; i < n; ++i) // get the total length
|
||||
sum_len += strlen(seq[i]);
|
||||
@@ -357,16 +423,21 @@ mm_idx_t *mm_idx_str(int w, int k, int is_hpc, int bucket_bits, int n, const cha
|
||||
mi->n_seq = n;
|
||||
mi->seq = (mm_idx_seq_t*)kcalloc(mi->km, n, sizeof(mm_idx_seq_t)); // ->seq is allocated from km
|
||||
mi->S = (uint32_t*)calloc((sum_len + 7) / 8, 4);
|
||||
mi->h = h = kh_init(str);
|
||||
for (i = 0, sum_len = 0; i < n; ++i) {
|
||||
const char *s = seq[i];
|
||||
mm_idx_seq_t *p = &mi->seq[i];
|
||||
uint32_t j;
|
||||
if (name && name[i]) {
|
||||
int absent;
|
||||
p->name = (char*)kmalloc(mi->km, strlen(name[i]) + 1);
|
||||
strcpy(p->name, name[i]);
|
||||
kh_put(str, h, p->name, &absent);
|
||||
assert(absent);
|
||||
}
|
||||
p->offset = sum_len;
|
||||
p->len = strlen(s);
|
||||
p->is_alt = 0;
|
||||
for (j = 0; j < p->len; ++j) {
|
||||
int c = seq_nt4_table[(uint8_t)s[j]];
|
||||
uint64_t o = sum_len + j;
|
||||
@@ -391,17 +462,20 @@ mm_idx_t *mm_idx_str(int w, int k, int is_hpc, int bucket_bits, int n, const cha
|
||||
void mm_idx_dump(FILE *fp, const mm_idx_t *mi)
|
||||
{
|
||||
uint64_t sum_len = 0;
|
||||
uint32_t x[5];
|
||||
int i;
|
||||
uint32_t x[5], i;
|
||||
|
||||
x[0] = mi->w, x[1] = mi->k, x[2] = mi->b, x[3] = mi->n_seq, x[4] = mi->flag;
|
||||
fwrite(MM_IDX_MAGIC, 1, 4, fp);
|
||||
fwrite(x, 4, 5, fp);
|
||||
for (i = 0; i < mi->n_seq; ++i) {
|
||||
uint8_t l;
|
||||
l = strlen(mi->seq[i].name);
|
||||
fwrite(&l, 1, 1, fp);
|
||||
fwrite(mi->seq[i].name, 1, l, fp);
|
||||
if (mi->seq[i].name) {
|
||||
uint8_t l = strlen(mi->seq[i].name);
|
||||
fwrite(&l, 1, 1, fp);
|
||||
fwrite(mi->seq[i].name, 1, l, fp);
|
||||
} else {
|
||||
uint8_t l = 0;
|
||||
fwrite(&l, 1, 1, fp);
|
||||
}
|
||||
fwrite(&mi->seq[i].len, 4, 1, fp);
|
||||
sum_len += mi->seq[i].len;
|
||||
}
|
||||
@@ -428,9 +502,8 @@ void mm_idx_dump(FILE *fp, const mm_idx_t *mi)
|
||||
|
||||
mm_idx_t *mm_idx_load(FILE *fp)
|
||||
{
|
||||
int i;
|
||||
char magic[4];
|
||||
uint32_t x[5];
|
||||
uint32_t x[5], i;
|
||||
uint64_t sum_len = 0;
|
||||
mm_idx_t *mi;
|
||||
|
||||
@@ -444,11 +517,14 @@ mm_idx_t *mm_idx_load(FILE *fp)
|
||||
uint8_t l;
|
||||
mm_idx_seq_t *s = &mi->seq[i];
|
||||
fread(&l, 1, 1, fp);
|
||||
s->name = (char*)kmalloc(mi->km, l + 1);
|
||||
fread(s->name, 1, l, fp);
|
||||
s->name[l] = 0;
|
||||
if (l) {
|
||||
s->name = (char*)kmalloc(mi->km, l + 1);
|
||||
fread(s->name, 1, l, fp);
|
||||
s->name[l] = 0;
|
||||
}
|
||||
fread(&s->len, 4, 1, fp);
|
||||
s->offset = sum_len;
|
||||
s->is_alt = 0;
|
||||
sum_len += s->len;
|
||||
}
|
||||
for (i = 0; i < 1<<mi->b; ++i) {
|
||||
@@ -482,14 +558,19 @@ mm_idx_t *mm_idx_load(FILE *fp)
|
||||
int64_t mm_idx_is_idx(const char *fn)
|
||||
{
|
||||
int fd, is_idx = 0;
|
||||
off_t ret, off_end;
|
||||
int64_t ret, off_end;
|
||||
char magic[4];
|
||||
|
||||
if (strcmp(fn, "-") == 0) return 0; // read from pipe; not an index
|
||||
fd = open(fn, O_RDONLY);
|
||||
if (fd < 0) return -1; // error
|
||||
#ifdef WIN32
|
||||
if ((off_end = _lseeki64(fd, 0, SEEK_END)) >= 4) {
|
||||
_lseeki64(fd, 0, SEEK_SET);
|
||||
#else
|
||||
if ((off_end = lseek(fd, 0, SEEK_END)) >= 4) {
|
||||
lseek(fd, 0, SEEK_SET);
|
||||
#endif // WIN32
|
||||
ret = read(fd, magic, 4);
|
||||
if (ret == 4 && strncmp(magic, MM_IDX_MAGIC, 4) == 0)
|
||||
is_idx = 1;
|
||||
@@ -535,7 +616,7 @@ mm_idx_t *mm_idx_reader_read(mm_idx_reader_t *r, int n_threads)
|
||||
mi = mm_idx_gen(r->fp.seq, r->opt.w, r->opt.k, r->opt.bucket_bits, r->opt.flag, r->opt.mini_batch_size, n_threads, r->opt.batch_size);
|
||||
if (mi) {
|
||||
if (r->fp_out) mm_idx_dump(r->fp_out, mi);
|
||||
++r->n_parts;
|
||||
mi->index = r->n_parts++;
|
||||
}
|
||||
return mi;
|
||||
}
|
||||
@@ -544,3 +625,151 @@ int mm_idx_reader_eof(const mm_idx_reader_t *r) // TODO: in extremely rare cases
|
||||
{
|
||||
return r->is_idx? (feof(r->fp.idx) || ftell(r->fp.idx) == r->idx_size) : mm_bseq_eof(r->fp.seq);
|
||||
}
|
||||
|
||||
#include <ctype.h>
|
||||
#include <zlib.h>
|
||||
#include "ksort.h"
|
||||
#include "kseq.h"
|
||||
KSTREAM_DECLARE(gzFile, gzread)
|
||||
|
||||
int mm_idx_alt_read(mm_idx_t *mi, const char *fn)
|
||||
{
|
||||
int n_alt = 0;
|
||||
gzFile fp;
|
||||
kstream_t *ks;
|
||||
kstring_t str = {0,0,0};
|
||||
fp = fn && strcmp(fn, "-")? gzopen(fn, "r") : gzdopen(fileno(stdin), "r");
|
||||
if (fp == 0) return -1;
|
||||
ks = ks_init(fp);
|
||||
if (mi->h == 0) mm_idx_index_name(mi);
|
||||
while (ks_getuntil(ks, KS_SEP_LINE, &str, 0) >= 0) {
|
||||
char *p;
|
||||
int id;
|
||||
for (p = str.s; *p && !isspace(*p); ++p) { }
|
||||
*p = 0;
|
||||
id = mm_idx_name2id(mi, str.s);
|
||||
if (id >= 0) mi->seq[id].is_alt = 1, ++n_alt;
|
||||
}
|
||||
mi->n_alt = n_alt;
|
||||
if (mm_verbose >= 3)
|
||||
fprintf(stderr, "[M::%s] found %d ALT contigs\n", __func__, n_alt);
|
||||
return n_alt;
|
||||
}
|
||||
|
||||
#define sort_key_bed(a) ((a).st)
|
||||
KRADIX_SORT_INIT(bed, mm_idx_intv1_t, sort_key_bed, 4)
|
||||
|
||||
mm_idx_intv_t *mm_idx_read_bed(const mm_idx_t *mi, const char *fn, int read_junc)
|
||||
{
|
||||
gzFile fp;
|
||||
kstream_t *ks;
|
||||
kstring_t str = {0,0,0};
|
||||
mm_idx_intv_t *I;
|
||||
|
||||
fp = fn && strcmp(fn, "-")? gzopen(fn, "r") : gzdopen(fileno(stdin), "r");
|
||||
if (fp == 0) return 0;
|
||||
I = (mm_idx_intv_t*)calloc(mi->n_seq, sizeof(*I));
|
||||
ks = ks_init(fp);
|
||||
while (ks_getuntil(ks, KS_SEP_LINE, &str, 0) >= 0) {
|
||||
mm_idx_intv_t *r;
|
||||
mm_idx_intv1_t t = {-1,-1,-1,-1,0};
|
||||
char *p, *q, *bl, *bs;
|
||||
int32_t i, id = -1, n_blk = 0;
|
||||
for (p = q = str.s, i = 0;; ++p) {
|
||||
if (*p == 0 || *p == '\t') {
|
||||
int32_t c = *p;
|
||||
*p = 0;
|
||||
if (i == 0) { // chr
|
||||
id = mm_idx_name2id(mi, q);
|
||||
if (id < 0) break; // unknown name; TODO: throw a warning
|
||||
} else if (i == 1) { // start
|
||||
t.st = atol(q); // TODO: watch out integer overflow!
|
||||
if (t.st < 0) break;
|
||||
} else if (i == 2) { // end
|
||||
t.en = atol(q);
|
||||
if (t.en < 0) break;
|
||||
} else if (i == 4) { // BED score
|
||||
t.score = atol(q);
|
||||
} else if (i == 5) { // strand
|
||||
t.strand = *q == '+'? 1 : *q == '-'? -1 : 0;
|
||||
} else if (i == 9) {
|
||||
if (!isdigit(*q)) break;
|
||||
n_blk = atol(q);
|
||||
} else if (i == 10) {
|
||||
bl = q;
|
||||
} else if (i == 11) {
|
||||
bs = q;
|
||||
break;
|
||||
}
|
||||
if (c == 0) break;
|
||||
++i, q = p + 1;
|
||||
}
|
||||
}
|
||||
if (id < 0 || t.st < 0 || t.st >= t.en) continue;
|
||||
r = &I[id];
|
||||
if (i >= 11 && read_junc) { // BED12
|
||||
int32_t st, sz, en;
|
||||
st = strtol(bs, &bs, 10); ++bs;
|
||||
sz = strtol(bl, &bl, 10); ++bl;
|
||||
en = t.st + st + sz;
|
||||
for (i = 1; i < n_blk; ++i) {
|
||||
mm_idx_intv1_t s = t;
|
||||
if (r->n == r->m) {
|
||||
r->m = r->m? r->m + (r->m>>1) : 16;
|
||||
r->a = (mm_idx_intv1_t*)realloc(r->a, sizeof(*r->a) * r->m);
|
||||
}
|
||||
st = strtol(bs, &bs, 10); ++bs;
|
||||
sz = strtol(bl, &bl, 10); ++bl;
|
||||
s.st = en, s.en = t.st + st;
|
||||
en = t.st + st + sz;
|
||||
if (s.en > s.st) r->a[r->n++] = s;
|
||||
}
|
||||
} else {
|
||||
if (r->n == r->m) {
|
||||
r->m = r->m? r->m + (r->m>>1) : 16;
|
||||
r->a = (mm_idx_intv1_t*)realloc(r->a, sizeof(*r->a) * r->m);
|
||||
}
|
||||
r->a[r->n++] = t;
|
||||
}
|
||||
}
|
||||
free(str.s);
|
||||
ks_destroy(ks);
|
||||
gzclose(fp);
|
||||
return I;
|
||||
}
|
||||
|
||||
int mm_idx_bed_read(mm_idx_t *mi, const char *fn, int read_junc)
|
||||
{
|
||||
int32_t i;
|
||||
if (mi->h == 0) mm_idx_index_name(mi);
|
||||
mi->I = mm_idx_read_bed(mi, fn, read_junc);
|
||||
if (mi->I == 0) return -1;
|
||||
for (i = 0; i < mi->n_seq; ++i) // TODO: eliminate redundant intervals
|
||||
radix_sort_bed(mi->I[i].a, mi->I[i].a + mi->I[i].n);
|
||||
return 0;
|
||||
}
|
||||
|
||||
int mm_idx_bed_junc(const mm_idx_t *mi, int32_t ctg, int32_t st, int32_t en, uint8_t *s)
|
||||
{
|
||||
int32_t i, left, right;
|
||||
mm_idx_intv_t *r;
|
||||
memset(s, 0, en - st);
|
||||
if (mi->I == 0 || ctg < 0 || ctg >= mi->n_seq) return -1;
|
||||
r = &mi->I[ctg];
|
||||
left = 0, right = r->n;
|
||||
while (right > left) {
|
||||
int32_t mid = left + ((right - left) >> 1);
|
||||
if (r->a[mid].st >= st) right = mid;
|
||||
else left = mid + 1;
|
||||
}
|
||||
for (i = left; i < r->n; ++i) {
|
||||
if (st <= r->a[i].st && en >= r->a[i].en && r->a[i].strand != 0) {
|
||||
if (r->a[i].strand > 0) {
|
||||
s[r->a[i].st - st] |= 1, s[r->a[i].en - 1 - st] |= 2;
|
||||
} else {
|
||||
s[r->a[i].st - st] |= 8, s[r->a[i].en - 1 - st] |= 4;
|
||||
}
|
||||
}
|
||||
}
|
||||
return left;
|
||||
}
|
||||
|
||||
@@ -18,15 +18,14 @@
|
||||
* | | | |
|
||||
* p=p->ptr->ptr->ptr->ptr p->ptr p->ptr->ptr p->ptr->ptr->ptr
|
||||
*/
|
||||
|
||||
#define MIN_CORE_SIZE 0x80000
|
||||
|
||||
typedef struct header_t {
|
||||
size_t size;
|
||||
struct header_t *ptr;
|
||||
} header_t;
|
||||
|
||||
typedef struct {
|
||||
void *par;
|
||||
size_t min_core_size;
|
||||
header_t base, *loop_head, *core_head; /* base is a zero-sized block always kept in the loop */
|
||||
} kmem_t;
|
||||
|
||||
@@ -36,31 +35,39 @@ static void panic(const char *s)
|
||||
abort();
|
||||
}
|
||||
|
||||
void *km_init(void)
|
||||
void *km_init2(void *km_par, size_t min_core_size)
|
||||
{
|
||||
return calloc(1, sizeof(kmem_t));
|
||||
kmem_t *km;
|
||||
km = (kmem_t*)kcalloc(km_par, 1, sizeof(kmem_t));
|
||||
km->par = km_par;
|
||||
km->min_core_size = min_core_size > 0? min_core_size : 0x80000;
|
||||
return (void*)km;
|
||||
}
|
||||
|
||||
void *km_init(void) { return km_init2(0, 0); }
|
||||
|
||||
void km_destroy(void *_km)
|
||||
{
|
||||
kmem_t *km = (kmem_t*)_km;
|
||||
void *km_par;
|
||||
header_t *p, *q;
|
||||
if (km == NULL) return;
|
||||
km_par = km->par;
|
||||
for (p = km->core_head; p != NULL;) {
|
||||
q = p->ptr;
|
||||
free(p);
|
||||
kfree(km_par, p);
|
||||
p = q;
|
||||
}
|
||||
free(km);
|
||||
kfree(km_par, km);
|
||||
}
|
||||
|
||||
static header_t *morecore(kmem_t *km, size_t nu)
|
||||
{
|
||||
header_t *q;
|
||||
size_t bytes, *p;
|
||||
nu = (nu + 1 + (MIN_CORE_SIZE - 1)) / MIN_CORE_SIZE * MIN_CORE_SIZE; /* the first +1 for core header */
|
||||
nu = (nu + 1 + (km->min_core_size - 1)) / km->min_core_size * km->min_core_size; /* the first +1 for core header */
|
||||
bytes = nu * sizeof(header_t);
|
||||
q = (header_t*)malloc(bytes);
|
||||
q = (header_t*)kmalloc(km->par, bytes);
|
||||
if (!q) panic("[morecore] insufficient memory");
|
||||
q->ptr = km->core_head, q->size = nu, km->core_head = q;
|
||||
p = (size_t*)(q + 1);
|
||||
@@ -125,7 +132,7 @@ void *kmalloc(void *_km, size_t n_bytes)
|
||||
|
||||
if (n_bytes == 0) return 0;
|
||||
if (km == NULL) return malloc(n_bytes);
|
||||
n_units = (n_bytes + sizeof(size_t) + sizeof(header_t) - 1) / sizeof(header_t) + 1;
|
||||
n_units = (n_bytes + sizeof(size_t) + sizeof(header_t) - 1) / sizeof(header_t); /* header+n_bytes requires at least this number of units */
|
||||
|
||||
if (!(q = km->loop_head)) /* the first time when kmalloc() is called, intialize it */
|
||||
q = km->loop_head = km->base.ptr = &km->base;
|
||||
@@ -160,18 +167,18 @@ void *kcalloc(void *_km, size_t count, size_t size)
|
||||
void *krealloc(void *_km, void *ap, size_t n_bytes) // TODO: this can be made more efficient in principle
|
||||
{
|
||||
kmem_t *km = (kmem_t*)_km;
|
||||
size_t n_units, *p, *q;
|
||||
size_t cap, *p, *q;
|
||||
|
||||
if (n_bytes == 0) {
|
||||
kfree(km, ap); return 0;
|
||||
}
|
||||
if (km == NULL) return realloc(ap, n_bytes);
|
||||
if (ap == NULL) return kmalloc(km, n_bytes);
|
||||
n_units = (n_bytes + sizeof(size_t) + sizeof(header_t) - 1) / sizeof(header_t);
|
||||
p = (size_t*)ap - 1;
|
||||
if (*p >= n_units) return ap; /* TODO: this prevents shrinking */
|
||||
cap = (*p) * sizeof(header_t) - sizeof(size_t);
|
||||
if (cap >= n_bytes) return ap; /* TODO: this prevents shrinking */
|
||||
q = (size_t*)kmalloc(km, n_bytes);
|
||||
memcpy(q, ap, (*p - 1) * sizeof(header_t));
|
||||
memcpy(q, ap, cap);
|
||||
kfree(km, ap);
|
||||
return q;
|
||||
}
|
||||
|
||||
@@ -17,6 +17,7 @@ void *kcalloc(void *km, size_t count, size_t size);
|
||||
void kfree(void *km, void *ptr);
|
||||
|
||||
void *km_init(void);
|
||||
void *km_init2(void *km_par, size_t min_core_size);
|
||||
void km_destroy(void *km);
|
||||
void km_stat(const void *_km, km_stat_t *s);
|
||||
|
||||
@@ -24,4 +25,52 @@ void km_stat(const void *_km, km_stat_t *s);
|
||||
}
|
||||
#endif
|
||||
|
||||
#define KMALLOC(km, ptr, len) ((ptr) = (__typeof__(ptr))kmalloc((km), (len) * sizeof(*(ptr))))
|
||||
#define KCALLOC(km, ptr, len) ((ptr) = (__typeof__(ptr))kcalloc((km), (len), sizeof(*(ptr))))
|
||||
#define KREALLOC(km, ptr, len) ((ptr) = (__typeof__(ptr))krealloc((km), (ptr), (len) * sizeof(*(ptr))))
|
||||
|
||||
#define KEXPAND(km, a, m) do { \
|
||||
(m) = (m) >= 4? (m) + ((m)>>1) : 16; \
|
||||
KREALLOC((km), (a), (m)); \
|
||||
} while (0)
|
||||
|
||||
#ifndef klib_unused
|
||||
#if (defined __clang__ && __clang_major__ >= 3) || (defined __GNUC__ && __GNUC__ >= 3)
|
||||
#define klib_unused __attribute__ ((__unused__))
|
||||
#else
|
||||
#define klib_unused
|
||||
#endif
|
||||
#endif /* klib_unused */
|
||||
|
||||
#define KALLOC_POOL_INIT2(SCOPE, name, kmptype_t) \
|
||||
typedef struct { \
|
||||
size_t cnt, n, max; \
|
||||
kmptype_t **buf; \
|
||||
void *km; \
|
||||
} kmp_##name##_t; \
|
||||
SCOPE kmp_##name##_t *kmp_init_##name(void *km) { \
|
||||
kmp_##name##_t *mp; \
|
||||
KCALLOC(km, mp, 1); \
|
||||
mp->km = km; \
|
||||
return mp; \
|
||||
} \
|
||||
SCOPE void kmp_destroy_##name(kmp_##name##_t *mp) { \
|
||||
size_t k; \
|
||||
for (k = 0; k < mp->n; ++k) kfree(mp->km, mp->buf[k]); \
|
||||
kfree(mp->km, mp->buf); kfree(mp->km, mp); \
|
||||
} \
|
||||
SCOPE kmptype_t *kmp_alloc_##name(kmp_##name##_t *mp) { \
|
||||
++mp->cnt; \
|
||||
if (mp->n == 0) return (kmptype_t*)kcalloc(mp->km, 1, sizeof(kmptype_t)); \
|
||||
return mp->buf[--mp->n]; \
|
||||
} \
|
||||
SCOPE void kmp_free_##name(kmp_##name##_t *mp, kmptype_t *p) { \
|
||||
--mp->cnt; \
|
||||
if (mp->n == mp->max) KEXPAND(mp->km, mp->buf, mp->max); \
|
||||
mp->buf[mp->n++] = p; \
|
||||
}
|
||||
|
||||
#define KALLOC_POOL_INIT(name, kmptype_t) \
|
||||
KALLOC_POOL_INIT2(static inline klib_unused, name, kmptype_t)
|
||||
|
||||
#endif
|
||||
|
||||
@@ -0,0 +1,120 @@
|
||||
#ifndef KETOPT_H
|
||||
#define KETOPT_H
|
||||
|
||||
#include <string.h> /* for strchr() and strncmp() */
|
||||
|
||||
#define ko_no_argument 0
|
||||
#define ko_required_argument 1
|
||||
#define ko_optional_argument 2
|
||||
|
||||
typedef struct {
|
||||
int ind; /* equivalent to optind */
|
||||
int opt; /* equivalent to optopt */
|
||||
char *arg; /* equivalent to optarg */
|
||||
int longidx; /* index of a long option; or -1 if short */
|
||||
/* private variables not intended for external uses */
|
||||
int i, pos, n_args;
|
||||
} ketopt_t;
|
||||
|
||||
typedef struct {
|
||||
char *name;
|
||||
int has_arg;
|
||||
int val;
|
||||
} ko_longopt_t;
|
||||
|
||||
static ketopt_t KETOPT_INIT = { 1, 0, 0, -1, 1, 0, 0 };
|
||||
|
||||
static void ketopt_permute(char *argv[], int j, int n) /* move argv[j] over n elements to the left */
|
||||
{
|
||||
int k;
|
||||
char *p = argv[j];
|
||||
for (k = 0; k < n; ++k)
|
||||
argv[j - k] = argv[j - k - 1];
|
||||
argv[j - k] = p;
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse command-line options and arguments
|
||||
*
|
||||
* This fuction has a similar interface to GNU's getopt_long(). Each call
|
||||
* parses one option and returns the option name. s->arg points to the option
|
||||
* argument if present. The function returns -1 when all command-line arguments
|
||||
* are parsed. In this case, s->ind is the index of the first non-option
|
||||
* argument.
|
||||
*
|
||||
* @param s status; shall be initialized to KETOPT_INIT on the first call
|
||||
* @param argc length of argv[]
|
||||
* @param argv list of command-line arguments; argv[0] is ignored
|
||||
* @param permute non-zero to move options ahead of non-option arguments
|
||||
* @param ostr option string
|
||||
* @param longopts long options
|
||||
*
|
||||
* @return ASCII for a short option; ko_longopt_t::val for a long option; -1 if
|
||||
* argv[] is fully processed; '?' for an unknown option or an ambiguous
|
||||
* long option; ':' if an option argument is missing
|
||||
*/
|
||||
static int ketopt(ketopt_t *s, int argc, char *argv[], int permute, const char *ostr, const ko_longopt_t *longopts)
|
||||
{
|
||||
int opt = -1, i0, j;
|
||||
if (permute) {
|
||||
while (s->i < argc && (argv[s->i][0] != '-' || argv[s->i][1] == '\0'))
|
||||
++s->i, ++s->n_args;
|
||||
}
|
||||
s->arg = 0, s->longidx = -1, i0 = s->i;
|
||||
if (s->i >= argc || argv[s->i][0] != '-' || argv[s->i][1] == '\0') {
|
||||
s->ind = s->i - s->n_args;
|
||||
return -1;
|
||||
}
|
||||
if (argv[s->i][0] == '-' && argv[s->i][1] == '-') { /* "--" or a long option */
|
||||
if (argv[s->i][2] == '\0') { /* a bare "--" */
|
||||
ketopt_permute(argv, s->i, s->n_args);
|
||||
++s->i, s->ind = s->i - s->n_args;
|
||||
return -1;
|
||||
}
|
||||
s->opt = 0, opt = '?', s->pos = -1;
|
||||
if (longopts) { /* parse long options */
|
||||
int k, n_exact = 0, n_partial = 0;
|
||||
const ko_longopt_t *o = 0, *o_exact = 0, *o_partial = 0;
|
||||
for (j = 2; argv[s->i][j] != '\0' && argv[s->i][j] != '='; ++j) {} /* find the end of the option name */
|
||||
for (k = 0; longopts[k].name != 0; ++k)
|
||||
if (strncmp(&argv[s->i][2], longopts[k].name, j - 2) == 0) {
|
||||
if (longopts[k].name[j - 2] == 0) ++n_exact, o_exact = &longopts[k];
|
||||
else ++n_partial, o_partial = &longopts[k];
|
||||
}
|
||||
if (n_exact > 1 || (n_exact == 0 && n_partial > 1)) return '?';
|
||||
o = n_exact == 1? o_exact : n_partial == 1? o_partial : 0;
|
||||
if (o) {
|
||||
s->opt = opt = o->val, s->longidx = o - longopts;
|
||||
if (argv[s->i][j] == '=') s->arg = &argv[s->i][j + 1];
|
||||
if (o->has_arg == 1 && argv[s->i][j] == '\0') {
|
||||
if (s->i < argc - 1) s->arg = argv[++s->i];
|
||||
else opt = ':'; /* missing option argument */
|
||||
}
|
||||
}
|
||||
}
|
||||
} else { /* a short option */
|
||||
char *p;
|
||||
if (s->pos == 0) s->pos = 1;
|
||||
opt = s->opt = argv[s->i][s->pos++];
|
||||
p = strchr((char*)ostr, opt);
|
||||
if (p == 0) {
|
||||
opt = '?'; /* unknown option */
|
||||
} else if (p[1] == ':') {
|
||||
if (argv[s->i][s->pos] == 0) {
|
||||
if (s->i < argc - 1) s->arg = argv[++s->i];
|
||||
else opt = ':'; /* missing option argument */
|
||||
} else s->arg = &argv[s->i][s->pos];
|
||||
s->pos = -1;
|
||||
}
|
||||
}
|
||||
if (s->pos < 0 || argv[s->i][s->pos] == 0) {
|
||||
++s->i, s->pos = 0;
|
||||
if (s->n_args > 0) /* permute */
|
||||
for (j = i0; j < s->i; ++j)
|
||||
ketopt_permute(argv, j, s->n_args);
|
||||
}
|
||||
s->ind = s->i - s->n_args;
|
||||
return opt;
|
||||
}
|
||||
|
||||
#endif
|
||||
@@ -0,0 +1,474 @@
|
||||
/* The MIT License
|
||||
|
||||
Copyright (c) 2019 by Attractive Chaos <attractor@live.co.uk>
|
||||
|
||||
Permission is hereby granted, free of charge, to any person obtaining
|
||||
a copy of this software and associated documentation files (the
|
||||
"Software"), to deal in the Software without restriction, including
|
||||
without limitation the rights to use, copy, modify, merge, publish,
|
||||
distribute, sublicense, and/or sell copies of the Software, and to
|
||||
permit persons to whom the Software is furnished to do so, subject to
|
||||
the following conditions:
|
||||
|
||||
The above copyright notice and this permission notice shall be
|
||||
included in all copies or substantial portions of the Software.
|
||||
|
||||
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,
|
||||
EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF
|
||||
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND
|
||||
NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS
|
||||
BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN
|
||||
ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN
|
||||
CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||
SOFTWARE.
|
||||
*/
|
||||
|
||||
/* An example:
|
||||
|
||||
#include <stdio.h>
|
||||
#include <string.h>
|
||||
#include <stdlib.h>
|
||||
#include "krmq.h"
|
||||
|
||||
struct my_node {
|
||||
char key;
|
||||
KRMQ_HEAD(struct my_node) head;
|
||||
};
|
||||
#define my_cmp(p, q) (((q)->key < (p)->key) - ((p)->key < (q)->key))
|
||||
KRMQ_INIT(my, struct my_node, head, my_cmp)
|
||||
|
||||
int main(void) {
|
||||
const char *str = "MNOLKQOPHIA"; // from wiki, except a duplicate
|
||||
struct my_node *root = 0;
|
||||
int i, l = strlen(str);
|
||||
for (i = 0; i < l; ++i) { // insert in the input order
|
||||
struct my_node *q, *p = malloc(sizeof(*p));
|
||||
p->key = str[i];
|
||||
q = krmq_insert(my, &root, p, 0);
|
||||
if (p != q) free(p); // if already present, free
|
||||
}
|
||||
krmq_itr_t(my) itr;
|
||||
krmq_itr_first(my, root, &itr); // place at first
|
||||
do { // traverse
|
||||
const struct my_node *p = krmq_at(&itr);
|
||||
putchar(p->key);
|
||||
free((void*)p); // free node
|
||||
} while (krmq_itr_next(my, &itr));
|
||||
putchar('\n');
|
||||
return 0;
|
||||
}
|
||||
*/
|
||||
|
||||
#ifndef KRMQ_H
|
||||
#define KRMQ_H
|
||||
|
||||
#ifdef __STRICT_ANSI__
|
||||
#define inline __inline__
|
||||
#endif
|
||||
|
||||
#define KRMQ_MAX_DEPTH 64
|
||||
|
||||
#define krmq_size(head, p) ((p)? (p)->head.size : 0)
|
||||
#define krmq_size_child(head, q, i) ((q)->head.p[(i)]? (q)->head.p[(i)]->head.size : 0)
|
||||
|
||||
#define KRMQ_HEAD(__type) \
|
||||
struct { \
|
||||
__type *p[2], *s; \
|
||||
signed char balance; /* balance factor */ \
|
||||
unsigned size; /* #elements in subtree */ \
|
||||
}
|
||||
|
||||
#define __KRMQ_FIND(suf, __scope, __type, __head, __cmp) \
|
||||
__scope __type *krmq_find_##suf(const __type *root, const __type *x, unsigned *cnt_) { \
|
||||
const __type *p = root; \
|
||||
unsigned cnt = 0; \
|
||||
while (p != 0) { \
|
||||
int cmp; \
|
||||
cmp = __cmp(x, p); \
|
||||
if (cmp >= 0) cnt += krmq_size_child(__head, p, 0) + 1; \
|
||||
if (cmp < 0) p = p->__head.p[0]; \
|
||||
else if (cmp > 0) p = p->__head.p[1]; \
|
||||
else break; \
|
||||
} \
|
||||
if (cnt_) *cnt_ = cnt; \
|
||||
return (__type*)p; \
|
||||
} \
|
||||
__scope __type *krmq_interval_##suf(const __type *root, const __type *x, __type **lower, __type **upper) { \
|
||||
const __type *p = root, *l = 0, *u = 0; \
|
||||
while (p != 0) { \
|
||||
int cmp; \
|
||||
cmp = __cmp(x, p); \
|
||||
if (cmp < 0) u = p, p = p->__head.p[0]; \
|
||||
else if (cmp > 0) l = p, p = p->__head.p[1]; \
|
||||
else { l = u = p; break; } \
|
||||
} \
|
||||
if (lower) *lower = (__type*)l; \
|
||||
if (upper) *upper = (__type*)u; \
|
||||
return (__type*)p; \
|
||||
}
|
||||
|
||||
#define __KRMQ_RMQ(suf, __scope, __type, __head, __cmp, __lt2) \
|
||||
__scope __type *krmq_rmq_##suf(const __type *root, const __type *lo, const __type *up) { /* CLOSED interval */ \
|
||||
const __type *p = root, *path[2][KRMQ_MAX_DEPTH], *min; \
|
||||
int plen[2] = {0, 0}, pcmp[2][KRMQ_MAX_DEPTH], i, cmp, lca; \
|
||||
if (root == 0) return 0; \
|
||||
while (p) { \
|
||||
cmp = __cmp(lo, p); \
|
||||
path[0][plen[0]] = p, pcmp[0][plen[0]++] = cmp; \
|
||||
if (cmp < 0) p = p->__head.p[0]; \
|
||||
else if (cmp > 0) p = p->__head.p[1]; \
|
||||
else break; \
|
||||
} \
|
||||
p = root; \
|
||||
while (p) { \
|
||||
cmp = __cmp(up, p); \
|
||||
path[1][plen[1]] = p, pcmp[1][plen[1]++] = cmp; \
|
||||
if (cmp < 0) p = p->__head.p[0]; \
|
||||
else if (cmp > 0) p = p->__head.p[1]; \
|
||||
else break; \
|
||||
} \
|
||||
for (i = 0; i < plen[0] && i < plen[1]; ++i) /* find the LCA */ \
|
||||
if (path[0][i] == path[1][i] && pcmp[0][i] <= 0 && pcmp[1][i] >= 0) \
|
||||
break; \
|
||||
if (i == plen[0] || i == plen[1]) return 0; /* no elements in the closed interval */ \
|
||||
lca = i, min = path[0][lca]; \
|
||||
for (i = lca + 1; i < plen[0]; ++i) { \
|
||||
if (pcmp[0][i] <= 0) { \
|
||||
if (__lt2(path[0][i], min)) min = path[0][i]; \
|
||||
if (path[0][i]->__head.p[1] && __lt2(path[0][i]->__head.p[1]->__head.s, min)) \
|
||||
min = path[0][i]->__head.p[1]->__head.s; \
|
||||
} \
|
||||
} \
|
||||
for (i = lca + 1; i < plen[1]; ++i) { \
|
||||
if (pcmp[1][i] >= 0) { \
|
||||
if (__lt2(path[1][i], min)) min = path[1][i]; \
|
||||
if (path[1][i]->__head.p[0] && __lt2(path[1][i]->__head.p[0]->__head.s, min)) \
|
||||
min = path[1][i]->__head.p[0]->__head.s; \
|
||||
} \
|
||||
} \
|
||||
return (__type*)min; \
|
||||
}
|
||||
|
||||
#define __KRMQ_ROTATE(suf, __type, __head, __lt2) \
|
||||
/* */ \
|
||||
static inline void krmq_update_min_##suf(__type *p, const __type *q, const __type *r) { \
|
||||
p->__head.s = !q || __lt2(p, q->__head.s)? p : q->__head.s; \
|
||||
p->__head.s = !r || __lt2(p->__head.s, r->__head.s)? p->__head.s : r->__head.s; \
|
||||
} \
|
||||
/* one rotation: (a,(b,c)q)p => ((a,b)p,c)q */ \
|
||||
static inline __type *krmq_rotate1_##suf(__type *p, int dir) { /* dir=0 to left; dir=1 to right */ \
|
||||
int opp = 1 - dir; /* opposite direction */ \
|
||||
__type *q = p->__head.p[opp], *s = p->__head.s; \
|
||||
unsigned size_p = p->__head.size; \
|
||||
p->__head.size -= q->__head.size - krmq_size_child(__head, q, dir); \
|
||||
q->__head.size = size_p; \
|
||||
krmq_update_min_##suf(p, p->__head.p[dir], q->__head.p[dir]); \
|
||||
q->__head.s = s; \
|
||||
p->__head.p[opp] = q->__head.p[dir]; \
|
||||
q->__head.p[dir] = p; \
|
||||
return q; \
|
||||
} \
|
||||
/* two consecutive rotations: (a,((b,c)r,d)q)p => ((a,b)p,(c,d)q)r */ \
|
||||
static inline __type *krmq_rotate2_##suf(__type *p, int dir) { \
|
||||
int b1, opp = 1 - dir; \
|
||||
__type *q = p->__head.p[opp], *r = q->__head.p[dir], *s = p->__head.s; \
|
||||
unsigned size_x_dir = krmq_size_child(__head, r, dir); \
|
||||
r->__head.size = p->__head.size; \
|
||||
p->__head.size -= q->__head.size - size_x_dir; \
|
||||
q->__head.size -= size_x_dir + 1; \
|
||||
krmq_update_min_##suf(p, p->__head.p[dir], r->__head.p[dir]); \
|
||||
krmq_update_min_##suf(q, q->__head.p[opp], r->__head.p[opp]); \
|
||||
r->__head.s = s; \
|
||||
p->__head.p[opp] = r->__head.p[dir]; \
|
||||
r->__head.p[dir] = p; \
|
||||
q->__head.p[dir] = r->__head.p[opp]; \
|
||||
r->__head.p[opp] = q; \
|
||||
b1 = dir == 0? +1 : -1; \
|
||||
if (r->__head.balance == b1) q->__head.balance = 0, p->__head.balance = -b1; \
|
||||
else if (r->__head.balance == 0) q->__head.balance = p->__head.balance = 0; \
|
||||
else q->__head.balance = b1, p->__head.balance = 0; \
|
||||
r->__head.balance = 0; \
|
||||
return r; \
|
||||
}
|
||||
|
||||
#define __KRMQ_INSERT(suf, __scope, __type, __head, __cmp, __lt2) \
|
||||
__scope __type *krmq_insert_##suf(__type **root_, __type *x, unsigned *cnt_) { \
|
||||
unsigned char stack[KRMQ_MAX_DEPTH]; \
|
||||
__type *path[KRMQ_MAX_DEPTH]; \
|
||||
__type *bp, *bq; \
|
||||
__type *p, *q, *r = 0; /* _r_ is potentially the new root */ \
|
||||
int i, which = 0, top, b1, path_len; \
|
||||
unsigned cnt = 0; \
|
||||
bp = *root_, bq = 0; \
|
||||
/* find the insertion location */ \
|
||||
for (p = bp, q = bq, top = path_len = 0; p; q = p, p = p->__head.p[which]) { \
|
||||
int cmp; \
|
||||
cmp = __cmp(x, p); \
|
||||
if (cmp >= 0) cnt += krmq_size_child(__head, p, 0) + 1; \
|
||||
if (cmp == 0) { \
|
||||
if (cnt_) *cnt_ = cnt; \
|
||||
return p; \
|
||||
} \
|
||||
if (p->__head.balance != 0) \
|
||||
bq = q, bp = p, top = 0; \
|
||||
stack[top++] = which = (cmp > 0); \
|
||||
path[path_len++] = p; \
|
||||
} \
|
||||
if (cnt_) *cnt_ = cnt; \
|
||||
x->__head.balance = 0, x->__head.size = 1, x->__head.p[0] = x->__head.p[1] = 0, x->__head.s = x; \
|
||||
if (q == 0) *root_ = x; \
|
||||
else q->__head.p[which] = x; \
|
||||
if (bp == 0) return x; \
|
||||
for (i = 0; i < path_len; ++i) ++path[i]->__head.size; \
|
||||
for (i = path_len - 1; i >= 0; --i) { \
|
||||
krmq_update_min_##suf(path[i], path[i]->__head.p[0], path[i]->__head.p[1]); \
|
||||
if (path[i]->__head.s != x) break; \
|
||||
} \
|
||||
for (p = bp, top = 0; p != x; p = p->__head.p[stack[top]], ++top) /* update balance factors */ \
|
||||
if (stack[top] == 0) --p->__head.balance; \
|
||||
else ++p->__head.balance; \
|
||||
if (bp->__head.balance > -2 && bp->__head.balance < 2) return x; /* no re-balance needed */ \
|
||||
/* re-balance */ \
|
||||
which = (bp->__head.balance < 0); \
|
||||
b1 = which == 0? +1 : -1; \
|
||||
q = bp->__head.p[1 - which]; \
|
||||
if (q->__head.balance == b1) { \
|
||||
r = krmq_rotate1_##suf(bp, which); \
|
||||
q->__head.balance = bp->__head.balance = 0; \
|
||||
} else r = krmq_rotate2_##suf(bp, which); \
|
||||
if (bq == 0) *root_ = r; \
|
||||
else bq->__head.p[bp != bq->__head.p[0]] = r; \
|
||||
return x; \
|
||||
}
|
||||
|
||||
#define __KRMQ_ERASE(suf, __scope, __type, __head, __cmp, __lt2) \
|
||||
__scope __type *krmq_erase_##suf(__type **root_, const __type *x, unsigned *cnt_) { \
|
||||
__type *p, *path[KRMQ_MAX_DEPTH], fake; \
|
||||
unsigned char dir[KRMQ_MAX_DEPTH]; \
|
||||
int i, d = 0, cmp; \
|
||||
unsigned cnt = 0; \
|
||||
fake = **root_, fake.__head.p[0] = *root_, fake.__head.p[1] = 0; \
|
||||
if (cnt_) *cnt_ = 0; \
|
||||
if (x) { \
|
||||
for (cmp = -1, p = &fake; cmp; cmp = __cmp(x, p)) { \
|
||||
int which = (cmp > 0); \
|
||||
if (cmp > 0) cnt += krmq_size_child(__head, p, 0) + 1; \
|
||||
dir[d] = which; \
|
||||
path[d++] = p; \
|
||||
p = p->__head.p[which]; \
|
||||
if (p == 0) { \
|
||||
if (cnt_) *cnt_ = 0; \
|
||||
return 0; \
|
||||
} \
|
||||
} \
|
||||
cnt += krmq_size_child(__head, p, 0) + 1; /* because p==x is not counted */ \
|
||||
} else { \
|
||||
for (p = &fake, cnt = 1; p; p = p->__head.p[0]) \
|
||||
dir[d] = 0, path[d++] = p; \
|
||||
p = path[--d]; \
|
||||
} \
|
||||
if (cnt_) *cnt_ = cnt; \
|
||||
for (i = 1; i < d; ++i) --path[i]->__head.size; \
|
||||
if (p->__head.p[1] == 0) { /* ((1,.)2,3)4 => (1,3)4; p=2 */ \
|
||||
path[d-1]->__head.p[dir[d-1]] = p->__head.p[0]; \
|
||||
} else { \
|
||||
__type *q = p->__head.p[1]; \
|
||||
if (q->__head.p[0] == 0) { /* ((1,2)3,4)5 => ((1)2,4)5; p=3,q=2 */ \
|
||||
q->__head.p[0] = p->__head.p[0]; \
|
||||
q->__head.balance = p->__head.balance; \
|
||||
path[d-1]->__head.p[dir[d-1]] = q; \
|
||||
path[d] = q, dir[d++] = 1; \
|
||||
q->__head.size = p->__head.size - 1; \
|
||||
} else { /* ((1,((.,2)3,4)5)6,7)8 => ((1,(2,4)5)3,7)8; p=6 */ \
|
||||
__type *r; \
|
||||
int e = d++; /* backup _d_ */\
|
||||
for (;;) { \
|
||||
dir[d] = 0; \
|
||||
path[d++] = q; \
|
||||
r = q->__head.p[0]; \
|
||||
if (r->__head.p[0] == 0) break; \
|
||||
q = r; \
|
||||
} \
|
||||
r->__head.p[0] = p->__head.p[0]; \
|
||||
q->__head.p[0] = r->__head.p[1]; \
|
||||
r->__head.p[1] = p->__head.p[1]; \
|
||||
r->__head.balance = p->__head.balance; \
|
||||
path[e-1]->__head.p[dir[e-1]] = r; \
|
||||
path[e] = r, dir[e] = 1; \
|
||||
for (i = e + 1; i < d; ++i) --path[i]->__head.size; \
|
||||
r->__head.size = p->__head.size - 1; \
|
||||
} \
|
||||
} \
|
||||
for (i = d - 1; i >= 0; --i) /* not sure why adding condition "path[i]->__head.s==p" doesn't work */ \
|
||||
krmq_update_min_##suf(path[i], path[i]->__head.p[0], path[i]->__head.p[1]); \
|
||||
while (--d > 0) { \
|
||||
__type *q = path[d]; \
|
||||
int which, other, b1 = 1, b2 = 2; \
|
||||
which = dir[d], other = 1 - which; \
|
||||
if (which) b1 = -b1, b2 = -b2; \
|
||||
q->__head.balance += b1; \
|
||||
if (q->__head.balance == b1) break; \
|
||||
else if (q->__head.balance == b2) { \
|
||||
__type *r = q->__head.p[other]; \
|
||||
if (r->__head.balance == -b1) { \
|
||||
path[d-1]->__head.p[dir[d-1]] = krmq_rotate2_##suf(q, which); \
|
||||
} else { \
|
||||
path[d-1]->__head.p[dir[d-1]] = krmq_rotate1_##suf(q, which); \
|
||||
if (r->__head.balance == 0) { \
|
||||
r->__head.balance = -b1; \
|
||||
q->__head.balance = b1; \
|
||||
break; \
|
||||
} else r->__head.balance = q->__head.balance = 0; \
|
||||
} \
|
||||
} \
|
||||
} \
|
||||
*root_ = fake.__head.p[0]; \
|
||||
return p; \
|
||||
}
|
||||
|
||||
#define krmq_free(__type, __head, __root, __free) do { \
|
||||
__type *_p, *_q; \
|
||||
for (_p = __root; _p; _p = _q) { \
|
||||
if (_p->__head.p[0] == 0) { \
|
||||
_q = _p->__head.p[1]; \
|
||||
__free(_p); \
|
||||
} else { \
|
||||
_q = _p->__head.p[0]; \
|
||||
_p->__head.p[0] = _q->__head.p[1]; \
|
||||
_q->__head.p[1] = _p; \
|
||||
} \
|
||||
} \
|
||||
} while (0)
|
||||
|
||||
#define __KRMQ_ITR(suf, __scope, __type, __head, __cmp) \
|
||||
struct krmq_itr_##suf { \
|
||||
const __type *stack[KRMQ_MAX_DEPTH], **top; \
|
||||
}; \
|
||||
__scope void krmq_itr_first_##suf(const __type *root, struct krmq_itr_##suf *itr) { \
|
||||
const __type *p; \
|
||||
for (itr->top = itr->stack - 1, p = root; p; p = p->__head.p[0]) \
|
||||
*++itr->top = p; \
|
||||
} \
|
||||
__scope int krmq_itr_find_##suf(const __type *root, const __type *x, struct krmq_itr_##suf *itr) { \
|
||||
const __type *p = root; \
|
||||
itr->top = itr->stack - 1; \
|
||||
while (p != 0) { \
|
||||
int cmp; \
|
||||
*++itr->top = p; \
|
||||
cmp = __cmp(x, p); \
|
||||
if (cmp < 0) p = p->__head.p[0]; \
|
||||
else if (cmp > 0) p = p->__head.p[1]; \
|
||||
else break; \
|
||||
} \
|
||||
return p? 1 : 0; \
|
||||
} \
|
||||
__scope int krmq_itr_next_bidir_##suf(struct krmq_itr_##suf *itr, int dir) { \
|
||||
const __type *p; \
|
||||
if (itr->top < itr->stack) return 0; \
|
||||
dir = !!dir; \
|
||||
p = (*itr->top)->__head.p[dir]; \
|
||||
if (p) { /* go down */ \
|
||||
for (; p; p = p->__head.p[!dir]) \
|
||||
*++itr->top = p; \
|
||||
return 1; \
|
||||
} else { /* go up */ \
|
||||
do { \
|
||||
p = *itr->top--; \
|
||||
} while (itr->top >= itr->stack && p == (*itr->top)->__head.p[dir]); \
|
||||
return itr->top < itr->stack? 0 : 1; \
|
||||
} \
|
||||
} \
|
||||
|
||||
/**
|
||||
* Insert a node to the tree
|
||||
*
|
||||
* @param suf name suffix used in KRMQ_INIT()
|
||||
* @param proot pointer to the root of the tree (in/out: root may change)
|
||||
* @param x node to insert (in)
|
||||
* @param cnt number of nodes smaller than or equal to _x_; can be NULL (out)
|
||||
*
|
||||
* @return _x_ if not present in the tree, or the node equal to x.
|
||||
*/
|
||||
#define krmq_insert(suf, proot, x, cnt) krmq_insert_##suf(proot, x, cnt)
|
||||
|
||||
/**
|
||||
* Find a node in the tree
|
||||
*
|
||||
* @param suf name suffix used in KRMQ_INIT()
|
||||
* @param root root of the tree
|
||||
* @param x node value to find (in)
|
||||
* @param cnt number of nodes smaller than or equal to _x_; can be NULL (out)
|
||||
*
|
||||
* @return node equal to _x_ if present, or NULL if absent
|
||||
*/
|
||||
#define krmq_find(suf, root, x, cnt) krmq_find_##suf(root, x, cnt)
|
||||
#define krmq_interval(suf, root, x, lower, upper) krmq_interval_##suf(root, x, lower, upper)
|
||||
#define krmq_rmq(suf, root, lo, up) krmq_rmq_##suf(root, lo, up)
|
||||
|
||||
/**
|
||||
* Delete a node from the tree
|
||||
*
|
||||
* @param suf name suffix used in KRMQ_INIT()
|
||||
* @param proot pointer to the root of the tree (in/out: root may change)
|
||||
* @param x node value to delete; if NULL, delete the first node (in)
|
||||
*
|
||||
* @return node removed from the tree if present, or NULL if absent
|
||||
*/
|
||||
#define krmq_erase(suf, proot, x, cnt) krmq_erase_##suf(proot, x, cnt)
|
||||
#define krmq_erase_first(suf, proot) krmq_erase_##suf(proot, 0, 0)
|
||||
|
||||
#define krmq_itr_t(suf) struct krmq_itr_##suf
|
||||
|
||||
/**
|
||||
* Place the iterator at the smallest object
|
||||
*
|
||||
* @param suf name suffix used in KRMQ_INIT()
|
||||
* @param root root of the tree
|
||||
* @param itr iterator
|
||||
*/
|
||||
#define krmq_itr_first(suf, root, itr) krmq_itr_first_##suf(root, itr)
|
||||
|
||||
/**
|
||||
* Place the iterator at the object equal to or greater than the query
|
||||
*
|
||||
* @param suf name suffix used in KRMQ_INIT()
|
||||
* @param root root of the tree
|
||||
* @param x query (in)
|
||||
* @param itr iterator (out)
|
||||
*
|
||||
* @return 1 if find; 0 otherwise. krmq_at(itr) is NULL if and only if query is
|
||||
* larger than all objects in the tree
|
||||
*/
|
||||
#define krmq_itr_find(suf, root, x, itr) krmq_itr_find_##suf(root, x, itr)
|
||||
|
||||
/**
|
||||
* Move to the next object in order
|
||||
*
|
||||
* @param itr iterator (modified)
|
||||
*
|
||||
* @return 1 if there is a next object; 0 otherwise
|
||||
*/
|
||||
#define krmq_itr_next(suf, itr) krmq_itr_next_bidir_##suf(itr, 1)
|
||||
#define krmq_itr_prev(suf, itr) krmq_itr_next_bidir_##suf(itr, 0)
|
||||
|
||||
/**
|
||||
* Return the pointer at the iterator
|
||||
*
|
||||
* @param itr iterator
|
||||
*
|
||||
* @return pointer if present; NULL otherwise
|
||||
*/
|
||||
#define krmq_at(itr) ((itr)->top < (itr)->stack? 0 : *(itr)->top)
|
||||
|
||||
#define KRMQ_INIT2(suf, __scope, __type, __head, __cmp, __lt2) \
|
||||
__KRMQ_FIND(suf, __scope, __type, __head, __cmp) \
|
||||
__KRMQ_RMQ(suf, __scope, __type, __head, __cmp, __lt2) \
|
||||
__KRMQ_ROTATE(suf, __type, __head, __lt2) \
|
||||
__KRMQ_INSERT(suf, __scope, __type, __head, __cmp, __lt2) \
|
||||
__KRMQ_ERASE(suf, __scope, __type, __head, __cmp, __lt2) \
|
||||
__KRMQ_ITR(suf, __scope, __type, __head, __cmp)
|
||||
|
||||
#define KRMQ_INIT(suf, __type, __head, __cmp, __lt2) \
|
||||
KRMQ_INIT2(suf,, __type, __head, __cmp, __lt2)
|
||||
|
||||
#endif
|
||||
@@ -37,6 +37,14 @@
|
||||
#define KS_SEP_LINE 2 // line separator: "\n" (Unix) or "\r\n" (Windows)
|
||||
#define KS_SEP_MAX 2
|
||||
|
||||
#ifndef klib_unused
|
||||
#if (defined __clang__ && __clang_major__ >= 3) || (defined __GNUC__ && __GNUC__ >= 3)
|
||||
#define klib_unused __attribute__ ((__unused__))
|
||||
#else
|
||||
#define klib_unused
|
||||
#endif
|
||||
#endif /* klib_unused */
|
||||
|
||||
#define __KS_TYPE(type_t) \
|
||||
typedef struct __kstream_t { \
|
||||
int begin, end; \
|
||||
@@ -64,7 +72,7 @@
|
||||
}
|
||||
|
||||
#define __KS_INLINED(__read) \
|
||||
static inline int ks_getc(kstream_t *ks) \
|
||||
static inline klib_unused int ks_getc(kstream_t *ks) \
|
||||
{ \
|
||||
if (ks->is_eof && ks->begin >= ks->end) return -1; \
|
||||
if (ks->begin >= ks->end) { \
|
||||
@@ -81,7 +89,7 @@
|
||||
#ifndef KSTRING_T
|
||||
#define KSTRING_T kstring_t
|
||||
typedef struct __kstring_t {
|
||||
unsigned l, m;
|
||||
size_t l, m;
|
||||
char *s;
|
||||
} kstring_t;
|
||||
#endif
|
||||
|
||||
@@ -37,7 +37,7 @@ typedef struct {
|
||||
int depth;
|
||||
} ks_isort_stack_t;
|
||||
|
||||
#define KSORT_SWAP(type_t, a, b) { register type_t t=(a); (a)=(b); (b)=t; }
|
||||
#define KSORT_SWAP(type_t, a, b) { type_t t=(a); (a)=(b); (b)=t; }
|
||||
|
||||
#define KSORT_INIT(name, type_t, __sort_lt) \
|
||||
void ks_heapdown_##name(size_t i, size_t n, type_t l[]) \
|
||||
|
||||
@@ -16,6 +16,13 @@
|
||||
#define KSW_EZ_SPLICE_REV 0x200
|
||||
#define KSW_EZ_SPLICE_FLANK 0x400
|
||||
|
||||
// The subset of CIGAR operators used by ksw code.
|
||||
// Use MM_CIGAR_* from minimap.h if you need the full list.
|
||||
#define KSW_CIGAR_MATCH 0
|
||||
#define KSW_CIGAR_INS 1
|
||||
#define KSW_CIGAR_DEL 2
|
||||
#define KSW_CIGAR_N_SKIP 3
|
||||
|
||||
#ifdef __cplusplus
|
||||
extern "C" {
|
||||
#endif
|
||||
@@ -61,7 +68,7 @@ void ksw_extd2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
int8_t gapo, int8_t gape, int8_t gapo2, int8_t gape2, int w, int zdrop, int end_bonus, int flag, ksw_extz_t *ez);
|
||||
|
||||
void ksw_exts2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uint8_t *target, int8_t m, const int8_t *mat,
|
||||
int8_t gapo, int8_t gape, int8_t gapo2, int8_t noncan, int zdrop, int flag, ksw_extz_t *ez);
|
||||
int8_t gapo, int8_t gape, int8_t gapo2, int8_t noncan, int zdrop, int8_t junc_bonus, int flag, const uint8_t *junc, ksw_extz_t *ez);
|
||||
|
||||
void ksw_extf2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uint8_t *target, int8_t mch, int8_t mis, int8_t e, int w, int xdrop, ksw_extz_t *ez);
|
||||
|
||||
@@ -127,23 +134,23 @@ static inline void ksw_backtrack(void *km, int is_rot, int is_rev, int min_intro
|
||||
r = i + j;
|
||||
if (i < off[r]) force_state = 2;
|
||||
if (off_end && i > off_end[r]) force_state = 1;
|
||||
tmp = force_state < 0? p[r * n_col + i - off[r]] : 0;
|
||||
tmp = force_state < 0? p[(size_t)r * n_col + i - off[r]] : 0;
|
||||
} else {
|
||||
if (j < off[i]) force_state = 2;
|
||||
if (off_end && j > off_end[i]) force_state = 1;
|
||||
tmp = force_state < 0? p[i * n_col + j - off[i]] : 0;
|
||||
tmp = force_state < 0? p[(size_t)i * n_col + j - off[i]] : 0;
|
||||
}
|
||||
if (state == 0) state = tmp & 7; // if requesting the H state, find state one maximizes it.
|
||||
else if (!(tmp >> (state + 2) & 1)) state = 0; // if requesting other states, _state_ stays the same if it is a continuation; otherwise, set to H
|
||||
if (state == 0) state = tmp & 7; // TODO: probably this line can be merged into the "else if" line right above; not 100% sure
|
||||
if (force_state >= 0) state = force_state;
|
||||
if (state == 0) cigar = ksw_push_cigar(km, &n_cigar, &m_cigar, cigar, 0, 1), --i, --j; // match
|
||||
else if (state == 1 || (state == 3 && min_intron_len <= 0)) cigar = ksw_push_cigar(km, &n_cigar, &m_cigar, cigar, 2, 1), --i; // deletion
|
||||
else if (state == 3 && min_intron_len > 0) cigar = ksw_push_cigar(km, &n_cigar, &m_cigar, cigar, 3, 1), --i; // intron
|
||||
else cigar = ksw_push_cigar(km, &n_cigar, &m_cigar, cigar, 1, 1), --j; // insertion
|
||||
if (state == 0) cigar = ksw_push_cigar(km, &n_cigar, &m_cigar, cigar, KSW_CIGAR_MATCH, 1), --i, --j;
|
||||
else if (state == 1 || (state == 3 && min_intron_len <= 0)) cigar = ksw_push_cigar(km, &n_cigar, &m_cigar, cigar, KSW_CIGAR_DEL, 1), --i;
|
||||
else if (state == 3 && min_intron_len > 0) cigar = ksw_push_cigar(km, &n_cigar, &m_cigar, cigar, KSW_CIGAR_N_SKIP, 1), --i;
|
||||
else cigar = ksw_push_cigar(km, &n_cigar, &m_cigar, cigar, KSW_CIGAR_INS, 1), --j;
|
||||
}
|
||||
if (i >= 0) cigar = ksw_push_cigar(km, &n_cigar, &m_cigar, cigar, min_intron_len > 0 && i >= min_intron_len? 3 : 2, i + 1); // first deletion
|
||||
if (j >= 0) cigar = ksw_push_cigar(km, &n_cigar, &m_cigar, cigar, 1, j + 1); // first insertion
|
||||
if (i >= 0) cigar = ksw_push_cigar(km, &n_cigar, &m_cigar, cigar, min_intron_len > 0 && i >= min_intron_len? KSW_CIGAR_N_SKIP : KSW_CIGAR_DEL, i + 1); // first deletion
|
||||
if (j >= 0) cigar = ksw_push_cigar(km, &n_cigar, &m_cigar, cigar, KSW_CIGAR_INS, j + 1); // first insertion
|
||||
if (!is_rev)
|
||||
for (i = 0; i < n_cigar>>1; ++i) // reverse CIGAR
|
||||
tmp = cigar[i], cigar[i] = cigar[n_cigar-1-i], cigar[n_cigar-1-i] = tmp;
|
||||
|
||||
+19
-20
@@ -17,18 +17,20 @@
|
||||
void __cpuidex(int cpuid[4], int func_id, int subfunc_id)
|
||||
{
|
||||
#if defined(__x86_64__)
|
||||
asm volatile ("cpuid"
|
||||
__asm__ volatile ("cpuid"
|
||||
: "=a" (cpuid[0]), "=b" (cpuid[1]), "=c" (cpuid[2]), "=d" (cpuid[3])
|
||||
: "0" (func_id), "2" (subfunc_id));
|
||||
#else // on 32bit, ebx can NOT be used as PIC code
|
||||
asm volatile ("xchgl %%ebx, %1; cpuid; xchgl %%ebx, %1"
|
||||
__asm__ volatile ("xchgl %%ebx, %1; cpuid; xchgl %%ebx, %1"
|
||||
: "=a" (cpuid[0]), "=r" (cpuid[1]), "=c" (cpuid[2]), "=d" (cpuid[3])
|
||||
: "0" (func_id), "2" (subfunc_id));
|
||||
#endif
|
||||
}
|
||||
#endif
|
||||
|
||||
int x86_simd(void)
|
||||
static int ksw_simd = -1;
|
||||
|
||||
static int x86_simd(void)
|
||||
{
|
||||
int flag = 0, cpuid[4], max_id;
|
||||
__cpuidex(cpuid, 0, 0);
|
||||
@@ -54,11 +56,10 @@ void ksw_extz2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
{
|
||||
extern void ksw_extz2_sse2(void *km, int qlen, const uint8_t *query, int tlen, const uint8_t *target, int8_t m, const int8_t *mat, int8_t q, int8_t e, int w, int zdrop, int end_bonus, int flag, ksw_extz_t *ez);
|
||||
extern void ksw_extz2_sse41(void *km, int qlen, const uint8_t *query, int tlen, const uint8_t *target, int8_t m, const int8_t *mat, int8_t q, int8_t e, int w, int zdrop, int end_bonus, int flag, ksw_extz_t *ez);
|
||||
unsigned simd;
|
||||
simd = x86_simd();
|
||||
if (simd & SIMD_SSE4_1)
|
||||
if (ksw_simd < 0) ksw_simd = x86_simd();
|
||||
if (ksw_simd & SIMD_SSE4_1)
|
||||
ksw_extz2_sse41(km, qlen, query, tlen, target, m, mat, q, e, w, zdrop, end_bonus, flag, ez);
|
||||
else if (simd & SIMD_SSE2)
|
||||
else if (ksw_simd & SIMD_SSE2)
|
||||
ksw_extz2_sse2(km, qlen, query, tlen, target, m, mat, q, e, w, zdrop, end_bonus, flag, ez);
|
||||
else abort();
|
||||
}
|
||||
@@ -70,28 +71,26 @@ void ksw_extd2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
int8_t q, int8_t e, int8_t q2, int8_t e2, int w, int zdrop, int end_bonus, int flag, ksw_extz_t *ez);
|
||||
extern void ksw_extd2_sse41(void *km, int qlen, const uint8_t *query, int tlen, const uint8_t *target, int8_t m, const int8_t *mat,
|
||||
int8_t q, int8_t e, int8_t q2, int8_t e2, int w, int zdrop, int end_bonus, int flag, ksw_extz_t *ez);
|
||||
unsigned simd;
|
||||
simd = x86_simd();
|
||||
if (simd & SIMD_SSE4_1)
|
||||
if (ksw_simd < 0) ksw_simd = x86_simd();
|
||||
if (ksw_simd & SIMD_SSE4_1)
|
||||
ksw_extd2_sse41(km, qlen, query, tlen, target, m, mat, q, e, q2, e2, w, zdrop, end_bonus, flag, ez);
|
||||
else if (simd & SIMD_SSE2)
|
||||
else if (ksw_simd & SIMD_SSE2)
|
||||
ksw_extd2_sse2(km, qlen, query, tlen, target, m, mat, q, e, q2, e2, w, zdrop, end_bonus, flag, ez);
|
||||
else abort();
|
||||
}
|
||||
|
||||
void ksw_exts2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uint8_t *target, int8_t m, const int8_t *mat,
|
||||
int8_t q, int8_t e, int8_t q2, int8_t noncan, int zdrop, int flag, ksw_extz_t *ez)
|
||||
int8_t q, int8_t e, int8_t q2, int8_t noncan, int zdrop, int8_t junc_bonus, int flag, const uint8_t *junc, ksw_extz_t *ez)
|
||||
{
|
||||
extern void ksw_exts2_sse2(void *km, int qlen, const uint8_t *query, int tlen, const uint8_t *target, int8_t m, const int8_t *mat,
|
||||
int8_t q, int8_t e, int8_t q2, int8_t noncan, int zdrop, int flag, ksw_extz_t *ez);
|
||||
int8_t q, int8_t e, int8_t q2, int8_t noncan, int zdrop, int8_t junc_bonus, int flag, const uint8_t *junc, ksw_extz_t *ez);
|
||||
extern void ksw_exts2_sse41(void *km, int qlen, const uint8_t *query, int tlen, const uint8_t *target, int8_t m, const int8_t *mat,
|
||||
int8_t q, int8_t e, int8_t q2, int8_t noncan, int zdrop, int flag, ksw_extz_t *ez);
|
||||
unsigned simd;
|
||||
simd = x86_simd();
|
||||
if (simd & SIMD_SSE4_1)
|
||||
ksw_exts2_sse41(km, qlen, query, tlen, target, m, mat, q, e, q2, noncan, zdrop, flag, ez);
|
||||
else if (simd & SIMD_SSE2)
|
||||
ksw_exts2_sse2(km, qlen, query, tlen, target, m, mat, q, e, q2, noncan, zdrop, flag, ez);
|
||||
int8_t q, int8_t e, int8_t q2, int8_t noncan, int zdrop, int8_t junc_bonus, int flag, const uint8_t *junc, ksw_extz_t *ez);
|
||||
if (ksw_simd < 0) ksw_simd = x86_simd();
|
||||
if (ksw_simd & SIMD_SSE4_1)
|
||||
ksw_exts2_sse41(km, qlen, query, tlen, target, m, mat, q, e, q2, noncan, zdrop, junc_bonus, flag, junc, ez);
|
||||
else if (ksw_simd & SIMD_SSE2)
|
||||
ksw_exts2_sse2(km, qlen, query, tlen, target, m, mat, q, e, q2, noncan, zdrop, junc_bonus, flag, junc, ez);
|
||||
else abort();
|
||||
}
|
||||
#endif
|
||||
|
||||
+13
-5
@@ -4,15 +4,23 @@
|
||||
#include "ksw2.h"
|
||||
|
||||
#ifdef __SSE2__
|
||||
#ifdef USE_SIMDE
|
||||
#include <simde/x86/sse2.h>
|
||||
#else
|
||||
#include <emmintrin.h>
|
||||
#endif
|
||||
|
||||
#ifdef KSW_SSE2_ONLY
|
||||
#undef __SSE4_1__
|
||||
#endif
|
||||
|
||||
#ifdef __SSE4_1__
|
||||
#ifdef USE_SIMDE
|
||||
#include <simde/x86/sse4.1.h>
|
||||
#else
|
||||
#include <smmintrin.h>
|
||||
#endif
|
||||
#endif
|
||||
|
||||
#ifdef KSW_CPU_DISPATCH
|
||||
#ifdef __SSE4_1__
|
||||
@@ -76,7 +84,7 @@ void ksw_extd2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
qe2_ = _mm_set1_epi8(q2 + e2);
|
||||
sc_mch_ = _mm_set1_epi8(mat[0]);
|
||||
sc_mis_ = _mm_set1_epi8(mat[1]);
|
||||
sc_N_ = _mm_set1_epi8(-e2);
|
||||
sc_N_ = mat[m*m-1] == 0? _mm_set1_epi8(-e2) : _mm_set1_epi8(mat[m*m-1]);
|
||||
m1_ = _mm_set1_epi8(m - 1); // wildcard
|
||||
|
||||
if (w < 0) w = tlen > qlen? tlen : qlen;
|
||||
@@ -111,7 +119,7 @@ void ksw_extd2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
for (t = 0; t < tlen_ * 16; ++t) H[t] = KSW_NEG_INF;
|
||||
}
|
||||
if (with_cigar) {
|
||||
mem2 = (uint8_t*)kmalloc(km, ((qlen + tlen - 1) * n_col_ + 1) * 16);
|
||||
mem2 = (uint8_t*)kmalloc(km, ((size_t)(qlen + tlen - 1) * n_col_ + 1) * 16);
|
||||
p = (__m128i*)(((size_t)mem2 + 15) >> 4 << 4);
|
||||
off = (int*)kmalloc(km, (qlen + tlen - 1) * sizeof(int) * 2);
|
||||
off_end = off + qlen + tlen - 1;
|
||||
@@ -218,7 +226,7 @@ void ksw_extd2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
#endif
|
||||
}
|
||||
} else if (!(flag&KSW_EZ_RIGHT)) { // gap left-alignment
|
||||
__m128i *pr = p + r * n_col_ - st_;
|
||||
__m128i *pr = p + (size_t)r * n_col_ - st_;
|
||||
off[r] = st, off_end[r] = en;
|
||||
for (t = st_; t <= en_; ++t) {
|
||||
__m128i d, z, a, b, a2, b2, xt1, x2t1, vt1, ut, tmp;
|
||||
@@ -265,7 +273,7 @@ void ksw_extd2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
_mm_store_si128(&pr[t], d);
|
||||
}
|
||||
} else { // gap right-alignment
|
||||
__m128i *pr = p + r * n_col_ - st_;
|
||||
__m128i *pr = p + (size_t)r * n_col_ - st_;
|
||||
off[r] = st, off_end[r] = en;
|
||||
for (t = st_; t <= en_; ++t) {
|
||||
__m128i d, z, a, b, a2, b2, xt1, x2t1, vt1, ut, tmp;
|
||||
@@ -382,7 +390,7 @@ void ksw_extd2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
int rev_cigar = !!(flag & KSW_EZ_REV_CIGAR);
|
||||
if (!ez->zdropped && !(flag&KSW_EZ_EXTZ_ONLY)) {
|
||||
ksw_backtrack(km, 1, rev_cigar, 0, (uint8_t*)p, off, off_end, n_col_*16, tlen-1, qlen-1, &ez->m_cigar, &ez->n_cigar, &ez->cigar);
|
||||
} else if (!ez->zdropped && (flag&KSW_EZ_EXTZ_ONLY) && ez->mqe + end_bonus > ez->max) {
|
||||
} else if (!ez->zdropped && (flag&KSW_EZ_EXTZ_ONLY) && ez->mqe + end_bonus > (int)ez->max) {
|
||||
ez->reach_end = 1;
|
||||
ksw_backtrack(km, 1, rev_cigar, 0, (uint8_t*)p, off, off_end, n_col_*16, ez->mqe_t, qlen-1, &ez->m_cigar, &ez->n_cigar, &ez->cigar);
|
||||
} else if (ez->max_t >= 0 && ez->max_q >= 0) {
|
||||
|
||||
+59
-19
@@ -4,27 +4,34 @@
|
||||
#include "ksw2.h"
|
||||
|
||||
#ifdef __SSE2__
|
||||
#ifdef USE_SIMDE
|
||||
#include <simde/x86/sse2.h>
|
||||
#else
|
||||
#include <emmintrin.h>
|
||||
|
||||
#endif
|
||||
#ifdef KSW_SSE2_ONLY
|
||||
#undef __SSE4_1__
|
||||
#endif
|
||||
|
||||
#ifdef __SSE4_1__
|
||||
#ifdef USE_SIMDE
|
||||
#include <simde/x86/sse4.1.h>
|
||||
#else
|
||||
#include <smmintrin.h>
|
||||
#endif
|
||||
#endif
|
||||
|
||||
#ifdef KSW_CPU_DISPATCH
|
||||
#ifdef __SSE4_1__
|
||||
void ksw_exts2_sse41(void *km, int qlen, const uint8_t *query, int tlen, const uint8_t *target, int8_t m, const int8_t *mat,
|
||||
int8_t q, int8_t e, int8_t q2, int8_t noncan, int zdrop, int flag, ksw_extz_t *ez)
|
||||
int8_t q, int8_t e, int8_t q2, int8_t noncan, int zdrop, int8_t junc_bonus, int flag, const uint8_t *junc, ksw_extz_t *ez)
|
||||
#else
|
||||
void ksw_exts2_sse2(void *km, int qlen, const uint8_t *query, int tlen, const uint8_t *target, int8_t m, const int8_t *mat,
|
||||
int8_t q, int8_t e, int8_t q2, int8_t noncan, int zdrop, int flag, ksw_extz_t *ez)
|
||||
int8_t q, int8_t e, int8_t q2, int8_t noncan, int zdrop, int8_t junc_bonus, int flag, const uint8_t *junc, ksw_extz_t *ez)
|
||||
#endif
|
||||
#else
|
||||
void ksw_exts2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uint8_t *target, int8_t m, const int8_t *mat,
|
||||
int8_t q, int8_t e, int8_t q2, int8_t noncan, int zdrop, int flag, ksw_extz_t *ez)
|
||||
int8_t q, int8_t e, int8_t q2, int8_t noncan, int zdrop, int8_t junc_bonus, int flag, const uint8_t *junc, ksw_extz_t *ez)
|
||||
#endif // ~KSW_CPU_DISPATCH
|
||||
{
|
||||
#define __dp_code_block1 \
|
||||
@@ -71,7 +78,7 @@ void ksw_exts2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
qe_ = _mm_set1_epi8(q + e);
|
||||
sc_mch_ = _mm_set1_epi8(mat[0]);
|
||||
sc_mis_ = _mm_set1_epi8(mat[1]);
|
||||
sc_N_ = _mm_set1_epi8(-e);
|
||||
sc_N_ = mat[m*m-1] == 0? _mm_set1_epi8(-e) : _mm_set1_epi8(mat[m*m-1]);
|
||||
m1_ = _mm_set1_epi8(m - 1); // wildcard
|
||||
|
||||
tlen_ = (tlen + 15) / 16;
|
||||
@@ -100,7 +107,7 @@ void ksw_exts2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
for (t = 0; t < tlen_ * 16; ++t) H[t] = KSW_NEG_INF;
|
||||
}
|
||||
if (with_cigar) {
|
||||
mem2 = (uint8_t*)kmalloc(km, ((qlen + tlen - 1) * n_col_ + 1) * 16);
|
||||
mem2 = (uint8_t*)kmalloc(km, ((size_t)(qlen + tlen - 1) * n_col_ + 1) * 16);
|
||||
p = (__m128i*)(((size_t)mem2 + 15) >> 4 << 4);
|
||||
off = (int*)kmalloc(km, (qlen + tlen - 1) * sizeof(int) * 2);
|
||||
off_end = off + qlen + tlen - 1;
|
||||
@@ -113,20 +120,53 @@ void ksw_exts2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
if (flag & (KSW_EZ_SPLICE_FOR|KSW_EZ_SPLICE_REV)) {
|
||||
int semi_cost = flag&KSW_EZ_SPLICE_FLANK? -noncan/2 : 0; // GTr or yAG is worth 0.5 bit; see PMID:18688272
|
||||
memset(donor, -noncan, tlen_ * 16);
|
||||
for (t = 0; t < tlen - 4; ++t) {
|
||||
int can_type = 0; // type of canonical site: 0=none, 1=GT/AG only, 2=GTr/yAG
|
||||
if ((flag & KSW_EZ_SPLICE_FOR) && target[t+1] == 2 && target[t+2] == 3) can_type = 1; // GTr...
|
||||
if ((flag & KSW_EZ_SPLICE_REV) && target[t+1] == 1 && target[t+2] == 3) can_type = 1; // CTr...
|
||||
if (can_type && (target[t+3] == 0 || target[t+3] == 2)) can_type = 2;
|
||||
if (can_type) ((int8_t*)donor)[t] = can_type == 2? 0 : semi_cost;
|
||||
}
|
||||
memset(acceptor, -noncan, tlen_ * 16);
|
||||
for (t = 2; t < tlen; ++t) {
|
||||
int can_type = 0;
|
||||
if ((flag & KSW_EZ_SPLICE_FOR) && target[t-1] == 0 && target[t] == 2) can_type = 1; // ...yAG
|
||||
if ((flag & KSW_EZ_SPLICE_REV) && target[t-1] == 0 && target[t] == 1) can_type = 1; // ...yAC
|
||||
if (can_type && (target[t-2] == 1 || target[t-2] == 3)) can_type = 2;
|
||||
if (can_type) ((int8_t*)acceptor)[t] = can_type == 2? 0 : semi_cost;
|
||||
if (!(flag & KSW_EZ_REV_CIGAR)) {
|
||||
for (t = 0; t < tlen - 4; ++t) {
|
||||
int can_type = 0; // type of canonical site: 0=none, 1=GT/AG only, 2=GTr/yAG
|
||||
if ((flag & KSW_EZ_SPLICE_FOR) && target[t+1] == 2 && target[t+2] == 3) can_type = 1; // GTr...
|
||||
if ((flag & KSW_EZ_SPLICE_REV) && target[t+1] == 1 && target[t+2] == 3) can_type = 1; // CTr...
|
||||
if (can_type && (target[t+3] == 0 || target[t+3] == 2)) can_type = 2;
|
||||
if (can_type) ((int8_t*)donor)[t] = can_type == 2? 0 : semi_cost;
|
||||
}
|
||||
if (junc)
|
||||
for (t = 0; t < tlen - 1; ++t)
|
||||
if (((flag & KSW_EZ_SPLICE_FOR) && (junc[t+1]&1)) || ((flag & KSW_EZ_SPLICE_REV) && (junc[t+1]&8)))
|
||||
((int8_t*)donor)[t] += junc_bonus;
|
||||
for (t = 2; t < tlen; ++t) {
|
||||
int can_type = 0;
|
||||
if ((flag & KSW_EZ_SPLICE_FOR) && target[t-1] == 0 && target[t] == 2) can_type = 1; // ...yAG
|
||||
if ((flag & KSW_EZ_SPLICE_REV) && target[t-1] == 0 && target[t] == 1) can_type = 1; // ...yAC
|
||||
if (can_type && (target[t-2] == 1 || target[t-2] == 3)) can_type = 2;
|
||||
if (can_type) ((int8_t*)acceptor)[t] = can_type == 2? 0 : semi_cost;
|
||||
}
|
||||
if (junc)
|
||||
for (t = 0; t < tlen; ++t)
|
||||
if (((flag & KSW_EZ_SPLICE_FOR) && (junc[t]&2)) || ((flag & KSW_EZ_SPLICE_REV) && (junc[t]&4)))
|
||||
((int8_t*)acceptor)[t] += junc_bonus;
|
||||
} else {
|
||||
for (t = 0; t < tlen - 4; ++t) {
|
||||
int can_type = 0; // type of canonical site: 0=none, 1=GT/AG only, 2=GTr/yAG
|
||||
if ((flag & KSW_EZ_SPLICE_FOR) && target[t+1] == 2 && target[t+2] == 0) can_type = 1; // GAy...
|
||||
if ((flag & KSW_EZ_SPLICE_REV) && target[t+1] == 1 && target[t+2] == 0) can_type = 1; // CAy...
|
||||
if (can_type && (target[t+3] == 1 || target[t+3] == 3)) can_type = 2;
|
||||
if (can_type) ((int8_t*)donor)[t] = can_type == 2? 0 : semi_cost;
|
||||
}
|
||||
if (junc)
|
||||
for (t = 0; t < tlen - 1; ++t)
|
||||
if (((flag & KSW_EZ_SPLICE_FOR) && (junc[t+1]&2)) || ((flag & KSW_EZ_SPLICE_REV) && (junc[t+1]&4)))
|
||||
((int8_t*)donor)[t] += junc_bonus;
|
||||
for (t = 2; t < tlen; ++t) {
|
||||
int can_type = 0;
|
||||
if ((flag & KSW_EZ_SPLICE_FOR) && target[t-1] == 3 && target[t] == 2) can_type = 1; // ...rTG
|
||||
if ((flag & KSW_EZ_SPLICE_REV) && target[t-1] == 3 && target[t] == 1) can_type = 1; // ...rTC
|
||||
if (can_type && (target[t-2] == 0 || target[t-2] == 2)) can_type = 2;
|
||||
if (can_type) ((int8_t*)acceptor)[t] = can_type == 2? 0 : semi_cost;
|
||||
}
|
||||
if (junc)
|
||||
for (t = 0; t < tlen; ++t)
|
||||
if (((flag & KSW_EZ_SPLICE_FOR) && (junc[t]&1)) || ((flag & KSW_EZ_SPLICE_REV) && (junc[t]&8)))
|
||||
((int8_t*)acceptor)[t] += junc_bonus;
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
+13
-5
@@ -3,15 +3,23 @@
|
||||
#include "ksw2.h"
|
||||
|
||||
#ifdef __SSE2__
|
||||
#ifdef USE_SIMDE
|
||||
#include <simde/x86/sse2.h>
|
||||
#else
|
||||
#include <emmintrin.h>
|
||||
#endif
|
||||
|
||||
#ifdef KSW_SSE2_ONLY
|
||||
#undef __SSE4_1__
|
||||
#endif
|
||||
|
||||
#ifdef __SSE4_1__
|
||||
#ifdef USE_SIMDE
|
||||
#include <simde/x86/sse4.1.h>
|
||||
#else
|
||||
#include <smmintrin.h>
|
||||
#endif
|
||||
#endif
|
||||
|
||||
#ifdef KSW_CPU_DISPATCH
|
||||
#ifdef __SSE4_1__
|
||||
@@ -65,7 +73,7 @@ void ksw_extz2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
flag16_ = _mm_set1_epi8(0x10);
|
||||
sc_mch_ = _mm_set1_epi8(mat[0]);
|
||||
sc_mis_ = _mm_set1_epi8(mat[1]);
|
||||
sc_N_ = _mm_set1_epi8(-e);
|
||||
sc_N_ = mat[m*m-1] == 0? _mm_set1_epi8(-e) : _mm_set1_epi8(mat[m*m-1]);
|
||||
m1_ = _mm_set1_epi8(m - 1); // wildcard
|
||||
max_sc_ = _mm_set1_epi8(mat[0] + (q + e) * 2);
|
||||
|
||||
@@ -89,7 +97,7 @@ void ksw_extz2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
for (t = 0; t < tlen_ * 16; ++t) H[t] = KSW_NEG_INF;
|
||||
}
|
||||
if (with_cigar) {
|
||||
mem2 = (uint8_t*)kmalloc(km, ((qlen + tlen - 1) * n_col_ + 1) * 16);
|
||||
mem2 = (uint8_t*)kmalloc(km, ((size_t)(qlen + tlen - 1) * n_col_ + 1) * 16);
|
||||
p = (__m128i*)(((size_t)mem2 + 15) >> 4 << 4);
|
||||
off = (int*)kmalloc(km, (qlen + tlen - 1) * sizeof(int) * 2);
|
||||
off_end = off + qlen + tlen - 1;
|
||||
@@ -169,7 +177,7 @@ void ksw_extz2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
#endif
|
||||
}
|
||||
} else if (!(flag&KSW_EZ_RIGHT)) { // gap left-alignment
|
||||
__m128i *pr = p + r * n_col_ - st_;
|
||||
__m128i *pr = p + (size_t)r * n_col_ - st_;
|
||||
off[r] = st, off_end[r] = en;
|
||||
for (t = st_; t <= en_; ++t) {
|
||||
__m128i d, z, a, b, xt1, vt1, ut, tmp;
|
||||
@@ -195,7 +203,7 @@ void ksw_extz2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
_mm_store_si128(&pr[t], d);
|
||||
}
|
||||
} else { // gap right-alignment
|
||||
__m128i *pr = p + r * n_col_ - st_;
|
||||
__m128i *pr = p + (size_t)r * n_col_ - st_;
|
||||
off[r] = st, off_end[r] = en;
|
||||
for (t = st_; t <= en_; ++t) {
|
||||
__m128i d, z, a, b, xt1, vt1, ut, tmp;
|
||||
@@ -293,7 +301,7 @@ void ksw_extz2_sse(void *km, int qlen, const uint8_t *query, int tlen, const uin
|
||||
int rev_cigar = !!(flag & KSW_EZ_REV_CIGAR);
|
||||
if (!ez->zdropped && !(flag&KSW_EZ_EXTZ_ONLY)) {
|
||||
ksw_backtrack(km, 1, rev_cigar, 0, (uint8_t*)p, off, off_end, n_col_*16, tlen-1, qlen-1, &ez->m_cigar, &ez->n_cigar, &ez->cigar);
|
||||
} else if (!ez->zdropped && (flag&KSW_EZ_EXTZ_ONLY) && ez->mqe + end_bonus > ez->max) {
|
||||
} else if (!ez->zdropped && (flag&KSW_EZ_EXTZ_ONLY) && ez->mqe + end_bonus > (int)ez->max) {
|
||||
ez->reach_end = 1;
|
||||
ksw_backtrack(km, 1, rev_cigar, 0, (uint8_t*)p, off, off_end, n_col_*16, ez->mqe_t, qlen-1, &ez->m_cigar, &ez->n_cigar, &ez->cigar);
|
||||
} else if (ez->max_t >= 0 && ez->max_q >= 0) {
|
||||
|
||||
+6
-1
@@ -1,9 +1,14 @@
|
||||
#include <stdlib.h>
|
||||
#include <stdint.h>
|
||||
#include <string.h>
|
||||
#include <emmintrin.h>
|
||||
#include "ksw2.h"
|
||||
|
||||
#ifdef USE_SIMDE
|
||||
#include <simde/x86/sse2.h>
|
||||
#else
|
||||
#include <emmintrin.h>
|
||||
#endif
|
||||
|
||||
#ifdef __GNUC__
|
||||
#define LIKELY(x) __builtin_expect((x),1)
|
||||
#define UNLIKELY(x) __builtin_expect((x),0)
|
||||
|
||||
@@ -0,0 +1,344 @@
|
||||
#include <stdint.h>
|
||||
#include <string.h>
|
||||
#include <stdio.h>
|
||||
#include <assert.h>
|
||||
#include "mmpriv.h"
|
||||
#include "kalloc.h"
|
||||
#include "krmq.h"
|
||||
|
||||
uint64_t *mg_chain_backtrack(void *km, int64_t n, const int32_t *f, const int64_t *p, int32_t *v, int32_t *t, int32_t min_cnt, int32_t min_sc, int32_t *n_u_, int32_t *n_v_)
|
||||
{
|
||||
mm128_t *z;
|
||||
uint64_t *u;
|
||||
int64_t i, k, n_z, n_v;
|
||||
int32_t n_u;
|
||||
|
||||
*n_u_ = *n_v_ = 0;
|
||||
for (i = 0, n_z = 0; i < n; ++i) // precompute n_z
|
||||
if (f[i] >= min_sc) ++n_z;
|
||||
if (n_z == 0) return 0;
|
||||
KMALLOC(km, z, n_z);
|
||||
for (i = 0, k = 0; i < n; ++i) // populate z[]
|
||||
if (f[i] >= min_sc) z[k].x = f[i], z[k++].y = i;
|
||||
radix_sort_128x(z, z + n_z);
|
||||
|
||||
memset(t, 0, n * 4);
|
||||
for (k = n_z - 1, n_v = n_u = 0; k >= 0; --k) { // precompute n_u
|
||||
int64_t n_v0 = n_v;
|
||||
int32_t sc;
|
||||
for (i = z[k].y; i >= 0 && t[i] == 0; i = p[i])
|
||||
++n_v, t[i] = 1;
|
||||
sc = i < 0? z[k].x : (int32_t)z[k].x - f[i];
|
||||
if (sc >= min_sc && n_v > n_v0 && n_v - n_v0 >= min_cnt)
|
||||
++n_u;
|
||||
else n_v = n_v0;
|
||||
}
|
||||
KMALLOC(km, u, n_u);
|
||||
memset(t, 0, n * 4);
|
||||
for (k = n_z - 1, n_v = n_u = 0; k >= 0; --k) { // populate u[]
|
||||
int64_t n_v0 = n_v;
|
||||
int32_t sc;
|
||||
for (i = z[k].y; i >= 0 && t[i] == 0; i = p[i])
|
||||
v[n_v++] = i, t[i] = 1;
|
||||
sc = i < 0? z[k].x : (int32_t)z[k].x - f[i];
|
||||
if (sc >= min_sc && n_v > n_v0 && n_v - n_v0 >= min_cnt)
|
||||
u[n_u++] = (uint64_t)sc << 32 | (n_v - n_v0);
|
||||
else n_v = n_v0;
|
||||
}
|
||||
kfree(km, z);
|
||||
assert(n_v < INT32_MAX);
|
||||
*n_u_ = n_u, *n_v_ = n_v;
|
||||
return u;
|
||||
}
|
||||
|
||||
static mm128_t *compact_a(void *km, int32_t n_u, uint64_t *u, int32_t n_v, int32_t *v, mm128_t *a)
|
||||
{
|
||||
mm128_t *b, *w;
|
||||
uint64_t *u2;
|
||||
int64_t i, j, k;
|
||||
|
||||
// write the result to b[]
|
||||
KMALLOC(km, b, n_v);
|
||||
for (i = 0, k = 0; i < n_u; ++i) {
|
||||
int32_t k0 = k, ni = (int32_t)u[i];
|
||||
for (j = 0; j < ni; ++j)
|
||||
b[k++] = a[v[k0 + (ni - j - 1)]];
|
||||
}
|
||||
kfree(km, v);
|
||||
|
||||
// sort u[] and a[] by the target position, such that adjacent chains may be joined
|
||||
KMALLOC(km, w, n_u);
|
||||
for (i = k = 0; i < n_u; ++i) {
|
||||
w[i].x = b[k].x, w[i].y = (uint64_t)k<<32|i;
|
||||
k += (int32_t)u[i];
|
||||
}
|
||||
radix_sort_128x(w, w + n_u);
|
||||
KMALLOC(km, u2, n_u);
|
||||
for (i = k = 0; i < n_u; ++i) {
|
||||
int32_t j = (int32_t)w[i].y, n = (int32_t)u[j];
|
||||
u2[i] = u[j];
|
||||
memcpy(&a[k], &b[w[i].y>>32], n * sizeof(mm128_t));
|
||||
k += n;
|
||||
}
|
||||
memcpy(u, u2, n_u * 8);
|
||||
memcpy(b, a, k * sizeof(mm128_t)); // write _a_ to _b_ and deallocate _a_ because _a_ is oversized, sometimes a lot
|
||||
kfree(km, a); kfree(km, w); kfree(km, u2);
|
||||
return b;
|
||||
}
|
||||
|
||||
static inline int32_t comput_sc(const mm128_t *ai, const mm128_t *aj, int32_t max_dist_x, int32_t max_dist_y, int32_t bw, float chn_pen_gap, float chn_pen_skip, int is_cdna, int n_seg)
|
||||
{
|
||||
int32_t dq = (int32_t)ai->y - (int32_t)aj->y, dr, dd, dg, q_span, sc;
|
||||
int32_t sidi = (ai->y & MM_SEED_SEG_MASK) >> MM_SEED_SEG_SHIFT;
|
||||
int32_t sidj = (aj->y & MM_SEED_SEG_MASK) >> MM_SEED_SEG_SHIFT;
|
||||
if (dq <= 0 || dq > max_dist_x) return INT32_MIN;
|
||||
dr = (int32_t)(ai->x - aj->x);
|
||||
if (sidi == sidj && (dr == 0 || dq > max_dist_y)) return INT32_MIN;
|
||||
dd = dr > dq? dr - dq : dq - dr;
|
||||
if (sidi == sidj && dd > bw) return INT32_MIN;
|
||||
if (n_seg > 1 && !is_cdna && sidi == sidj && dr > max_dist_y) return INT32_MIN;
|
||||
dg = dr < dq? dr : dq;
|
||||
q_span = aj->y>>32&0xff;
|
||||
sc = q_span < dg? q_span : dg;
|
||||
if (dd || dg > q_span) {
|
||||
float lin_pen, log_pen;
|
||||
lin_pen = chn_pen_gap * (float)dd + chn_pen_skip * (float)dg;
|
||||
log_pen = dd >= 1? mg_log2(dd + 1) : 0.0f; // mg_log2() only works for dd>=2
|
||||
if (is_cdna || sidi != sidj) {
|
||||
if (sidi != sidj && dr == 0) ++sc; // possibly due to overlapping paired ends; give a minor bonus
|
||||
else if (dr > dq || sidi != sidj) sc -= (int)(lin_pen < log_pen? lin_pen : log_pen); // deletion or jump between paired ends
|
||||
else sc -= (int)(lin_pen + .5f * log_pen);
|
||||
} else sc -= (int)(lin_pen + .5f * log_pen);
|
||||
}
|
||||
return sc;
|
||||
}
|
||||
|
||||
/* Input:
|
||||
* a[].x: tid<<33 | rev<<32 | tpos
|
||||
* a[].y: flags<<40 | q_span<<32 | q_pos
|
||||
* Output:
|
||||
* n_u: #chains
|
||||
* u[]: score<<32 | #anchors (sum of lower 32 bits of u[] is the returned length of a[])
|
||||
* input a[] is deallocated on return
|
||||
*/
|
||||
mm128_t *mg_lchain_dp(int max_dist_x, int max_dist_y, int bw, int max_skip, int max_iter, int min_cnt, int min_sc, float chn_pen_gap, float chn_pen_skip,
|
||||
int is_cdna, int n_seg, int64_t n, mm128_t *a, int *n_u_, uint64_t **_u, void *km)
|
||||
{ // TODO: make sure this works when n has more than 32 bits
|
||||
int32_t *f, *t, *v, n_u, n_v, mmax_f = 0;
|
||||
int64_t *p, i, j, max_ii, st = 0, n_iter = 0;
|
||||
uint64_t *u;
|
||||
|
||||
if (_u) *_u = 0, *n_u_ = 0;
|
||||
if (n == 0 || a == 0) {
|
||||
kfree(km, a);
|
||||
return 0;
|
||||
}
|
||||
if (max_dist_x < bw) max_dist_x = bw;
|
||||
if (max_dist_y < bw && !is_cdna) max_dist_y = bw;
|
||||
KMALLOC(km, p, n);
|
||||
KMALLOC(km, f, n);
|
||||
KMALLOC(km, v, n);
|
||||
KCALLOC(km, t, n);
|
||||
|
||||
// fill the score and backtrack arrays
|
||||
for (i = 0, max_ii = -1; i < n; ++i) {
|
||||
int64_t max_j = -1, end_j;
|
||||
int32_t max_f = a[i].y>>32&0xff, n_skip = 0;
|
||||
while (st < i && (a[i].x>>32 != a[st].x>>32 || a[i].x > a[st].x + max_dist_x)) ++st;
|
||||
if (i - st > max_iter) st = i - max_iter;
|
||||
for (j = i - 1; j >= st; --j) {
|
||||
int32_t sc;
|
||||
sc = comput_sc(&a[i], &a[j], max_dist_x, max_dist_y, bw, chn_pen_gap, chn_pen_skip, is_cdna, n_seg);
|
||||
++n_iter;
|
||||
if (sc == INT32_MIN) continue;
|
||||
sc += f[j];
|
||||
if (sc > max_f) {
|
||||
max_f = sc, max_j = j;
|
||||
if (n_skip > 0) --n_skip;
|
||||
} else if (t[j] == (int32_t)i) {
|
||||
if (++n_skip > max_skip)
|
||||
break;
|
||||
}
|
||||
if (p[j] >= 0) t[p[j]] = i;
|
||||
}
|
||||
end_j = j;
|
||||
if (max_ii < 0 || a[i].x - a[max_ii].x > (int64_t)max_dist_x) {
|
||||
int32_t max = INT32_MIN;
|
||||
max_ii = -1;
|
||||
for (j = i - 1; j >= st; --j)
|
||||
if (max < f[j]) max = f[j], max_ii = j;
|
||||
}
|
||||
if (max_ii >= 0 && max_ii < end_j) {
|
||||
int32_t tmp;
|
||||
tmp = comput_sc(&a[i], &a[max_ii], max_dist_x, max_dist_y, bw, chn_pen_gap, chn_pen_skip, is_cdna, n_seg);
|
||||
if (tmp != INT32_MIN && max_f < tmp + f[max_ii])
|
||||
max_f = tmp + f[max_ii], max_j = max_ii;
|
||||
}
|
||||
f[i] = max_f, p[i] = max_j;
|
||||
v[i] = max_j >= 0 && v[max_j] > max_f? v[max_j] : max_f; // v[] keeps the peak score up to i; f[] is the score ending at i, not always the peak
|
||||
if (max_ii < 0 || (a[i].x - a[max_ii].x <= (int64_t)max_dist_x && f[max_ii] < f[i]))
|
||||
max_ii = i;
|
||||
if (mmax_f < max_f) mmax_f = max_f;
|
||||
}
|
||||
|
||||
u = mg_chain_backtrack(km, n, f, p, v, t, min_cnt, min_sc, &n_u, &n_v);
|
||||
*n_u_ = n_u, *_u = u; // NB: note that u[] may not be sorted by score here
|
||||
kfree(km, p); kfree(km, f); kfree(km, t);
|
||||
if (n_u == 0) {
|
||||
kfree(km, a); kfree(km, v);
|
||||
return 0;
|
||||
}
|
||||
return compact_a(km, n_u, u, n_v, v, a);
|
||||
}
|
||||
|
||||
typedef struct lc_elem_s {
|
||||
int32_t y;
|
||||
int64_t i;
|
||||
double pri;
|
||||
KRMQ_HEAD(struct lc_elem_s) head;
|
||||
} lc_elem_t;
|
||||
|
||||
#define lc_elem_cmp(a, b) ((a)->y < (b)->y? -1 : (a)->y > (b)->y? 1 : ((a)->i > (b)->i) - ((a)->i < (b)->i))
|
||||
#define lc_elem_lt2(a, b) ((a)->pri < (b)->pri)
|
||||
KRMQ_INIT(lc_elem, lc_elem_t, head, lc_elem_cmp, lc_elem_lt2)
|
||||
|
||||
KALLOC_POOL_INIT(rmq, lc_elem_t)
|
||||
|
||||
static inline int32_t comput_sc_simple(const mm128_t *ai, const mm128_t *aj, float chn_pen_gap, float chn_pen_skip, int32_t *exact, int32_t *width)
|
||||
{
|
||||
int32_t dq = (int32_t)ai->y - (int32_t)aj->y, dr, dd, dg, q_span, sc;
|
||||
dr = (int32_t)(ai->x - aj->x);
|
||||
*width = dd = dr > dq? dr - dq : dq - dr;
|
||||
dg = dr < dq? dr : dq;
|
||||
q_span = aj->y>>32&0xff;
|
||||
sc = q_span < dg? q_span : dg;
|
||||
if (exact) *exact = (dd == 0 && dg <= q_span);
|
||||
if (dd || dq > q_span) {
|
||||
float lin_pen, log_pen;
|
||||
lin_pen = chn_pen_gap * (float)dd + chn_pen_skip * (float)dg;
|
||||
log_pen = dd >= 1? mg_log2(dd + 1) : 0.0f; // mg_log2() only works for dd>=2
|
||||
sc -= (int)(lin_pen + .5f * log_pen);
|
||||
}
|
||||
return sc;
|
||||
}
|
||||
|
||||
mm128_t *mg_lchain_rmq(int max_dist, int max_dist_inner, int bw, int max_chn_skip, int cap_rmq_size, int min_cnt, int min_sc, float chn_pen_gap, float chn_pen_skip,
|
||||
int64_t n, mm128_t *a, int *n_u_, uint64_t **_u, void *km)
|
||||
{
|
||||
int32_t *f,*t, *v, n_u, n_v, mmax_f = 0, max_rmq_size = 0;
|
||||
int64_t *p, i, i0, st = 0, st_inner = 0, n_iter = 0;
|
||||
uint64_t *u;
|
||||
lc_elem_t *root = 0, *root_inner = 0;
|
||||
void *mem_mp = 0;
|
||||
kmp_rmq_t *mp;
|
||||
|
||||
if (_u) *_u = 0, *n_u_ = 0;
|
||||
if (n == 0 || a == 0) {
|
||||
kfree(km, a);
|
||||
return 0;
|
||||
}
|
||||
if (max_dist < bw) max_dist = bw;
|
||||
if (max_dist_inner <= 0 || max_dist_inner >= max_dist) max_dist_inner = 0;
|
||||
KMALLOC(km, p, n);
|
||||
KMALLOC(km, f, n);
|
||||
KCALLOC(km, t, n);
|
||||
KMALLOC(km, v, n);
|
||||
mem_mp = km_init2(km, 0x10000);
|
||||
mp = kmp_init_rmq(mem_mp);
|
||||
|
||||
// fill the score and backtrack arrays
|
||||
for (i = i0 = 0; i < n; ++i) {
|
||||
int64_t max_j = -1;
|
||||
int32_t q_span = a[i].y>>32&0xff, max_f = q_span;
|
||||
lc_elem_t s, *q, *r, lo, hi;
|
||||
// add in-range anchors
|
||||
if (i0 < i && a[i0].x != a[i].x) {
|
||||
int64_t j;
|
||||
for (j = i0; j < i; ++j) {
|
||||
q = kmp_alloc_rmq(mp);
|
||||
q->y = (int32_t)a[j].y, q->i = j, q->pri = -(f[j] + 0.5 * chn_pen_gap * ((int32_t)a[j].x + (int32_t)a[j].y));
|
||||
krmq_insert(lc_elem, &root, q, 0);
|
||||
if (max_dist_inner > 0) {
|
||||
r = kmp_alloc_rmq(mp);
|
||||
*r = *q;
|
||||
krmq_insert(lc_elem, &root_inner, r, 0);
|
||||
}
|
||||
}
|
||||
i0 = i;
|
||||
}
|
||||
// get rid of active chains out of range
|
||||
while (st < i && (a[i].x>>32 != a[st].x>>32 || a[i].x > a[st].x + max_dist || krmq_size(head, root) > cap_rmq_size)) {
|
||||
s.y = (int32_t)a[st].y, s.i = st;
|
||||
if ((q = krmq_find(lc_elem, root, &s, 0)) != 0) {
|
||||
q = krmq_erase(lc_elem, &root, q, 0);
|
||||
kmp_free_rmq(mp, q);
|
||||
}
|
||||
++st;
|
||||
}
|
||||
if (max_dist_inner > 0) { // similar to the block above, but applied to the inner tree
|
||||
while (st_inner < i && (a[i].x>>32 != a[st_inner].x>>32 || a[i].x > a[st_inner].x + max_dist_inner || krmq_size(head, root_inner) > cap_rmq_size)) {
|
||||
s.y = (int32_t)a[st_inner].y, s.i = st_inner;
|
||||
if ((q = krmq_find(lc_elem, root_inner, &s, 0)) != 0) {
|
||||
q = krmq_erase(lc_elem, &root_inner, q, 0);
|
||||
kmp_free_rmq(mp, q);
|
||||
}
|
||||
++st_inner;
|
||||
}
|
||||
}
|
||||
// RMQ
|
||||
lo.i = INT32_MAX, lo.y = (int32_t)a[i].y - max_dist;
|
||||
hi.i = 0, hi.y = (int32_t)a[i].y;
|
||||
if ((q = krmq_rmq(lc_elem, root, &lo, &hi)) != 0) {
|
||||
int32_t sc, exact, width, n_skip = 0;
|
||||
int64_t j = q->i;
|
||||
assert(q->y >= lo.y && q->y <= hi.y);
|
||||
sc = f[j] + comput_sc_simple(&a[i], &a[j], chn_pen_gap, chn_pen_skip, &exact, &width);
|
||||
if (width <= bw && sc > max_f) max_f = sc, max_j = j;
|
||||
if (!exact && root_inner && (int32_t)a[i].y > 0) {
|
||||
lc_elem_t *lo, *hi;
|
||||
s.y = (int32_t)a[i].y - 1, s.i = n;
|
||||
krmq_interval(lc_elem, root_inner, &s, &lo, &hi);
|
||||
if (lo) {
|
||||
const lc_elem_t *q;
|
||||
int32_t width, n_rmq_iter = 0;
|
||||
krmq_itr_t(lc_elem) itr;
|
||||
krmq_itr_find(lc_elem, root_inner, lo, &itr);
|
||||
while ((q = krmq_at(&itr)) != 0) {
|
||||
if (q->y < (int32_t)a[i].y - max_dist_inner) break;
|
||||
++n_rmq_iter;
|
||||
j = q->i;
|
||||
sc = f[j] + comput_sc_simple(&a[i], &a[j], chn_pen_gap, chn_pen_skip, 0, &width);
|
||||
if (width <= bw) {
|
||||
if (sc > max_f) {
|
||||
max_f = sc, max_j = j;
|
||||
if (n_skip > 0) --n_skip;
|
||||
} else if (t[j] == (int32_t)i) {
|
||||
if (++n_skip > max_chn_skip)
|
||||
break;
|
||||
}
|
||||
if (p[j] >= 0) t[p[j]] = i;
|
||||
}
|
||||
if (!krmq_itr_prev(lc_elem, &itr)) break;
|
||||
}
|
||||
n_iter += n_rmq_iter;
|
||||
}
|
||||
}
|
||||
}
|
||||
// set max
|
||||
assert(max_j < 0 || (a[max_j].x < a[i].x && (int32_t)a[max_j].y < (int32_t)a[i].y));
|
||||
f[i] = max_f, p[i] = max_j;
|
||||
v[i] = max_j >= 0 && v[max_j] > max_f? v[max_j] : max_f; // v[] keeps the peak score up to i; f[] is the score ending at i, not always the peak
|
||||
if (mmax_f < max_f) mmax_f = max_f;
|
||||
if (max_rmq_size < krmq_size(head, root)) max_rmq_size = krmq_size(head, root);
|
||||
}
|
||||
km_destroy(mem_mp);
|
||||
|
||||
u = mg_chain_backtrack(km, n, f, p, v, t, min_cnt, min_sc, &n_u, &n_v);
|
||||
*n_u_ = n_u, *_u = u; // NB: note that u[] may not be sorted by score here
|
||||
kfree(km, p); kfree(km, f); kfree(km, t);
|
||||
if (n_u == 0) {
|
||||
kfree(km, a); kfree(km, v);
|
||||
return 0;
|
||||
}
|
||||
return compact_a(km, n_u, u, n_v, v, a);
|
||||
}
|
||||
Submodule
+1
Submodule lib/simde added at b30129b3b4
@@ -1,12 +1,13 @@
|
||||
#include <stdlib.h>
|
||||
#include <stdio.h>
|
||||
#include <string.h>
|
||||
#include <errno.h>
|
||||
#include "bseq.h"
|
||||
#include "minimap.h"
|
||||
#include "mmpriv.h"
|
||||
#include "getopt.h"
|
||||
#include "ketopt.h"
|
||||
|
||||
#define MM_VERSION "2.8-r672"
|
||||
#define MM_VERSION "2.21-dev-r1094-dirty"
|
||||
|
||||
#ifdef __linux__
|
||||
#include <sys/resource.h>
|
||||
@@ -22,57 +23,86 @@ void liftrlimit()
|
||||
void liftrlimit() {}
|
||||
#endif
|
||||
|
||||
static struct option long_options[] = {
|
||||
{ "bucket-bits", required_argument, 0, 0 },
|
||||
{ "mb-size", required_argument, 0, 'K' },
|
||||
{ "seed", required_argument, 0, 0 },
|
||||
{ "no-kalloc", no_argument, 0, 0 },
|
||||
{ "print-qname", no_argument, 0, 0 },
|
||||
{ "no-self", no_argument, 0, 'D' },
|
||||
{ "print-seeds", no_argument, 0, 0 },
|
||||
{ "max-chain-skip", required_argument, 0, 0 },
|
||||
{ "min-dp-len", required_argument, 0, 0 },
|
||||
{ "print-aln-seq", no_argument, 0, 0 },
|
||||
{ "splice", no_argument, 0, 0 },
|
||||
{ "cost-non-gt-ag", required_argument, 0, 'C' },
|
||||
{ "no-long-join", no_argument, 0, 0 },
|
||||
{ "sr", no_argument, 0, 0 },
|
||||
{ "frag", required_argument, 0, 0 },
|
||||
{ "secondary", required_argument, 0, 0 },
|
||||
{ "cs", optional_argument, 0, 0 },
|
||||
{ "end-bonus", required_argument, 0, 0 },
|
||||
{ "no-pairing", no_argument, 0, 0 },
|
||||
{ "splice-flank", required_argument, 0, 0 },
|
||||
{ "idx-no-seq", no_argument, 0, 0 },
|
||||
{ "end-seed-pen", required_argument, 0, 0 }, // 21
|
||||
{ "for-only", no_argument, 0, 0 }, // 22
|
||||
{ "rev-only", no_argument, 0, 0 }, // 23
|
||||
{ "heap-sort", required_argument, 0, 0 }, // 24
|
||||
{ "all-chain", no_argument, 0, 'P' },
|
||||
{ "dual", required_argument, 0, 0 }, // 26
|
||||
{ "help", no_argument, 0, 'h' },
|
||||
{ "max-intron-len", required_argument, 0, 'G' },
|
||||
{ "version", no_argument, 0, 'V' },
|
||||
{ "min-count", required_argument, 0, 'n' },
|
||||
{ "min-chain-score",required_argument, 0, 'm' },
|
||||
{ "mask-level", required_argument, 0, 'M' },
|
||||
{ "min-dp-score", required_argument, 0, 's' },
|
||||
{ "sam", no_argument, 0, 'a' },
|
||||
{ 0, 0, 0, 0}
|
||||
static ko_longopt_t long_options[] = {
|
||||
{ "bucket-bits", ko_required_argument, 300 },
|
||||
{ "mb-size", ko_required_argument, 'K' },
|
||||
{ "seed", ko_required_argument, 302 },
|
||||
{ "no-kalloc", ko_no_argument, 303 },
|
||||
{ "print-qname", ko_no_argument, 304 },
|
||||
{ "no-self", ko_no_argument, 'D' },
|
||||
{ "print-seeds", ko_no_argument, 306 },
|
||||
{ "max-chain-skip", ko_required_argument, 307 },
|
||||
{ "min-dp-len", ko_required_argument, 308 },
|
||||
{ "print-aln-seq", ko_no_argument, 309 },
|
||||
{ "splice", ko_no_argument, 310 },
|
||||
{ "cost-non-gt-ag", ko_required_argument, 'C' },
|
||||
{ "no-long-join", ko_no_argument, 312 },
|
||||
{ "sr", ko_no_argument, 313 },
|
||||
{ "frag", ko_required_argument, 314 },
|
||||
{ "secondary", ko_required_argument, 315 },
|
||||
{ "cs", ko_optional_argument, 316 },
|
||||
{ "end-bonus", ko_required_argument, 317 },
|
||||
{ "no-pairing", ko_no_argument, 318 },
|
||||
{ "splice-flank", ko_required_argument, 319 },
|
||||
{ "idx-no-seq", ko_no_argument, 320 },
|
||||
{ "end-seed-pen", ko_required_argument, 321 },
|
||||
{ "for-only", ko_no_argument, 322 },
|
||||
{ "rev-only", ko_no_argument, 323 },
|
||||
{ "heap-sort", ko_required_argument, 324 },
|
||||
{ "all-chain", ko_no_argument, 'P' },
|
||||
{ "dual", ko_required_argument, 326 },
|
||||
{ "max-clip-ratio", ko_required_argument, 327 },
|
||||
{ "min-occ-floor", ko_required_argument, 328 },
|
||||
{ "MD", ko_no_argument, 329 },
|
||||
{ "lj-min-ratio", ko_required_argument, 330 },
|
||||
{ "score-N", ko_required_argument, 331 },
|
||||
{ "eqx", ko_no_argument, 332 },
|
||||
{ "paf-no-hit", ko_no_argument, 333 },
|
||||
{ "split-prefix", ko_required_argument, 334 },
|
||||
{ "no-end-flt", ko_no_argument, 335 },
|
||||
{ "hard-mask-level",ko_no_argument, 336 },
|
||||
{ "cap-sw-mem", ko_required_argument, 337 },
|
||||
{ "max-qlen", ko_required_argument, 338 },
|
||||
{ "max-chain-iter", ko_required_argument, 339 },
|
||||
{ "junc-bed", ko_required_argument, 340 },
|
||||
{ "junc-bonus", ko_required_argument, 341 },
|
||||
{ "sam-hit-only", ko_no_argument, 342 },
|
||||
{ "chain-gap-scale",ko_required_argument, 343 },
|
||||
{ "alt", ko_required_argument, 344 },
|
||||
{ "alt-drop", ko_required_argument, 345 },
|
||||
{ "mask-len", ko_required_argument, 346 },
|
||||
{ "rmq", ko_optional_argument, 347 },
|
||||
{ "qstrand", ko_no_argument, 348 },
|
||||
{ "cap-kalloc", ko_required_argument, 349 },
|
||||
{ "help", ko_no_argument, 'h' },
|
||||
{ "max-intron-len", ko_required_argument, 'G' },
|
||||
{ "version", ko_no_argument, 'V' },
|
||||
{ "min-count", ko_required_argument, 'n' },
|
||||
{ "min-chain-score",ko_required_argument, 'm' },
|
||||
{ "mask-level", ko_required_argument, 'M' },
|
||||
{ "min-dp-score", ko_required_argument, 's' },
|
||||
{ "sam", ko_no_argument, 'a' },
|
||||
{ 0, 0, 0 }
|
||||
};
|
||||
|
||||
static inline int64_t mm_parse_num(const char *str)
|
||||
static inline int64_t mm_parse_num2(const char *str, char **q)
|
||||
{
|
||||
double x;
|
||||
char *p;
|
||||
x = strtod(optarg, &p);
|
||||
if (*p == 'G' || *p == 'g') x *= 1e9;
|
||||
else if (*p == 'M' || *p == 'm') x *= 1e6;
|
||||
else if (*p == 'K' || *p == 'k') x *= 1e3;
|
||||
x = strtod(str, &p);
|
||||
if (*p == 'G' || *p == 'g') x *= 1e9, ++p;
|
||||
else if (*p == 'M' || *p == 'm') x *= 1e6, ++p;
|
||||
else if (*p == 'K' || *p == 'k') x *= 1e3, ++p;
|
||||
if (q) *q = p;
|
||||
return (int64_t)(x + .499);
|
||||
}
|
||||
|
||||
static inline void yes_or_no(mm_mapopt_t *opt, int flag, int long_idx, const char *arg, int yes_to_set)
|
||||
static inline int64_t mm_parse_num(const char *str)
|
||||
{
|
||||
return mm_parse_num2(str, 0);
|
||||
}
|
||||
|
||||
static inline void yes_or_no(mm_mapopt_t *opt, int64_t flag, int long_idx, const char *arg, int yes_to_set)
|
||||
{
|
||||
if (yes_to_set) {
|
||||
if (strcmp(arg, "yes") == 0 || strcmp(arg, "y") == 0) opt->flag |= flag;
|
||||
@@ -87,11 +117,12 @@ static inline void yes_or_no(mm_mapopt_t *opt, int flag, int long_idx, const cha
|
||||
|
||||
int main(int argc, char *argv[])
|
||||
{
|
||||
const char *opt_str = "2aSDw:k:K:t:r:f:Vv:g:G:I:d:XT:s:x:Hcp:M:n:z:A:B:O:E:m:N:Qu:R:hF:LC:";
|
||||
const char *opt_str = "2aSDw:k:K:t:r:f:Vv:g:G:I:d:XT:s:x:Hcp:M:n:z:A:B:O:E:m:N:Qu:R:hF:LC:yYPo:e:U:";
|
||||
ketopt_t o = KETOPT_INIT;
|
||||
mm_mapopt_t opt;
|
||||
mm_idxopt_t ipt;
|
||||
int i, c, n_threads = 3, long_idx;
|
||||
char *fnw = 0, *rg = 0, *s;
|
||||
int i, c, n_threads = 3, n_parts, old_best_n = -1;
|
||||
char *fnw = 0, *rg = 0, *junc_bed = 0, *s, *alt_list = 0;
|
||||
FILE *fp_help = stderr;
|
||||
mm_idx_reader_t *idx_rdr;
|
||||
mm_idx_t *mi;
|
||||
@@ -101,30 +132,35 @@ int main(int argc, char *argv[])
|
||||
mm_realtime0 = realtime();
|
||||
mm_set_opt(0, &ipt, &opt);
|
||||
|
||||
while ((c = getopt_long(argc, argv, opt_str, long_options, &long_idx)) >= 0) // apply option -x/preset first
|
||||
while ((c = ketopt(&o, argc, argv, 1, opt_str, long_options)) >= 0) { // test command line options and apply option -x/preset first
|
||||
if (c == 'x') {
|
||||
if (mm_set_opt(optarg, &ipt, &opt) < 0) {
|
||||
fprintf(stderr, "[ERROR] unknown preset '%s'\n", optarg);
|
||||
if (mm_set_opt(o.arg, &ipt, &opt) < 0) {
|
||||
fprintf(stderr, "[ERROR] unknown preset '%s'\n", o.arg);
|
||||
return 1;
|
||||
}
|
||||
break;
|
||||
} else if (c == ':') {
|
||||
fprintf(stderr, "[ERROR] missing option argument\n");
|
||||
return 1;
|
||||
} else if (c == '?') {
|
||||
fprintf(stderr, "[ERROR] unknown option in \"%s\"\n", argv[o.i - 1]);
|
||||
return 1;
|
||||
}
|
||||
optreset = 1;
|
||||
}
|
||||
o = KETOPT_INIT;
|
||||
|
||||
while ((c = getopt_long(argc, argv, opt_str, long_options, &long_idx)) >= 0) {
|
||||
if (c == 'w') ipt.w = atoi(optarg);
|
||||
else if (c == 'k') ipt.k = atoi(optarg);
|
||||
while ((c = ketopt(&o, argc, argv, 1, opt_str, long_options)) >= 0) {
|
||||
if (c == 'w') ipt.w = atoi(o.arg);
|
||||
else if (c == 'k') ipt.k = atoi(o.arg);
|
||||
else if (c == 'H') ipt.flag |= MM_I_HPC;
|
||||
else if (c == 'd') fnw = optarg; // the above are indexing related options, except -I
|
||||
else if (c == 'r') opt.bw = (int)mm_parse_num(optarg);
|
||||
else if (c == 't') n_threads = atoi(optarg);
|
||||
else if (c == 'v') mm_verbose = atoi(optarg);
|
||||
else if (c == 'g') opt.max_gap = (int)mm_parse_num(optarg);
|
||||
else if (c == 'G') mm_mapopt_max_intron_len(&opt, (int)mm_parse_num(optarg));
|
||||
else if (c == 'F') opt.max_frag_len = (int)mm_parse_num(optarg);
|
||||
else if (c == 'N') opt.best_n = atoi(optarg);
|
||||
else if (c == 'p') opt.pri_ratio = atof(optarg);
|
||||
else if (c == 'M') opt.mask_level = atof(optarg);
|
||||
else if (c == 'd') fnw = o.arg; // the above are indexing related options, except -I
|
||||
else if (c == 't') n_threads = atoi(o.arg);
|
||||
else if (c == 'v') mm_verbose = atoi(o.arg);
|
||||
else if (c == 'g') opt.max_gap = (int)mm_parse_num(o.arg);
|
||||
else if (c == 'G') mm_mapopt_max_intron_len(&opt, (int)mm_parse_num(o.arg));
|
||||
else if (c == 'F') opt.max_frag_len = (int)mm_parse_num(o.arg);
|
||||
else if (c == 'N') old_best_n = opt.best_n, opt.best_n = atoi(o.arg);
|
||||
else if (c == 'p') opt.pri_ratio = atof(o.arg);
|
||||
else if (c == 'M') opt.mask_level = atof(o.arg);
|
||||
else if (c == 'c') opt.flag |= MM_F_OUT_CG | MM_F_CIGAR;
|
||||
else if (c == 'D') opt.flag |= MM_F_NO_DIAG;
|
||||
else if (c == 'P') opt.flag |= MM_F_ALL_CHAINS;
|
||||
@@ -133,57 +169,91 @@ int main(int argc, char *argv[])
|
||||
else if (c == 'Q') opt.flag |= MM_F_NO_QUAL;
|
||||
else if (c == 'Y') opt.flag |= MM_F_SOFTCLIP;
|
||||
else if (c == 'L') opt.flag |= MM_F_LONG_CIGAR;
|
||||
else if (c == 'T') opt.sdust_thres = atoi(optarg);
|
||||
else if (c == 'n') opt.min_cnt = atoi(optarg);
|
||||
else if (c == 'm') opt.min_chain_score = atoi(optarg);
|
||||
else if (c == 'A') opt.a = atoi(optarg);
|
||||
else if (c == 'B') opt.b = atoi(optarg);
|
||||
else if (c == 'z') opt.zdrop = atoi(optarg);
|
||||
else if (c == 's') opt.min_dp_max = atoi(optarg);
|
||||
else if (c == 'C') opt.noncan = atoi(optarg);
|
||||
else if (c == 'I') ipt.batch_size = mm_parse_num(optarg);
|
||||
else if (c == 'K') opt.mini_batch_size = (int)mm_parse_num(optarg);
|
||||
else if (c == 'R') rg = optarg;
|
||||
else if (c == 'y') opt.flag |= MM_F_COPY_COMMENT;
|
||||
else if (c == 'T') opt.sdust_thres = atoi(o.arg);
|
||||
else if (c == 'n') opt.min_cnt = atoi(o.arg);
|
||||
else if (c == 'm') opt.min_chain_score = atoi(o.arg);
|
||||
else if (c == 'A') opt.a = atoi(o.arg);
|
||||
else if (c == 'B') opt.b = atoi(o.arg);
|
||||
else if (c == 's') opt.min_dp_max = atoi(o.arg);
|
||||
else if (c == 'C') opt.noncan = atoi(o.arg);
|
||||
else if (c == 'I') ipt.batch_size = mm_parse_num(o.arg);
|
||||
else if (c == 'K') opt.mini_batch_size = mm_parse_num(o.arg);
|
||||
else if (c == 'e') opt.occ_dist = mm_parse_num(o.arg);
|
||||
else if (c == 'R') rg = o.arg;
|
||||
else if (c == 'h') fp_help = stdout;
|
||||
else if (c == '2') opt.flag |= MM_F_2_IO_THREADS;
|
||||
else if (c == 0 && long_idx == 0) ipt.bucket_bits = atoi(optarg); // --bucket-bits
|
||||
else if (c == 0 && long_idx == 2) opt.seed = atoi(optarg); // --seed
|
||||
else if (c == 0 && long_idx == 3) mm_dbg_flag |= MM_DBG_NO_KALLOC; // --no-kalloc
|
||||
else if (c == 0 && long_idx == 4) mm_dbg_flag |= MM_DBG_PRINT_QNAME; // --print-qname
|
||||
else if (c == 0 && long_idx == 6) mm_dbg_flag |= MM_DBG_PRINT_QNAME | MM_DBG_PRINT_SEED, n_threads = 1; // --print-seed
|
||||
else if (c == 0 && long_idx == 7) opt.max_chain_skip = atoi(optarg); // --max-chain-skip
|
||||
else if (c == 0 && long_idx == 8) opt.min_ksw_len = atoi(optarg); // --min-dp-len
|
||||
else if (c == 0 && long_idx == 9) mm_dbg_flag |= MM_DBG_PRINT_QNAME | MM_DBG_PRINT_ALN_SEQ, n_threads = 1; // --print-aln-seq
|
||||
else if (c == 0 && long_idx ==10) opt.flag |= MM_F_SPLICE; // --splice
|
||||
else if (c == 0 && long_idx ==12) opt.flag |= MM_F_NO_LJOIN; // --no-long-join
|
||||
else if (c == 0 && long_idx ==13) opt.flag |= MM_F_SR; // --sr
|
||||
else if (c == 0 && long_idx ==17) opt.end_bonus = atoi(optarg); // --end-bonus
|
||||
else if (c == 0 && long_idx ==18) opt.flag |= MM_F_INDEPEND_SEG; // --no-pairing
|
||||
else if (c == 0 && long_idx ==20) ipt.flag |= MM_I_NO_SEQ; // --idx-no-seq
|
||||
else if (c == 0 && long_idx ==21) opt.anchor_ext_shift = atoi(optarg); // --end-seed-pen
|
||||
else if (c == 0 && long_idx ==22) opt.flag |= MM_F_FOR_ONLY; // --for-only
|
||||
else if (c == 0 && long_idx ==23) opt.flag |= MM_F_REV_ONLY; // --rev-only
|
||||
else if (c == 0 && long_idx == 14) { // --frag
|
||||
yes_or_no(&opt, MM_F_FRAG_MODE, long_idx, optarg, 1);
|
||||
} else if (c == 0 && long_idx == 15) { // --secondary
|
||||
yes_or_no(&opt, MM_F_NO_PRINT_2ND, long_idx, optarg, 0);
|
||||
} else if (c == 0 && long_idx == 16) { // --cs
|
||||
else if (c == 'o') {
|
||||
if (strcmp(o.arg, "-") != 0) {
|
||||
if (freopen(o.arg, "wb", stdout) == NULL) {
|
||||
fprintf(stderr, "[ERROR]\033[1;31m failed to write the output to file '%s'\033[0m: %s\n", o.arg, strerror(errno));
|
||||
exit(1);
|
||||
}
|
||||
}
|
||||
}
|
||||
else if (c == 300) ipt.bucket_bits = atoi(o.arg); // --bucket-bits
|
||||
else if (c == 302) opt.seed = atoi(o.arg); // --seed
|
||||
else if (c == 303) mm_dbg_flag |= MM_DBG_NO_KALLOC; // --no-kalloc
|
||||
else if (c == 304) mm_dbg_flag |= MM_DBG_PRINT_QNAME; // --print-qname
|
||||
else if (c == 306) mm_dbg_flag |= MM_DBG_PRINT_QNAME | MM_DBG_PRINT_SEED, n_threads = 1; // --print-seed
|
||||
else if (c == 307) opt.max_chain_skip = atoi(o.arg); // --max-chain-skip
|
||||
else if (c == 339) opt.max_chain_iter = atoi(o.arg); // --max-chain-iter
|
||||
else if (c == 308) opt.min_ksw_len = atoi(o.arg); // --min-dp-len
|
||||
else if (c == 309) mm_dbg_flag |= MM_DBG_PRINT_QNAME | MM_DBG_PRINT_ALN_SEQ, n_threads = 1; // --print-aln-seq
|
||||
else if (c == 310) opt.flag |= MM_F_SPLICE; // --splice
|
||||
else if (c == 312) opt.flag |= MM_F_NO_LJOIN; // --no-long-join
|
||||
else if (c == 313) opt.flag |= MM_F_SR; // --sr
|
||||
else if (c == 317) opt.end_bonus = atoi(o.arg); // --end-bonus
|
||||
else if (c == 318) opt.flag |= MM_F_INDEPEND_SEG; // --no-pairing
|
||||
else if (c == 320) ipt.flag |= MM_I_NO_SEQ; // --idx-no-seq
|
||||
else if (c == 321) opt.anchor_ext_shift = atoi(o.arg); // --end-seed-pen
|
||||
else if (c == 322) opt.flag |= MM_F_FOR_ONLY; // --for-only
|
||||
else if (c == 323) opt.flag |= MM_F_REV_ONLY; // --rev-only
|
||||
else if (c == 327) opt.max_clip_ratio = atof(o.arg); // --max-clip-ratio
|
||||
else if (c == 328) opt.min_mid_occ = atoi(o.arg); // --min-occ-floor
|
||||
else if (c == 329) opt.flag |= MM_F_OUT_MD; // --MD
|
||||
else if (c == 331) opt.sc_ambi = atoi(o.arg); // --score-N
|
||||
else if (c == 332) opt.flag |= MM_F_EQX; // --eqx
|
||||
else if (c == 333) opt.flag |= MM_F_PAF_NO_HIT; // --paf-no-hit
|
||||
else if (c == 334) opt.split_prefix = o.arg; // --split-prefix
|
||||
else if (c == 335) opt.flag |= MM_F_NO_END_FLT; // --no-end-flt
|
||||
else if (c == 336) opt.flag |= MM_F_HARD_MLEVEL; // --hard-mask-level
|
||||
else if (c == 337) opt.max_sw_mat = mm_parse_num(o.arg); // --cap-sw-mat
|
||||
else if (c == 338) opt.max_qlen = mm_parse_num(o.arg); // --max-qlen
|
||||
else if (c == 340) junc_bed = o.arg; // --junc-bed
|
||||
else if (c == 341) opt.junc_bonus = atoi(o.arg); // --junc-bonus
|
||||
else if (c == 342) opt.flag |= MM_F_SAM_HIT_ONLY; // --sam-hit-only
|
||||
else if (c == 343) opt.chain_gap_scale = atof(o.arg); // --chain-gap-scale
|
||||
else if (c == 344) alt_list = o.arg; // --alt
|
||||
else if (c == 345) opt.alt_drop = atof(o.arg); // --alt-drop
|
||||
else if (c == 346) opt.mask_len = mm_parse_num(o.arg); // --mask-len
|
||||
else if (c == 348) opt.flag |= MM_F_QSTRAND | MM_F_NO_INV; // --qstrand
|
||||
else if (c == 349) opt.cap_kalloc = mm_parse_num(o.arg); // --cap-kalloc
|
||||
else if (c == 330) {
|
||||
fprintf(stderr, "[WARNING] \033[1;31m --lj-min-ratio has been deprecated.\033[0m\n");
|
||||
} else if (c == 314) { // --frag
|
||||
yes_or_no(&opt, MM_F_FRAG_MODE, o.longidx, o.arg, 1);
|
||||
} else if (c == 315) { // --secondary
|
||||
yes_or_no(&opt, MM_F_NO_PRINT_2ND, o.longidx, o.arg, 0);
|
||||
} else if (c == 316) { // --cs
|
||||
opt.flag |= MM_F_OUT_CS | MM_F_CIGAR;
|
||||
if (optarg == 0 || strcmp(optarg, "short") == 0) {
|
||||
if (o.arg == 0 || strcmp(o.arg, "short") == 0) {
|
||||
opt.flag &= ~MM_F_OUT_CS_LONG;
|
||||
} else if (strcmp(optarg, "long") == 0) {
|
||||
} else if (strcmp(o.arg, "long") == 0) {
|
||||
opt.flag |= MM_F_OUT_CS_LONG;
|
||||
} else if (strcmp(optarg, "none") == 0) {
|
||||
} else if (strcmp(o.arg, "none") == 0) {
|
||||
opt.flag &= ~MM_F_OUT_CS;
|
||||
} else if (mm_verbose >= 2) {
|
||||
fprintf(stderr, "[WARNING]\033[1;31m --cs only takes 'short' or 'long'. Invalid values are assumed to be 'short'.\033[0m\n");
|
||||
}
|
||||
} else if (c == 0 && long_idx == 19) { // --splice-flank
|
||||
yes_or_no(&opt, MM_F_SPLICE_FLANK, long_idx, optarg, 1);
|
||||
} else if (c == 0 && long_idx == 24) { // --heap-sort
|
||||
yes_or_no(&opt, MM_F_HEAP_SORT, long_idx, optarg, 1);
|
||||
} else if (c == 0 && long_idx == 26) { // --dual
|
||||
yes_or_no(&opt, MM_F_NO_DUAL, long_idx, optarg, 0);
|
||||
} else if (c == 319) { // --splice-flank
|
||||
yes_or_no(&opt, MM_F_SPLICE_FLANK, o.longidx, o.arg, 1);
|
||||
} else if (c == 324) { // --heap-sort
|
||||
yes_or_no(&opt, MM_F_HEAP_SORT, o.longidx, o.arg, 1);
|
||||
} else if (c == 326) { // --dual
|
||||
yes_or_no(&opt, MM_F_NO_DUAL, o.longidx, o.arg, 0);
|
||||
} else if (c == 347) { // --rmq
|
||||
yes_or_no(&opt, MM_F_RMQ, o.longidx, o.arg, 1);
|
||||
} else if (c == 'S') {
|
||||
opt.flag |= MM_F_OUT_CS | MM_F_CIGAR | MM_F_OUT_CS_LONG;
|
||||
if (mm_verbose >= 2)
|
||||
@@ -191,27 +261,36 @@ int main(int argc, char *argv[])
|
||||
} else if (c == 'V') {
|
||||
puts(MM_VERSION);
|
||||
return 0;
|
||||
} else if (c == 'r') {
|
||||
opt.bw = (int)mm_parse_num2(o.arg, &s);
|
||||
if (*s == ',') opt.bw_long = (int)mm_parse_num2(s + 1, &s);
|
||||
} else if (c == 'U') {
|
||||
opt.min_mid_occ = strtol(o.arg, &s, 10);
|
||||
if (*s == ',') opt.max_mid_occ = strtol(s + 1, &s, 10);
|
||||
} else if (c == 'f') {
|
||||
double x;
|
||||
char *p;
|
||||
x = strtod(optarg, &p);
|
||||
x = strtod(o.arg, &p);
|
||||
if (x < 1.0) opt.mid_occ_frac = x, opt.mid_occ = 0;
|
||||
else opt.mid_occ = (int)(x + .499);
|
||||
if (*p == ',') opt.max_occ = (int)(strtod(p+1, &p) + .499);
|
||||
} else if (c == 'u') {
|
||||
if (*optarg == 'b') opt.flag |= MM_F_SPLICE_FOR|MM_F_SPLICE_REV; // both strands
|
||||
else if (*optarg == 'f') opt.flag |= MM_F_SPLICE_FOR, opt.flag &= ~MM_F_SPLICE_REV; // match GT-AG
|
||||
else if (*optarg == 'r') opt.flag |= MM_F_SPLICE_REV, opt.flag &= ~MM_F_SPLICE_FOR; // match CT-AC (reverse complement of GT-AG)
|
||||
else if (*optarg == 'n') opt.flag &= ~(MM_F_SPLICE_FOR|MM_F_SPLICE_REV); // don't try to match the GT-AG signal
|
||||
if (*o.arg == 'b') opt.flag |= MM_F_SPLICE_FOR|MM_F_SPLICE_REV; // both strands
|
||||
else if (*o.arg == 'f') opt.flag |= MM_F_SPLICE_FOR, opt.flag &= ~MM_F_SPLICE_REV; // match GT-AG
|
||||
else if (*o.arg == 'r') opt.flag |= MM_F_SPLICE_REV, opt.flag &= ~MM_F_SPLICE_FOR; // match CT-AC (reverse complement of GT-AG)
|
||||
else if (*o.arg == 'n') opt.flag &= ~(MM_F_SPLICE_FOR|MM_F_SPLICE_REV); // don't try to match the GT-AG signal
|
||||
else {
|
||||
fprintf(stderr, "[ERROR]\033[1;31m unrecognized cDNA direction\033[0m\n");
|
||||
return 1;
|
||||
}
|
||||
} else if (c == 'z') {
|
||||
opt.zdrop = opt.zdrop_inv = strtol(o.arg, &s, 10);
|
||||
if (*s == ',') opt.zdrop_inv = strtol(s + 1, &s, 10);
|
||||
} else if (c == 'O') {
|
||||
opt.q = opt.q2 = strtol(optarg, &s, 10);
|
||||
opt.q = opt.q2 = strtol(o.arg, &s, 10);
|
||||
if (*s == ',') opt.q2 = strtol(s + 1, &s, 10);
|
||||
} else if (c == 'E') {
|
||||
opt.e = opt.e2 = strtol(optarg, &s, 10);
|
||||
opt.e = opt.e2 = strtol(o.arg, &s, 10);
|
||||
if (*s == ',') opt.e2 = strtol(s + 1, &s, 10);
|
||||
}
|
||||
}
|
||||
@@ -223,14 +302,18 @@ int main(int argc, char *argv[])
|
||||
ipt.flag |= MM_I_NO_SEQ;
|
||||
if (mm_check_opt(&ipt, &opt) < 0)
|
||||
return 1;
|
||||
if (opt.best_n == 0) {
|
||||
fprintf(stderr, "[WARNING]\033[1;31m changed '-N 0' to '-N %d --secondary=no'.\033[0m\n", old_best_n);
|
||||
opt.best_n = old_best_n, opt.flag |= MM_F_NO_PRINT_2ND;
|
||||
}
|
||||
|
||||
if (argc == optind || fp_help == stdout) {
|
||||
if (argc == o.ind || fp_help == stdout) {
|
||||
fprintf(fp_help, "Usage: minimap2 [options] <target.fa>|<target.idx> [query.fa] [...]\n");
|
||||
fprintf(fp_help, "Options:\n");
|
||||
fprintf(fp_help, " Indexing:\n");
|
||||
fprintf(fp_help, " -H use homopolymer-compressed k-mer\n");
|
||||
fprintf(fp_help, " -H use homopolymer-compressed k-mer (preferrable for PacBio)\n");
|
||||
fprintf(fp_help, " -k INT k-mer size (no larger than 28) [%d]\n", ipt.k);
|
||||
fprintf(fp_help, " -w INT minizer window size [%d]\n", ipt.w);
|
||||
fprintf(fp_help, " -w INT minimizer window size [%d]\n", ipt.w);
|
||||
fprintf(fp_help, " -I NUM split index for every ~NUM input bases [4G]\n");
|
||||
fprintf(fp_help, " -d FILE dump index to FILE []\n");
|
||||
fprintf(fp_help, " Mapping:\n");
|
||||
@@ -238,7 +321,7 @@ int main(int argc, char *argv[])
|
||||
fprintf(fp_help, " -g NUM stop chain enlongation if there are no minimizers in INT-bp [%d]\n", opt.max_gap);
|
||||
fprintf(fp_help, " -G NUM max intron length (effective with -xsplice; changing -r) [200k]\n");
|
||||
fprintf(fp_help, " -F NUM max fragment length (effective with -xsr or in the fragment mode) [800]\n");
|
||||
fprintf(fp_help, " -r NUM bandwidth used in chaining and DP-based alignment [%d]\n", opt.bw);
|
||||
fprintf(fp_help, " -r NUM[,NUM] chaining/alignment bandwidth and long-join bandwidth [%d,%d]\n", opt.bw, opt.bw_long);
|
||||
fprintf(fp_help, " -n INT minimal number of minimizers on a chain [%d]\n", opt.min_cnt);
|
||||
fprintf(fp_help, " -m INT minimal chaining score (matching bases minus log gap penalty) [%d]\n", opt.min_chain_score);
|
||||
// fprintf(fp_help, " -T INT SDUST threshold; 0 to disable SDUST [%d]\n", opt.sdust_thres); // TODO: this option is never used; might be buggy
|
||||
@@ -247,48 +330,48 @@ int main(int argc, char *argv[])
|
||||
fprintf(fp_help, " -N INT retain at most INT secondary alignments [%d]\n", opt.best_n);
|
||||
fprintf(fp_help, " Alignment:\n");
|
||||
fprintf(fp_help, " -A INT matching score [%d]\n", opt.a);
|
||||
fprintf(fp_help, " -B INT mismatch penalty [%d]\n", opt.b);
|
||||
fprintf(fp_help, " -B INT mismatch penalty (larger value for lower divergence) [%d]\n", opt.b);
|
||||
fprintf(fp_help, " -O INT[,INT] gap open penalty [%d,%d]\n", opt.q, opt.q2);
|
||||
fprintf(fp_help, " -E INT[,INT] gap extension penalty; a k-long gap costs min{O1+k*E1,O2+k*E2} [%d,%d]\n", opt.e, opt.e2);
|
||||
fprintf(fp_help, " -z INT Z-drop score [%d]\n", opt.zdrop);
|
||||
fprintf(fp_help, " -z INT[,INT] Z-drop score and inversion Z-drop score [%d,%d]\n", opt.zdrop, opt.zdrop_inv);
|
||||
fprintf(fp_help, " -s INT minimal peak DP alignment score [%d]\n", opt.min_dp_max);
|
||||
fprintf(fp_help, " -u CHAR how to find GT-AG. f:transcript strand, b:both strands, n:don't match GT-AG [n]\n");
|
||||
fprintf(fp_help, " Input/Output:\n");
|
||||
fprintf(fp_help, " -a output in the SAM format (PAF by default)\n");
|
||||
fprintf(fp_help, " -Q don't output base quality in SAM\n");
|
||||
fprintf(fp_help, " -o FILE output alignments to FILE [stdout]\n");
|
||||
fprintf(fp_help, " -L write CIGAR with >65535 ops at the CG tag\n");
|
||||
fprintf(fp_help, " -R STR SAM read group line in a format like '@RG\\tID:foo\\tSM:bar' []\n");
|
||||
fprintf(fp_help, " -c output CIGAR in PAF\n");
|
||||
fprintf(fp_help, " --cs[=STR] output the cs tag; STR is 'short' (if absent) or 'long' [none]\n");
|
||||
fprintf(fp_help, " --MD output the MD tag\n");
|
||||
fprintf(fp_help, " --eqx write =/X CIGAR operators\n");
|
||||
fprintf(fp_help, " -Y use soft clipping for supplementary alignments\n");
|
||||
fprintf(fp_help, " -t INT number of threads [%d]\n", n_threads);
|
||||
fprintf(fp_help, " -K NUM minibatch size for mapping [500M]\n");
|
||||
// fprintf(fp_help, " -v INT verbose level [%d]\n", mm_verbose);
|
||||
fprintf(fp_help, " --version show version number\n");
|
||||
fprintf(fp_help, " Preset:\n");
|
||||
fprintf(fp_help, " -x STR preset (always applied before other options) []\n");
|
||||
fprintf(fp_help, " map-pb: -Hk19 (PacBio vs reference mapping)\n");
|
||||
fprintf(fp_help, " map-ont: -k15 (Oxford Nanopore vs reference mapping)\n");
|
||||
fprintf(fp_help, " asm5: -k19 -w19 -A1 -B19 -O39,81 -E3,1 -s200 -z200 (asm to ref mapping; break at 5%% div.)\n");
|
||||
fprintf(fp_help, " asm10: -k19 -w19 -A1 -B9 -O16,41 -E2,1 -s200 -z200 (asm to ref mapping; break at 10%% div.)\n");
|
||||
fprintf(fp_help, " ava-pb: -Hk19 -Xw5 -m100 -g10000 --max-chain-skip 25 (PacBio read overlap)\n");
|
||||
fprintf(fp_help, " ava-ont: -k15 -Xw5 -m100 -g10000 --max-chain-skip 25 (ONT read overlap)\n");
|
||||
fprintf(fp_help, " splice: long-read spliced alignment (see minimap2.1 for details)\n");
|
||||
fprintf(fp_help, " sr: short single-end reads without splicing (see minimap2.1 for details)\n");
|
||||
fprintf(fp_help, "\nSee `man ./minimap2.1' for detailed description of command-line options.\n");
|
||||
fprintf(fp_help, " -x STR preset (always applied before other options; see minimap2.1 for details) []\n");
|
||||
fprintf(fp_help, " - map-pb/map-ont - PacBio CLR/Nanopore vs reference mapping\n");
|
||||
fprintf(fp_help, " - map-hifi - PacBio HiFi reads vs reference mapping\n");
|
||||
fprintf(fp_help, " - ava-pb/ava-ont - PacBio/Nanopore read overlap\n");
|
||||
fprintf(fp_help, " - asm5/asm10/asm20 - asm-to-ref mapping, for ~0.1/1/5%% sequence divergence\n");
|
||||
fprintf(fp_help, " - splice/splice:hq - long-read/Pacbio-CCS spliced alignment\n");
|
||||
fprintf(fp_help, " - sr - genomic short-read mapping\n");
|
||||
fprintf(fp_help, "\nSee `man ./minimap2.1' for detailed description of these and other advanced command-line options.\n");
|
||||
return fp_help == stdout? 0 : 1;
|
||||
}
|
||||
|
||||
if ((opt.flag & MM_F_SR) && argc - optind > 3) {
|
||||
if ((opt.flag & MM_F_SR) && argc - o.ind > 3) {
|
||||
fprintf(stderr, "[ERROR] incorrect input: in the sr mode, please specify no more than two query files.\n");
|
||||
return 1;
|
||||
}
|
||||
idx_rdr = mm_idx_reader_open(argv[optind], &ipt, fnw);
|
||||
idx_rdr = mm_idx_reader_open(argv[o.ind], &ipt, fnw);
|
||||
if (idx_rdr == 0) {
|
||||
fprintf(stderr, "[ERROR] failed to open file '%s'\n", argv[optind]);
|
||||
fprintf(stderr, "[ERROR] failed to open file '%s': %s\n", argv[o.ind], strerror(errno));
|
||||
return 1;
|
||||
}
|
||||
if (!idx_rdr->is_idx && fnw == 0 && argc - optind < 2) {
|
||||
if (!idx_rdr->is_idx && fnw == 0 && argc - o.ind < 2) {
|
||||
fprintf(stderr, "[ERROR] missing input: please specify a query file to map or option -d to keep the index\n");
|
||||
mm_idx_reader_close(idx_rdr);
|
||||
return 1;
|
||||
@@ -296,6 +379,7 @@ int main(int argc, char *argv[])
|
||||
if (opt.best_n == 0 && (opt.flag&MM_F_CIGAR) && mm_verbose >= 2)
|
||||
fprintf(stderr, "[WARNING]\033[1;31m `-N 0' reduces alignment accuracy. Please use --secondary=no to suppress secondary alignments.\033[0m\n");
|
||||
while ((mi = mm_idx_reader_read(idx_rdr, n_threads)) != 0) {
|
||||
int ret;
|
||||
if ((opt.flag & MM_F_CIGAR) && (mi->flag & MM_I_NO_SEQ)) {
|
||||
fprintf(stderr, "[ERROR] the prebuilt index doesn't contain sequences.\n");
|
||||
mm_idx_destroy(mi);
|
||||
@@ -304,32 +388,61 @@ int main(int argc, char *argv[])
|
||||
}
|
||||
if ((opt.flag & MM_F_OUT_SAM) && idx_rdr->n_parts == 1) {
|
||||
if (mm_idx_reader_eof(idx_rdr)) {
|
||||
mm_write_sam_hdr(mi, rg, MM_VERSION, argc, argv);
|
||||
if (opt.split_prefix == 0)
|
||||
ret = mm_write_sam_hdr(mi, rg, MM_VERSION, argc, argv);
|
||||
else
|
||||
ret = mm_write_sam_hdr(0, rg, MM_VERSION, argc, argv);
|
||||
} else {
|
||||
mm_write_sam_hdr(0, rg, MM_VERSION, argc, argv);
|
||||
if (mm_verbose >= 2)
|
||||
fprintf(stderr, "[WARNING]\033[1;31m For a multi-part index, no @SQ lines will be outputted.\033[0m\n");
|
||||
ret = mm_write_sam_hdr(0, rg, MM_VERSION, argc, argv);
|
||||
if (opt.split_prefix == 0 && mm_verbose >= 2)
|
||||
fprintf(stderr, "[WARNING]\033[1;31m For a multi-part index, no @SQ lines will be outputted. Please use --split-prefix.\033[0m\n");
|
||||
}
|
||||
if (ret != 0) {
|
||||
mm_idx_destroy(mi);
|
||||
mm_idx_reader_close(idx_rdr);
|
||||
return 1;
|
||||
}
|
||||
}
|
||||
if (mm_verbose >= 3)
|
||||
fprintf(stderr, "[M::%s::%.3f*%.2f] loaded/built the index for %d target sequence(s)\n",
|
||||
__func__, realtime() - mm_realtime0, cputime() / (realtime() - mm_realtime0), mi->n_seq);
|
||||
if (argc != optind + 1) mm_mapopt_update(&opt, mi);
|
||||
if (argc != o.ind + 1) mm_mapopt_update(&opt, mi);
|
||||
if (mm_verbose >= 3) mm_idx_stat(mi);
|
||||
if (junc_bed) mm_idx_bed_read(mi, junc_bed, 1);
|
||||
if (alt_list) mm_idx_alt_read(mi, alt_list);
|
||||
if (argc - (o.ind + 1) == 0) continue; // no query files
|
||||
ret = 0;
|
||||
if (!(opt.flag & MM_F_FRAG_MODE)) {
|
||||
for (i = optind + 1; i < argc; ++i)
|
||||
mm_map_file(mi, argv[i], &opt, n_threads);
|
||||
for (i = o.ind + 1; i < argc; ++i) {
|
||||
ret = mm_map_file(mi, argv[i], &opt, n_threads);
|
||||
if (ret < 0) break;
|
||||
}
|
||||
} else {
|
||||
mm_map_file_frag(mi, argc - (optind + 1), (const char**)&argv[optind + 1], &opt, n_threads);
|
||||
ret = mm_map_file_frag(mi, argc - (o.ind + 1), (const char**)&argv[o.ind + 1], &opt, n_threads);
|
||||
}
|
||||
mm_idx_destroy(mi);
|
||||
if (ret < 0) {
|
||||
fprintf(stderr, "ERROR: failed to map the query file\n");
|
||||
exit(EXIT_FAILURE);
|
||||
}
|
||||
}
|
||||
n_parts = idx_rdr->n_parts;
|
||||
mm_idx_reader_close(idx_rdr);
|
||||
|
||||
fprintf(stderr, "[M::%s] Version: %s\n", __func__, MM_VERSION);
|
||||
fprintf(stderr, "[M::%s] CMD:", __func__);
|
||||
for (i = 0; i < argc; ++i)
|
||||
fprintf(stderr, " %s", argv[i]);
|
||||
fprintf(stderr, "\n[M::%s] Real time: %.3f sec; CPU: %.3f sec\n", __func__, realtime() - mm_realtime0, cputime());
|
||||
if (opt.split_prefix)
|
||||
mm_split_merge(argc - (o.ind + 1), (const char**)&argv[o.ind + 1], &opt, n_parts);
|
||||
|
||||
if (fflush(stdout) == EOF) {
|
||||
perror("[ERROR] failed to write the results");
|
||||
exit(EXIT_FAILURE);
|
||||
}
|
||||
|
||||
if (mm_verbose >= 3) {
|
||||
fprintf(stderr, "[M::%s] Version: %s\n", __func__, MM_VERSION);
|
||||
fprintf(stderr, "[M::%s] CMD:", __func__);
|
||||
for (i = 0; i < argc; ++i)
|
||||
fprintf(stderr, " %s", argv[i]);
|
||||
fprintf(stderr, "\n[M::%s] Real time: %.3f sec; CPU: %.3f sec; Peak RSS: %.3f GB\n", __func__, realtime() - mm_realtime0, cputime(), peakrss() / 1024.0 / 1024.0 / 1024.0);
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <assert.h>
|
||||
#include <errno.h>
|
||||
#include "kthread.h"
|
||||
#include "kvec.h"
|
||||
#include "kalloc.h"
|
||||
@@ -11,6 +12,7 @@
|
||||
|
||||
struct mm_tbuf_s {
|
||||
void *km;
|
||||
int rep_len, frag_gap;
|
||||
};
|
||||
|
||||
mm_tbuf_t *mm_tbuf_init(void)
|
||||
@@ -28,6 +30,11 @@ void mm_tbuf_destroy(mm_tbuf_t *b)
|
||||
free(b);
|
||||
}
|
||||
|
||||
void *mm_tbuf_get_km(mm_tbuf_t *b)
|
||||
{
|
||||
return b->km;
|
||||
}
|
||||
|
||||
static int mm_dust_minier(void *km, int n, mm128_t *a, int l_seq, const char *seq, int sdust_thres)
|
||||
{
|
||||
int n_dreg, j, k, u = 0;
|
||||
@@ -39,12 +46,12 @@ static int mm_dust_minier(void *km, int n, mm128_t *a, int l_seq, const char *se
|
||||
for (j = k = 0; j < n; ++j) { // squeeze out minimizers that significantly overlap with LCRs
|
||||
int32_t qpos = (uint32_t)a[j].y>>1, span = a[j].x&0xff;
|
||||
int32_t s = qpos - (span - 1), e = s + span;
|
||||
while (u < n_dreg && (uint32_t)dreg[u] <= s) ++u;
|
||||
if (u < n_dreg && dreg[u]>>32 < e) {
|
||||
while (u < n_dreg && (int32_t)dreg[u] <= s) ++u;
|
||||
if (u < n_dreg && (int32_t)(dreg[u]>>32) < e) {
|
||||
int v, l = 0;
|
||||
for (v = u; v < n_dreg && dreg[v]>>32 < e; ++v) { // iterate over LCRs overlapping this minimizer
|
||||
int ss = s > dreg[v]>>32? s : dreg[v]>>32;
|
||||
int ee = e < (uint32_t)dreg[v]? e : (uint32_t)dreg[v];
|
||||
for (v = u; v < n_dreg && (int32_t)(dreg[v]>>32) < e; ++v) { // iterate over LCRs overlapping this minimizer
|
||||
int ss = s > (int32_t)(dreg[v]>>32)? s : dreg[v]>>32;
|
||||
int ee = e < (int32_t)dreg[v]? e : (uint32_t)dreg[v];
|
||||
l += ee - ss;
|
||||
}
|
||||
if (l <= span>>1) a[k++] = a[j]; // keep the minimizer if less than half of it falls in masked region
|
||||
@@ -56,9 +63,10 @@ static int mm_dust_minier(void *km, int n, mm128_t *a, int l_seq, const char *se
|
||||
|
||||
static void collect_minimizers(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int n_segs, const int *qlens, const char **seqs, mm128_v *mv)
|
||||
{
|
||||
int i, j, n, sum = 0;
|
||||
int i, n, sum = 0;
|
||||
mv->n = 0;
|
||||
for (i = n = 0; i < n_segs; ++i) {
|
||||
size_t j;
|
||||
mm_sketch(km, seqs[i], qlens[i], mi->w, mi->k, i, mi->flag&MM_I_HPC, mv);
|
||||
for (j = n; j < mv->n; ++j)
|
||||
mv->a[j].y += sum << 1;
|
||||
@@ -72,55 +80,14 @@ static void collect_minimizers(void *km, const mm_mapopt_t *opt, const mm_idx_t
|
||||
#define heap_lt(a, b) ((a).x > (b).x)
|
||||
KSORT_INIT(heap, mm128_t, heap_lt)
|
||||
|
||||
typedef struct {
|
||||
uint32_t n;
|
||||
uint32_t q_pos, q_span;
|
||||
uint32_t seg_id:31, is_tandem:1;
|
||||
const uint64_t *cr;
|
||||
} mm_match_t;
|
||||
|
||||
static mm_match_t *collect_matches(void *km, int *_n_m, int max_occ, const mm_idx_t *mi, const mm128_v *mv, int64_t *n_a, int *rep_len, int *n_mini_pos, uint64_t **mini_pos)
|
||||
{
|
||||
int i, rep_st = 0, rep_en = 0, n_m;
|
||||
mm_match_t *m;
|
||||
*n_mini_pos = 0;
|
||||
*mini_pos = (uint64_t*)kmalloc(km, mv->n * sizeof(uint64_t));
|
||||
m = (mm_match_t*)kmalloc(km, mv->n * sizeof(mm_match_t));
|
||||
for (i = n_m = 0, *rep_len = 0, *n_a = 0; i < mv->n; ++i) {
|
||||
const uint64_t *cr;
|
||||
mm128_t *p = &mv->a[i];
|
||||
uint32_t q_pos = (uint32_t)p->y, q_span = p->x & 0xff;
|
||||
int t;
|
||||
cr = mm_idx_get(mi, p->x>>8, &t);
|
||||
if (t >= max_occ) {
|
||||
int en = (q_pos >> 1) + 1, st = en - q_span;
|
||||
if (st > rep_en) {
|
||||
*rep_len += rep_en - rep_st;
|
||||
rep_st = st, rep_en = en;
|
||||
} else rep_en = en;
|
||||
} else {
|
||||
mm_match_t *q = &m[n_m++];
|
||||
q->q_pos = q_pos, q->q_span = q_span, q->cr = cr, q->n = t, q->seg_id = p->y >> 32;
|
||||
q->is_tandem = 0;
|
||||
if (i > 0 && p->x>>8 == mv->a[i - 1].x>>8) q->is_tandem = 1;
|
||||
if (i < mv->n - 1 && p->x>>8 == mv->a[i + 1].x>>8) q->is_tandem = 1;
|
||||
*n_a += q->n;
|
||||
(*mini_pos)[(*n_mini_pos)++] = (uint64_t)q_span<<32 | q_pos>>1;
|
||||
}
|
||||
}
|
||||
*rep_len += rep_en - rep_st;
|
||||
*_n_m = n_m;
|
||||
return m;
|
||||
}
|
||||
|
||||
static inline int skip_seed(int flag, uint64_t r, const mm_match_t *q, const char *qname, int qlen, const mm_idx_t *mi, int *is_self)
|
||||
static inline int skip_seed(int flag, uint64_t r, const mm_seed_t *q, const char *qname, int qlen, const mm_idx_t *mi, int *is_self)
|
||||
{
|
||||
*is_self = 0;
|
||||
if (qname && (flag & (MM_F_NO_DIAG|MM_F_NO_DUAL))) {
|
||||
const mm_idx_seq_t *s = &mi->seq[r>>32];
|
||||
int cmp;
|
||||
cmp = strcmp(qname, s->name);
|
||||
if ((flag&MM_F_NO_DIAG) && cmp == 0 && s->len == qlen) {
|
||||
if ((flag&MM_F_NO_DIAG) && cmp == 0 && (int)s->len == qlen) {
|
||||
if ((uint32_t)r>>1 == (q->q_pos>>1)) return 1; // avoid the diagnonal anchors
|
||||
if ((r&1) == (q->q_pos&1)) *is_self = 1; // this flag is used to avoid spurious extension on self chain
|
||||
}
|
||||
@@ -142,10 +109,10 @@ static mm128_t *collect_seed_hits_heap(void *km, const mm_mapopt_t *opt, int max
|
||||
{
|
||||
int i, n_m, heap_size = 0;
|
||||
int64_t j, n_for = 0, n_rev = 0;
|
||||
mm_match_t *m;
|
||||
mm_seed_t *m;
|
||||
mm128_t *a, *heap;
|
||||
|
||||
m = collect_matches(km, &n_m, max_occ, mi, mv, n_a, rep_len, n_mini_pos, mini_pos);
|
||||
m = mm_collect_matches(km, &n_m, qlen, max_occ, opt->max_max_occ, opt->occ_dist, mi, mv, n_a, rep_len, n_mini_pos, mini_pos);
|
||||
|
||||
heap = (mm128_t*)kmalloc(km, n_m * sizeof(mm128_t));
|
||||
a = (mm128_t*)kmalloc(km, *n_a * sizeof(mm128_t));
|
||||
@@ -159,23 +126,24 @@ static mm128_t *collect_seed_hits_heap(void *km, const mm_mapopt_t *opt, int max
|
||||
}
|
||||
ks_heapmake_heap(heap_size, heap);
|
||||
while (heap_size > 0) {
|
||||
mm_match_t *q = &m[heap->y>>32];
|
||||
mm_seed_t *q = &m[heap->y>>32];
|
||||
mm128_t *p;
|
||||
uint64_t r = heap->x;
|
||||
int32_t is_self, rpos = (uint32_t)r >> 1;
|
||||
if (skip_seed(opt->flag, r, q, qname, qlen, mi, &is_self)) continue;
|
||||
if ((r&1) == (q->q_pos&1)) { // forward strand
|
||||
p = &a[n_for++];
|
||||
p->x = (r&0xffffffff00000000ULL) | rpos;
|
||||
p->y = (uint64_t)q->q_span << 32 | q->q_pos >> 1;
|
||||
} else { // reverse strand
|
||||
p = &a[(*n_a) - (++n_rev)];
|
||||
p->x = 1ULL<<63 | (r&0xffffffff00000000ULL) | rpos;
|
||||
p->y = (uint64_t)q->q_span << 32 | (qlen - ((q->q_pos>>1) + 1 - q->q_span) - 1);
|
||||
if (!skip_seed(opt->flag, r, q, qname, qlen, mi, &is_self)) {
|
||||
if ((r&1) == (q->q_pos&1)) { // forward strand
|
||||
p = &a[n_for++];
|
||||
p->x = (r&0xffffffff00000000ULL) | rpos;
|
||||
p->y = (uint64_t)q->q_span << 32 | q->q_pos >> 1;
|
||||
} else { // reverse strand
|
||||
p = &a[(*n_a) - (++n_rev)];
|
||||
p->x = 1ULL<<63 | (r&0xffffffff00000000ULL) | rpos;
|
||||
p->y = (uint64_t)q->q_span << 32 | (qlen - ((q->q_pos>>1) + 1 - q->q_span) - 1);
|
||||
}
|
||||
p->y |= (uint64_t)q->seg_id << MM_SEED_SEG_SHIFT;
|
||||
if (q->is_tandem) p->y |= MM_SEED_TANDEM;
|
||||
if (is_self) p->y |= MM_SEED_SELF;
|
||||
}
|
||||
p->y |= (uint64_t)q->seg_id << MM_SEED_SEG_SHIFT;
|
||||
if (q->is_tandem) p->y |= MM_SEED_TANDEM;
|
||||
if (is_self) p->y |= MM_SEED_SELF;
|
||||
// update the heap
|
||||
if ((uint32_t)heap->y < q->n - 1) {
|
||||
++heap[0].y;
|
||||
@@ -205,14 +173,15 @@ static mm128_t *collect_seed_hits_heap(void *km, const mm_mapopt_t *opt, int max
|
||||
static mm128_t *collect_seed_hits(void *km, const mm_mapopt_t *opt, int max_occ, const mm_idx_t *mi, const char *qname, const mm128_v *mv, int qlen, int64_t *n_a, int *rep_len,
|
||||
int *n_mini_pos, uint64_t **mini_pos)
|
||||
{
|
||||
int i, k, n_m;
|
||||
mm_match_t *m;
|
||||
int i, n_m;
|
||||
mm_seed_t *m;
|
||||
mm128_t *a;
|
||||
m = collect_matches(km, &n_m, max_occ, mi, mv, n_a, rep_len, n_mini_pos, mini_pos);
|
||||
m = mm_collect_matches(km, &n_m, qlen, max_occ, opt->max_max_occ, opt->occ_dist, mi, mv, n_a, rep_len, n_mini_pos, mini_pos);
|
||||
a = (mm128_t*)kmalloc(km, *n_a * sizeof(mm128_t));
|
||||
for (i = 0, *n_a = 0; i < n_m; ++i) {
|
||||
mm_match_t *q = &m[i];
|
||||
mm_seed_t *q = &m[i];
|
||||
const uint64_t *r = q->cr;
|
||||
uint32_t k;
|
||||
for (k = 0; k < q->n; ++k) {
|
||||
int32_t is_self, rpos = (uint32_t)r[k] >> 1;
|
||||
mm128_t *p;
|
||||
@@ -221,9 +190,13 @@ static mm128_t *collect_seed_hits(void *km, const mm_mapopt_t *opt, int max_occ,
|
||||
if ((r[k]&1) == (q->q_pos&1)) { // forward strand
|
||||
p->x = (r[k]&0xffffffff00000000ULL) | rpos;
|
||||
p->y = (uint64_t)q->q_span << 32 | q->q_pos >> 1;
|
||||
} else { // reverse strand
|
||||
} else if (!(opt->flag & MM_F_QSTRAND)) { // reverse strand and not in the query-strand mode
|
||||
p->x = 1ULL<<63 | (r[k]&0xffffffff00000000ULL) | rpos;
|
||||
p->y = (uint64_t)q->q_span << 32 | (qlen - ((q->q_pos>>1) + 1 - q->q_span) - 1);
|
||||
} else { // reverse strand; query-strand
|
||||
int32_t len = mi->seq[r[k]>>32].len;
|
||||
p->x = 1ULL<<63 | (r[k]&0xffffffff00000000ULL) | (len - (rpos + 1 - q->q_span) - 1); // coordinate only accurate for non-HPC seeds
|
||||
p->y = (uint64_t)q->q_span << 32 | q->q_pos >> 1;
|
||||
}
|
||||
p->y |= (uint64_t)q->seg_id << MM_SEED_SEG_SHIFT;
|
||||
if (q->is_tandem) p->y |= MM_SEED_TANDEM;
|
||||
@@ -238,11 +211,9 @@ static mm128_t *collect_seed_hits(void *km, const mm_mapopt_t *opt, int max_occ,
|
||||
static void chain_post(const mm_mapopt_t *opt, int max_chain_gap_ref, const mm_idx_t *mi, void *km, int qlen, int n_segs, const int *qlens, int *n_regs, mm_reg1_t *regs, mm128_t *a)
|
||||
{
|
||||
if (!(opt->flag & MM_F_ALL_CHAINS)) { // don't choose primary mapping(s)
|
||||
mm_set_parent(km, opt->mask_level, *n_regs, regs, opt->a * 2 + opt->b);
|
||||
mm_set_parent(km, opt->mask_level, opt->mask_len, *n_regs, regs, opt->a * 2 + opt->b, opt->flag&MM_F_HARD_MLEVEL, opt->alt_drop);
|
||||
if (n_segs <= 1) mm_select_sub(km, opt->pri_ratio, mi->k*2, opt->best_n, n_regs, regs);
|
||||
else mm_select_sub_multi(km, opt->pri_ratio, 0.2f, 0.7f, max_chain_gap_ref, mi->k*2, opt->best_n, n_segs, qlens, n_regs, regs);
|
||||
if (!(opt->flag & (MM_F_SPLICE|MM_F_SR|MM_F_NO_LJOIN))) // long join not working well without primary chains
|
||||
mm_join_long(km, opt, qlen, n_regs, regs, a);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -251,7 +222,7 @@ static mm_reg1_t *align_regs(const mm_mapopt_t *opt, const mm_idx_t *mi, void *k
|
||||
if (!(opt->flag & MM_F_CIGAR)) return regs;
|
||||
regs = mm_align_skeleton(km, opt, mi, qlen, seq, n_regs, regs, a); // this calls mm_filter_regs()
|
||||
if (!(opt->flag & MM_F_ALL_CHAINS)) { // don't choose primary mapping(s)
|
||||
mm_set_parent(km, opt->mask_level, *n_regs, regs, opt->a * 2 + opt->b);
|
||||
mm_set_parent(km, opt->mask_level, opt->mask_len, *n_regs, regs, opt->a * 2 + opt->b, opt->flag&MM_F_HARD_MLEVEL, opt->alt_drop);
|
||||
mm_select_sub(km, opt->pri_ratio, mi->k*2, opt->best_n, n_regs, regs);
|
||||
mm_set_sam_pri(*n_regs, regs);
|
||||
}
|
||||
@@ -274,6 +245,7 @@ void mm_map_frag(const mm_idx_t *mi, int n_segs, const int *qlens, const char **
|
||||
qlen_sum += qlens[i], n_regs[i] = 0, regs[i] = 0;
|
||||
|
||||
if (qlen_sum == 0 || n_segs <= 0 || n_segs > MM_MAX_SEG) return;
|
||||
if (opt->max_qlen > 0 && qlen_sum > opt->max_qlen) return;
|
||||
|
||||
hash = qname? __ac_X31_hash_string(qname) : 0;
|
||||
hash ^= __ac_Wang_hash(qlen_sum) + __ac_Wang_hash(opt->seed);
|
||||
@@ -301,17 +273,33 @@ void mm_map_frag(const mm_idx_t *mi, int n_segs, const int *qlens, const char **
|
||||
if (max_chain_gap_ref < opt->max_gap) max_chain_gap_ref = opt->max_gap;
|
||||
} else max_chain_gap_ref = opt->max_gap;
|
||||
|
||||
a = mm_chain_dp(max_chain_gap_ref, max_chain_gap_qry, opt->bw, opt->max_chain_skip, opt->min_cnt, opt->min_chain_score, is_splice, n_segs, n_a, a, &n_regs0, &u, b->km);
|
||||
if (opt->flag & MM_F_RMQ) {
|
||||
a = mg_lchain_rmq(opt->max_gap, opt->rmq_inner_dist, opt->bw, opt->max_chain_skip, opt->rmq_size_cap, opt->min_cnt, opt->min_chain_score,
|
||||
opt->chain_gap_scale * 0.01 * mi->k, 0.0f, n_a, a, &n_regs0, &u, b->km);
|
||||
} else {
|
||||
a = mg_lchain_dp(max_chain_gap_ref, max_chain_gap_qry, opt->bw, opt->max_chain_skip, opt->max_chain_iter, opt->min_cnt, opt->min_chain_score,
|
||||
opt->chain_gap_scale * 0.01 * mi->k, 0.0f, is_splice, n_segs, n_a, a, &n_regs0, &u, b->km);
|
||||
}
|
||||
|
||||
if (opt->max_occ > opt->mid_occ && rep_len > 0) {
|
||||
if (opt->bw_long > opt->bw && (opt->flag & (MM_F_SPLICE|MM_F_SR|MM_F_NO_LJOIN)) == 0 && n_segs == 1 && n_regs0 > 1) { // re-chain/long-join for long sequences
|
||||
int32_t st = (int32_t)a[0].y, en = (int32_t)a[(int32_t)u[0] - 1].y;
|
||||
if (qlen_sum - (en - st) > opt->rmq_rescue_size || en - st > qlen_sum * opt->rmq_rescue_ratio) {
|
||||
int32_t i;
|
||||
for (i = 0, n_a = 0; i < n_regs0; ++i) n_a += (int32_t)u[i];
|
||||
kfree(b->km, u);
|
||||
radix_sort_128x(a, a + n_a);
|
||||
a = mg_lchain_rmq(opt->max_gap, opt->rmq_inner_dist, opt->bw_long, opt->max_chain_skip, opt->rmq_size_cap, opt->min_cnt, opt->min_chain_score,
|
||||
opt->chain_gap_scale * 0.01 * mi->k, 0.0f, n_a, a, &n_regs0, &u, b->km);
|
||||
}
|
||||
} else if (opt->max_occ > opt->mid_occ && rep_len > 0 && !(opt->flag & MM_F_RMQ)) { // re-chain, mostly for short reads
|
||||
int rechain = 0;
|
||||
if (n_regs0 > 0) { // test if the best chain has all the segments
|
||||
int n_chained_segs = 1, max = 0, max_i = -1, max_off = -1, off = 0;
|
||||
for (i = 0; i < n_regs0; ++i) { // find the best chain
|
||||
if (max < u[i]>>32) max = u[i]>>32, max_i = i, max_off = off;
|
||||
if (max < (int)(u[i]>>32)) max = u[i]>>32, max_i = i, max_off = off;
|
||||
off += (uint32_t)u[i];
|
||||
}
|
||||
for (i = 1; i < (uint32_t)u[max_i]; ++i) // count the number of segments in the best chain
|
||||
for (i = 1; i < (int32_t)u[max_i]; ++i) // count the number of segments in the best chain
|
||||
if ((a[max_off+i].y&MM_SEED_SEG_MASK) != (a[max_off+i-1].y&MM_SEED_SEG_MASK))
|
||||
++n_chained_segs;
|
||||
if (n_chained_segs < n_segs)
|
||||
@@ -323,11 +311,18 @@ void mm_map_frag(const mm_idx_t *mi, int n_segs, const int *qlens, const char **
|
||||
kfree(b->km, mini_pos);
|
||||
if (opt->flag & MM_F_HEAP_SORT) a = collect_seed_hits_heap(b->km, opt, opt->max_occ, mi, qname, &mv, qlen_sum, &n_a, &rep_len, &n_mini_pos, &mini_pos);
|
||||
else a = collect_seed_hits(b->km, opt, opt->max_occ, mi, qname, &mv, qlen_sum, &n_a, &rep_len, &n_mini_pos, &mini_pos);
|
||||
a = mm_chain_dp(max_chain_gap_ref, max_chain_gap_qry, opt->bw, opt->max_chain_skip, opt->min_cnt, opt->min_chain_score, is_splice, n_segs, n_a, a, &n_regs0, &u, b->km);
|
||||
a = mg_lchain_dp(max_chain_gap_ref, max_chain_gap_qry, opt->bw, opt->max_chain_skip, opt->max_chain_iter, opt->min_cnt, opt->min_chain_score,
|
||||
opt->chain_gap_scale * 0.01 * mi->k, 0.0f, is_splice, n_segs, n_a, a, &n_regs0, &u, b->km);
|
||||
}
|
||||
}
|
||||
b->frag_gap = max_chain_gap_ref;
|
||||
b->rep_len = rep_len;
|
||||
|
||||
regs0 = mm_gen_regs(b->km, hash, qlen_sum, n_regs0, u, a);
|
||||
regs0 = mm_gen_regs(b->km, hash, qlen_sum, n_regs0, u, a, !!(opt->flag&MM_F_QSTRAND));
|
||||
if (mi->n_alt) {
|
||||
mm_mark_alt(mi, n_regs0, regs0);
|
||||
mm_hit_sort(b->km, &n_regs0, regs0, opt->alt_drop); // this step can be merged into mm_gen_regs(); will do if this shows up in profile
|
||||
}
|
||||
|
||||
if (mm_dbg_flag & MM_DBG_PRINT_SEED)
|
||||
for (j = 0; j < n_regs0; ++j)
|
||||
@@ -336,20 +331,22 @@ void mm_map_frag(const mm_idx_t *mi, int n_segs, const int *qlens, const char **
|
||||
i == regs0[j].as? 0 : ((int32_t)a[i].y - (int32_t)a[i-1].y) - ((int32_t)a[i].x - (int32_t)a[i-1].x));
|
||||
|
||||
chain_post(opt, max_chain_gap_ref, mi, b->km, qlen_sum, n_segs, qlens, &n_regs0, regs0, a);
|
||||
if (!is_sr) mm_est_err(mi, qlen_sum, n_regs0, regs0, a, n_mini_pos, mini_pos);
|
||||
if (!is_sr && !(opt->flag&MM_F_QSTRAND))
|
||||
mm_est_err(mi, qlen_sum, n_regs0, regs0, a, n_mini_pos, mini_pos);
|
||||
|
||||
if (n_segs == 1) { // uni-segment
|
||||
regs0 = align_regs(opt, mi, b->km, qlens[0], seqs[0], &n_regs0, regs0, a);
|
||||
mm_set_mapq(n_regs0, regs0, opt->min_chain_score, opt->a, rep_len, is_sr);
|
||||
regs0 = (mm_reg1_t*)realloc(regs0, sizeof(*regs0) * n_regs0);
|
||||
mm_set_mapq(b->km, n_regs0, regs0, opt->min_chain_score, opt->a, rep_len, is_sr);
|
||||
n_regs[0] = n_regs0, regs[0] = regs0;
|
||||
} else { // multi-segment
|
||||
mm_seg_t *seg;
|
||||
seg = mm_seg_gen(b->km, hash, n_segs, qlens, n_regs0, regs0, n_regs, regs, a); // split fragment chain to separate segment chains
|
||||
free(regs0);
|
||||
for (i = 0; i < n_segs; ++i) {
|
||||
mm_set_parent(b->km, opt->mask_level, n_regs[i], regs[i], opt->a * 2 + opt->b); // update mm_reg1_t::parent
|
||||
mm_set_parent(b->km, opt->mask_level, opt->mask_len, n_regs[i], regs[i], opt->a * 2 + opt->b, opt->flag&MM_F_HARD_MLEVEL, opt->alt_drop); // update mm_reg1_t::parent
|
||||
regs[i] = align_regs(opt, mi, b->km, qlens[i], seqs[i], &n_regs[i], regs[i], seg[i].a);
|
||||
mm_set_mapq(n_regs[i], regs[i], opt->min_chain_score, opt->a, rep_len, is_sr);
|
||||
mm_set_mapq(b->km, n_regs[i], regs[i], opt->min_chain_score, opt->a, rep_len, is_sr);
|
||||
}
|
||||
mm_seg_free(b->km, n_segs, seg);
|
||||
if (n_segs == 2 && opt->pe_ori >= 0 && (opt->flag&MM_F_CIGAR))
|
||||
@@ -366,7 +363,9 @@ void mm_map_frag(const mm_idx_t *mi, int n_segs, const int *qlens, const char **
|
||||
if (mm_dbg_flag & MM_DBG_PRINT_QNAME)
|
||||
fprintf(stderr, "QM\t%s\t%d\tcap=%ld,nCore=%ld,largest=%ld\n", qname, qlen_sum, kmst.capacity, kmst.n_cores, kmst.largest);
|
||||
assert(kmst.n_blocks == kmst.n_cores); // otherwise, there is a memory leak
|
||||
if (kmst.largest > 1U<<28) {
|
||||
if (kmst.largest > 1U<<28 || (opt->cap_kalloc > 0 && kmst.capacity > opt->cap_kalloc)) {
|
||||
if (mm_dbg_flag & MM_DBG_PRINT_QNAME)
|
||||
fprintf(stderr, "[W::%s] reset thread-local memory after read %s\n", __func__, qname);
|
||||
km_destroy(b->km);
|
||||
b->km = km_init();
|
||||
}
|
||||
@@ -385,18 +384,23 @@ mm_reg1_t *mm_map(const mm_idx_t *mi, int qlen, const char *seq, int *n_regs, mm
|
||||
**************************/
|
||||
|
||||
typedef struct {
|
||||
int mini_batch_size, n_processed, n_threads, n_fp;
|
||||
int n_processed, n_threads, n_fp;
|
||||
int64_t mini_batch_size;
|
||||
const mm_mapopt_t *opt;
|
||||
mm_bseq_file_t **fp;
|
||||
const mm_idx_t *mi;
|
||||
kstring_t str;
|
||||
|
||||
int n_parts;
|
||||
uint32_t *rid_shift;
|
||||
FILE *fp_split, **fp_parts;
|
||||
} pipeline_t;
|
||||
|
||||
typedef struct {
|
||||
const pipeline_t *p;
|
||||
int n_seq, n_frag;
|
||||
mm_bseq1_t *seq;
|
||||
int *n_reg, *seg_off, *n_seg;
|
||||
int *n_reg, *seg_off, *n_seg, *rep_len, *frag_gap;
|
||||
mm_reg1_t **reg;
|
||||
mm_tbuf_t **buf;
|
||||
} step_t;
|
||||
@@ -406,10 +410,13 @@ static void worker_for(void *_data, long i, int tid) // kt_for() callback
|
||||
step_t *s = (step_t*)_data;
|
||||
int qlens[MM_MAX_SEG], j, off = s->seg_off[i], pe_ori = s->p->opt->pe_ori;
|
||||
const char *qseqs[MM_MAX_SEG];
|
||||
double t = 0.0;
|
||||
mm_tbuf_t *b = s->buf[tid];
|
||||
assert(s->n_seg[i] <= MM_MAX_SEG);
|
||||
if (mm_dbg_flag & MM_DBG_PRINT_QNAME)
|
||||
if (mm_dbg_flag & MM_DBG_PRINT_QNAME) {
|
||||
fprintf(stderr, "QR\t%s\t%d\t%d\n", s->seq[off].name, tid, s->seq[off].l_seq);
|
||||
t = realtime();
|
||||
}
|
||||
for (j = 0; j < s->n_seg[i]; ++j) {
|
||||
if (s->n_seg[i] == 2 && ((j == 0 && (pe_ori>>1&1)) || (j == 1 && (pe_ori&1))))
|
||||
mm_revcomp_bseq(&s->seq[off + j]);
|
||||
@@ -417,10 +424,17 @@ static void worker_for(void *_data, long i, int tid) // kt_for() callback
|
||||
qseqs[j] = s->seq[off + j].seq;
|
||||
}
|
||||
if (s->p->opt->flag & MM_F_INDEPEND_SEG) {
|
||||
for (j = 0; j < s->n_seg[i]; ++j)
|
||||
for (j = 0; j < s->n_seg[i]; ++j) {
|
||||
mm_map_frag(s->p->mi, 1, &qlens[j], &qseqs[j], &s->n_reg[off+j], &s->reg[off+j], b, s->p->opt, s->seq[off+j].name);
|
||||
s->rep_len[off + j] = b->rep_len;
|
||||
s->frag_gap[off + j] = b->frag_gap;
|
||||
}
|
||||
} else {
|
||||
mm_map_frag(s->p->mi, s->n_seg[i], qlens, qseqs, &s->n_reg[off], &s->reg[off], b, s->p->opt, s->seq[off].name);
|
||||
for (j = 0; j < s->n_seg[i]; ++j) {
|
||||
s->rep_len[off + j] = b->rep_len;
|
||||
s->frag_gap[off + j] = b->frag_gap;
|
||||
}
|
||||
}
|
||||
for (j = 0; j < s->n_seg[i]; ++j) // flip the query strand and coordinate to the original read strand
|
||||
if (s->n_seg[i] == 2 && ((j == 0 && (pe_ori>>1&1)) || (j == 1 && (pe_ori&1)))) {
|
||||
@@ -434,6 +448,73 @@ static void worker_for(void *_data, long i, int tid) // kt_for() callback
|
||||
r->rev = !r->rev;
|
||||
}
|
||||
}
|
||||
if (mm_dbg_flag & MM_DBG_PRINT_QNAME)
|
||||
fprintf(stderr, "QT\t%s\t%d\t%.6f\n", s->seq[off].name, tid, realtime() - t);
|
||||
}
|
||||
|
||||
static void merge_hits(step_t *s)
|
||||
{
|
||||
int f, i, k0, k, max_seg = 0, *n_reg_part, *rep_len_part, *frag_gap_part, *qlens;
|
||||
void *km;
|
||||
FILE **fp = s->p->fp_parts;
|
||||
const mm_mapopt_t *opt = s->p->opt;
|
||||
|
||||
km = km_init();
|
||||
for (f = 0; f < s->n_frag; ++f)
|
||||
max_seg = max_seg > s->n_seg[f]? max_seg : s->n_seg[f];
|
||||
qlens = CALLOC(int, max_seg + s->p->n_parts * 3);
|
||||
n_reg_part = qlens + max_seg;
|
||||
rep_len_part = n_reg_part + s->p->n_parts;
|
||||
frag_gap_part = rep_len_part + s->p->n_parts;
|
||||
for (f = 0, k = k0 = 0; f < s->n_frag; ++f) {
|
||||
k0 = k;
|
||||
for (i = 0; i < s->n_seg[f]; ++i, ++k) {
|
||||
int j, l, t, rep_len = 0;
|
||||
qlens[i] = s->seq[k].l_seq;
|
||||
for (j = 0, s->n_reg[k] = 0; j < s->p->n_parts; ++j) {
|
||||
mm_err_fread(&n_reg_part[j], sizeof(int), 1, fp[j]);
|
||||
mm_err_fread(&rep_len_part[j], sizeof(int), 1, fp[j]);
|
||||
mm_err_fread(&frag_gap_part[j], sizeof(int), 1, fp[j]);
|
||||
s->n_reg[k] += n_reg_part[j];
|
||||
if (rep_len < rep_len_part[j])
|
||||
rep_len = rep_len_part[j];
|
||||
}
|
||||
s->reg[k] = CALLOC(mm_reg1_t, s->n_reg[k]);
|
||||
for (j = 0, l = 0; j < s->p->n_parts; ++j) {
|
||||
for (t = 0; t < n_reg_part[j]; ++t, ++l) {
|
||||
mm_reg1_t *r = &s->reg[k][l];
|
||||
uint32_t capacity;
|
||||
mm_err_fread(r, sizeof(mm_reg1_t), 1, fp[j]);
|
||||
r->rid += s->p->rid_shift[j];
|
||||
if (opt->flag & MM_F_CIGAR) {
|
||||
mm_err_fread(&capacity, 4, 1, fp[j]);
|
||||
r->p = (mm_extra_t*)calloc(capacity, 4);
|
||||
r->p->capacity = capacity;
|
||||
mm_err_fread(r->p, r->p->capacity, 4, fp[j]);
|
||||
}
|
||||
}
|
||||
}
|
||||
if (!(opt->flag&MM_F_SR) && s->seq[k].l_seq >= opt->rank_min_len)
|
||||
mm_update_dp_max(s->seq[k].l_seq, s->n_reg[k], s->reg[k], opt->rank_frac, opt->a, opt->b);
|
||||
for (j = 0; j < s->n_reg[k]; ++j) {
|
||||
mm_reg1_t *r = &s->reg[k][j];
|
||||
if (r->p) r->p->dp_max2 = 0; // reset ->dp_max2 as mm_set_parent() doesn't clear it; necessary with mm_update_dp_max()
|
||||
r->subsc = 0; // this may not be necessary
|
||||
r->n_sub = 0; // n_sub will be an underestimate as we don't see all the chains now, but it can't be accurate anyway
|
||||
}
|
||||
mm_hit_sort(km, &s->n_reg[k], s->reg[k], opt->alt_drop);
|
||||
mm_set_parent(km, opt->mask_level, opt->mask_len, s->n_reg[k], s->reg[k], opt->a * 2 + opt->b, opt->flag&MM_F_HARD_MLEVEL, opt->alt_drop);
|
||||
if (!(opt->flag & MM_F_ALL_CHAINS)) {
|
||||
mm_select_sub(km, opt->pri_ratio, s->p->mi->k*2, opt->best_n, &s->n_reg[k], s->reg[k]);
|
||||
mm_set_sam_pri(s->n_reg[k], s->reg[k]);
|
||||
}
|
||||
mm_set_mapq(km, s->n_reg[k], s->reg[k], opt->min_chain_score, opt->a, rep_len, !!(opt->flag & MM_F_SR));
|
||||
}
|
||||
if (s->n_seg[f] == 2 && opt->pe_ori >= 0 && (opt->flag&MM_F_CIGAR))
|
||||
mm_pair(km, frag_gap_part[0], opt->pe_bonus, opt->a * 2 + opt->b, opt->a, qlens, &s->n_reg[k0], &s->reg[k0]);
|
||||
}
|
||||
free(qlens);
|
||||
km_destroy(km);
|
||||
}
|
||||
|
||||
static void *worker_pipeline(void *shared, int step, void *in)
|
||||
@@ -442,11 +523,12 @@ static void *worker_pipeline(void *shared, int step, void *in)
|
||||
pipeline_t *p = (pipeline_t*)shared;
|
||||
if (step == 0) { // step 0: read sequences
|
||||
int with_qual = (!!(p->opt->flag & MM_F_OUT_SAM) && !(p->opt->flag & MM_F_NO_QUAL));
|
||||
int with_comment = !!(p->opt->flag & MM_F_COPY_COMMENT);
|
||||
int frag_mode = (p->n_fp > 1 || !!(p->opt->flag & MM_F_FRAG_MODE));
|
||||
step_t *s;
|
||||
s = (step_t*)calloc(1, sizeof(step_t));
|
||||
if (p->n_fp > 1) s->seq = mm_bseq_read_frag(p->n_fp, p->fp, p->mini_batch_size, with_qual, &s->n_seq);
|
||||
else s->seq = mm_bseq_read2(p->fp[0], p->mini_batch_size, with_qual, frag_mode, &s->n_seq);
|
||||
if (p->n_fp > 1) s->seq = mm_bseq_read_frag2(p->n_fp, p->fp, p->mini_batch_size, with_qual, with_comment, &s->n_seq);
|
||||
else s->seq = mm_bseq_read3(p->fp[0], p->mini_batch_size, with_qual, with_comment, frag_mode, &s->n_seq);
|
||||
if (s->seq) {
|
||||
s->p = p;
|
||||
for (i = 0; i < s->n_seq; ++i)
|
||||
@@ -454,9 +536,11 @@ static void *worker_pipeline(void *shared, int step, void *in)
|
||||
s->buf = (mm_tbuf_t**)calloc(p->n_threads, sizeof(mm_tbuf_t*));
|
||||
for (i = 0; i < p->n_threads; ++i)
|
||||
s->buf[i] = mm_tbuf_init();
|
||||
s->n_reg = (int*)calloc(3 * s->n_seq, sizeof(int));
|
||||
s->seg_off = s->n_reg + s->n_seq; // seg_off and n_seg are allocated together with n_reg
|
||||
s->n_reg = (int*)calloc(5 * s->n_seq, sizeof(int));
|
||||
s->seg_off = s->n_reg + s->n_seq; // seg_off, n_seg, rep_len and frag_gap are allocated together with n_reg
|
||||
s->n_seg = s->seg_off + s->n_seq;
|
||||
s->rep_len = s->n_seg + s->n_seq;
|
||||
s->frag_gap = s->rep_len + s->n_seq;
|
||||
s->reg = (mm_reg1_t**)calloc(s->n_seq, sizeof(mm_reg1_t*));
|
||||
for (i = 1, j = 0; i <= s->n_seq; ++i)
|
||||
if (i == s->n_seq || !frag_mode || !mm_qname_same(s->seq[i-1].name, s->seq[i].name)) {
|
||||
@@ -467,7 +551,8 @@ static void *worker_pipeline(void *shared, int step, void *in)
|
||||
return s;
|
||||
} else free(s);
|
||||
} else if (step == 1) { // step 1: map
|
||||
kt_for(p->n_threads, worker_for, in, ((step_t*)in)->n_frag);
|
||||
if (p->n_parts > 0) merge_hits((step_t*)in);
|
||||
else kt_for(p->n_threads, worker_for, in, ((step_t*)in)->n_frag);
|
||||
return in;
|
||||
} else if (step == 2) { // step 2: output
|
||||
void *km = 0;
|
||||
@@ -480,20 +565,36 @@ static void *worker_pipeline(void *shared, int step, void *in)
|
||||
int seg_st = s->seg_off[k], seg_en = s->seg_off[k] + s->n_seg[k];
|
||||
for (i = seg_st; i < seg_en; ++i) {
|
||||
mm_bseq1_t *t = &s->seq[i];
|
||||
for (j = 0; j < s->n_reg[i]; ++j) {
|
||||
mm_reg1_t *r = &s->reg[i][j];
|
||||
assert(!r->sam_pri || r->id == r->parent);
|
||||
if ((p->opt->flag & MM_F_NO_PRINT_2ND) && r->id != r->parent)
|
||||
continue;
|
||||
if (p->opt->split_prefix && p->n_parts == 0) { // then write to temporary files
|
||||
mm_err_fwrite(&s->n_reg[i], sizeof(int), 1, p->fp_split);
|
||||
mm_err_fwrite(&s->rep_len[i], sizeof(int), 1, p->fp_split);
|
||||
mm_err_fwrite(&s->frag_gap[i], sizeof(int), 1, p->fp_split);
|
||||
for (j = 0; j < s->n_reg[i]; ++j) {
|
||||
mm_reg1_t *r = &s->reg[i][j];
|
||||
mm_err_fwrite(r, sizeof(mm_reg1_t), 1, p->fp_split);
|
||||
if (p->opt->flag & MM_F_CIGAR) {
|
||||
mm_err_fwrite(&r->p->capacity, 4, 1, p->fp_split);
|
||||
mm_err_fwrite(r->p, r->p->capacity, 4, p->fp_split);
|
||||
}
|
||||
}
|
||||
} else if (s->n_reg[i] > 0) { // the query has at least one hit
|
||||
for (j = 0; j < s->n_reg[i]; ++j) {
|
||||
mm_reg1_t *r = &s->reg[i][j];
|
||||
assert(!r->sam_pri || r->id == r->parent);
|
||||
if ((p->opt->flag & MM_F_NO_PRINT_2ND) && r->id != r->parent)
|
||||
continue;
|
||||
if (p->opt->flag & MM_F_OUT_SAM)
|
||||
mm_write_sam3(&p->str, mi, t, i - seg_st, j, s->n_seg[k], &s->n_reg[seg_st], (const mm_reg1_t*const*)&s->reg[seg_st], km, p->opt->flag, s->rep_len[i]);
|
||||
else
|
||||
mm_write_paf3(&p->str, mi, t, r, km, p->opt->flag, s->rep_len[i]);
|
||||
mm_err_puts(p->str.s);
|
||||
}
|
||||
} else if ((p->opt->flag & MM_F_PAF_NO_HIT) || ((p->opt->flag & MM_F_OUT_SAM) && !(p->opt->flag & MM_F_SAM_HIT_ONLY))) { // output an empty hit, if requested
|
||||
if (p->opt->flag & MM_F_OUT_SAM)
|
||||
mm_write_sam2(&p->str, mi, t, i - seg_st, j, s->n_seg[k], &s->n_reg[seg_st], (const mm_reg1_t*const*)&s->reg[seg_st], km, p->opt->flag);
|
||||
mm_write_sam3(&p->str, mi, t, i - seg_st, -1, s->n_seg[k], &s->n_reg[seg_st], (const mm_reg1_t*const*)&s->reg[seg_st], km, p->opt->flag, s->rep_len[i]);
|
||||
else
|
||||
mm_write_paf(&p->str, mi, t, r, km, p->opt->flag);
|
||||
puts(p->str.s);
|
||||
}
|
||||
if (s->n_reg[i] == 0 && (p->opt->flag & MM_F_OUT_SAM)) {
|
||||
mm_write_sam2(&p->str, mi, t, i - seg_st, -1, s->n_seg[k], &s->n_reg[seg_st], (const mm_reg1_t*const*)&s->reg[seg_st], km, p->opt->flag);
|
||||
puts(p->str.s);
|
||||
mm_write_paf3(&p->str, mi, t, 0, 0, p->opt->flag, s->rep_len[i]);
|
||||
mm_err_puts(p->str.s);
|
||||
}
|
||||
}
|
||||
for (i = seg_st; i < seg_en; ++i) {
|
||||
@@ -501,9 +602,10 @@ static void *worker_pipeline(void *shared, int step, void *in)
|
||||
free(s->reg[i]);
|
||||
free(s->seq[i].seq); free(s->seq[i].name);
|
||||
if (s->seq[i].qual) free(s->seq[i].qual);
|
||||
if (s->seq[i].comment) free(s->seq[i].comment);
|
||||
}
|
||||
}
|
||||
free(s->reg); free(s->n_reg); free(s->seq); // seg_off and n_seg were allocated with reg; no memory leak here
|
||||
free(s->reg); free(s->n_reg); free(s->seq); // seg_off, n_seg, rep_len and frag_gap were allocated with reg; no memory leak here
|
||||
km_destroy(km);
|
||||
if (mm_verbose >= 3)
|
||||
fprintf(stderr, "[M::%s::%.3f*%.2f] mapped %d sequences\n", __func__, realtime() - mm_realtime0, cputime() / (realtime() - mm_realtime0), s->n_seq);
|
||||
@@ -512,32 +614,44 @@ static void *worker_pipeline(void *shared, int step, void *in)
|
||||
return 0;
|
||||
}
|
||||
|
||||
static mm_bseq_file_t **open_bseqs(int n, const char **fn)
|
||||
{
|
||||
mm_bseq_file_t **fp;
|
||||
int i, j;
|
||||
fp = (mm_bseq_file_t**)calloc(n, sizeof(mm_bseq_file_t*));
|
||||
for (i = 0; i < n; ++i) {
|
||||
if ((fp[i] = mm_bseq_open(fn[i])) == 0) {
|
||||
if (mm_verbose >= 1)
|
||||
fprintf(stderr, "ERROR: failed to open file '%s': %s\n", fn[i], strerror(errno));
|
||||
for (j = 0; j < i; ++j)
|
||||
mm_bseq_close(fp[j]);
|
||||
free(fp);
|
||||
return 0;
|
||||
}
|
||||
}
|
||||
return fp;
|
||||
}
|
||||
|
||||
int mm_map_file_frag(const mm_idx_t *idx, int n_segs, const char **fn, const mm_mapopt_t *opt, int n_threads)
|
||||
{
|
||||
int i, j, pl_threads;
|
||||
int i, pl_threads;
|
||||
pipeline_t pl;
|
||||
if (n_segs < 1) return -1;
|
||||
memset(&pl, 0, sizeof(pipeline_t));
|
||||
pl.n_fp = n_segs;
|
||||
pl.fp = (mm_bseq_file_t**)calloc(n_segs, sizeof(mm_bseq_file_t*));
|
||||
for (i = 0; i < n_segs; ++i) {
|
||||
pl.fp[i] = mm_bseq_open(fn[i]);
|
||||
if (pl.fp[i] == 0) {
|
||||
if (mm_verbose >= 1)
|
||||
fprintf(stderr, "ERROR: failed to open file '%s'\n", fn[i]);
|
||||
for (j = 0; j < i; ++j)
|
||||
mm_bseq_close(pl.fp[j]);
|
||||
free(pl.fp);
|
||||
return -1;
|
||||
}
|
||||
}
|
||||
pl.fp = open_bseqs(pl.n_fp, fn);
|
||||
if (pl.fp == 0) return -1;
|
||||
pl.opt = opt, pl.mi = idx;
|
||||
pl.n_threads = n_threads > 1? n_threads : 1;
|
||||
pl.mini_batch_size = opt->mini_batch_size;
|
||||
if (opt->split_prefix)
|
||||
pl.fp_split = mm_split_init(opt->split_prefix, idx);
|
||||
pl_threads = n_threads == 1? 1 : (opt->flag&MM_F_2_IO_THREADS)? 3 : 2;
|
||||
kt_pipeline(pl_threads, worker_pipeline, &pl, 3);
|
||||
|
||||
free(pl.str.s);
|
||||
for (i = 0; i < n_segs; ++i)
|
||||
if (pl.fp_split) fclose(pl.fp_split);
|
||||
for (i = 0; i < pl.n_fp; ++i)
|
||||
mm_bseq_close(pl.fp[i]);
|
||||
free(pl.fp);
|
||||
return 0;
|
||||
@@ -547,3 +661,48 @@ int mm_map_file(const mm_idx_t *idx, const char *fn, const mm_mapopt_t *opt, int
|
||||
{
|
||||
return mm_map_file_frag(idx, 1, &fn, opt, n_threads);
|
||||
}
|
||||
|
||||
int mm_split_merge(int n_segs, const char **fn, const mm_mapopt_t *opt, int n_split_idx)
|
||||
{
|
||||
int i;
|
||||
pipeline_t pl;
|
||||
mm_idx_t *mi;
|
||||
if (n_segs < 1 || n_split_idx < 1) return -1;
|
||||
memset(&pl, 0, sizeof(pipeline_t));
|
||||
pl.n_fp = n_segs;
|
||||
pl.fp = open_bseqs(pl.n_fp, fn);
|
||||
if (pl.fp == 0) return -1;
|
||||
pl.opt = opt;
|
||||
pl.mini_batch_size = opt->mini_batch_size;
|
||||
|
||||
pl.n_parts = n_split_idx;
|
||||
pl.fp_parts = CALLOC(FILE*, pl.n_parts);
|
||||
pl.rid_shift = CALLOC(uint32_t, pl.n_parts);
|
||||
pl.mi = mi = mm_split_merge_prep(opt->split_prefix, n_split_idx, pl.fp_parts, pl.rid_shift);
|
||||
if (pl.mi == 0) {
|
||||
free(pl.fp_parts);
|
||||
free(pl.rid_shift);
|
||||
return -1;
|
||||
}
|
||||
for (i = n_split_idx - 1; i > 0; --i)
|
||||
pl.rid_shift[i] = pl.rid_shift[i - 1];
|
||||
for (pl.rid_shift[0] = 0, i = 1; i < n_split_idx; ++i)
|
||||
pl.rid_shift[i] += pl.rid_shift[i - 1];
|
||||
if (opt->flag & MM_F_OUT_SAM)
|
||||
for (i = 0; i < (int32_t)pl.mi->n_seq; ++i)
|
||||
printf("@SQ\tSN:%s\tLN:%d\n", pl.mi->seq[i].name, pl.mi->seq[i].len);
|
||||
|
||||
kt_pipeline(2, worker_pipeline, &pl, 3);
|
||||
|
||||
free(pl.str.s);
|
||||
mm_idx_destroy(mi);
|
||||
free(pl.rid_shift);
|
||||
for (i = 0; i < n_split_idx; ++i)
|
||||
fclose(pl.fp_parts[i]);
|
||||
free(pl.fp_parts);
|
||||
for (i = 0; i < pl.n_fp; ++i)
|
||||
mm_bseq_close(pl.fp[i]);
|
||||
free(pl.fp);
|
||||
mm_split_rm_tmp(opt->split_prefix, n_split_idx);
|
||||
return 0;
|
||||
}
|
||||
|
||||
@@ -29,6 +29,16 @@
|
||||
#define MM_F_REV_ONLY 0x200000
|
||||
#define MM_F_HEAP_SORT 0x400000
|
||||
#define MM_F_ALL_CHAINS 0x800000
|
||||
#define MM_F_OUT_MD 0x1000000
|
||||
#define MM_F_COPY_COMMENT 0x2000000
|
||||
#define MM_F_EQX 0x4000000 // use =/X instead of M
|
||||
#define MM_F_PAF_NO_HIT 0x8000000 // output unmapped reads to PAF
|
||||
#define MM_F_NO_END_FLT 0x10000000
|
||||
#define MM_F_HARD_MLEVEL 0x20000000
|
||||
#define MM_F_SAM_HIT_ONLY 0x40000000
|
||||
#define MM_F_RMQ (0x80000000LL)
|
||||
#define MM_F_QSTRAND (0x100000000LL)
|
||||
#define MM_F_NO_INV (0x200000000LL)
|
||||
|
||||
#define MM_I_HPC 0x1
|
||||
#define MM_I_NO_SEQ 0x2
|
||||
@@ -38,6 +48,18 @@
|
||||
|
||||
#define MM_MAX_SEG 255
|
||||
|
||||
#define MM_CIGAR_MATCH 0
|
||||
#define MM_CIGAR_INS 1
|
||||
#define MM_CIGAR_DEL 2
|
||||
#define MM_CIGAR_N_SKIP 3
|
||||
#define MM_CIGAR_SOFTCLIP 4
|
||||
#define MM_CIGAR_HARDCLIP 5
|
||||
#define MM_CIGAR_PADDING 6
|
||||
#define MM_CIGAR_EQ_MATCH 7
|
||||
#define MM_CIGAR_X_MISMATCH 8
|
||||
|
||||
#define MM_CIGAR_STR "MIDNSHP=XB"
|
||||
|
||||
#ifdef __cplusplus
|
||||
extern "C" {
|
||||
#endif
|
||||
@@ -51,15 +73,19 @@ typedef struct {
|
||||
char *name; // name of the db sequence
|
||||
uint64_t offset; // offset in mm_idx_t::S
|
||||
uint32_t len; // length
|
||||
uint32_t is_alt;
|
||||
} mm_idx_seq_t;
|
||||
|
||||
typedef struct {
|
||||
int32_t b, w, k, flag;
|
||||
uint32_t n_seq; // number of reference sequences
|
||||
int32_t index;
|
||||
int32_t n_alt;
|
||||
mm_idx_seq_t *seq; // sequence name, length and offset
|
||||
uint32_t *S; // 4-bit packed sequence
|
||||
struct mm_idx_bucket_s *B; // index (hidden)
|
||||
void *km;
|
||||
struct mm_idx_intv_s *I; // intervals (hidden)
|
||||
void *km, *h;
|
||||
} mm_idx_t;
|
||||
|
||||
// minimap2 alignment
|
||||
@@ -82,7 +108,7 @@ typedef struct {
|
||||
int32_t mlen, blen; // seeded exact match length; seeded alignment block length
|
||||
int32_t n_sub; // number of suboptimal mappings
|
||||
int32_t score0; // initial chaining score (before chain merging/spliting)
|
||||
uint32_t mapq:8, split:2, rev:1, inv:1, sam_pri:1, proper_frag:1, pe_thru:1, seg_split:1, seg_id:8, dummy:8;
|
||||
uint32_t mapq:8, split:2, rev:1, inv:1, sam_pri:1, proper_frag:1, pe_thru:1, seg_split:1, seg_id:8, split_inv:1, is_alt:1, dummy:6;
|
||||
uint32_t hash;
|
||||
float div;
|
||||
mm_extra_t *p;
|
||||
@@ -91,43 +117,60 @@ typedef struct {
|
||||
// indexing and mapping options
|
||||
typedef struct {
|
||||
short k, w, flag, bucket_bits;
|
||||
int mini_batch_size;
|
||||
int64_t mini_batch_size;
|
||||
uint64_t batch_size;
|
||||
} mm_idxopt_t;
|
||||
|
||||
typedef struct {
|
||||
int64_t flag; // see MM_F_* macros
|
||||
int seed;
|
||||
int sdust_thres; // score threshold for SDUST; 0 to disable
|
||||
int flag; // see MM_F_* macros
|
||||
|
||||
int bw; // bandwidth
|
||||
int max_qlen; // max query length
|
||||
|
||||
int bw, bw_long; // bandwidth
|
||||
int max_gap, max_gap_ref; // break a chain if there are no minimizers in a max_gap window
|
||||
int max_frag_len;
|
||||
int max_chain_skip;
|
||||
int max_chain_skip, max_chain_iter;
|
||||
int min_cnt; // min number of minimizers on each chain
|
||||
int min_chain_score; // min chaining score
|
||||
float chain_gap_scale;
|
||||
int rmq_size_cap, rmq_inner_dist;
|
||||
int rmq_rescue_size;
|
||||
float rmq_rescue_ratio;
|
||||
|
||||
float mask_level;
|
||||
int mask_len;
|
||||
float pri_ratio;
|
||||
int best_n; // top best_n chains are subjected to DP alignment
|
||||
|
||||
int max_join_long, max_join_short;
|
||||
int min_join_flank_sc;
|
||||
float alt_drop;
|
||||
|
||||
int a, b, q, e, q2, e2; // matching score, mismatch, gap-open and gap-ext penalties
|
||||
int sc_ambi; // score when one or both bases are "N"
|
||||
int noncan; // cost of non-canonical splicing sites
|
||||
int zdrop; // break alignment if alignment score drops too fast along the diagonal
|
||||
int junc_bonus;
|
||||
int zdrop, zdrop_inv; // break alignment if alignment score drops too fast along the diagonal
|
||||
int end_bonus;
|
||||
int min_dp_max; // drop an alignment if the score of the max scoring segment is below this threshold
|
||||
int min_ksw_len;
|
||||
int anchor_ext_len, anchor_ext_shift;
|
||||
float max_clip_ratio; // drop an alignment if BOTH ends are clipped above this ratio
|
||||
|
||||
int rank_min_len;
|
||||
float rank_frac;
|
||||
|
||||
int pe_ori, pe_bonus;
|
||||
|
||||
float mid_occ_frac; // only used by mm_mapopt_update(); see below
|
||||
int32_t min_mid_occ, max_mid_occ;
|
||||
int32_t mid_occ; // ignore seeds with occurrences above this threshold
|
||||
int32_t max_occ;
|
||||
int mini_batch_size; // size of a batch of query bases to process in parallel
|
||||
int32_t max_occ, max_max_occ, occ_dist;
|
||||
int64_t mini_batch_size; // size of a batch of query bases to process in parallel
|
||||
int64_t max_sw_mat;
|
||||
int64_t cap_kalloc;
|
||||
|
||||
const char *split_prefix;
|
||||
} mm_mapopt_t;
|
||||
|
||||
// index reader
|
||||
@@ -212,6 +255,36 @@ void mm_idx_reader_close(mm_idx_reader_t *r);
|
||||
|
||||
int mm_idx_reader_eof(const mm_idx_reader_t *r);
|
||||
|
||||
/**
|
||||
* Check whether the file contains a minimap2 index
|
||||
*
|
||||
* @param fn file name
|
||||
*
|
||||
* @return the file size if fn is an index file; 0 if fn is not.
|
||||
*/
|
||||
int64_t mm_idx_is_idx(const char *fn);
|
||||
|
||||
/**
|
||||
* Load a part of an index
|
||||
*
|
||||
* Given a uni-part index, this function loads the entire index into memory.
|
||||
* Given a multi-part index, it loads one part only and places the file pointer
|
||||
* at the end of that part.
|
||||
*
|
||||
* @param fp pointer to FILE object
|
||||
*
|
||||
* @return minimap2 index read from fp
|
||||
*/
|
||||
mm_idx_t *mm_idx_load(FILE *fp);
|
||||
|
||||
/**
|
||||
* Append an index (or one part of a full index) to file
|
||||
*
|
||||
* @param fp pointer to FILE object
|
||||
* @param mi minimap2 index
|
||||
*/
|
||||
void mm_idx_dump(FILE *fp, const mm_idx_t *mi);
|
||||
|
||||
/**
|
||||
* Create an index from strings in memory
|
||||
*
|
||||
@@ -260,6 +333,8 @@ mm_tbuf_t *mm_tbuf_init(void);
|
||||
*/
|
||||
void mm_tbuf_destroy(mm_tbuf_t *b);
|
||||
|
||||
void *mm_tbuf_get_km(mm_tbuf_t *b);
|
||||
|
||||
/**
|
||||
* Align a query sequence against an index
|
||||
*
|
||||
@@ -296,6 +371,31 @@ int mm_map_file(const mm_idx_t *idx, const char *fn, const mm_mapopt_t *opt, int
|
||||
|
||||
int mm_map_file_frag(const mm_idx_t *idx, int n_segs, const char **fn, const mm_mapopt_t *opt, int n_threads);
|
||||
|
||||
/**
|
||||
* Generate the cs tag (new in 2.12)
|
||||
*
|
||||
* @param km memory blocks; set to NULL if unsure
|
||||
* @param buf buffer to write the cs/MD tag; typicall NULL on the first call
|
||||
* @param max_len max length of the buffer; typically set to 0 on the first call
|
||||
* @param mi index
|
||||
* @param r alignment
|
||||
* @param seq query sequence
|
||||
* @param no_iden true to use : instead of =
|
||||
*
|
||||
* @return the length of cs
|
||||
*/
|
||||
int mm_gen_cs(void *km, char **buf, int *max_len, const mm_idx_t *mi, const mm_reg1_t *r, const char *seq, int no_iden);
|
||||
int mm_gen_MD(void *km, char **buf, int *max_len, const mm_idx_t *mi, const mm_reg1_t *r, const char *seq);
|
||||
|
||||
// query sequence name and sequence in the minimap2 index
|
||||
int mm_idx_index_name(mm_idx_t *mi);
|
||||
int mm_idx_name2id(const mm_idx_t *mi, const char *name);
|
||||
int mm_idx_getseq(const mm_idx_t *mi, uint32_t rid, uint32_t st, uint32_t en, uint8_t *seq);
|
||||
|
||||
int mm_idx_alt_read(mm_idx_t *mi, const char *fn);
|
||||
int mm_idx_bed_read(mm_idx_t *mi, const char *fn, int read_junc);
|
||||
int mm_idx_bed_junc(const mm_idx_t *mi, int32_t ctg, int32_t st, int32_t en, uint8_t *s);
|
||||
|
||||
// deprecated APIs for backward compatibility
|
||||
void mm_mapopt_init(mm_mapopt_t *opt);
|
||||
mm_idx_t *mm_idx_build(const char *fn, int w, int k, int flag, int n_threads);
|
||||
|
||||
+197
-51
@@ -1,4 +1,4 @@
|
||||
.TH minimap2 1 "1 February 2018" "minimap2-2.8 (r672)" "Bioinformatics tools"
|
||||
.TH minimap2 1 "6 July 2021" "minimap2-2.21 (r1071)" "Bioinformatics tools"
|
||||
.SH NAME
|
||||
.PP
|
||||
minimap2 - mapping and alignment between collections of DNA sequences
|
||||
@@ -121,21 +121,56 @@ provided as the target sequences, options
|
||||
.BR -w ,
|
||||
.B -I
|
||||
will be effectively overridden by the options stored in the index file.
|
||||
.TP
|
||||
.BI --alt \ FILE
|
||||
List of ALT contigs [null]
|
||||
.TP
|
||||
.BI --alt-drop \ FLOAT
|
||||
Drop ALT hits by
|
||||
.I FLOAT
|
||||
fraction when ranking and computing mapping quality [0.15]
|
||||
.SS Mapping options
|
||||
.TP 10
|
||||
.BI -f \ FLOAT
|
||||
Ignore top
|
||||
.BI -f \ FLOAT | INT1 [, INT2 ]
|
||||
If fraction, ignore top
|
||||
.I FLOAT
|
||||
fraction of most frequent minimizers [0.0002]
|
||||
fraction of most frequent minimizers [0.0002]. If integer,
|
||||
ignore minimizers occuring more than
|
||||
.I INT1
|
||||
times.
|
||||
.I INT2
|
||||
is only effective in the
|
||||
.B --sr
|
||||
or
|
||||
.B -xsr
|
||||
mode, which sets the threshold for a second round of seeding.
|
||||
.TP
|
||||
.BI -g \ INT
|
||||
.BI -U \ INT1 [, INT2 ]
|
||||
Lower and upper bounds of k-mer occurrences [10,1000000]. The final k-mer occurrence threshold is
|
||||
.RI max{ INT1 ,\ min{ INT2 ,
|
||||
.BR -f }}.
|
||||
This option prevents excessively small or large
|
||||
.B -f
|
||||
estimated from the input reference. It deprecates
|
||||
.B --min-occ-floor
|
||||
in earlier versions of minimap2.
|
||||
.TP
|
||||
.BI -e \ INT
|
||||
Sample a high-frequency minimizer every
|
||||
.I INT
|
||||
basepairs [500].
|
||||
.TP
|
||||
.BI -g \ NUM
|
||||
Stop chain enlongation if there are no minimizers within
|
||||
.IR INT -bp
|
||||
[10000].
|
||||
.IR NUM -bp
|
||||
[10k].
|
||||
.TP
|
||||
.BI -r \ INT
|
||||
Bandwidth used in chaining and DP-based alignment [500]. This option
|
||||
approximately controls the maximum gap size.
|
||||
.BI -r \ NUM1 [, NUM2 ]
|
||||
Bandwidth for chaining and base alignment [500,20k].
|
||||
.I NUM1
|
||||
is used for initial chaining and alignment extension;
|
||||
.I NUM2
|
||||
for RMQ-based re-chaining and closing gaps in alignments.
|
||||
.TP
|
||||
.BI -n \ INT
|
||||
Discard chains consisting of
|
||||
@@ -161,8 +196,10 @@ and
|
||||
have no effect when this option is in use.
|
||||
.TP
|
||||
.BR --dual = yes | no
|
||||
During chaining, whether to skip pairs wherein the query name is
|
||||
lexicographically greater than the target name [yes]
|
||||
If
|
||||
.BR no ,
|
||||
skip query-target pairs wherein the query name is lexicographically greater
|
||||
than the target name [yes]
|
||||
.TP
|
||||
.B -X
|
||||
Equivalent to
|
||||
@@ -174,7 +211,7 @@ Primarily used for all-vs-all read overlapping.
|
||||
.BI -p \ FLOAT
|
||||
Minimal secondary-to-primary score ratio to output secondary mappings [0.8].
|
||||
Between two chains overlaping over half of the shorter chain (controlled by
|
||||
.BR --mask-level ),
|
||||
.BR -M ),
|
||||
the chain with a lower score is secondary to the chain with a higher score.
|
||||
If the ratio of the scores is below
|
||||
.IR FLOAT ,
|
||||
@@ -207,14 +244,39 @@ Mark as secondary a chain that overlaps with a better chain by
|
||||
.I FLOAT
|
||||
or more of the shorter chain [0.5]
|
||||
.TP
|
||||
.BR --rmq = no | yes
|
||||
Use the minigraph chaining algorithm [no]. The minigraph algorithm is better
|
||||
for aligning contigs through long INDELs.
|
||||
.TP
|
||||
.B --hard-mask-level
|
||||
Honor option
|
||||
.B -M
|
||||
and disable a heurstic to save unmapped subsequences and disables
|
||||
.BR --mask-len .
|
||||
.TP
|
||||
.BI --mask-len \ NUM
|
||||
Keep an alignment if dropping it leaves an unaligned region on query longer than
|
||||
.IR INT
|
||||
[inf]. Effective without
|
||||
.BR --hard-mask-level .
|
||||
.TP
|
||||
.BI --max-chain-skip \ INT
|
||||
A heuristics that stops chaining early [50]. Minimap2 uses dynamic programming
|
||||
A heuristics that stops chaining early [25]. Minimap2 uses dynamic programming
|
||||
for chaining. The time complexity is quadratic in the number of seeds. This
|
||||
option makes minimap2 exits the inner loop if it repeatedly sees seeds already
|
||||
on chains. Set
|
||||
.I INT
|
||||
to a large number to switch off this heurstics.
|
||||
.TP
|
||||
.BI --max-chain-iter \ INT
|
||||
Check up to
|
||||
.I INT
|
||||
partial chains during chaining [5000]. This is a heuristic to avoid quadratic
|
||||
time complexity in the worst case.
|
||||
.TP
|
||||
.BI --chain-gap-scale \ FLOAT
|
||||
Scale of gap cost during chaining [1.0]
|
||||
.TP
|
||||
.B --no-long-join
|
||||
Disable the long gap patching heuristic. When this option is applied, the
|
||||
maximum alignment gap is mostly controlled by
|
||||
@@ -229,6 +291,9 @@ applies a second round of chaining with a higher minimizer occurrence threshold
|
||||
if no good chain is found. In addition, minimap2 attempts to patch gaps between
|
||||
seeds with ungapped alignment.
|
||||
.TP
|
||||
.BI --split-prefix \ STR
|
||||
Prefix to create temporary files. Typically used for a multi-part index.
|
||||
.TP
|
||||
.BR --frag = no | yes
|
||||
Whether to enable the fragment mode [no]
|
||||
.TP
|
||||
@@ -243,6 +308,10 @@ Only map to the reverse complement strand of the reference sequences.
|
||||
.BR --heap-sort = no | yes
|
||||
If yes, sort anchors with heap merge, instead of radix sort. Heap merge is
|
||||
faster for short reads, but slower for long reads. [no]
|
||||
.TP
|
||||
.B --no-pairing
|
||||
Treat two reads in a pair as independent reads. The mate related fields in SAM
|
||||
are still properly populated.
|
||||
.SS Alignment options
|
||||
.TP 10
|
||||
.BI -A \ INT
|
||||
@@ -269,11 +338,22 @@ Cost for a non-canonical GT-AG splicing (effective with
|
||||
.BR --splice )
|
||||
[0]
|
||||
.TP
|
||||
.BI -z \ INT
|
||||
Break an alignment if the running score drops too quickly along the diagonal of
|
||||
the DP matrix (diagonal X-drop, or Z-drop) [400]. Increasing the value improves
|
||||
the contiguity of the alignment at the cost of poor alignment in the middle
|
||||
(e.g. caused by a long inversion).
|
||||
.BI -z \ INT1[,INT2]
|
||||
Truncate an alignment if the running alignment score drops too quickly along
|
||||
the diagonal of the DP matrix (diagonal X-drop, or Z-drop) [400,200]. If the
|
||||
drop of score is above
|
||||
.IR INT2 ,
|
||||
minimap2 will reverse complement the query in the related region and align
|
||||
again to test small inversions. Minimap2 truncates alignment if there is an
|
||||
inversion or the drop of score is greater than
|
||||
.IR INT1 .
|
||||
Decrease
|
||||
.I INT2
|
||||
to find small inversions at the cost of performance and false positives.
|
||||
Increase
|
||||
.I INT1
|
||||
to improves the contiguity of alignment at the cost of poor alignment in the
|
||||
middle.
|
||||
.TP
|
||||
.BI -s \ INT
|
||||
Minimal peak DP alignment score to output [40]. The peak score is computed from
|
||||
@@ -292,6 +372,9 @@ no attempt to match GT-AG [n]
|
||||
.BI --end-bonus \ INT
|
||||
Score bonus when alignment extends to the end of the query sequence [0].
|
||||
.TP
|
||||
.BI --score-N \ INT
|
||||
Score of a mismatch involving ambiguous bases [1].
|
||||
.TP
|
||||
.BR --splice-flank = yes | no
|
||||
Assume the next base to a
|
||||
.B GT
|
||||
@@ -309,6 +392,17 @@ on SIRV data, please add
|
||||
.B --splice-flank=no
|
||||
to the command line.
|
||||
.TP
|
||||
.BR --junc-bed \ FILE
|
||||
Gene annotations in the BED12 format (aka 12-column BED), or intron positions
|
||||
in 5-column BED. With this option, minimap2 prefers splicing in annotations.
|
||||
BED12 file can be converted from GTF/GFF3 with `paftools.js gff2bed anno.gtf'
|
||||
[].
|
||||
.TP
|
||||
.BR --junc-bonus \ INT
|
||||
Score bonus for a splice donor or acceptor found in annotation (effective with
|
||||
.BR --junc-bed )
|
||||
[9].
|
||||
.TP
|
||||
.BI --end-seed-pen \ INT
|
||||
Drop a terminal anchor if
|
||||
.IR s <log( g )+ INT ,
|
||||
@@ -320,12 +414,26 @@ the length of the terminal gap in the chain. This option is only effective
|
||||
with
|
||||
.BR --splice .
|
||||
It helps to avoid tiny terminal exons. [6]
|
||||
.TP
|
||||
.B --no-end-flt
|
||||
Don't filter seeds towards the ends of chains before performing base-level
|
||||
alignment.
|
||||
.TP
|
||||
.BI --cap-sw-mem \ NUM
|
||||
Skip alignment if the DP matrix size is above
|
||||
.IR NUM .
|
||||
Set 0 to disable [100m].
|
||||
.SS Input/output options
|
||||
.TP 10
|
||||
.B -a
|
||||
Generate CIGAR and output alignments in the SAM format. Minimap2 outputs in PAF
|
||||
by default.
|
||||
.TP
|
||||
.BI -o \ FILE
|
||||
Output alignments to
|
||||
.I FILE
|
||||
[stdout].
|
||||
.TP
|
||||
.B -Q
|
||||
Ignore base quality in the input file.
|
||||
.TP
|
||||
@@ -340,6 +448,9 @@ SAM read group line in a format like
|
||||
.B @RG\\\\tID:foo\\\\tSM:bar
|
||||
[].
|
||||
.TP
|
||||
.B -y
|
||||
Copy input FASTA/Q comments to output.
|
||||
.TP
|
||||
.B -c
|
||||
Generate CIGAR. In PAF, the CIGAR is written to the `cg' custom tag.
|
||||
.TP
|
||||
@@ -358,6 +469,12 @@ is given,
|
||||
.I short
|
||||
is assumed. [none]
|
||||
.TP
|
||||
.B --MD
|
||||
Output the MD tag (see the SAM spec).
|
||||
.TP
|
||||
.B --eqx
|
||||
Output =/X CIGAR operators for sequence match/mismatch.
|
||||
.TP
|
||||
.B -Y
|
||||
In SAM output, use soft clipping for supplementary alignments.
|
||||
.TP
|
||||
@@ -391,6 +508,18 @@ memory.
|
||||
.BR --secondary = yes | no
|
||||
Whether to output secondary alignments [yes]
|
||||
.TP
|
||||
.BI --max-qlen \ NUM
|
||||
Filter out query sequences longer than
|
||||
.IR NUM .
|
||||
.TP
|
||||
.B --paf-no-hit
|
||||
In PAF, output unmapped queries; the strand and the reference name fields are
|
||||
set to `*'. Warning: some paftools.js commands may not work with such output
|
||||
for the moment.
|
||||
.TP
|
||||
.B --sam-hit-only
|
||||
In SAM, don't output unmapped reads.
|
||||
.TP
|
||||
.B --version
|
||||
Print version number to stdout
|
||||
.SS Preset options
|
||||
@@ -404,53 +533,47 @@ Available
|
||||
.I STR
|
||||
are:
|
||||
.RS
|
||||
.TP 8
|
||||
.B map-pb
|
||||
PacBio/Oxford Nanopore read to reference mapping
|
||||
.RB ( -Hk19 )
|
||||
.TP
|
||||
.TP 10
|
||||
.B map-ont
|
||||
Slightly more sensitive for Oxford Nanopore to reference mapping
|
||||
.RB ( -k15 ).
|
||||
For PacBio reads, HPC minimizers consistently leads to faster performance and
|
||||
more sensitive results in comparison to normal minimizers. For Oxford Nanopore
|
||||
data, normal minimizers are better, though not much. The effectiveness of HPC
|
||||
is determined by the sequencing error mode.
|
||||
Align noisy long reads of ~10% error rate to a reference genome. This is the
|
||||
default mode.
|
||||
.TP
|
||||
.B map-hifi
|
||||
Align PacBio high-fidelity (HiFi) reads to a reference genome
|
||||
.RB ( -k19
|
||||
.B -w19 -U50,500 -g10k -A1 -B4 -O6,26 -E2,1
|
||||
.BR -s200 ).
|
||||
.TP
|
||||
.B map-pb
|
||||
Align older PacBio continuous long (CLR) reads to a reference genome
|
||||
.RB ( -Hk19 ).
|
||||
.TP
|
||||
.B asm5
|
||||
Long assembly to reference mapping
|
||||
.RB ( -k19
|
||||
.B -w19 -A1 -B19 -O39,81 -E3,1 -s200
|
||||
.BR -z200 ).
|
||||
.B -w19 -U50,500 --rmq -r100k -g10k -A1 -B19 -O39,81 -E3,1 -s200 -z200
|
||||
.BR -N50 ).
|
||||
Typically, the alignment will not extend to regions with 5% or higher sequence
|
||||
divergence. Only use this preset if the average divergence is far below 5%.
|
||||
.TP
|
||||
.B asm10
|
||||
Long assembly to reference mapping
|
||||
.RB ( -k19
|
||||
.B -w19 -A1 -B9 -O16,41 -E2,1 -s200
|
||||
.BR -z200 ).
|
||||
.B -w19 -U50,500 --rmq -r100k -g10k -A1 -B9 -O16,41 -E2,1 -s200 -z200
|
||||
.BR -N50 ).
|
||||
Up to 10% sequence divergence.
|
||||
.TP
|
||||
.B ava-pb
|
||||
PacBio all-vs-all overlap mapping
|
||||
.RB ( -Hk19
|
||||
.B -Xw5 -m100 -g10000 --max-chain-skip
|
||||
.BR 25 ).
|
||||
.TP
|
||||
.B ava-ont
|
||||
Oxford Nanopore all-vs-all overlap mapping
|
||||
.RB ( -k15
|
||||
.B -Xw5 -m100 -g10000 --max-chain-skip
|
||||
.BR 25 ).
|
||||
Similarly, the major difference from
|
||||
.B ava-pb
|
||||
is that this preset is not using HPC minimizers.
|
||||
.B asm20
|
||||
Long assembly to reference mapping
|
||||
.RB ( -k19
|
||||
.B -w10 -U50,500 --rmq -r100k -g10k -A1 -B4 -O6,26 -E2,1 -s200 -z200
|
||||
.BR -N50 ).
|
||||
Up to 20% sequence divergence.
|
||||
.TP
|
||||
.B splice
|
||||
Long-read spliced alignment
|
||||
.RB ( -k15
|
||||
.B -w5 --splice -g2000 -G200k -A1 -B2 -O2,32 -E1,0 -C9 -z200 -ub
|
||||
.B -w5 --splice -g2k -G200k -A1 -B2 -O2,32 -E1,0 -b0 -C9 -z200 -ub --junc-bonus=9 --cap-sw-mem=0
|
||||
.BR --splice-flank=yes ).
|
||||
In the splice mode, 1) long deletions are taken as introns and represented as
|
||||
the
|
||||
@@ -460,12 +583,30 @@ costs are different during chaining; 4) the computation of the
|
||||
.RB ` ms '
|
||||
tag ignores introns to demote hits to pseudogenes.
|
||||
.TP
|
||||
.B splice:hq
|
||||
Long-read splice alignment for PacBio CCS reads
|
||||
.RB ( -xsplice
|
||||
.B -C5 -O6,24
|
||||
.BR -B4 ).
|
||||
.TP
|
||||
.B sr
|
||||
Short single-end reads without splicing
|
||||
.RB ( -k21
|
||||
.B -w11 --sr --frag=yes -A2 -B8 -O12,32 -E2,1 -r50 -p.5 -N20 -f1000,5000 -n2 -m20
|
||||
.B -s40 -g200 -2K50m --heap-sort=yes
|
||||
.B -w11 --sr --frag=yes -A2 -B8 -O12,32 -E2,1 -b0 -r100 -p.5 -N20 -f1000,5000 -n2 -m20
|
||||
.B -s40 -g100 -2K50m --heap-sort=yes
|
||||
.BR --secondary=no ).
|
||||
.TP
|
||||
.B ava-pb
|
||||
PacBio CLR all-vs-all overlap mapping
|
||||
.RB ( -Hk19
|
||||
.B -Xw5 -e0
|
||||
.BR -m100 ).
|
||||
.TP
|
||||
.B ava-ont
|
||||
Oxford Nanopore all-vs-all overlap mapping
|
||||
.RB ( -k15
|
||||
.B -Xw5 -e0 -m100
|
||||
.BR -r2k ).
|
||||
.RE
|
||||
.SS Miscellaneous options
|
||||
.TP 10
|
||||
@@ -522,12 +663,17 @@ cm i Number of minimizers on the chain
|
||||
s1 i Chaining score
|
||||
s2 i Chaining score of the best secondary chain
|
||||
NM i Total number of mismatches and gaps in the alignment
|
||||
MD Z To generate the ref sequence in the alignment
|
||||
AS i DP alignment score
|
||||
SA Z List of other supplementary alignments
|
||||
ms i DP score of the max scoring segment in the alignment
|
||||
nn i Number of ambiguous bases in the alignment
|
||||
ts A Transcript strand (splice mode only)
|
||||
cg Z CIGAR string (only in PAF)
|
||||
cs Z Difference string
|
||||
dv f Approximate per-base sequence divergence
|
||||
de f Gap-compressed per-base sequence divergence
|
||||
rl i Length of query regions harboring repetitive seeds
|
||||
.TE
|
||||
|
||||
.PP
|
||||
|
||||
@@ -1,3 +1,4 @@
|
||||
#include <stdlib.h>
|
||||
#include "mmpriv.h"
|
||||
|
||||
int mm_verbose = 1;
|
||||
@@ -86,6 +87,8 @@ double cputime()
|
||||
|
||||
return kernelModeTime + userModeTime;
|
||||
}
|
||||
|
||||
long peakrss(void) { return 0; }
|
||||
#else
|
||||
#include <sys/resource.h>
|
||||
#include <sys/time.h>
|
||||
@@ -96,16 +99,57 @@ double cputime(void)
|
||||
getrusage(RUSAGE_SELF, &r);
|
||||
return r.ru_utime.tv_sec + r.ru_stime.tv_sec + 1e-6 * (r.ru_utime.tv_usec + r.ru_stime.tv_usec);
|
||||
}
|
||||
|
||||
long peakrss(void)
|
||||
{
|
||||
struct rusage r;
|
||||
getrusage(RUSAGE_SELF, &r);
|
||||
#ifdef __linux__
|
||||
return r.ru_maxrss * 1024;
|
||||
#else
|
||||
return r.ru_maxrss;
|
||||
#endif
|
||||
}
|
||||
|
||||
#endif /* WIN32 || _WIN32 */
|
||||
|
||||
double realtime(void)
|
||||
{
|
||||
struct timeval tp;
|
||||
struct timezone tzp;
|
||||
gettimeofday(&tp, &tzp);
|
||||
gettimeofday(&tp, NULL);
|
||||
return tp.tv_sec + tp.tv_usec * 1e-6;
|
||||
}
|
||||
|
||||
void mm_err_puts(const char *str)
|
||||
{
|
||||
int ret;
|
||||
ret = puts(str);
|
||||
if (ret == EOF) {
|
||||
perror("[ERROR] failed to write the results");
|
||||
exit(EXIT_FAILURE);
|
||||
}
|
||||
}
|
||||
|
||||
void mm_err_fwrite(const void *p, size_t size, size_t nitems, FILE *fp)
|
||||
{
|
||||
int ret;
|
||||
ret = fwrite(p, size, nitems, fp);
|
||||
if (ret == EOF) {
|
||||
perror("[ERROR] failed to write data");
|
||||
exit(EXIT_FAILURE);
|
||||
}
|
||||
}
|
||||
|
||||
void mm_err_fread(void *p, size_t size, size_t nitems, FILE *fp)
|
||||
{
|
||||
int ret;
|
||||
ret = fread(p, size, nitems, fp);
|
||||
if (ret == EOF) {
|
||||
perror("[ERROR] failed to read data");
|
||||
exit(EXIT_FAILURE);
|
||||
}
|
||||
}
|
||||
|
||||
#include "ksort.h"
|
||||
|
||||
#define sort_key_128x(a) ((a).x)
|
||||
@@ -115,3 +159,4 @@ KRADIX_SORT_INIT(128x, mm128_t, sort_key_128x, 8)
|
||||
KRADIX_SORT_INIT(64, uint64_t, sort_key_64, 8)
|
||||
|
||||
KSORT_INIT_GENERIC(uint32_t)
|
||||
KSORT_INIT_GENERIC(uint64_t)
|
||||
|
||||
+170
-19
@@ -1,28 +1,179 @@
|
||||
The [K8 Javascript shell][k8] is needed to run Javascripts in this directory.
|
||||
Precompiled k8 binaries for Mac and Linux can be found at the [K8 release
|
||||
page][k8bin].
|
||||
## <a name="started"></a>Getting Started
|
||||
|
||||
* [paf2aln.js](paf2aln.js): convert PAF to [MAF][maf] or BLAST-like output for
|
||||
eyeballing. PAF has to be generated with minimap2 option `-S`, which writes
|
||||
the aligned sequences to the `cs` tag. An example:
|
||||
```sh
|
||||
../minimap2 -S ../test/MT-*.fa | k8 paf2aln.js /dev/stdin
|
||||
```
|
||||
```sh
|
||||
# install minimap2
|
||||
git clone https://github.com/lh3/minimap2
|
||||
cd minimap2 && make
|
||||
# install the k8 javascript shell
|
||||
curl -L https://github.com/attractivechaos/k8/releases/download/v0.2.4/k8-0.2.4.tar.bz2 | tar -jxf -
|
||||
cp k8-0.2.4/k8-`uname -s` k8 # or copy it to a directory on your $PATH
|
||||
# export PATH="$PATH:`pwd`:`pwd`/misc" # run this if k8, minimap2 or paftools.js not on your $PATH
|
||||
minimap2 --cs test/MT-human.fa test/MT-orang.fa | paftools.js view - # view alignment
|
||||
minimap2 -c test/MT-human.fa test/MT-orang.fa | paftools.js stat - # basic alignment statistics
|
||||
minimap2 -c --cs test/MT-human.fa test/MT-orang.fa \
|
||||
| sort -k6,6 -k8,8n | paftools.js call -L15000 - # calling variants from asm-to-ref alignment
|
||||
minimap2 -c test/MT-human.fa test/MT-orang.fa \
|
||||
| paftools.js liftover -l10000 - <(echo -e "MT_orang\t2000\t5000") # liftOver
|
||||
# no test data for the following examples
|
||||
paftools.js junceval -e anno.gtf splice.sam > out.txt # compare splice junctions to annotations
|
||||
paftools.js splice2bed anno.gtf > anno.bed # convert GTF/GFF3 to BED12
|
||||
```
|
||||
|
||||
* [mapstat.js](mapstat.js): output basic statistics such as the number of
|
||||
non-redundant mapped bases, number of split and secondary alignments and
|
||||
number of long gaps. This scripts seamlessly works with both SAM and PAF.
|
||||
## Table of Contents
|
||||
|
||||
* [sim-pbsim.js](sim-pbsim.js): convert reads simulated with [PBSIM][pbsim] to
|
||||
FASTA and encode the true mapping positions to read names in a format like
|
||||
`S1_33!chr1!225258409!225267761!-`.
|
||||
- [Getting Started](#started)
|
||||
- [Introduction](#intro)
|
||||
- [Evaluation](#eval)
|
||||
- [Evaluating mapping accuracy with simulated reads](#mapeval)
|
||||
- [Evaluating read overlap sensitivity](#oveval)
|
||||
- [Calling Variants from Assemblies](#asmvar)
|
||||
|
||||
* [sim-eval.js](sim-eval.js): evaluate mapping accuracy for FASTA generated
|
||||
with [sim-pbsim.js](sim-pbsim.js) or [sim-mason2.js](sim-mason2.js).
|
||||
## <a name="intro"></a>Introduction
|
||||
|
||||
* [sam2paf.js](sam2paf.js): convert SAM to PAF.
|
||||
paftools.js is a script that processes alignments in the [PAF format][paf],
|
||||
such as converting between formats, evaluating mapping accuracy, lifting over
|
||||
BED files based on alignment, and calling variants from assembly-to-assembly
|
||||
alignment. This script *requires* the [k8 Javascript shell][k8] to run. On
|
||||
Linux or Mac, you can download the precompiled k8 binary with:
|
||||
|
||||
```sh
|
||||
curl -L https://github.com/attractivechaos/k8/releases/download/v0.2.4/k8-0.2.4.tar.bz2 | tar -jxf -
|
||||
cp k8-0.2.4/k8-`uname -s` $HOME/bin/k8 # assuming $HOME/bin in your $PATH
|
||||
```
|
||||
|
||||
It is highly recommended to copy the executable `k8` to a directory on your
|
||||
`$PATH` such as `/usr/bin/env` can find it. Like python scripts, once you
|
||||
install `k8`, you can launch paftools.js in one of the two ways:
|
||||
|
||||
```sh
|
||||
path/to/paftools.js # only if k8 is on your $PATH
|
||||
k8 path/to/paftools.js
|
||||
```
|
||||
|
||||
In a nutshell, paftools.js has the following commands:
|
||||
|
||||
```
|
||||
Usage: paftools.js <command> [arguments]
|
||||
Commands:
|
||||
view convert PAF to BLAST-like (for eyeballing) or MAF
|
||||
splice2bed convert spliced alignment in PAF/SAM to BED12
|
||||
sam2paf convert SAM to PAF
|
||||
delta2paf convert MUMmer's delta to PAF
|
||||
gff2bed convert GTF/GFF3 to BED12
|
||||
|
||||
stat collect basic mapping information in PAF/SAM
|
||||
liftover simplistic liftOver
|
||||
call call variants from asm-to-ref alignment with the cs tag
|
||||
bedcov compute the number of bases covered
|
||||
|
||||
mapeval evaluate mapping accuracy using mason2/PBSIM-simulated FASTQ
|
||||
mason2fq convert mason2-simulated SAM to FASTQ
|
||||
pbsim2fq convert PBSIM-simulated MAF to FASTQ
|
||||
junceval evaluate splice junction consistency with known annotations
|
||||
ov-eval evaluate read overlap sensitivity using read-to-ref mapping
|
||||
```
|
||||
|
||||
paftools.js seamlessly reads both plain text files and gzip'd text files.
|
||||
|
||||
## <a name="eval"></a>Evaluation
|
||||
|
||||
### <a name="mapeval"></a>Evaluating mapping accuracy with simulated reads
|
||||
|
||||
The **pbsim2fq** command of paftools.js converts the MAF output of [pbsim][pbsim]
|
||||
to FASTQ and encodes the true mapping position in the read name in a format like
|
||||
`S1_33!chr1!225258409!225267761!-`. Similarly, the **mason2fq** command
|
||||
converts [mason2][mason2] simulated SAM to FASTQ.
|
||||
|
||||
Command **mapeval** evaluates mapped SAM/PAF. Here is example output:
|
||||
|
||||
```
|
||||
Q 60 32478 0 0.000000000 32478
|
||||
Q 22 16 1 0.000030775 32494
|
||||
Q 21 43 1 0.000061468 32537
|
||||
Q 19 73 1 0.000091996 32610
|
||||
Q 14 66 1 0.000122414 32676
|
||||
Q 10 27 3 0.000214048 32703
|
||||
Q 8 14 1 0.000244521 32717
|
||||
Q 7 13 2 0.000305530 32730
|
||||
Q 6 46 1 0.000335611 32776
|
||||
Q 3 10 1 0.000366010 32786
|
||||
Q 2 20 2 0.000426751 32806
|
||||
Q 1 248 94 0.003267381 33054
|
||||
Q 0 31 17 0.003778147 33085
|
||||
U 3
|
||||
```
|
||||
|
||||
where each Q-line gives the quality threshold, the number of reads mapped with
|
||||
mapping quality equal to or greater than the threshold, number of wrong
|
||||
mappings, accumulative mapping error rate and the accumulative number of
|
||||
mapped reads. The U-line, if present, gives the number of unmapped reads if
|
||||
they are present in the SAM file.
|
||||
|
||||
Suppose the reported mapping coordinate overlap with the true coordinate like
|
||||
the following:
|
||||
|
||||
```
|
||||
truth: --------------------
|
||||
mapper: ----------------------
|
||||
|<- l1 ->|<-- o -->|<-- l2 -->|
|
||||
```
|
||||
|
||||
Let `r=o/(l1+o+l2)`. The reported mapping is considered correct if `r>0.1` by
|
||||
default.
|
||||
|
||||
### <a name="oveval"></a>Evaluating read overlap sensitivity
|
||||
|
||||
Command **ov-eval** takes *sorted* read-to-reference alignment and read
|
||||
overlaps in PAF as input, and evaluates the sensitivity. For example:
|
||||
|
||||
```sh
|
||||
minimap2 -cx map-pb ref.fa reads.fq.gz | sort -k6,6 -k8,8n > reads-to-ref.paf
|
||||
minimap2 -x ava-pb reads.fq.gz reads.fq.gz > ovlp.paf
|
||||
k8 ov-eval.js reads-to-ref.paf ovlp.paf
|
||||
```
|
||||
|
||||
## <a name="asmvar"></a>Calling Variants from Haploid Assemblies
|
||||
|
||||
The **call** command of paftools.js calls variants from coordinate-sorted
|
||||
assembly-to-reference alignment. It calls variants from the [cs tag][cs] and
|
||||
identifies confident/callable regions as those covered by exactly one contig.
|
||||
Here are example command lines:
|
||||
|
||||
```sh
|
||||
minimap2 -cx asm5 -t8 --cs ref.fa asm.fa > asm.paf # keeping this file is recommended; --cs required!
|
||||
sort -k6,6 -k8,8n asm.paf > asm.srt.paf # sort by reference start coordinate
|
||||
k8 paftools.js call asm.srt.paf > asm.var.txt
|
||||
```
|
||||
|
||||
Here is sample output:
|
||||
|
||||
```
|
||||
V chr1 2276040 2276041 1 60 c g LJII01000171.1 1217409 1217410 +
|
||||
V chr1 2280409 2280410 1 60 a g LJII01000171.1 1221778 1221779 +
|
||||
V chr1 2280504 2280505 1 60 a g LJII01000171.1 1221873 1221874 +
|
||||
R chr1 2325140 2436340
|
||||
V chr1 2325287 2325287 1 60 - ct LJII01000171.1 1272894 1272896 +
|
||||
V chr1 2325642 2325644 1 60 tt - LJII01000171.1 1273251 1273251 +
|
||||
V chr1 2326051 2326052 1 60 c t LJII01000171.1 1273658 1273659 +
|
||||
V chr1 2326287 2326288 1 60 c t LJII01000171.1 1273894 1273895 +
|
||||
```
|
||||
|
||||
where a line starting with `R` gives regions covered by one query contig, and a
|
||||
V-line encodes a variant in the following format: chr, start, end, query depth,
|
||||
mapping quality, REF allele, ALT allele, query name, query start, end and the
|
||||
query orientation. Generally, you should only look at variants where column 5
|
||||
is one.
|
||||
|
||||
By default, when calling variants, "paftools.js call" ignores alignments 50kb
|
||||
or shorter; when deriving callable regions, it ignores alignments 10kb or
|
||||
shorter. It uses two thresholds to avoid edge effects. These defaults are
|
||||
designed for long-read assemblies. For short reads, both should be reduced.
|
||||
|
||||
|
||||
|
||||
[paf]: https://github.com/lh3/miniasm/blob/master/PAF.md
|
||||
[cs]: https://github.com/lh3/minimap2#cs
|
||||
[k8]: https://github.com/attractivechaos/k8
|
||||
[k8bin]: https://github.com/attractivechaos/k8/releases
|
||||
[maf]: https://genome.ucsc.edu/FAQ/FAQformat#format5
|
||||
[pbsim]: https://github.com/pfaucon/PBSIM-PacBio-Simulator
|
||||
[mason2]: https://github.com/seqan/seqan/tree/master/apps/mason2
|
||||
|
||||
@@ -1,258 +0,0 @@
|
||||
/*******************************
|
||||
* Command line option parsing *
|
||||
*******************************/
|
||||
|
||||
var getopt = function(args, ostr) {
|
||||
var oli; // option letter list index
|
||||
if (typeof(getopt.place) == 'undefined')
|
||||
getopt.ind = 0, getopt.arg = null, getopt.place = -1;
|
||||
if (getopt.place == -1) { // update scanning pointer
|
||||
if (getopt.ind >= args.length || args[getopt.ind].charAt(getopt.place = 0) != '-') {
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
if (getopt.place + 1 < args[getopt.ind].length && args[getopt.ind].charAt(++getopt.place) == '-') { // found "--"
|
||||
++getopt.ind;
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
}
|
||||
var optopt = args[getopt.ind].charAt(getopt.place++); // character checked for validity
|
||||
if (optopt == ':' || (oli = ostr.indexOf(optopt)) < 0) {
|
||||
if (optopt == '-') return null; // if the user didn't specify '-' as an option, assume it means null.
|
||||
if (getopt.place < 0) ++getopt.ind;
|
||||
return '?';
|
||||
}
|
||||
if (oli+1 >= ostr.length || ostr.charAt(++oli) != ':') { // don't need argument
|
||||
getopt.arg = null;
|
||||
if (getopt.place < 0 || getopt.place >= args[getopt.ind].length) ++getopt.ind, getopt.place = -1;
|
||||
} else { // need an argument
|
||||
if (getopt.place >= 0 && getopt.place < args[getopt.ind].length)
|
||||
getopt.arg = args[getopt.ind].substr(getopt.place);
|
||||
else if (args.length <= ++getopt.ind) { // no arg
|
||||
getopt.place = -1;
|
||||
if (ostr.length > 0 && ostr.charAt(0) == ':') return ':';
|
||||
return '?';
|
||||
} else getopt.arg = args[getopt.ind]; // white space
|
||||
getopt.place = -1;
|
||||
++getopt.ind;
|
||||
}
|
||||
return optopt;
|
||||
}
|
||||
|
||||
/***********************
|
||||
* Interval operations *
|
||||
***********************/
|
||||
|
||||
Interval = {};
|
||||
|
||||
Interval.sort = function(a)
|
||||
{
|
||||
if (typeof a[0] == 'number')
|
||||
a.sort(function(x, y) { return x - y });
|
||||
else a.sort(function(x, y) { return x[0] != y[0]? x[0] - y[0] : x[1] - y[1] });
|
||||
}
|
||||
|
||||
Interval.merge = function(a, sorted)
|
||||
{
|
||||
if (typeof sorted == 'undefined') sorted = true;
|
||||
if (!sorted) Interval.sort(a);
|
||||
var k = 0;
|
||||
for (var i = 1; i < a.length; ++i) {
|
||||
if (a[k][1] >= a[i][0])
|
||||
a[k][1] = a[k][1] > a[i][1]? a[k][1] : a[i][1];
|
||||
else a[++k] = a[i].slice(0);
|
||||
}
|
||||
a.length = k + 1;
|
||||
}
|
||||
|
||||
Interval.dedup = function(a, sorted)
|
||||
{
|
||||
if (typeof sorted == 'undefined') sorted = true;
|
||||
if (!sorted) Interval.sort(a);
|
||||
var k = 0;
|
||||
for (var i = 1; i < a.length; ++i)
|
||||
if (a[k][0] != a[i][0] || a[k][1] != a[i][1])
|
||||
a[++k] = a[i].slice(0);
|
||||
a.length = k + 1;
|
||||
}
|
||||
|
||||
Interval.index_end = function(a, sorted)
|
||||
{
|
||||
if (a.length == 0) return;
|
||||
if (typeof sorted == 'undefined') sorted = true;
|
||||
if (!sorted) Interval.sort(a);
|
||||
a[0].push(0);
|
||||
var k = 0, k_en = a[0][1];
|
||||
for (var i = 1; i < a.length; ++i) {
|
||||
if (k_en <= a[i][0]) {
|
||||
for (++k; k < i; ++k)
|
||||
if (a[k][1] > a[i][0])
|
||||
break;
|
||||
k_en = a[k][1];
|
||||
}
|
||||
a[i].push(k);
|
||||
}
|
||||
}
|
||||
|
||||
Interval.find_intv = function(a, x)
|
||||
{
|
||||
var left = -1, right = a.length;
|
||||
if (typeof a[0] == 'number') {
|
||||
while (right - left > 1) {
|
||||
var mid = left + ((right - left) >> 1);
|
||||
if (a[mid] > x) right = mid;
|
||||
else if (a[mid] < x) left = mid;
|
||||
else return mid;
|
||||
}
|
||||
} else {
|
||||
while (right - left > 1) {
|
||||
var mid = left + ((right - left) >> 1);
|
||||
if (a[mid][0] > x) right = mid;
|
||||
else if (a[mid][0] < x) left = mid;
|
||||
else return mid;
|
||||
}
|
||||
}
|
||||
return left;
|
||||
}
|
||||
|
||||
Interval.find_ovlp = function(a, st, en)
|
||||
{
|
||||
if (a.length == 0 || st >= en) return [];
|
||||
var l = Interval.find_intv(a, st);
|
||||
var k = l < 0? 0 : a[l][a[l].length - 1];
|
||||
var b = [];
|
||||
for (var i = k; i < a.length; ++i) {
|
||||
if (a[i][0] >= en) break;
|
||||
else if (st < a[i][1])
|
||||
b.push(a[i]);
|
||||
}
|
||||
return b;
|
||||
}
|
||||
|
||||
/*****************
|
||||
* Main function *
|
||||
*****************/
|
||||
|
||||
function read_bed(fn, to_merge, to_dedup)
|
||||
{
|
||||
var file = new File(fn);
|
||||
var buf = new Bytes();
|
||||
var h = {};
|
||||
while (file.readline(buf) >= 0) {
|
||||
var t = buf.toString().split("\t");
|
||||
if (h[t[0]] == null)
|
||||
h[t[0]] = [];
|
||||
var bst = parseInt(t[1]);
|
||||
var ben = parseInt(t[2]);
|
||||
if (t.length >= 12 && /^\d+$/.test(t[9])) {
|
||||
t[9] = parseInt(t[9]);
|
||||
var sz = t[10].split(",");
|
||||
var st = t[11].split(",");
|
||||
for (var i = 0; i < t[9]; ++i) {
|
||||
st[i] = parseInt(st[i]);
|
||||
sz[i] = parseInt(sz[i]);
|
||||
h[t[0]].push([bst + st[i], bst + st[i] + sz[i], 0, 0, 0]);
|
||||
}
|
||||
} else {
|
||||
h[t[0]].push([bst, ben, 0, 0, 0]);
|
||||
}
|
||||
}
|
||||
buf.destroy();
|
||||
file.close();
|
||||
for (var chr in h) {
|
||||
if (to_merge) Interval.merge(h[chr], false);
|
||||
else if (to_dedup) Interval.dedup(h[chr], false);
|
||||
else Interval.sort(h[chr]);
|
||||
Interval.index_end(h[chr]);
|
||||
}
|
||||
return h;
|
||||
}
|
||||
|
||||
function main(args)
|
||||
{
|
||||
var c, print_len = false, to_merge = true, to_dedup = false, fn_excl = null;
|
||||
while ((c = getopt(args, "pde:")) != null) {
|
||||
if (c == 'p') print_len = true;
|
||||
else if (c == 'd') to_dedup = true, to_merge = false;
|
||||
else if (c == 'e') fn_excl = getopt.arg;
|
||||
}
|
||||
|
||||
if (args.length - getopt.ind < 2) {
|
||||
print("Usage: k8 cnt-feat.js [options] <target.bed> <feature.bed>");
|
||||
print("Options:");
|
||||
print(" -e FILE exclude features overlapping regions in BED FILE []");
|
||||
print(" -p print number of covered bases for each feature");
|
||||
exit(1);
|
||||
}
|
||||
|
||||
var excl = fn_excl != null? read_bed(fn_excl, true, false) : null;
|
||||
var target = read_bed(args[getopt.ind], to_merge, to_dedup);
|
||||
|
||||
var file, buf = new Bytes();
|
||||
var tot_len = 0, hit_len = 0;
|
||||
file = args[getopt.ind+1] != '-'? new File(args[getopt.ind+1]) : new File();
|
||||
while (file.readline(buf) >= 0) {
|
||||
var t = buf.toString().split("\t");
|
||||
var a = [];
|
||||
var bst = parseInt(t[1]);
|
||||
var ben = parseInt(t[2]);
|
||||
if (t.length >= 12 && /^\d+$/.test(t[9])) { // BED12
|
||||
t[9] = parseInt(t[9]);
|
||||
var sz = t[10].split(",");
|
||||
var st = t[11].split(",");
|
||||
for (var i = 0; i < t[9]; ++i) {
|
||||
st[i] = parseInt(st[i]);
|
||||
sz[i] = parseInt(sz[i]);
|
||||
a.push([bst + st[i], bst + st[i] + sz[i], false]);
|
||||
}
|
||||
} else a.push([bst, ben, false]); // 3-column BED
|
||||
var feat_len = 0;
|
||||
for (var i = 0; i < a.length; ++i) {
|
||||
if (excl != null && excl[t[0]] != null) {
|
||||
var oe = Interval.find_ovlp(excl[t[0]], a[i][0], a[i][1]);
|
||||
if (oe.length > 0)
|
||||
continue;
|
||||
}
|
||||
a[i][2] = true;
|
||||
feat_len += a[i][1] - a[i][0];
|
||||
}
|
||||
tot_len += feat_len;
|
||||
if (target[t[0]] == null) continue;
|
||||
var b = [];
|
||||
for (var i = 0; i < a.length; ++i) {
|
||||
if (!a[i][2]) continue;
|
||||
var o = Interval.find_ovlp(target[t[0]], a[i][0], a[i][1]);
|
||||
for (var j = 0; j < o.length; ++j) {
|
||||
var max_st = o[j][0] > a[i][0]? o[j][0] : a[i][0];
|
||||
var min_en = o[j][1] < a[i][1]? o[j][1] : a[i][1];
|
||||
b.push([max_st, min_en]);
|
||||
o[j][2] += min_en - max_st;
|
||||
++o[j][3];
|
||||
if (max_st == o[j][0] && min_en == o[j][1])
|
||||
++o[j][4];
|
||||
}
|
||||
}
|
||||
// find the length covered
|
||||
var feat_hit_len = 0;
|
||||
if (b.length > 0) {
|
||||
b.sort(function(a,b) {return a[0]-b[0]});
|
||||
var st = b[0][0], en = b[0][1];
|
||||
for (var i = 1; i < b.length; ++i) {
|
||||
if (b[i][0] <= en) en = en > b[i][1]? en : b[i][1];
|
||||
else feat_hit_len += en - st, st = b[i][0], en = b[i][1];
|
||||
}
|
||||
feat_hit_len += en - st;
|
||||
}
|
||||
hit_len += feat_hit_len;
|
||||
if (print_len) print('F', t.slice(0, 4).join("\t"), feat_len, feat_hit_len);
|
||||
}
|
||||
file.close();
|
||||
|
||||
buf.destroy();
|
||||
|
||||
warn("# feature bases: " + tot_len);
|
||||
warn("# feature bases overlapping targets: " + hit_len + ' (' + (100.0 * hit_len / tot_len).toFixed(2) + '%)');
|
||||
}
|
||||
|
||||
main(arguments);
|
||||
-150
@@ -1,150 +0,0 @@
|
||||
var getopt = function(args, ostr) {
|
||||
var oli; // option letter list index
|
||||
if (typeof(getopt.place) == 'undefined')
|
||||
getopt.ind = 0, getopt.arg = null, getopt.place = -1;
|
||||
if (getopt.place == -1) { // update scanning pointer
|
||||
if (getopt.ind >= args.length || args[getopt.ind].charAt(getopt.place = 0) != '-') {
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
if (getopt.place + 1 < args[getopt.ind].length && args[getopt.ind].charAt(++getopt.place) == '-') { // found "--"
|
||||
++getopt.ind;
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
}
|
||||
var optopt = args[getopt.ind].charAt(getopt.place++); // character checked for validity
|
||||
if (optopt == ':' || (oli = ostr.indexOf(optopt)) < 0) {
|
||||
if (optopt == '-') return null; // if the user didn't specify '-' as an option, assume it means null.
|
||||
if (getopt.place < 0) ++getopt.ind;
|
||||
return '?';
|
||||
}
|
||||
if (oli+1 >= ostr.length || ostr.charAt(++oli) != ':') { // don't need argument
|
||||
getopt.arg = null;
|
||||
if (getopt.place < 0 || getopt.place >= args[getopt.ind].length) ++getopt.ind, getopt.place = -1;
|
||||
} else { // need an argument
|
||||
if (getopt.place >= 0 && getopt.place < args[getopt.ind].length)
|
||||
getopt.arg = args[getopt.ind].substr(getopt.place);
|
||||
else if (args.length <= ++getopt.ind) { // no arg
|
||||
getopt.place = -1;
|
||||
if (ostr.length > 0 && ostr.charAt(0) == ':') return ':';
|
||||
return '?';
|
||||
} else getopt.arg = args[getopt.ind]; // white space
|
||||
getopt.place = -1;
|
||||
++getopt.ind;
|
||||
}
|
||||
return optopt;
|
||||
}
|
||||
|
||||
var c, fn_ucsc_fai = null, is_short = false;
|
||||
while ((c = getopt(arguments, "u:s")) != null) {
|
||||
if (c == 'u') fn_ucsc_fai = getopt.arg;
|
||||
else if (c == 's') is_short = true;
|
||||
}
|
||||
|
||||
if (getopt.ind == arguments.length) {
|
||||
print("Usage: k8 gff2bed.js [-u ucsc-genome.fa.fai] <in.gff>");
|
||||
exit(1);
|
||||
}
|
||||
|
||||
var ens2ucsc = {};
|
||||
if (fn_ucsc_fai != null) {
|
||||
var buf = new Bytes();
|
||||
var file = new File(fn_ucsc_fai);
|
||||
while (file.readline(buf) >= 0) {
|
||||
var t = buf.toString().split("\t");
|
||||
var s = t[0];
|
||||
if (/_(random|alt|decoy)$/.test(s)) {
|
||||
s = s.replace(/_(random|alt|decoy)$/, '');
|
||||
s = s.replace(/^chr\S+_/, '');
|
||||
} else {
|
||||
s = s.replace(/^chrUn_/, '');
|
||||
}
|
||||
s = s.replace(/v(\d+)/, ".$1");
|
||||
if (s != t[0]) ens2ucsc[s] = t[0];
|
||||
}
|
||||
file.close();
|
||||
buf.destroy();
|
||||
}
|
||||
|
||||
var colors = {
|
||||
'protein_coding':'0,128,255',
|
||||
'lincRNA':'0,192,0',
|
||||
'snRNA':'0,192,0',
|
||||
'miRNA':'0,192,0',
|
||||
'misc_RNA':'0,192,0'
|
||||
};
|
||||
|
||||
function print_bed12(exons, cds_st, cds_en, is_short)
|
||||
{
|
||||
if (exons.length == 0) return;
|
||||
var name = is_short? exons[0][7] + "|" + exons[0][5] : exons[0].slice(4, 7).join("|");
|
||||
var a = exons.sort(function(a,b) {return a[1]-b[1]});
|
||||
var sizes = [], starts = [], st, en;
|
||||
st = a[0][1];
|
||||
en = a[a.length - 1][2];
|
||||
if (cds_st == 1<<30) cds_st = st;
|
||||
if (cds_en == 0) cds_en = en;
|
||||
if (cds_st < st || cds_en > en)
|
||||
throw Error("inconsistent thick start or end for transcript " + a[0][4]);
|
||||
for (var i = 0; i < a.length; ++i) {
|
||||
sizes.push(a[i][2] - a[i][1]);
|
||||
starts.push(a[i][1] - st);
|
||||
}
|
||||
var color = colors[a[0][5]];
|
||||
if (color == null) color = '196,196,196';
|
||||
print(a[0][0], st, en, name, 1000, a[0][3], cds_st, cds_en, color, a.length, sizes.join(",") + ",", starts.join(",") + ",");
|
||||
}
|
||||
|
||||
var re_gtf = /(transcript_id|transcript_type|transcript_biotype|gene_name|transcript_name) "([^"]+)";/g;
|
||||
var re_gff3 = /(transcript_id|transcript_type|transcript_biotype|gene_name|transcript_name)=([^;]+)/g;
|
||||
var buf = new Bytes();
|
||||
var file = new File(arguments[getopt.ind]);
|
||||
|
||||
var exons = [], cds_st = 1<<30, cds_en = 0, last_id = null;
|
||||
while (file.readline(buf) >= 0) {
|
||||
var t = buf.toString().split("\t");
|
||||
if (t[0].charAt(0) == '#') continue;
|
||||
if (t[2] != "CDS" && t[2] != "exon") continue;
|
||||
t[3] = parseInt(t[3]) - 1;
|
||||
t[4] = parseInt(t[4]);
|
||||
var id = null, type = "", gname = "N/A", biotype = "", m, tname = "N/A";
|
||||
while ((m = re_gtf.exec(t[8])) != null) {
|
||||
if (m[1] == "transcript_id") id = m[2];
|
||||
else if (m[1] == "transcript_type") type = m[2];
|
||||
else if (m[1] == "transcript_biotype") biotype = m[2];
|
||||
else if (m[1] == "gene_name") name = m[2];
|
||||
else if (m[1] == "transcript_name") tname = m[2];
|
||||
}
|
||||
while ((m = re_gff3.exec(t[8])) != null) {
|
||||
if (m[1] == "transcript_id") id = m[2];
|
||||
else if (m[1] == "transcript_type") type = m[2];
|
||||
else if (m[1] == "transcript_biotype") biotype = m[2];
|
||||
else if (m[1] == "gene_name") name = m[2];
|
||||
else if (m[1] == "transcript_name") tname = m[2];
|
||||
}
|
||||
if (type == "" && biotype != "") type = biotype;
|
||||
if (id == null) throw Error("No transcript_id");
|
||||
if (id != last_id) {
|
||||
print_bed12(exons, cds_st, cds_en, is_short);
|
||||
exons = [], cds_st = 1<<30, cds_en = 0;
|
||||
last_id = id;
|
||||
}
|
||||
if (t[2] == "CDS") {
|
||||
cds_st = cds_st < t[3]? cds_st : t[3];
|
||||
cds_en = cds_en > t[4]? cds_en : t[4];
|
||||
} else if (t[2] == "exon") {
|
||||
if (fn_ucsc_fai != null) {
|
||||
if (ens2ucsc[t[0]] != null)
|
||||
t[0] = ens2ucsc[t[0]];
|
||||
else if (/^[A-Z]+\d+\.\d+$/.test(t[0]))
|
||||
t[0] = t[0].replace(/([A-Z]+\d+)\.(\d+)/, "chrUn_$1v$2");
|
||||
}
|
||||
exons.push([t[0], t[3], t[4], t[6], id, type, name, tname]);
|
||||
}
|
||||
}
|
||||
if (last_id != null)
|
||||
print_bed12(exons, cds_st, cds_en, is_short);
|
||||
|
||||
file.close();
|
||||
buf.destroy();
|
||||
@@ -1,267 +0,0 @@
|
||||
/*******************************
|
||||
* Command line option parsing *
|
||||
*******************************/
|
||||
|
||||
var getopt = function(args, ostr) {
|
||||
var oli; // option letter list index
|
||||
if (typeof(getopt.place) == 'undefined')
|
||||
getopt.ind = 0, getopt.arg = null, getopt.place = -1;
|
||||
if (getopt.place == -1) { // update scanning pointer
|
||||
if (getopt.ind >= args.length || args[getopt.ind].charAt(getopt.place = 0) != '-') {
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
if (getopt.place + 1 < args[getopt.ind].length && args[getopt.ind].charAt(++getopt.place) == '-') { // found "--"
|
||||
++getopt.ind;
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
}
|
||||
var optopt = args[getopt.ind].charAt(getopt.place++); // character checked for validity
|
||||
if (optopt == ':' || (oli = ostr.indexOf(optopt)) < 0) {
|
||||
if (optopt == '-') return null; // if the user didn't specify '-' as an option, assume it means null.
|
||||
if (getopt.place < 0) ++getopt.ind;
|
||||
return '?';
|
||||
}
|
||||
if (oli+1 >= ostr.length || ostr.charAt(++oli) != ':') { // don't need argument
|
||||
getopt.arg = null;
|
||||
if (getopt.place < 0 || getopt.place >= args[getopt.ind].length) ++getopt.ind, getopt.place = -1;
|
||||
} else { // need an argument
|
||||
if (getopt.place >= 0 && getopt.place < args[getopt.ind].length)
|
||||
getopt.arg = args[getopt.ind].substr(getopt.place);
|
||||
else if (args.length <= ++getopt.ind) { // no arg
|
||||
getopt.place = -1;
|
||||
if (ostr.length > 0 && ostr.charAt(0) == ':') return ':';
|
||||
return '?';
|
||||
} else getopt.arg = args[getopt.ind]; // white space
|
||||
getopt.place = -1;
|
||||
++getopt.ind;
|
||||
}
|
||||
return optopt;
|
||||
}
|
||||
|
||||
/***********************
|
||||
* Interval operations *
|
||||
***********************/
|
||||
|
||||
Interval = {};
|
||||
|
||||
Interval.sort = function(a)
|
||||
{
|
||||
if (typeof a[0] == 'number')
|
||||
a.sort(function(x, y) { return x - y });
|
||||
else a.sort(function(x, y) { return x[0] != y[0]? x[0] - y[0] : x[1] - y[1] });
|
||||
}
|
||||
|
||||
Interval.merge = function(a, sorted)
|
||||
{
|
||||
if (typeof sorted == 'undefined') sorted = true;
|
||||
if (!sorted) Interval.sort(a);
|
||||
var k = 0;
|
||||
for (var i = 1; i < a.length; ++i) {
|
||||
if (a[k][1] >= a[i][0])
|
||||
a[k][1] = a[k][1] > a[i][1]? a[k][1] : a[i][1];
|
||||
else a[++k] = a[i].slice(0);
|
||||
}
|
||||
a.length = k + 1;
|
||||
}
|
||||
|
||||
Interval.index_end = function(a, sorted)
|
||||
{
|
||||
if (a.length == 0) return;
|
||||
if (typeof sorted == 'undefined') sorted = true;
|
||||
if (!sorted) Interval.sort(a);
|
||||
a[0].push(0);
|
||||
var k = 0, k_en = a[0][1];
|
||||
for (var i = 1; i < a.length; ++i) {
|
||||
if (k_en <= a[i][0]) {
|
||||
for (++k; k < i; ++k)
|
||||
if (a[k][1] > a[i][0])
|
||||
break;
|
||||
k_en = a[k][1];
|
||||
}
|
||||
a[i].push(k);
|
||||
}
|
||||
}
|
||||
|
||||
Interval.find_intv = function(a, x)
|
||||
{
|
||||
var left = -1, right = a.length;
|
||||
if (typeof a[0] == 'number') {
|
||||
while (right - left > 1) {
|
||||
var mid = left + ((right - left) >> 1);
|
||||
if (a[mid] > x) right = mid;
|
||||
else if (a[mid] < x) left = mid;
|
||||
else return mid;
|
||||
}
|
||||
} else {
|
||||
while (right - left > 1) {
|
||||
var mid = left + ((right - left) >> 1);
|
||||
if (a[mid][0] > x) right = mid;
|
||||
else if (a[mid][0] < x) left = mid;
|
||||
else return mid;
|
||||
}
|
||||
}
|
||||
return left;
|
||||
}
|
||||
|
||||
Interval.find_ovlp = function(a, st, en)
|
||||
{
|
||||
if (a.length == 0 || st >= en) return [];
|
||||
var l = Interval.find_intv(a, st);
|
||||
var k = l < 0? 0 : a[l][a[l].length - 1];
|
||||
var b = [];
|
||||
for (var i = k; i < a.length; ++i) {
|
||||
if (a[i][0] >= en) break;
|
||||
else if (st < a[i][1])
|
||||
b.push(a[i]);
|
||||
}
|
||||
return b;
|
||||
}
|
||||
|
||||
/*****************
|
||||
* Main function *
|
||||
*****************/
|
||||
|
||||
var c, l_fuzzy = 0, print_ovlp = false, print_err_only = false, first_only = false;
|
||||
while ((c = getopt(arguments, "l:ep")) != null) {
|
||||
if (c == 'l') l_fuzzy = parseInt(getopt.arg);
|
||||
else if (c == 'e') print_err_only = print_ovlp = true;
|
||||
else if (c == 'p') print_ovlp = true;
|
||||
}
|
||||
|
||||
if (arguments.length - getopt.ind < 2) {
|
||||
print("Usage: k8 intron-eval.js [options] <gene.gtf> <aln.sam>");
|
||||
exit(1);
|
||||
}
|
||||
|
||||
var file, buf = new Bytes();
|
||||
|
||||
var tr = {};
|
||||
file = new File(arguments[getopt.ind]);
|
||||
while (file.readline(buf) >= 0) {
|
||||
var m, t = buf.toString().split("\t");
|
||||
if (t[0].charAt(0) == '#') continue;
|
||||
if (t[2] != 'exon') continue;
|
||||
var st = parseInt(t[3]) - 1;
|
||||
var en = parseInt(t[4]);
|
||||
if ((m = /transcript_id "(\S+)"/.exec(t[8])) == null) continue;
|
||||
var tid = m[1];
|
||||
if (tr[tid] == null) tr[tid] = [t[0], t[6], 0, 0, []];
|
||||
tr[tid][4].push([st, en]);
|
||||
}
|
||||
file.close();
|
||||
|
||||
var anno = {};
|
||||
for (var tid in tr) {
|
||||
var t = tr[tid];
|
||||
Interval.sort(t[4]);
|
||||
t[2] = t[4][0][0];
|
||||
t[3] = t[4][t[4].length - 1][1];
|
||||
if (anno[t[0]] == null) anno[t[0]] = [];
|
||||
var s = t[4];
|
||||
for (var i = 0; i < s.length - 1; ++i) {
|
||||
if (s[i][1] >= s[i+1][0])
|
||||
warn("WARNING: incorrect annotation for transcript "+tid+" ("+s[i][1]+" >= "+s[i+1][0]+")")
|
||||
anno[t[0]].push([s[i][1], s[i+1][0]]);
|
||||
}
|
||||
}
|
||||
tr = null;
|
||||
|
||||
for (var chr in anno) {
|
||||
var e = anno[chr];
|
||||
if (e.length == 0) continue;
|
||||
Interval.sort(e);
|
||||
var k = 0;
|
||||
for (var i = 1; i < e.length; ++i) // dedup
|
||||
if (e[i][0] != e[k][0] || e[i][1] != e[k][1])
|
||||
e[++k] = e[i].slice(0);
|
||||
e.length = k + 1;
|
||||
Interval.index_end(e);
|
||||
}
|
||||
|
||||
var n_pri = 0, n_unmapped = 0, n_mapped = 0;
|
||||
var n_sgl = 0, n_splice = 0, n_splice_hit = 0, n_splice_novel = 0;
|
||||
|
||||
file = new File(arguments[getopt.ind+1]);
|
||||
var last_qname = null;
|
||||
var re_cigar = /(\d+)([MIDNSHX=])/g;
|
||||
while (file.readline(buf) >= 0) {
|
||||
var m, t = buf.toString().split("\t");
|
||||
|
||||
if (t[0].charAt(0) == '@') continue;
|
||||
var flag = parseInt(t[1]);
|
||||
if (flag&0x100) continue;
|
||||
if (first_only && last_qname == t[0]) continue;
|
||||
if (t[2] == '*') {
|
||||
++n_unmapped;
|
||||
continue;
|
||||
} else {
|
||||
++n_pri;
|
||||
if (last_qname != t[0]) ++n_mapped;
|
||||
}
|
||||
|
||||
var pos = parseInt(t[3]) - 1, intron = [];
|
||||
while ((m = re_cigar.exec(t[5])) != null) {
|
||||
var len = parseInt(m[1]), op = m[2];
|
||||
if (op == 'N') {
|
||||
intron.push([pos, pos + len]);
|
||||
pos += len;
|
||||
} else if (op == 'M' || op == 'X' || op == '=' || op == 'D') pos += len;
|
||||
}
|
||||
if (intron.length == 0) {
|
||||
++n_sgl;
|
||||
continue;
|
||||
}
|
||||
n_splice += intron.length;
|
||||
|
||||
var chr = anno[t[2]];
|
||||
if (chr != null) {
|
||||
for (var i = 0; i < intron.length; ++i) {
|
||||
var o = Interval.find_ovlp(chr, intron[i][0], intron[i][1]);
|
||||
if (o.length > 0) {
|
||||
var hit = false;
|
||||
for (var j = 0; j < o.length; ++j) {
|
||||
var st_diff = intron[i][0] - o[j][0];
|
||||
var en_diff = intron[i][1] - o[j][1];
|
||||
if (st_diff < 0) st_diff = -st_diff;
|
||||
if (en_diff < 0) en_diff = -en_diff;
|
||||
if (st_diff <= l_fuzzy && en_diff <= l_fuzzy)
|
||||
++n_splice_hit, hit = true;
|
||||
if (hit) break;
|
||||
}
|
||||
if (print_ovlp) {
|
||||
var type = hit? 'C' : 'P';
|
||||
if (hit && print_err_only) continue;
|
||||
var x = '[';
|
||||
for (var j = 0; j < o.length; ++j) {
|
||||
if (j) x += ', ';
|
||||
x += '(' + o[j][0] + "," + o[j][1] + ')';
|
||||
}
|
||||
x += ']';
|
||||
print(type, t[0], i+1, t[2], intron[i][0], intron[i][1], x);
|
||||
}
|
||||
} else {
|
||||
++n_splice_novel;
|
||||
if (print_ovlp)
|
||||
print('N', t[0], i+1, t[2], intron[i][0], intron[i][1]);
|
||||
}
|
||||
}
|
||||
} else {
|
||||
n_splice_novel += intron.length;
|
||||
}
|
||||
last_qname = t[0];
|
||||
}
|
||||
file.close();
|
||||
|
||||
buf.destroy();
|
||||
|
||||
if (!print_ovlp) {
|
||||
print("# unmapped reads: " + n_unmapped);
|
||||
print("# mapped reads: " + n_mapped);
|
||||
print("# primary alignments: " + n_pri);
|
||||
print("# singletons: " + n_sgl);
|
||||
print("# predicted introns: " + n_splice);
|
||||
print("# non-overlapping introns: " + n_splice_novel);
|
||||
print("# correct introns: " + n_splice_hit + " (" + (n_splice_hit / n_splice * 100).toFixed(2) + "%)");
|
||||
}
|
||||
-183
@@ -1,183 +0,0 @@
|
||||
var getopt = function(args, ostr) {
|
||||
var oli; // option letter list index
|
||||
if (typeof(getopt.place) == 'undefined')
|
||||
getopt.ind = 0, getopt.arg = null, getopt.place = -1;
|
||||
if (getopt.place == -1) { // update scanning pointer
|
||||
if (getopt.ind >= args.length || args[getopt.ind].charAt(getopt.place = 0) != '-') {
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
if (getopt.place + 1 < args[getopt.ind].length && args[getopt.ind].charAt(++getopt.place) == '-') { // found "--"
|
||||
++getopt.ind;
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
}
|
||||
var optopt = args[getopt.ind].charAt(getopt.place++); // character checked for validity
|
||||
if (optopt == ':' || (oli = ostr.indexOf(optopt)) < 0) {
|
||||
if (optopt == '-') return null; // if the user didn't specify '-' as an option, assume it means null.
|
||||
if (getopt.place < 0) ++getopt.ind;
|
||||
return '?';
|
||||
}
|
||||
if (oli+1 >= ostr.length || ostr.charAt(++oli) != ':') { // don't need argument
|
||||
getopt.arg = null;
|
||||
if (getopt.place < 0 || getopt.place >= args[getopt.ind].length) ++getopt.ind, getopt.place = -1;
|
||||
} else { // need an argument
|
||||
if (getopt.place >= 0 && getopt.place < args[getopt.ind].length)
|
||||
getopt.arg = args[getopt.ind].substr(getopt.place);
|
||||
else if (args.length <= ++getopt.ind) { // no arg
|
||||
getopt.place = -1;
|
||||
if (ostr.length > 0 && ostr.charAt(0) == ':') return ':';
|
||||
return '?';
|
||||
} else getopt.arg = args[getopt.ind]; // white space
|
||||
getopt.place = -1;
|
||||
++getopt.ind;
|
||||
}
|
||||
return optopt;
|
||||
}
|
||||
|
||||
var c, gap_out_len = null;
|
||||
while ((c = getopt(arguments, "l:")) != null)
|
||||
if (c == 'l') gap_out_len = parseInt(getopt.arg);
|
||||
|
||||
if (getopt.ind == arguments.length) {
|
||||
print("Usage: k8 mapstat.js [-l gapOutLen] <in.sam>|<in.paf>");
|
||||
exit(1);
|
||||
}
|
||||
|
||||
var buf = new Bytes();
|
||||
var file = new File(arguments[getopt.ind]);
|
||||
var re = /(\d+)([MIDSHNX=])/g;
|
||||
|
||||
var lineno = 0, n_pri = 0, n_2nd = 0, n_seq = 0, n_cigar_64k = 0, l_tot = 0, l_cov = 0;
|
||||
var n_gap = [[0, 0, 0, 0, 0, 0], [0, 0, 0, 0, 0, 0]];
|
||||
|
||||
function cov_len(regs)
|
||||
{
|
||||
regs.sort(function(a,b) {return a[0]-b[0]});
|
||||
var st = regs[0][0], en = regs[0][1], l = 0;
|
||||
for (var i = 1; i < regs.length; ++i) {
|
||||
if (regs[i][0] < en)
|
||||
en = en > regs[i][1]? en : regs[i][1];
|
||||
else l += en - st, st = regs[i][0], en = regs[i][1];
|
||||
}
|
||||
l += en - st;
|
||||
return l;
|
||||
}
|
||||
|
||||
var last = null, last_qlen = null, regs = [];
|
||||
while (file.readline(buf) >= 0) {
|
||||
var line = buf.toString();
|
||||
++lineno;
|
||||
if (line.charAt(0) != '@') {
|
||||
var t = line.split("\t", 12);
|
||||
var m, rs, cigar = null, is_pri = false, is_sam = false, is_rev = false, tname = null;
|
||||
var atlen = null, aqlen, qs, qe, mapq, ori_qlen;
|
||||
if (t[4] == '+' || t[4] == '-') { // PAF
|
||||
if (!/\ts2:i:\d+/.test(line)) {
|
||||
++n_2nd;
|
||||
continue;
|
||||
}
|
||||
if ((m = /\tcg:Z:(\S+)/.exec(line)) != null)
|
||||
cigar = m[1];
|
||||
if (cigar == null) {
|
||||
warn("WARNING: no CIGAR at line " + lineno);
|
||||
continue;
|
||||
}
|
||||
tname = t[5];
|
||||
qs = parseInt(t[2]), qe = parseInt(t[3]);
|
||||
aqlen = qe - qs;
|
||||
is_rev = t[4] == '+'? false : true;
|
||||
rs = parseInt(t[7]);
|
||||
atlen = parseInt(t[8]) - rs;
|
||||
mapq = parseInt(t[11]);
|
||||
ori_qlen = parseInt(t[1]);
|
||||
} else { // SAM
|
||||
var flag = parseInt(t[1]);
|
||||
if ((flag & 4) || t[2] == '*' || t[5] == '*') continue;
|
||||
if (flag & 0x100) {
|
||||
++n_2nd;
|
||||
continue;
|
||||
}
|
||||
cigar = t[5];
|
||||
tname = t[2];
|
||||
rs = parseInt(t[3]) - 1;
|
||||
mapq = parseInt(t[4]);
|
||||
aqlen = t[9].length;
|
||||
is_sam = true;
|
||||
is_rev = !!(flag&0x10);
|
||||
}
|
||||
++n_pri;
|
||||
if (last != t[0]) {
|
||||
if (last != null) {
|
||||
l_tot += last_qlen;
|
||||
l_cov += cov_len(regs);
|
||||
}
|
||||
regs = [];
|
||||
++n_seq, last = t[0];
|
||||
}
|
||||
var M = 0, tl = 0, ql = 0, clip = [0, 0], n_cigar = 0, sclip = 0;
|
||||
while ((m = re.exec(cigar)) != null) {
|
||||
var l = parseInt(m[1]);
|
||||
++n_cigar;
|
||||
if (m[2] == 'M' || m[2] == '=' || m[2] == 'X') {
|
||||
tl += l, ql += l, M += l;
|
||||
} else if (m[2] == 'I' || m[2] == 'D') {
|
||||
var type;
|
||||
if (l < 50) type = 0;
|
||||
else if (l < 100) type = 1;
|
||||
else if (l < 300) type = 2;
|
||||
else if (l < 400) type = 3;
|
||||
else if (l < 1000) type = 4;
|
||||
else type = 5;
|
||||
if (m[2] == 'I') ql += l, ++n_gap[0][type];
|
||||
else tl += l, ++n_gap[1][type];
|
||||
if (gap_out_len != null && l >= gap_out_len)
|
||||
print(t[0], ql, is_rev? '-' : '+', tname, rs + tl, m[2], l);
|
||||
} else if (m[2] == 'N') {
|
||||
tl += l;
|
||||
} else if (m[2] == 'S') {
|
||||
clip[M == 0? 0 : 1] = l, sclip += l;
|
||||
} else if (m[2] == 'H') {
|
||||
clip[M == 0? 0 : 1] = l;
|
||||
}
|
||||
}
|
||||
if (n_cigar > 65535) ++n_cigar_64k;
|
||||
if (ql + sclip != aqlen)
|
||||
warn("WARNING: aligned query length is inconsistent with CIGAR at line " + lineno + " (" + (ql+sclip) + " != " + aqlen + ")");
|
||||
if (atlen != null && atlen != tl)
|
||||
warn("WARNING: aligned reference length is inconsistent with CIGAR at line " + lineno);
|
||||
if (is_sam) {
|
||||
qs = clip[is_rev? 1 : 0], qe = qs + ql;
|
||||
ori_qlen = clip[0] + ql + clip[1];
|
||||
}
|
||||
regs.push([qs, qe]);
|
||||
last_qlen = ori_qlen;
|
||||
}
|
||||
}
|
||||
l_tot += last_qlen;
|
||||
l_cov += cov_len(regs);
|
||||
|
||||
file.close();
|
||||
buf.destroy();
|
||||
|
||||
if (gap_out_len == null) {
|
||||
print("Number of mapped sequences: " + n_seq);
|
||||
print("Number of primary alignments: " + n_pri);
|
||||
print("Number of secondary alignments: " + n_2nd);
|
||||
print("Number of primary alignments with >65535 CIGAR operations: " + n_cigar_64k);
|
||||
print("Number of bases in mapped sequences: " + l_tot);
|
||||
print("Number of mapped bases: " + l_cov);
|
||||
print("Number of insertions in [0,50): " + n_gap[0][0]);
|
||||
print("Number of insertions in [50,100): " + n_gap[0][1]);
|
||||
print("Number of insertions in [100,300): " + n_gap[0][2]);
|
||||
print("Number of insertions in [300,400): " + n_gap[0][3]);
|
||||
print("Number of insertions in [400,1000): " + n_gap[0][4]);
|
||||
print("Number of insertions in [1000,inf): " + n_gap[0][5]);
|
||||
print("Number of deletions in [0,50): " + n_gap[1][0]);
|
||||
print("Number of deletions in [50,100): " + n_gap[1][1]);
|
||||
print("Number of deletions in [100,300): " + n_gap[1][2]);
|
||||
print("Number of deletions in [300,400): " + n_gap[1][3]);
|
||||
print("Number of deletions in [400,1000): " + n_gap[1][4]);
|
||||
print("Number of deletions in [1000,inf): " + n_gap[1][5]);
|
||||
}
|
||||
Executable
+335
@@ -0,0 +1,335 @@
|
||||
#!/usr/bin/env k8
|
||||
|
||||
var getopt = function(args, ostr) {
|
||||
var oli; // option letter list index
|
||||
if (typeof(getopt.place) == 'undefined')
|
||||
getopt.ind = 0, getopt.arg = null, getopt.place = -1;
|
||||
if (getopt.place == -1) { // update scanning pointer
|
||||
if (getopt.ind >= args.length || args[getopt.ind].charAt(getopt.place = 0) != '-') {
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
if (getopt.place + 1 < args[getopt.ind].length && args[getopt.ind].charAt(++getopt.place) == '-') { // found "--"
|
||||
++getopt.ind;
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
}
|
||||
var optopt = args[getopt.ind].charAt(getopt.place++); // character checked for validity
|
||||
if (optopt == ':' || (oli = ostr.indexOf(optopt)) < 0) {
|
||||
if (optopt == '-') return null; // if the user didn't specify '-' as an option, assume it means null.
|
||||
if (getopt.place < 0) ++getopt.ind;
|
||||
return '?';
|
||||
}
|
||||
if (oli+1 >= ostr.length || ostr.charAt(++oli) != ':') { // don't need argument
|
||||
getopt.arg = null;
|
||||
if (getopt.place < 0 || getopt.place >= args[getopt.ind].length) ++getopt.ind, getopt.place = -1;
|
||||
} else { // need an argument
|
||||
if (getopt.place >= 0 && getopt.place < args[getopt.ind].length)
|
||||
getopt.arg = args[getopt.ind].substr(getopt.place);
|
||||
else if (args.length <= ++getopt.ind) { // no arg
|
||||
getopt.place = -1;
|
||||
if (ostr.length > 0 && ostr.charAt(0) == ':') return ':';
|
||||
return '?';
|
||||
} else getopt.arg = args[getopt.ind]; // white space
|
||||
getopt.place = -1;
|
||||
++getopt.ind;
|
||||
}
|
||||
return optopt;
|
||||
}
|
||||
|
||||
function read_fastx(file, buf)
|
||||
{
|
||||
if (file.readline(buf) < 0) return null;
|
||||
var m, line = buf.toString();
|
||||
if ((m = /^([>@])(\S+)/.exec(line)) == null)
|
||||
throw Error("wrong fastx format");
|
||||
var is_fq = (m[1] == '@');
|
||||
var name = m[2];
|
||||
if (file.readline(buf) < 0)
|
||||
throw Error("missing sequence line");
|
||||
var seq = buf.toString();
|
||||
if (is_fq) { // skip quality
|
||||
file.readline(buf);
|
||||
file.readline(buf);
|
||||
}
|
||||
return [name, seq];
|
||||
}
|
||||
|
||||
function filter_paf(a, opt)
|
||||
{
|
||||
if (a.length == 0) return;
|
||||
var k = 0;
|
||||
for (var i = 0; i < a.length; ++i) {
|
||||
var ai = a[i];
|
||||
if (ai[10] < opt.min_blen) continue;
|
||||
if (ai[9] < ai[10] * opt.min_iden) continue;
|
||||
var clip = [0, 0];
|
||||
if (ai[4] == '+') {
|
||||
clip[0] = ai[2] < ai[7]? ai[2] : ai[7];
|
||||
clip[1] = ai[1] - ai[3] < ai[6] - ai[8]? ai[1] - ai[3] : ai[6] - ai[8];
|
||||
} else {
|
||||
clip[0] = ai[2] < ai[6] - ai[8]? ai[2] : ai[6] - ai[8];
|
||||
clip[1] = ai[1] - ai[3] < ai[7]? ai[1] - ai[3] : ai[7];
|
||||
}
|
||||
if (clip[0] > opt.max_clip_len || clip[1] > opt.max_clip_len) continue;
|
||||
a[k++] = ai;
|
||||
}
|
||||
a.length = k;
|
||||
}
|
||||
|
||||
function parse_events(t, ev, id, buf)
|
||||
{
|
||||
var re = /(:(\d+))|(([\+\-\*])([a-z]+))/g;
|
||||
var m, cs = null;
|
||||
for (var j = 12; j < t.length; ++j) {
|
||||
if ((m = /^cs:Z:(\S+)/.exec(t[j])) != null) {
|
||||
cs = m[1].toLowerCase();
|
||||
break;
|
||||
}
|
||||
}
|
||||
if (cs == null) {
|
||||
warn("Warning: no cs tag for read '" + t[0] + "'");
|
||||
return;
|
||||
}
|
||||
var st = t[2], en = t[3];
|
||||
var x = st;
|
||||
while ((m = re.exec(cs)) != null) {
|
||||
var l;
|
||||
if (m[2] != null) { // an identitcal match ":\d+"
|
||||
l = parseInt(m[2]);
|
||||
// [start, end, type, index, changed_base]
|
||||
ev.push([x, x + l, 0, id]);
|
||||
} else {
|
||||
if (m[4] == '*') {
|
||||
l = 1;
|
||||
ev.push([x, x + 1, 1, id, m[5][0]]);
|
||||
} else if (m[4] == '+') {
|
||||
l = m[5].length;
|
||||
ev.push([x, x + l, 2, id]);
|
||||
} else if (m[4] == '-') {
|
||||
l = 0;
|
||||
ev.push([x, x, -1, id, m[5]]);
|
||||
}
|
||||
}
|
||||
x += l;
|
||||
}
|
||||
if (x != en)
|
||||
throw Error("inconsistent cs for read '" + t[0] + "'");
|
||||
}
|
||||
|
||||
function find_het_sub(ev, a, opt)
|
||||
{
|
||||
var n = a.length, last0_i = -1, h = [], d = [];
|
||||
for (var i = 0; i < n; ++i) h[i] = [], d[i] = [];
|
||||
for (var i = 0; i < ev.length; ++i) {
|
||||
if (ev[i][2] == 0) {
|
||||
if (last0_i < 0 || ev[i][0] != ev[last0_i][0]) last0_i = i;
|
||||
else if (ev[i][1] > ev[last0_i][1])
|
||||
last0_i = i;
|
||||
} else if (ev[i][2] == 1 && last0_i >= 0 && ev[i][0] < ev[last0_i][1]) {
|
||||
if (ev[last0_i][1] - ev[last0_i][0] >= opt.min_mlen) {
|
||||
if (opt.dbg_ev) print("EV", ev[last0_i].join("\t"), "|", ev[i].join("\t"));
|
||||
var e0 = ev[last0_i], hl = h[e0[3]];
|
||||
if (hl.length == 0 || hl[hl.length-1][0] != e0[0])
|
||||
hl.push([e0[0], e0[1]]);
|
||||
d[ev[i][3]].push([ev[i][0], e0[1] - e0[0]]);
|
||||
}
|
||||
}
|
||||
}
|
||||
var b = [];
|
||||
for (var i = 0; i < n; ++i) {
|
||||
var sh = 0, dh = 0;
|
||||
for (var j = 0; j < h[i].length; ++j)
|
||||
sh += h[i][j][1] - h[i][j][0];
|
||||
for (var j = 0; j < d[i].length; ++j)
|
||||
dh += d[i][j][1];
|
||||
// [start, end, index, #consistent, lenConsistent, #conflictive, lenConflictive, identity, mlen]
|
||||
b[i] = [a[i][2], a[i][3], i, h[i].length, sh, d[i].length, dh, a[i][9] / a[i][10], a[i][9]];
|
||||
}
|
||||
return b;
|
||||
}
|
||||
|
||||
function flt_utg_for_ec(b, opt)
|
||||
{
|
||||
var k = 0;
|
||||
for (var i = 0; i < b.length; ++i) {
|
||||
var bi = b[i];
|
||||
if (bi[4] == 0 && bi[6] == 0) b[k++] = bi; // entirely ambiguous
|
||||
else if (bi[6] < (bi[4] + bi[6]) * opt.max_ratio0) b[k++] = bi;
|
||||
}
|
||||
b.length = k;
|
||||
if (b.length == 0) return;
|
||||
// find the longest contiguous segment
|
||||
b.sort(function(x,y) { return x[0]-y[0] });
|
||||
var st = b[0][0], en = b[0][1], max_st = 0, max_en = 0, max_max_en = en;
|
||||
for (var i = 1; i < b.length; ++i) {
|
||||
if (b[i][0] > en) {
|
||||
if (en - st > max_en - max_st)
|
||||
max_st = st, max_en = en;
|
||||
st = b[i][0], en = b[i][1];
|
||||
} else {
|
||||
en = en > b[i][1]? en : b[i][1];
|
||||
}
|
||||
max_max_en = max_max_en > b[i][1]? max_max_en : b[i][1];
|
||||
}
|
||||
if (en - st > max_en - max_st)
|
||||
max_st = st, max_en = en;
|
||||
if (max_max_en != en || st != b[0][0]) {
|
||||
var k = 0;
|
||||
for (var i = 0; i < b.length; ++i)
|
||||
if (b[i][0] < max_en && b[i][1] > max_st)
|
||||
b[k++] = b[i];
|
||||
b.length = k;
|
||||
}
|
||||
}
|
||||
|
||||
function flt_utg_for_bin(b, opt) // filter out alignments clearly on the wrong phase
|
||||
{
|
||||
var k = 0;
|
||||
for (var i = 0; i < b.length; ++i) {
|
||||
var bi = b[i];
|
||||
if (bi[4] + bi[6] == 0 || bi[4] >= (bi[4] + bi[6]) * opt.max_ratio0) b[k++] = bi;
|
||||
}
|
||||
b.length = k;
|
||||
}
|
||||
|
||||
function ec_core(b, n_a, ev, buf, ecb) // error correction
|
||||
{
|
||||
var intv = [];
|
||||
for (var i = 0; i < n_a; ++i)
|
||||
intv[i] = null;
|
||||
intv[b[0][2]] = [b[0][0], b[0][1]];
|
||||
var en = b[0][1];
|
||||
for (var i = 1; i < b.length; ++i) {
|
||||
if (b[i][1] <= en) continue;
|
||||
intv[b[i][2]] = [en, b[i][1]];
|
||||
en = b[i][1];
|
||||
}
|
||||
var k = 0;
|
||||
ecb.capacity = buf.capacity;
|
||||
ecb.length = 0;
|
||||
for (var i = 0; i < ev.length; ++i) {
|
||||
var e = ev[i], I = intv[e[3]];
|
||||
if (I == null) continue;
|
||||
if (e[0] >= I[0] && e[0] < I[1]) { // this is to reduce duplicated events around junctions
|
||||
//print("X", e.join("\t"));
|
||||
if (e[2] == 0) {
|
||||
ecb.length += e[1] - e[0];
|
||||
for (var j = e[0]; j < e[1]; ++j)
|
||||
ecb[k++] = buf[j];
|
||||
} else if (e[2] == 1) {
|
||||
++ecb.length;
|
||||
ecb[k++] = e[4].charCodeAt(0);
|
||||
} else if (e[2] < 0) {
|
||||
ecb.length += e[4].length;
|
||||
for (var j = 0; j < e[4].length; ++j)
|
||||
ecb[k++] = e[4].charCodeAt(j);
|
||||
} // else, skip e[2] == 2
|
||||
}
|
||||
}
|
||||
if (ecb.length != k) throw Error("BUG!");
|
||||
}
|
||||
|
||||
function process_paf(a, opt, fp_seq, buf, ecb)
|
||||
{
|
||||
if (a.length == 0) return;
|
||||
var len = a[0][1], name = a[0][0], seq = null;
|
||||
if (len < opt.min_rlen) return;
|
||||
if (fp_seq) {
|
||||
var ret;
|
||||
while ((ret = read_fastx(fp_seq, buf)) != null)
|
||||
if (ret[0] == a[0][0])
|
||||
break;
|
||||
if (ret == null)
|
||||
throw Error("failed to find sequence for read '" + a[0][0] + "'");
|
||||
name = ret[0], seq = ret[1];
|
||||
if (seq.length != len)
|
||||
throw Error("inconsistent length for read '" + name + "'");
|
||||
}
|
||||
filter_paf(a, opt);
|
||||
if (a.length == 0) return;
|
||||
var ev = [];
|
||||
for (var i = 0; i < a.length; ++i)
|
||||
parse_events(a[i], ev, i, buf);
|
||||
ev.sort(function(x,y) { return x[0]!=y[0]? x[0]-y[0] : x[2]-y[2] });
|
||||
if (seq == null) print("SQ", name, a[0][1], a.length);
|
||||
var b = find_het_sub(ev, a, opt);
|
||||
if (opt.ec) flt_utg_for_ec(b, opt);
|
||||
else flt_utg_for_bin(b, opt);
|
||||
if (seq == null) {
|
||||
for (var i = 0; i < b.length; ++i) {
|
||||
var m, ai = a[b[i][2]], score = 0;
|
||||
for (var j = 10; j < ai.length; ++j)
|
||||
if ((m = /^AS:i:(\d+)/.exec(ai[j])) != null)
|
||||
score = m[1];
|
||||
print("TS", b[i][2], b[i][0], b[i][1], ai.slice(5, 9).join("\t"), b[i].slice(3, 7).join("\t"), score);
|
||||
}
|
||||
print("//");
|
||||
} else { // error correction
|
||||
if (b.length == 0) return;
|
||||
buf.set(seq, 0);
|
||||
ec_core(b, a.length, ev, buf, ecb);
|
||||
print(">" + name);
|
||||
print(ecb);
|
||||
}
|
||||
}
|
||||
|
||||
function main(args)
|
||||
{
|
||||
var c, opt = { min_rlen:5000, min_blen:5000, min_iden:0.8, min_mlen:5, max_clip_len:500, max_ratio0:0.25, dbg_ev:false };
|
||||
while ((c = getopt(args, "l:b:d:m:c:r:E")) != null) {
|
||||
if (c == 'l') opt.min_rlen = parseInt(getopt.arg);
|
||||
else if (c == 'b') opt.min_blen = parseInt(getopt.arg);
|
||||
else if (c == 'd') opt.min_iden = parseFloat(getopt.arg);
|
||||
else if (c == 'm') opt.min_slen = parseInt(getopt.arg);
|
||||
else if (c == 'c') opt.max_clip_len = parseInt(getopt.arg);
|
||||
else if (c == 'r') opt.max_ratio0 = parseFloat(getopt.arg);
|
||||
else if (c == 'E') opt.dbg_ev = true;
|
||||
}
|
||||
if (args.length - getopt.ind < 1) {
|
||||
print("Usage: mmphase.js [options] <map-with-cs.paf> [reads.fa]");
|
||||
print("Options:");
|
||||
print(" -l INT min read length [" + opt.min_rlen + "]");
|
||||
print(" -b INT min alignment length [" + opt.min_blen + "]");
|
||||
print(" -d FLOAT min identity [" + opt.min_iden + "]");
|
||||
print(" -s INT min match length [" + opt.min_mlen + "]");
|
||||
print(" -c INT max clip length [" + opt.max_clip_len + "]");
|
||||
print(" -r FLOAT initial ratio for haplotype filtering [" + opt.max_ratio0 + "]");
|
||||
return 0;
|
||||
}
|
||||
|
||||
opt.ec = args.length - getopt.ind < 2? false : true;
|
||||
if (!opt.ec) {
|
||||
print("CC");
|
||||
print("CC", "SQ qName qLen nHits");
|
||||
print("CC", "TS index qStart qEnd tName tLen tStart tEnd nConsistent lCons nConflictive lConf score");
|
||||
print("CC");
|
||||
}
|
||||
|
||||
var buf = new Bytes(), ecb = new Bytes();
|
||||
var fp_paf = new File(args[getopt.ind]);
|
||||
var fp_seq = args.length - getopt.ind >= 2? new File(args[getopt.ind+1]) : null;
|
||||
var a = [];
|
||||
while (fp_paf.readline(buf) >= 0) {
|
||||
var t = buf.toString().split("\t");
|
||||
if (a.length > 0 && a[0][0] != t[0]) {
|
||||
process_paf(a, opt, fp_seq, buf, ecb);
|
||||
a.length = 0;
|
||||
}
|
||||
for (var i = 1; i <= 3; ++i) t[i] = parseInt(t[i]);
|
||||
if (t[1] < opt.min_rlen) continue;
|
||||
for (var i = 6; i <= 10; ++i) t[i] = parseInt(t[i]);
|
||||
if (t[10] < opt.min_blen) continue;
|
||||
a.push(t);
|
||||
}
|
||||
if (a.length >= 0)
|
||||
process_paf(a, opt, fp_seq, buf, ecb);
|
||||
if (fp_seq) fp_seq.close();
|
||||
fp_paf.close();
|
||||
ecb.destroy();
|
||||
buf.destroy();
|
||||
}
|
||||
|
||||
var ret = main(arguments)
|
||||
exit(ret)
|
||||
-105
@@ -1,105 +0,0 @@
|
||||
var getopt = function(args, ostr) {
|
||||
var oli; // option letter list index
|
||||
if (typeof(getopt.place) == 'undefined')
|
||||
getopt.ind = 0, getopt.arg = null, getopt.place = -1;
|
||||
if (getopt.place == -1) { // update scanning pointer
|
||||
if (getopt.ind >= args.length || args[getopt.ind].charAt(getopt.place = 0) != '-') {
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
if (getopt.place + 1 < args[getopt.ind].length && args[getopt.ind].charAt(++getopt.place) == '-') { // found "--"
|
||||
++getopt.ind;
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
}
|
||||
var optopt = args[getopt.ind].charAt(getopt.place++); // character checked for validity
|
||||
if (optopt == ':' || (oli = ostr.indexOf(optopt)) < 0) {
|
||||
if (optopt == '-') return null; // if the user didn't specify '-' as an option, assume it means null.
|
||||
if (getopt.place < 0) ++getopt.ind;
|
||||
return '?';
|
||||
}
|
||||
if (oli+1 >= ostr.length || ostr.charAt(++oli) != ':') { // don't need argument
|
||||
getopt.arg = null;
|
||||
if (getopt.place < 0 || getopt.place >= args[getopt.ind].length) ++getopt.ind, getopt.place = -1;
|
||||
} else { // need an argument
|
||||
if (getopt.place >= 0 && getopt.place < args[getopt.ind].length)
|
||||
getopt.arg = args[getopt.ind].substr(getopt.place);
|
||||
else if (args.length <= ++getopt.ind) { // no arg
|
||||
getopt.place = -1;
|
||||
if (ostr.length > 0 && ostr.charAt(0) == ':') return ':';
|
||||
return '?';
|
||||
} else getopt.arg = args[getopt.ind]; // white space
|
||||
getopt.place = -1;
|
||||
++getopt.ind;
|
||||
}
|
||||
return optopt;
|
||||
}
|
||||
|
||||
var c, min_ovlp = 2000, min_frac = 0.95, min_mapq = 10;
|
||||
while ((c = getopt(arguments, "q:l:f:")) != null) {
|
||||
if (c == 'q') min_mapq = parseInt(getopt.arg);
|
||||
else if (c == 'l') min_ovlp = parseInt(getopt.arg);
|
||||
else if (c == 'f') min_frac = parseFloat(getopt.arg);
|
||||
}
|
||||
if (arguments.length - getopt.ind < 2) {
|
||||
print("Usage: sort -k6,6 -k8,8n to-ref.paf | k8 ov-eval.js [options] - <ovlp.paf>");
|
||||
print("Options:");
|
||||
print(" -l INT min overlap length [2000]");
|
||||
print(" -q INT min mapping quality [10]");
|
||||
print(" -f FLOAT min fraction of mapped length [0.95]");
|
||||
exit(1);
|
||||
}
|
||||
|
||||
var buf = new Bytes();
|
||||
var file = arguments[getopt.ind] == '-'? new File() : new File(arguments[getopt.ind]);
|
||||
var a = [], h = {};
|
||||
while (file.readline(buf) >= 0) {
|
||||
var t = buf.toString().split("\t");
|
||||
var is_pri = false;
|
||||
if (parseInt(t[11]) < min_mapq) continue;
|
||||
for (var i = 12; i < t.length; ++i)
|
||||
if (t[i] == 'tp:A:P')
|
||||
is_pri = true;
|
||||
if (!is_pri) continue;
|
||||
for (var i = 1; i <= 3; ++i)
|
||||
t[i] = parseInt(t[i]);
|
||||
for (var i = 6; i <= 8; ++i)
|
||||
t[i] = parseInt(t[i]);
|
||||
if (t[3] - t[2] < min_ovlp || t[8] - t[7] < min_ovlp || (t[3] - t[2]) / t[1] < min_frac)
|
||||
continue;
|
||||
var ctg = t[5], st = t[7], en = t[8];
|
||||
while (a.length > 0) {
|
||||
if (a[0][0] == ctg && a[0][2] > st)
|
||||
break;
|
||||
else a.shift();
|
||||
}
|
||||
for (var j = 0; j < a.length; ++j) {
|
||||
if (a[j][3] == t[0]) continue;
|
||||
var len = (en > a[j][2]? a[j][2] : en) - st;
|
||||
if (len >= min_ovlp) {
|
||||
var key = a[j][3] < t[0]? a[j][3] + "\t" + t[0] : t[0] + "\t" + a[j][3];
|
||||
h[key] = len;
|
||||
}
|
||||
}
|
||||
a.push([ctg, st, en, t[0]]);
|
||||
}
|
||||
file.close();
|
||||
|
||||
file = new File(arguments[getopt.ind + 1]);
|
||||
while (file.readline(buf) >= 0) {
|
||||
var t = buf.toString().split("\t");
|
||||
var key = t[0] < t[5]? t[0] + "\t" + t[5] : t[5] + "\t" + t[0];
|
||||
if (h[key] > 0) h[key] = -h[key];
|
||||
}
|
||||
file.close();
|
||||
buf.destroy();
|
||||
|
||||
var n_ovlp = 0, n_missing = 0;
|
||||
for (var key in h) {
|
||||
++n_ovlp;
|
||||
if (h[key] > 0) ++n_missing;
|
||||
}
|
||||
print(n_ovlp + " overlaps inferred from the reference mapping");
|
||||
print(n_missing + " missed by the read overlapper");
|
||||
print((100 * (1 - n_missing / n_ovlp)).toFixed(2) + "% sensitivity");
|
||||
-196
@@ -1,196 +0,0 @@
|
||||
var getopt = function(args, ostr) {
|
||||
var oli; // option letter list index
|
||||
if (typeof(getopt.place) == 'undefined')
|
||||
getopt.ind = 0, getopt.arg = null, getopt.place = -1;
|
||||
if (getopt.place == -1) { // update scanning pointer
|
||||
if (getopt.ind >= args.length || args[getopt.ind].charAt(getopt.place = 0) != '-') {
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
if (getopt.place + 1 < args[getopt.ind].length && args[getopt.ind].charAt(++getopt.place) == '-') { // found "--"
|
||||
++getopt.ind;
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
}
|
||||
var optopt = args[getopt.ind].charAt(getopt.place++); // character checked for validity
|
||||
if (optopt == ':' || (oli = ostr.indexOf(optopt)) < 0) {
|
||||
if (optopt == '-') return null; // if the user didn't specify '-' as an option, assume it means null.
|
||||
if (getopt.place < 0) ++getopt.ind;
|
||||
return '?';
|
||||
}
|
||||
if (oli+1 >= ostr.length || ostr.charAt(++oli) != ':') { // don't need argument
|
||||
getopt.arg = null;
|
||||
if (getopt.place < 0 || getopt.place >= args[getopt.ind].length) ++getopt.ind, getopt.place = -1;
|
||||
} else { // need an argument
|
||||
if (getopt.place >= 0 && getopt.place < args[getopt.ind].length)
|
||||
getopt.arg = args[getopt.ind].substr(getopt.place);
|
||||
else if (args.length <= ++getopt.ind) { // no arg
|
||||
getopt.place = -1;
|
||||
if (ostr.length > 0 && ostr.charAt(0) == ':') return ':';
|
||||
return '?';
|
||||
} else getopt.arg = args[getopt.ind]; // white space
|
||||
getopt.place = -1;
|
||||
++getopt.ind;
|
||||
}
|
||||
return optopt;
|
||||
}
|
||||
|
||||
var c, line_len = 80, fmt = "aln";
|
||||
while ((c = getopt(arguments, "f:l:")) != null) {
|
||||
if (c == 'f') {
|
||||
fmt = getopt.arg;
|
||||
if (fmt != "aln" && fmt != "lastz-cigar" && fmt != "maf")
|
||||
throw Error("format must be one of aln, lastz-cigar and maf");
|
||||
} else if (c == 'l') line_len = parseInt(getopt.arg);
|
||||
}
|
||||
if (line_len == 0) line_len = 0x7fffffff;
|
||||
|
||||
if (getopt.ind == arguments.length) {
|
||||
print("Usage: k8 paf2aln.js [options] <in.paf>");
|
||||
print("Options:");
|
||||
print(" -f STR output format: aln (BLAST-like), maf or lastz-cigar [aln]");
|
||||
print(" -l INT line length in BLAST-like output [80]");
|
||||
exit(1);
|
||||
}
|
||||
|
||||
function padding_str(x, len, right)
|
||||
{
|
||||
var s = x.toString();
|
||||
if (s.length < len) {
|
||||
if (right) s += Array(len - s.length + 1).join(" ");
|
||||
else s = Array(len - s.length + 1).join(" ") + s;
|
||||
}
|
||||
return s;
|
||||
}
|
||||
|
||||
function update_aln(s_ref, s_qry, s_mid, type, seq, slen)
|
||||
{
|
||||
var l = type == '*'? 1 : seq.length;
|
||||
if (type == '=' || type == ':') {
|
||||
s_ref.set(seq);
|
||||
s_qry.set(seq);
|
||||
s_mid.set(Array(l+1).join("|"));
|
||||
slen[0] += l, slen[1] += l;
|
||||
} else if (type == '*') {
|
||||
s_ref.set(seq.charAt(0));
|
||||
s_qry.set(seq.charAt(1));
|
||||
s_mid.set(' ');
|
||||
slen[0] += 1, slen[1] += 1;
|
||||
} else if (type == '+') {
|
||||
s_ref.set(Array(l+1).join("-"));
|
||||
s_qry.set(seq);
|
||||
s_mid.set(Array(l+1).join(" "));
|
||||
slen[1] += l;
|
||||
} else if (type == '-') {
|
||||
s_ref.set(seq);
|
||||
s_qry.set(Array(l+1).join("-"));
|
||||
s_mid.set(Array(l+1).join(" "));
|
||||
slen[0] += l;
|
||||
}
|
||||
}
|
||||
|
||||
function print_aln(rs, qs, strand, slen, elen, s_ref, s_qry, s_mid)
|
||||
{
|
||||
print(["Ref+:", padding_str(rs + slen[0] + 1, 10, false), s_ref.toString(), padding_str(rs + elen[0], 10, true)].join(" "));
|
||||
print(" " + s_mid.toString());
|
||||
var st, en;
|
||||
if (strand == '+') st = qs + slen[1] + 1, en = qs + elen[1];
|
||||
else st = qs - slen[1], en = qs - elen[1] + 1;
|
||||
print(["Qry" + strand + ":", padding_str(st, 10, false), s_qry.toString(), padding_str(en, 10, true)].join(" "));
|
||||
}
|
||||
|
||||
var s_ref = new Bytes(), s_qry = new Bytes(), s_mid = new Bytes(); // these are used to show padded alignment
|
||||
var re_cs = /([:=\-\+\*])(\d+|[A-Za-z]+)/g;
|
||||
var re_cg = /(\d+)([MIDNSH])/g;
|
||||
|
||||
var buf = new Bytes();
|
||||
var file = arguments[getopt.ind] == "-"? new File() : new File(arguments[getopt.ind]);
|
||||
var lineno = 0;
|
||||
if (fmt == "maf") print("##maf version=1\n");
|
||||
while (file.readline(buf) >= 0) {
|
||||
var m, line = buf.toString();
|
||||
var t = line.split("\t", 12);
|
||||
++lineno;
|
||||
s_ref.length = s_qry.length = s_mid.length = 0;
|
||||
var slen = [0, 0], elen = [0, 0];
|
||||
if (fmt == "lastz-cigar") { // LASTZ-cigar output
|
||||
var cg = (m = /\tcg:Z:(\S+)/.exec(line)) != null? m[1] : null;
|
||||
if (cg == null) {
|
||||
warn("WARNING: converting to LASTZ-cigar format requires the 'cg' tag, which is absent on line " + lineno);
|
||||
continue;
|
||||
}
|
||||
var score = (m = /\tAS:i:(\d+)/.exec(line)) != null? m[1] : 0;
|
||||
var out = ['cigar:', t[0], t[2], t[3], t[4], t[5], t[7], t[8], '+', score];
|
||||
while ((m = re_cg.exec(cg)) != null)
|
||||
out.push(m[2], m[1]);
|
||||
print(out.join(" "));
|
||||
} else if (fmt == "maf") { // MAF output
|
||||
var cs = (m = /\tcs:Z:(\S+)/.exec(line)) != null? m[1] : null;
|
||||
if (cs == null) {
|
||||
warn("WARNING: converting to MAF requires the 'cs' tag, which is absent on line " + lineno);
|
||||
continue;
|
||||
}
|
||||
while ((m = re_cs.exec(cs)) != null) {
|
||||
if (m[1] == ':')
|
||||
throw Error("converting to MAF only works with 'minimap2 --cs=long'");
|
||||
update_aln(s_ref, s_qry, s_mid, m[1], m[2], elen);
|
||||
}
|
||||
var score = (m = /\tAS:i:(\d+)/.exec(line)) != null? parseInt(m[1]) : 0;
|
||||
var len = t[0].length > t[5].length? t[0].length : t[5].length;
|
||||
print("a " + score);
|
||||
print(["s", padding_str(t[5], len, true), padding_str(t[7], 10, false), padding_str(parseInt(t[8]) - parseInt(t[7]), 10, false),
|
||||
"+", padding_str(t[6], 10, false), s_ref.toString()].join(" "));
|
||||
var qs, qe, ql = parseInt(t[1]);
|
||||
if (t[4] == '+') {
|
||||
qs = parseInt(t[2]);
|
||||
qe = parseInt(t[3]);
|
||||
} else {
|
||||
qs = ql - parseInt(t[3]);
|
||||
qe = ql - parseInt(t[2]);
|
||||
}
|
||||
print(["s", padding_str(t[0], len, true), padding_str(qs, 10, false), padding_str(qe - qs, 10, false),
|
||||
t[4], padding_str(ql, 10, false), s_qry.toString()].join(" "));
|
||||
print("");
|
||||
} else { // BLAST-like output
|
||||
var cs = (m = /\tcs:Z:(\S+)/.exec(line)) != null? m[1] : null;
|
||||
if (cs == null) {
|
||||
warn("WARNING: converting to BLAST-like alignment requires the 'cs' tag, which is absent on line " + lineno);
|
||||
continue;
|
||||
}
|
||||
line = line.replace(/\tc[sg]:Z:\S+/g, ""); // get rid of cs or cg tags
|
||||
print('>' + line);
|
||||
var rs = parseInt(t[7]), qs = t[4] == '+'? parseInt(t[2]) : parseInt(t[3]);
|
||||
var n_blocks = 0;
|
||||
while ((m = re_cs.exec(cs)) != null) {
|
||||
if (m[1] == ':') m[2] = Array(parseInt(m[2]) + 1).join("=");
|
||||
var start = 0, rest = m[1] == '*'? 1 : m[2].length;
|
||||
while (rest > 0) {
|
||||
var l_proc;
|
||||
if (s_ref.length + rest >= line_len) {
|
||||
l_proc = line_len - s_ref.length;
|
||||
update_aln(s_ref, s_qry, s_mid, m[1], m[1] == '*'? m[2] : m[2].substr(start, l_proc), elen);
|
||||
if (n_blocks > 0) print("");
|
||||
print_aln(rs, qs, t[4], slen, elen, s_ref, s_qry, s_mid);
|
||||
++n_blocks;
|
||||
s_ref.length = s_qry.length = s_mid.length = 0;
|
||||
slen[0] = elen[0], slen[1] = elen[1];
|
||||
} else {
|
||||
l_proc = rest;
|
||||
update_aln(s_ref, s_qry, s_mid, m[1], m[1] == '*'? m[2] : m[2].substr(start, l_proc), elen);
|
||||
}
|
||||
rest -= l_proc, start += l_proc;
|
||||
}
|
||||
}
|
||||
if (s_ref.length > 0) {
|
||||
if (n_blocks > 0) print("");
|
||||
print_aln(rs, qs, t[4], slen, elen, s_ref, s_qry, s_mid);
|
||||
++n_blocks;
|
||||
}
|
||||
print("//");
|
||||
}
|
||||
}
|
||||
file.close();
|
||||
buf.destroy();
|
||||
|
||||
s_ref.destroy(); s_qry.destroy(); s_mid.destroy();
|
||||
@@ -1,188 +0,0 @@
|
||||
var getopt = function(args, ostr) {
|
||||
var oli; // option letter list index
|
||||
if (typeof(getopt.place) == 'undefined')
|
||||
getopt.ind = 0, getopt.arg = null, getopt.place = -1;
|
||||
if (getopt.place == -1) { // update scanning pointer
|
||||
if (getopt.ind >= args.length || args[getopt.ind].charAt(getopt.place = 0) != '-') {
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
if (getopt.place + 1 < args[getopt.ind].length && args[getopt.ind].charAt(++getopt.place) == '-') { // found "--"
|
||||
++getopt.ind;
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
}
|
||||
var optopt = args[getopt.ind].charAt(getopt.place++); // character checked for validity
|
||||
if (optopt == ':' || (oli = ostr.indexOf(optopt)) < 0) {
|
||||
if (optopt == '-') return null; // if the user didn't specify '-' as an option, assume it means null.
|
||||
if (getopt.place < 0) ++getopt.ind;
|
||||
return '?';
|
||||
}
|
||||
if (oli+1 >= ostr.length || ostr.charAt(++oli) != ':') { // don't need argument
|
||||
getopt.arg = null;
|
||||
if (getopt.place < 0 || getopt.place >= args[getopt.ind].length) ++getopt.ind, getopt.place = -1;
|
||||
} else { // need an argument
|
||||
if (getopt.place >= 0 && getopt.place < args[getopt.ind].length)
|
||||
getopt.arg = args[getopt.ind].substr(getopt.place);
|
||||
else if (args.length <= ++getopt.ind) { // no arg
|
||||
getopt.place = -1;
|
||||
if (ostr.length > 0 && ostr.charAt(0) == ':') return ':';
|
||||
return '?';
|
||||
} else getopt.arg = args[getopt.ind]; // white space
|
||||
getopt.place = -1;
|
||||
++getopt.ind;
|
||||
}
|
||||
return optopt;
|
||||
}
|
||||
|
||||
var re_cs = /([:=*+-])(\d+|[A-Za-z]+)/g;
|
||||
var c, min_cov_len = 10000, min_var_len = 50000, gap_thres = 50, min_mapq = 5;
|
||||
while ((c = getopt(arguments, "l:L:g:q:")) != null) {
|
||||
if (c == 'l') min_cov_len = parseInt(getopt.arg);
|
||||
else if (c == 'L') min_var_len = parseInt(getopt.arg);
|
||||
else if (c == 'g') gap_thres = parseInt(getopt.arg);
|
||||
else if (c == 'q') min_mapq = parseInt(getopt.arg);
|
||||
}
|
||||
|
||||
if (arguments.length == getopt.ind) {
|
||||
print("Usage: k8 paf2diff.js [options] <with-cs.paf>");
|
||||
print("Options:");
|
||||
print(" -l INT min alignment length to compute coverage ["+min_cov_len+"]");
|
||||
print(" -L INT min alignment length to call variants ["+min_var_len+"]");
|
||||
print(" -q INT min mapping quality ["+min_mapq+"]");
|
||||
print(" -g INT short/long gap threshold (for statistics only) ["+gap_thres+"]");
|
||||
exit(1);
|
||||
}
|
||||
|
||||
var file = arguments[getopt.ind] == '-'? new File() : new File(arguments[getopt.ind]);
|
||||
var buf = new Bytes();
|
||||
var tot_len = 0, n_sub = [0, 0, 0], n_ins = [0, 0, 0, 0], n_del = [0, 0, 0, 0];
|
||||
|
||||
function count_var(o)
|
||||
{
|
||||
if (o[3] > 1) return;
|
||||
if (o[5] == '-' && o[6] == '-') return;
|
||||
if (o[5] == '-') { // insertion
|
||||
var l = o[6].length;
|
||||
if (l == 1) ++n_ins[0];
|
||||
else if (l == 2) ++n_ins[1];
|
||||
else if (l < gap_thres) ++n_ins[2];
|
||||
else ++n_ins[3];
|
||||
} else if (o[6] == '-') { // deletion
|
||||
var l = o[5].length;
|
||||
if (l == 1) ++n_del[0];
|
||||
else if (l == 2) ++n_del[1];
|
||||
else if (l < gap_thres) ++n_del[2];
|
||||
else ++n_del[3];
|
||||
} else {
|
||||
++n_sub[0];
|
||||
var s = o[5] + o[6];
|
||||
if (s == 'ag' || s == 'ga' || s == 'ct' || s == 'tc')
|
||||
++n_sub[1];
|
||||
else ++n_sub[2];
|
||||
}
|
||||
}
|
||||
|
||||
var a = [], out = [];
|
||||
var c1_ctg = null, c1_start = 0, c1_end = 0, c1_counted = false, c1_len = 0;
|
||||
while (file.readline(buf) >= 0) {
|
||||
var line = buf.toString();
|
||||
if (!/\ts2:i:/.test(line)) continue; // skip secondary alignments
|
||||
var m, t = line.split("\t", 12);
|
||||
for (var i = 6; i <= 11; ++i)
|
||||
t[i] = parseInt(t[i]);
|
||||
if (t[10] < min_cov_len || t[11] < min_mapq) continue;
|
||||
var ctg = t[5], x = t[7], end = t[8];
|
||||
// compute regions covered by 1 contig
|
||||
if (ctg != c1_ctg || x >= c1_end) {
|
||||
if (c1_counted && c1_end > c1_start) {
|
||||
c1_len += c1_end - c1_start;
|
||||
print('R', c1_ctg, c1_start, c1_end);
|
||||
}
|
||||
c1_ctg = ctg, c1_start = x, c1_end = end;
|
||||
c1_counted = (t[10] >= min_var_len);
|
||||
} else if (end > c1_end) { // overlap
|
||||
if (c1_counted && x > c1_start) {
|
||||
c1_len += x - c1_start;
|
||||
print('R', c1_ctg, c1_start, x);
|
||||
}
|
||||
c1_start = c1_end, c1_end = end;
|
||||
c1_counted = (t[10] >= min_var_len);
|
||||
} else { // contained
|
||||
if (c1_counted && x > c1_start) {
|
||||
c1_len += x - c1_start;
|
||||
print('R', c1_ctg, c1_start, x);
|
||||
}
|
||||
c1_start = end;
|
||||
}
|
||||
// output variants ahead of this alignment
|
||||
while (out.length) {
|
||||
if (out[0][0] != ctg || out[0][2] <= x) {
|
||||
count_var(out[0]);
|
||||
print('V', out[0].join("\t"));
|
||||
out.shift();
|
||||
} else break;
|
||||
}
|
||||
// update coverage
|
||||
for (var i = 0; i < out.length; ++i)
|
||||
if (out[i][1] >= x && out[i][2] <= end)
|
||||
++out[i][3];
|
||||
// drop alignments that don't overlap with the current one
|
||||
var k = 0;
|
||||
for (var i = 0; i < a.length; ++i)
|
||||
if (a[0][0] == ctg && a[0][2] > x)
|
||||
a[k++] = a[i];
|
||||
a.length = k;
|
||||
// core loop
|
||||
if (t[10] >= min_var_len) {
|
||||
if ((m = /\tcs:Z:(\S+)/.exec(line)) == null) continue; // no cs tag
|
||||
var cs = m[1];
|
||||
var blen = 0, n_diff = 0;
|
||||
tot_len += t[10];
|
||||
while ((m = re_cs.exec(cs)) != null) {
|
||||
var cov = 1;
|
||||
if (m[1] == '*' || m[1] == '+' || m[1] == '-')
|
||||
for (var i = 0; i < a.length; ++i)
|
||||
if (a[0][2] > x) ++cov;
|
||||
if (m[1] == '=' || m[1] == ':') {
|
||||
var l = m[1] == '='? m[2].length : parseInt(m[2]);
|
||||
x += l, blen += l;
|
||||
} else if (m[1] == '*') {
|
||||
out.push([t[5], x, x+1, cov, t[11], m[2].charAt(0), m[2].charAt(1)]);
|
||||
++x, ++blen, ++n_diff;
|
||||
} else if (m[1] == '+') {
|
||||
out.push([t[5], x, x, cov, t[11], '-', m[2]]);
|
||||
++blen, ++n_diff;
|
||||
} else if (m[1] == '-') {
|
||||
out.push([t[5], x, x + m[2].length, cov, t[11], m[2], '-']);
|
||||
x += m[2].length, ++blen, ++n_diff;
|
||||
}
|
||||
}
|
||||
}
|
||||
a.push([t[5], t[7], t[8]]);
|
||||
}
|
||||
if (c1_counted && c1_end > c1_start) {
|
||||
c1_len += c1_end - c1_start;
|
||||
print('R', c1_ctg, c1_start, c1_end);
|
||||
}
|
||||
while (out.length) {
|
||||
count_var(out[0]);
|
||||
print('V', out[0].join("\t"));
|
||||
out.shift();
|
||||
}
|
||||
|
||||
//warn(tot_len + " alignment columns considered in calling");
|
||||
warn(c1_len + " reference bases covered by exactly one contig");
|
||||
warn(n_sub[0] + " substitutions; ts/tv = " + (n_sub[1]/n_sub[2]).toFixed(3));
|
||||
warn(n_del[0] + " 1bp deletions");
|
||||
warn(n_ins[0] + " 1bp insertions");
|
||||
warn(n_del[1] + " 2bp deletions");
|
||||
warn(n_ins[1] + " 2bp insertions");
|
||||
warn(n_del[2] + " [3,"+gap_thres+") deletions");
|
||||
warn(n_ins[2] + " [3,"+gap_thres+") insertions");
|
||||
warn(n_del[3] + " >="+gap_thres+" deletions");
|
||||
warn(n_ins[3] + " >="+gap_thres+" insertions");
|
||||
|
||||
buf.destroy();
|
||||
file.close();
|
||||
Executable
+3126
File diff suppressed because it is too large
Load Diff
-114
@@ -1,114 +0,0 @@
|
||||
var getopt = function(args, ostr) {
|
||||
var oli; // option letter list index
|
||||
if (typeof(getopt.place) == 'undefined')
|
||||
getopt.ind = 0, getopt.arg = null, getopt.place = -1;
|
||||
if (getopt.place == -1) { // update scanning pointer
|
||||
if (getopt.ind >= args.length || args[getopt.ind].charAt(getopt.place = 0) != '-') {
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
if (getopt.place + 1 < args[getopt.ind].length && args[getopt.ind].charAt(++getopt.place) == '-') { // found "--"
|
||||
++getopt.ind;
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
}
|
||||
var optopt = args[getopt.ind].charAt(getopt.place++); // character checked for validity
|
||||
if (optopt == ':' || (oli = ostr.indexOf(optopt)) < 0) {
|
||||
if (optopt == '-') return null; // if the user didn't specify '-' as an option, assume it means null.
|
||||
if (getopt.place < 0) ++getopt.ind;
|
||||
return '?';
|
||||
}
|
||||
if (oli+1 >= ostr.length || ostr.charAt(++oli) != ':') { // don't need argument
|
||||
getopt.arg = null;
|
||||
if (getopt.place < 0 || getopt.place >= args[getopt.ind].length) ++getopt.ind, getopt.place = -1;
|
||||
} else { // need an argument
|
||||
if (getopt.place >= 0 && getopt.place < args[getopt.ind].length)
|
||||
getopt.arg = args[getopt.ind].substr(getopt.place);
|
||||
else if (args.length <= ++getopt.ind) { // no arg
|
||||
getopt.place = -1;
|
||||
if (ostr.length > 0 && ostr.charAt(0) == ':') return ':';
|
||||
return '?';
|
||||
} else getopt.arg = args[getopt.ind]; // white space
|
||||
getopt.place = -1;
|
||||
++getopt.ind;
|
||||
}
|
||||
return optopt;
|
||||
}
|
||||
|
||||
var c, pri_only = false;
|
||||
while ((c = getopt(arguments, "p")) != null)
|
||||
if (c == 'p') pri_only = true;
|
||||
|
||||
var file = arguments.length == getopt.ind? new File() : new File(arguments[getopt.ind]);
|
||||
var buf = new Bytes();
|
||||
var re = /(\d+)([MIDSHNX=])/g;
|
||||
|
||||
var len = {}, lineno = 0;
|
||||
while (file.readline(buf) >= 0) {
|
||||
var m, n_cigar = 0, line = buf.toString();
|
||||
++lineno;
|
||||
if (line.charAt(0) == '@') {
|
||||
if (/^@SQ/.test(line)) {
|
||||
var name = (m = /\tSN:(\S+)/.exec(line)) != null? m[1] : null;
|
||||
var l = (m = /\tLN:(\d+)/.exec(line)) != null? parseInt(m[1]) : null;
|
||||
if (name != null && l != null) len[name] = l;
|
||||
}
|
||||
continue;
|
||||
}
|
||||
var t = line.split("\t");
|
||||
var flag = parseInt(t[1]);
|
||||
if (t[9] != '*' && t[10] != '*' && t[9].length != t[10].length) throw Error("ERROR at line " + lineno + ": inconsistent SEQ and QUAL lengths - " + t[9].length + " != " + t[10].length);
|
||||
if (t[2] == '*' || (flag&4)) continue;
|
||||
if (pri_only && (flag&0x100)) continue;
|
||||
var tlen = len[t[2]];
|
||||
if (tlen == null) throw Error("ERROR at line " + lineno + ": can't find the length of contig " + t[2]);
|
||||
var nn = (m = /\tnn:i:(\d+)/.exec(line)) != null? parseInt(m[1]) : 0;
|
||||
var NM = (m = /\tNM:i:(\d+)/.exec(line)) != null? parseInt(m[1]) : null;
|
||||
var have_NM = NM == null? false : true;
|
||||
NM += nn;
|
||||
var clip = [0, 0], I = [0, 0], D = [0, 0], M = 0, N = 0, ql = 0, tl = 0, mm = 0, ext_cigar = false;
|
||||
while ((m = re.exec(t[5])) != null) {
|
||||
var l = parseInt(m[1]);
|
||||
if (m[2] == 'M') M += l, ql += l, tl += l, ext_cigar = false;
|
||||
else if (m[2] == 'I') ++I[0], I[1] += l, ql += l;
|
||||
else if (m[2] == 'D') ++D[0], D[1] += l, tl += l;
|
||||
else if (m[2] == 'N') N += l, tl += l;
|
||||
else if (m[2] == 'S') clip[M == 0? 0 : 1] = l, ql += l;
|
||||
else if (m[2] == 'H') clip[M == 0? 0 : 1] = l;
|
||||
else if (m[2] == '=') M += l, ql += l, tl += l, ext_cigar = true;
|
||||
else if (m[2] == 'X') M += l, ql += l, tl += l, mm += l, ext_cigar = true;
|
||||
++n_cigar;
|
||||
}
|
||||
if (n_cigar > 65535)
|
||||
warn("WARNING at line " + lineno + ": " + n_cigar + " CIGAR operations");
|
||||
if (tl + parseInt(t[3]) - 1 > tlen) {
|
||||
warn("WARNING at line " + lineno + ": alignment end position larger than ref length; skipped");
|
||||
continue;
|
||||
}
|
||||
if (t[9] != '*' && t[9].length != ql) {
|
||||
warn("WARNING at line " + lineno + ": SEQ length inconsistent with CIGAR (" + t[9].length + " != " + ql + "); skipped");
|
||||
continue;
|
||||
}
|
||||
if (!have_NM || ext_cigar) NM = I[1] + D[1] + mm;
|
||||
if (NM < I[1] + D[1] + mm) {
|
||||
warn("WARNING at line " + lineno + ": NM is less than the total number of gaps (" + NM + " < " + (I[1]+D[1]+mm) + ")");
|
||||
NM = I[1] + D[1] + mm;
|
||||
}
|
||||
var extra = ["mm:i:"+(NM-I[1]-D[1]), "io:i:"+I[0], "in:i:"+I[1], "do:i:"+D[0], "dn:i:"+D[1]];
|
||||
var match = M - (NM - I[1] - D[1]);
|
||||
var blen = M + I[1] + D[1];
|
||||
var qlen = M + I[1] + clip[0] + clip[1];
|
||||
var qs, qe;
|
||||
if (flag&16) qs = clip[1], qe = qlen - clip[0];
|
||||
else qs = clip[0], qe = qlen - clip[1];
|
||||
var ts = parseInt(t[3]) - 1, te = ts + M + D[1] + N;
|
||||
var qname = t[0];
|
||||
if ((flag&1) && (flag&0x40)) qname += '/1';
|
||||
if ((flag&1) && (flag&0x80)) qname += '/2';
|
||||
var a = [qname, qlen, qs, qe, flag&16? '-' : '+', t[2], tlen, ts, te, match, blen, t[4]];
|
||||
print(a.join("\t"), extra.join("\t"));
|
||||
}
|
||||
|
||||
buf.destroy();
|
||||
file.close();
|
||||
@@ -1,193 +0,0 @@
|
||||
var getopt = function(args, ostr) {
|
||||
var oli; // option letter list index
|
||||
if (typeof(getopt.place) == 'undefined')
|
||||
getopt.ind = 0, getopt.arg = null, getopt.place = -1;
|
||||
if (getopt.place == -1) { // update scanning pointer
|
||||
if (getopt.ind >= args.length || args[getopt.ind].charAt(getopt.place = 0) != '-') {
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
if (getopt.place + 1 < args[getopt.ind].length && args[getopt.ind].charAt(++getopt.place) == '-') { // found "--"
|
||||
++getopt.ind;
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
}
|
||||
var optopt = args[getopt.ind].charAt(getopt.place++); // character checked for validity
|
||||
if (optopt == ':' || (oli = ostr.indexOf(optopt)) < 0) {
|
||||
if (optopt == '-') return null; // if the user didn't specify '-' as an option, assume it means null.
|
||||
if (getopt.place < 0) ++getopt.ind;
|
||||
return '?';
|
||||
}
|
||||
if (oli+1 >= ostr.length || ostr.charAt(++oli) != ':') { // don't need argument
|
||||
getopt.arg = null;
|
||||
if (getopt.place < 0 || getopt.place >= args[getopt.ind].length) ++getopt.ind, getopt.place = -1;
|
||||
} else { // need an argument
|
||||
if (getopt.place >= 0 && getopt.place < args[getopt.ind].length)
|
||||
getopt.arg = args[getopt.ind].substr(getopt.place);
|
||||
else if (args.length <= ++getopt.ind) { // no arg
|
||||
getopt.place = -1;
|
||||
if (ostr.length > 0 && ostr.charAt(0) == ':') return ':';
|
||||
return '?';
|
||||
} else getopt.arg = args[getopt.ind]; // white space
|
||||
getopt.place = -1;
|
||||
++getopt.ind;
|
||||
}
|
||||
return optopt;
|
||||
}
|
||||
|
||||
var c, max_mapq = 60, mode = 0, err_out_q = 256, print_err = false, ovlp_ratio = 0.1, cap_short_mapq = false;
|
||||
while ((c = getopt(arguments, "Q:r:m:c")) != null) {
|
||||
if (c == 'Q') err_out_q = parseInt(getopt.arg), print_err = true;
|
||||
else if (c == 'r') ovlp_ratio = parseFloat(getopt.arg);
|
||||
else if (c == 'm') mode = parseInt(getopt.arg);
|
||||
else if (c == 'c') cap_short_mapq = true;
|
||||
}
|
||||
|
||||
var file = arguments.length == getopt.ind? new File() : new File(arguments[getopt.ind]);
|
||||
var buf = new Bytes();
|
||||
|
||||
var tot = [], err = [];
|
||||
for (var q = 0; q <= max_mapq; ++q)
|
||||
tot[q] = err[q] = 0;
|
||||
|
||||
function is_correct(s, b)
|
||||
{
|
||||
if (s[0] != b[0] || s[3] != b[3]) return false;
|
||||
var o, l;
|
||||
if (s[1] < b[1]) {
|
||||
if (s[2] <= b[1]) return false;
|
||||
o = (s[2] < b[2]? s[2] : b[2]) - b[1];
|
||||
l = (s[2] > b[2]? s[2] : b[2]) - s[1];
|
||||
} else {
|
||||
if (b[2] <= s[1]) return false;
|
||||
o = (s[2] < b[2]? s[2] : b[2]) - s[1];
|
||||
l = (s[2] > b[2]? s[2] : b[2]) - b[1];
|
||||
}
|
||||
return o/l > ovlp_ratio? true : false;
|
||||
}
|
||||
|
||||
function count_err(qname, a, tot, err, mode)
|
||||
{
|
||||
if (a.length == 0) return;
|
||||
|
||||
var m, s;
|
||||
if ((m = /^(\S+)!(\S+)!(\d+)!(\d+)!([\+\-])$/.exec(qname)) != null) { // pbsim single-end reads
|
||||
s = [m[1], m[2], parseInt(m[3]), parseInt(m[4]), m[5]];
|
||||
} else if ((m = /^(\S+)!(\S+)!(\d+)_(\d+)!(\d+)_(\d+)!([\+\-])([\+\-])\/([12])$/.exec(qname)) != null) { // mason2 paired-end reads
|
||||
if (m[9] == '1') {
|
||||
s = [m[1], m[2], parseInt(m[3]), parseInt(m[5]), m[7]];
|
||||
} else {
|
||||
s = [m[1], m[2], parseInt(m[4]), parseInt(m[6]), m[8]];
|
||||
}
|
||||
} else throw Error("Failed to parse simulated read names '" + qname + "'");
|
||||
s.shift(); // skip the orginal read name
|
||||
|
||||
if (mode == 0 || mode == 1) { // longest only or first only
|
||||
var max_i = 0;
|
||||
if (mode == 0) { // longest only
|
||||
var max = 0;
|
||||
for (var i = 0; i < a.length; ++i)
|
||||
if (a[i][5] > max)
|
||||
max = a[i][5], max_i = i;
|
||||
}
|
||||
var mapq = a[max_i][4];
|
||||
++tot[mapq];
|
||||
if (!is_correct(s, a[max_i])) {
|
||||
if (mapq >= err_out_q)
|
||||
print('E', qname, a[max_i].join("\t"));
|
||||
++err[mapq];
|
||||
}
|
||||
} else if (mode == 2) { // all primary mode
|
||||
var max_err_mapq = -1, max_mapq = 0, max_err_i = -1;
|
||||
if (cap_short_mapq) {
|
||||
var max = 0, max_q = 0;
|
||||
for (var i = 0; i < a.length; ++i)
|
||||
if (a[i][5] > max)
|
||||
max = a[i][5], max_q = a[i][4];
|
||||
for (var i = 0; i < a.length; ++i)
|
||||
a[i][4] = max_q < a[i][4]? max_q : a[i][4];
|
||||
}
|
||||
for (var i = 0; i < a.length; ++i) {
|
||||
max_mapq = max_mapq > a[i][4]? max_mapq : a[i][4];
|
||||
if (!is_correct(s, a[i]))
|
||||
if (a[i][4] > max_err_mapq)
|
||||
max_err_mapq = a[i][4], max_err_i = i;
|
||||
}
|
||||
if (max_err_mapq >= 0) {
|
||||
++tot[max_err_mapq], ++err[max_err_mapq];
|
||||
if (max_err_mapq >= err_out_q)
|
||||
print('E', qname, a[max_err_i].join("\t"));
|
||||
} else ++tot[max_mapq];
|
||||
}
|
||||
}
|
||||
|
||||
var lineno = 0, last = null, a = [], n_unmapped = null;
|
||||
var re_cigar = /(\d+)([MIDSHN])/g;
|
||||
while (file.readline(buf) >= 0) {
|
||||
var m, line = buf.toString();
|
||||
++lineno;
|
||||
if (line[0] != '@') {
|
||||
var t = line.split("\t");
|
||||
if (t[4] == '+' || t[4] == '-') { // PAF
|
||||
if (last != t[0]) {
|
||||
if (last != null) count_err(last, a, tot, err, mode);
|
||||
a = [], last = t[0];
|
||||
}
|
||||
if (/\ts1:i:\d+/.test(line) && !/\ts2:i:\d+/.test(line)) // secondary alignment in minimap2 PAF
|
||||
continue;
|
||||
var mapq = parseInt(t[11]);
|
||||
if (mapq > max_mapq) mapq = max_mapq;
|
||||
a.push([t[5], parseInt(t[7]), parseInt(t[8]), t[4], mapq, parseInt(t[9])]);
|
||||
} else { // SAM
|
||||
var flag = parseInt(t[1]);
|
||||
var read_no = flag>>6&0x3;
|
||||
var qname = t[0];
|
||||
if (!/\/[12]$/.test(qname))
|
||||
qname = read_no == 1 || read_no == 2? t[0] + '/' + read_no : t[0];
|
||||
if (last != qname) {
|
||||
if (last != null) count_err(last, a, tot, err, mode);
|
||||
a = [], last = qname;
|
||||
}
|
||||
if (flag&0x100) continue; // secondary alignment
|
||||
if ((flag&0x4) || t[2] == '*') { // unmapped
|
||||
if (n_unmapped == null) n_unmapped = 0;
|
||||
++n_unmapped;
|
||||
continue;
|
||||
}
|
||||
var mapq = parseInt(t[4]);
|
||||
if (mapq > max_mapq) mapq = max_mapq;
|
||||
var pos = parseInt(t[3]) - 1, pos_end = pos;
|
||||
var n_gap = 0, mlen = 0;
|
||||
while ((m = re_cigar.exec(t[5])) != null) {
|
||||
var len = parseInt(m[1]);
|
||||
if (m[2] == 'M') pos_end += len, mlen += len;
|
||||
else if (m[2] == 'I') n_gap += len;
|
||||
else if (m[2] == 'D') n_gap += len, pos_end += len;
|
||||
}
|
||||
var score = pos_end - pos;
|
||||
if ((m = /\tNM:i:(\d+)/.exec(line)) != null) {
|
||||
var NM = parseInt(m[1]);
|
||||
if (NM >= n_gap) score = mlen - (NM - n_gap);
|
||||
}
|
||||
a.push([t[2], pos, pos_end, (flag&16)? '-' : '+', mapq, score]);
|
||||
}
|
||||
}
|
||||
}
|
||||
if (last != null) count_err(last, a, tot, err, mode);
|
||||
|
||||
buf.destroy();
|
||||
file.close();
|
||||
|
||||
var sum_tot = 0, sum_err = 0, q_out = -1, sum_tot2 = 0, sum_err2 = 0;
|
||||
for (var q = max_mapq; q >= 0; --q) {
|
||||
if (tot[q] == 0) continue;
|
||||
if (q_out < 0 || err[q] > 0) {
|
||||
if (q_out >= 0) print('Q', q_out, sum_tot, sum_err, (sum_err2/sum_tot2).toFixed(9), sum_tot2);
|
||||
sum_tot = sum_err = 0, q_out = q;
|
||||
}
|
||||
sum_tot += tot[q], sum_err += err[q];
|
||||
sum_tot2 += tot[q], sum_err2 += err[q];
|
||||
}
|
||||
print('Q', q_out, sum_tot, sum_err, (sum_err2/sum_tot2).toFixed(9), sum_tot2);
|
||||
if (n_unmapped != null) print('U', n_unmapped);
|
||||
@@ -1,105 +0,0 @@
|
||||
Bytes.prototype.reverse = function()
|
||||
{
|
||||
for (var i = 0; i < this.length>>1; ++i) {
|
||||
var tmp = this[i];
|
||||
this[i] = this[this.length - i - 1];
|
||||
this[this.length - i - 1] = tmp;
|
||||
}
|
||||
}
|
||||
|
||||
// reverse complement a DNA string
|
||||
Bytes.prototype.revcomp = function()
|
||||
{
|
||||
if (Bytes.rctab == null) {
|
||||
var s1 = 'WSATUGCYRKMBDHVNwsatugcyrkmbdhvn';
|
||||
var s2 = 'WSTAACGRYMKVHDBNwstaacgrymkvhdbn';
|
||||
Bytes.rctab = [];
|
||||
for (var i = 0; i < 256; ++i) Bytes.rctab[i] = 0;
|
||||
for (var i = 0; i < s1.length; ++i)
|
||||
Bytes.rctab[s1.charCodeAt(i)] = s2.charCodeAt(i);
|
||||
}
|
||||
for (var i = 0; i < this.length>>1; ++i) {
|
||||
var tmp = this[this.length - i - 1];
|
||||
this[this.length - i - 1] = Bytes.rctab[this[i]];
|
||||
this[i] = Bytes.rctab[tmp];
|
||||
}
|
||||
if (this.length&1)
|
||||
this[this.length>>1] = Bytes.rctab[this[this.length>>1]];
|
||||
}
|
||||
|
||||
if (arguments.length == 0) {
|
||||
print("Usage: k8 sim-mason2.js <mason.sam>");
|
||||
exit(1);
|
||||
}
|
||||
|
||||
function print_se(a)
|
||||
{
|
||||
print('@' + a.slice(0, 5).join("!") + " " + a[8]);
|
||||
print(a[5]);
|
||||
print("+");
|
||||
print(a[6]);
|
||||
}
|
||||
|
||||
var buf = new Bytes(), buf2 = new Bytes();
|
||||
var file = new File(arguments[0]);
|
||||
var re = /(\d+)([MIDSHN])/g;
|
||||
var last = null;
|
||||
while (file.readline(buf) >= 0) {
|
||||
var t = buf.toString().split("\t");
|
||||
if (t[0].charAt(0) == '@') continue;
|
||||
var m, l_ref = 0;
|
||||
while ((m = re.exec(t[5])) != null)
|
||||
if (m[2] == 'D' || m[2] == 'M' || m[2] == 'N')
|
||||
l_ref += parseInt(m[1]);
|
||||
var flag = parseInt(t[1]);
|
||||
var rev = !!(flag&16);
|
||||
var seq, qual;
|
||||
if (rev) {
|
||||
buf2.length = 0;
|
||||
buf2.set(t[9], 0);
|
||||
buf2.revcomp();
|
||||
seq = buf2.toString();
|
||||
buf2.set(t[10], 0);
|
||||
buf2.reverse();
|
||||
qual = buf2.toString();
|
||||
} else seq = t[9], qual = t[10];
|
||||
var qname = t[0];
|
||||
qname = qname.replace(/^simulated./, "");
|
||||
var chr = t[2];
|
||||
var pos = parseInt(t[3]) - 1;
|
||||
var strand = (flag&16)? '-' : '+';
|
||||
var read_no = flag&0xc0;
|
||||
if (read_no == 0x40) read_no = 1;
|
||||
else if (read_no == 0x80) read_no = 2;
|
||||
else read_no = 0;
|
||||
var err = 0, snp = 0, indel = 0;
|
||||
for (var i = 11; i < t.length; ++i) {
|
||||
if ((m = /^XE:i:(\d+)/.exec(t[i])) != null) err = m[1];
|
||||
else if ((m = /^XS:i:(\d+)/.exec(t[i])) != null) snp = m[1];
|
||||
else if ((m = /^XI:i:(\d+)/.exec(t[i])) != null) indel = m[1];
|
||||
}
|
||||
var comment = [err, snp, indel].join(":");
|
||||
if (last == null) {
|
||||
last = [qname, chr, pos, pos + l_ref, strand, seq, qual, read_no, comment];
|
||||
} else if (last[0] != qname) {
|
||||
print_se(last);
|
||||
last = [qname, chr, pos, pos + l_ref, strand, seq, qual, read_no, comment];
|
||||
} else {
|
||||
if (read_no == 2) { // last[] is the first read
|
||||
if (last[7] != 1) throw Error("ERROR: can't find read1");
|
||||
var name = [qname, chr, last[2] + "_" + pos, last[3] + "_" + (pos + l_ref), last[4] + strand].join("!");
|
||||
print('@' + name + '/1' + ' ' + last[8]); print(last[5]); print("+"); print(last[6]);
|
||||
print('@' + name + '/2' + ' ' + comment); print(seq); print("+"); print(qual);
|
||||
} else {
|
||||
if (last[7] != 2) throw Error("ERROR: can't find read2");
|
||||
var name = [qname, chr, pos + "_" + last[2], (pos + l_ref) + "_" + last[3], strand + last[4]].join("!");
|
||||
print('@' + name + '/1' + ' ' + comment); print(seq); print("+"); print(qual);
|
||||
print('@' + name + '/2' + ' ' + last[8]); print(last[5]); print("+"); print(last[6]);
|
||||
}
|
||||
last = null;
|
||||
}
|
||||
}
|
||||
if (last != null) print_se(last);
|
||||
file.close();
|
||||
buf.destroy();
|
||||
buf2.destroy();
|
||||
@@ -1,81 +0,0 @@
|
||||
Bytes.prototype.reverse = function()
|
||||
{
|
||||
for (var i = 0; i < this.length>>1; ++i) {
|
||||
var tmp = this[i];
|
||||
this[i] = this[this.length - i - 1];
|
||||
this[this.length - i - 1] = tmp;
|
||||
}
|
||||
}
|
||||
|
||||
// reverse complement a DNA string
|
||||
Bytes.prototype.revcomp = function()
|
||||
{
|
||||
if (Bytes.rctab == null) {
|
||||
var s1 = 'WSATUGCYRKMBDHVNwsatugcyrkmbdhvn';
|
||||
var s2 = 'WSTAACGRYMKVHDBNwstaacgrymkvhdbn';
|
||||
Bytes.rctab = [];
|
||||
for (var i = 0; i < 256; ++i) Bytes.rctab[i] = 0;
|
||||
for (var i = 0; i < s1.length; ++i)
|
||||
Bytes.rctab[s1.charCodeAt(i)] = s2.charCodeAt(i);
|
||||
}
|
||||
for (var i = 0; i < this.length>>1; ++i) {
|
||||
var tmp = this[this.length - i - 1];
|
||||
this[this.length - i - 1] = Bytes.rctab[this[i]];
|
||||
this[i] = Bytes.rctab[tmp];
|
||||
}
|
||||
if (this.length&1)
|
||||
this[this.length>>1] = Bytes.rctab[this[this.length>>1]];
|
||||
}
|
||||
|
||||
if (arguments.length < 2) {
|
||||
print("Usage: k8 sim-pbsim.js <ref.fa.fai> <pbsim1.maf> [[pbsim2.maf] ...]");
|
||||
exit(1);
|
||||
}
|
||||
|
||||
var file, buf = new Bytes(), buf2 = new Bytes();
|
||||
file = new File(arguments[0]);
|
||||
var chr_list = [];
|
||||
while (file.readline(buf) >= 0) {
|
||||
var t = buf.toString().split(/\s+/);
|
||||
chr_list.push(t[0]);
|
||||
}
|
||||
file.close();
|
||||
|
||||
for (var k = 1; k < arguments.length; ++k) {
|
||||
var fn = arguments[k];
|
||||
file = new File(fn);
|
||||
var state = 0, reg;
|
||||
while (file.readline(buf) >= 0) {
|
||||
var line = buf.toString();
|
||||
if (state == 0 && line.charAt(0) == 'a') {
|
||||
state = 1;
|
||||
} else if (state == 1 && line.charAt(0) == 's') {
|
||||
var t = line.split(/\s+/);
|
||||
var st = parseInt(t[2]);
|
||||
reg = [st, st + parseInt(t[3])];
|
||||
state = 2;
|
||||
} else if (state == 2 && line.charAt(0) == 's') {
|
||||
var m, t = line.split(/\s+/);
|
||||
if ((m = /S(\d+)_\d+/.exec(t[1])) == null) throw Error("Failed to parse the read name");
|
||||
var chr_id = parseInt(m[1]) - 1;
|
||||
if (chr_id >= chr_list.length) throw Error("Index outside the chr list");
|
||||
var name = [t[1], chr_list[chr_id], reg[0], reg[1], t[4]].join("!");
|
||||
var seq = t[6].replace(/\-/g, "");
|
||||
if (seq.length != parseInt(t[5])) throw Error("Inconsistent read length");
|
||||
if (seq.indexOf("NN") < 0) {
|
||||
if (t[4] == '-') {
|
||||
buf2.set(seq, 0);
|
||||
buf2.length = seq.length;
|
||||
buf2.revcomp();
|
||||
seq = buf2.toString();
|
||||
}
|
||||
print(">" + name);
|
||||
print(seq);
|
||||
}
|
||||
state = 0;
|
||||
}
|
||||
}
|
||||
file.close();
|
||||
}
|
||||
buf.destroy();
|
||||
buf2.destroy();
|
||||
@@ -1,148 +0,0 @@
|
||||
var getopt = function(args, ostr) {
|
||||
var oli; // option letter list index
|
||||
if (typeof(getopt.place) == 'undefined')
|
||||
getopt.ind = 0, getopt.arg = null, getopt.place = -1;
|
||||
if (getopt.place == -1) { // update scanning pointer
|
||||
if (getopt.ind >= args.length || args[getopt.ind].charAt(getopt.place = 0) != '-') {
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
if (getopt.place + 1 < args[getopt.ind].length && args[getopt.ind].charAt(++getopt.place) == '-') { // found "--"
|
||||
++getopt.ind;
|
||||
getopt.place = -1;
|
||||
return null;
|
||||
}
|
||||
}
|
||||
var optopt = args[getopt.ind].charAt(getopt.place++); // character checked for validity
|
||||
if (optopt == ':' || (oli = ostr.indexOf(optopt)) < 0) {
|
||||
if (optopt == '-') return null; // if the user didn't specify '-' as an option, assume it means null.
|
||||
if (getopt.place < 0) ++getopt.ind;
|
||||
return '?';
|
||||
}
|
||||
if (oli+1 >= ostr.length || ostr.charAt(++oli) != ':') { // don't need argument
|
||||
getopt.arg = null;
|
||||
if (getopt.place < 0 || getopt.place >= args[getopt.ind].length) ++getopt.ind, getopt.place = -1;
|
||||
} else { // need an argument
|
||||
if (getopt.place >= 0 && getopt.place < args[getopt.ind].length)
|
||||
getopt.arg = args[getopt.ind].substr(getopt.place);
|
||||
else if (args.length <= ++getopt.ind) { // no arg
|
||||
getopt.place = -1;
|
||||
if (ostr.length > 0 && ostr.charAt(0) == ':') return ':';
|
||||
return '?';
|
||||
} else getopt.arg = args[getopt.ind]; // white space
|
||||
getopt.place = -1;
|
||||
++getopt.ind;
|
||||
}
|
||||
return optopt;
|
||||
}
|
||||
|
||||
var colors = ["0,128,255", "255,0,0", "0,192,0"];
|
||||
|
||||
function print_lines(a, fmt) {
|
||||
if (a.length == 0) return;
|
||||
if (fmt == "bed") {
|
||||
var n_pri = 0;
|
||||
for (var i = 0; i < a.length; ++i)
|
||||
if (a[i][8] == 0) ++n_pri;
|
||||
if (n_pri > 1) {
|
||||
for (var i = 0; i < a.length; ++i)
|
||||
if (a[i][8] == 0) a[i][8] = 1;
|
||||
} else if (n_pri == 0) {
|
||||
warn("Warning: " + a[0][3] + " doesn't have a primary alignment");
|
||||
}
|
||||
for (var i = 0; i < a.length; ++i) {
|
||||
a[i][8] = colors[a[i][8]];
|
||||
print(a[i].join("\t"));
|
||||
}
|
||||
}
|
||||
a.length = 0;
|
||||
}
|
||||
|
||||
function main(args) {
|
||||
var re = /(\d+)([MIDNSH])/g;
|
||||
var c, fmt = "bed", fn_name_conv = null;
|
||||
while ((c = getopt(args, "f:n:")) != null) {
|
||||
if (c == 'f') fmt = getopt.arg;
|
||||
else if (c == 'n') fn_name_conv = getopt.arg;
|
||||
}
|
||||
if (getopt.ind == args.length) {
|
||||
warn("Usage: k8 splice2bed.js <in.paf>");
|
||||
exit(1);
|
||||
}
|
||||
|
||||
var conv = null;
|
||||
if (fn_name_conv != null) {
|
||||
conv = new Map();
|
||||
var file = new File(fn_name_conv);
|
||||
var buf = new Bytes();
|
||||
while (file.readline(buf) >= 0) {
|
||||
var t = buf.toString().split("\t");
|
||||
conv.put(t[0], t[1]);
|
||||
}
|
||||
buf.destroy();
|
||||
file.close();
|
||||
}
|
||||
|
||||
var file = new File(args[getopt.ind]);
|
||||
var buf = new Bytes();
|
||||
var a = [];
|
||||
while (file.readline(buf) >= 0) {
|
||||
var line = buf.toString();
|
||||
if (line.charAt(0) == '@') continue; // skip SAM header lines
|
||||
var t = line.split("\t");
|
||||
var is_pri = false, cigar = null, a1;
|
||||
var qname = conv != null? conv.get(t[0]) : null;
|
||||
if (qname != null) t[0] = qname;
|
||||
if (t.length >= 10 && t[4] != '+' && t[4] != '-' && /^\d+/.test(t[1])) { // SAM
|
||||
var flag = parseInt(t[1]);
|
||||
if (flag&1) t[0] += '/' + (flag>>6&3);
|
||||
}
|
||||
if (a.length && a[0][3] != t[0]) {
|
||||
print_lines(a, fmt);
|
||||
a = [];
|
||||
}
|
||||
if (t.length >= 12 && (t[4] == '+' || t[4] == '-')) { // PAF
|
||||
for (var i = 12; i < t.length; ++i) {
|
||||
if (t[i].substr(0, 5) == 'cg:Z:') {
|
||||
cigar = t[i].substr(5);
|
||||
} else if (t[i].substr(0, 5) == 's2:i:') {
|
||||
is_pri = true;
|
||||
}
|
||||
}
|
||||
a1 = [t[5], t[7], t[8], t[0], Math.floor(t[9]/t[10]*1000), t[4]];
|
||||
} else if (t.length >= 10) { // SAM
|
||||
var flag = parseInt(t[1]);
|
||||
if ((flag&4) || a[2] == '*') continue;
|
||||
cigar = t[5];
|
||||
is_pri = (flag&0x100)? false : true;
|
||||
a1 = [t[2], parseInt(t[3])-1, null, t[0], 1000, (flag&16)? '-' : '+'];
|
||||
} else {
|
||||
throw Error("unrecognized input format");
|
||||
}
|
||||
if (cigar == null) throw Error("missing CIGAR");
|
||||
var m, x0 = 0, x = 0, bs = [], bl = [];
|
||||
while ((m = re.exec(cigar)) != null) {
|
||||
if (m[2] == 'M' || m[2] == 'D') {
|
||||
x += parseInt(m[1]);
|
||||
} else if (m[2] == 'N') {
|
||||
bs.push(x0);
|
||||
bl.push(x - x0);
|
||||
x += parseInt(m[1]);
|
||||
x0 = x;
|
||||
}
|
||||
}
|
||||
bs.push(x0);
|
||||
bl.push(x - x0);
|
||||
// write the BED12 line
|
||||
if (a1[2] == null) a1[2] = a1[1] + x;
|
||||
a1.push(a1[1], a1[2]); // thick start/end is the same as start/end
|
||||
a1.push(is_pri? 0 : 2, bs.length, bl.join(",")+",", bs.join(",")+",");
|
||||
a.push(a1);
|
||||
}
|
||||
print_lines(a, fmt);
|
||||
buf.destroy();
|
||||
file.close();
|
||||
if (conv != null) conv.destroy();
|
||||
}
|
||||
|
||||
main(arguments);
|
||||
@@ -4,6 +4,7 @@
|
||||
#include <assert.h>
|
||||
#include "minimap.h"
|
||||
#include "bseq.h"
|
||||
#include "kseq.h"
|
||||
|
||||
#define MM_PARENT_UNSET (-1)
|
||||
#define MM_PARENT_TMP_PRI (-2)
|
||||
@@ -28,17 +29,20 @@
|
||||
#define mm_seq4_set(s, i, c) ((s)[(i)>>3] |= (uint32_t)(c) << (((i)&7)<<2))
|
||||
#define mm_seq4_get(s, i) ((s)[(i)>>3] >> (((i)&7)<<2) & 0xf)
|
||||
|
||||
#define MALLOC(type, len) ((type*)malloc((len) * sizeof(type)))
|
||||
#define CALLOC(type, len) ((type*)calloc((len), sizeof(type)))
|
||||
|
||||
#ifdef __cplusplus
|
||||
extern "C" {
|
||||
#endif
|
||||
|
||||
#ifndef KSTRING_T
|
||||
#define KSTRING_T kstring_t
|
||||
typedef struct __kstring_t {
|
||||
unsigned l, m;
|
||||
char *s;
|
||||
} kstring_t;
|
||||
#endif
|
||||
typedef struct {
|
||||
uint32_t n;
|
||||
uint32_t q_pos;
|
||||
uint32_t q_span:31, flt:1;
|
||||
uint32_t seg_id:31, is_tandem:1;
|
||||
const uint64_t *cr;
|
||||
} mm_seed_t;
|
||||
|
||||
typedef struct {
|
||||
int n_u, n_a;
|
||||
@@ -48,6 +52,7 @@ typedef struct {
|
||||
|
||||
double cputime(void);
|
||||
double realtime(void);
|
||||
long peakrss(void);
|
||||
|
||||
void radix_sort_128x(mm128_t *beg, mm128_t *end);
|
||||
void radix_sort_64(uint64_t *beg, uint64_t *end);
|
||||
@@ -55,30 +60,42 @@ uint32_t ks_ksmall_uint32_t(size_t n, uint32_t arr[], size_t kk);
|
||||
|
||||
void mm_sketch(void *km, const char *str, int len, int w, int k, uint32_t rid, int is_hpc, mm128_v *p);
|
||||
|
||||
void mm_write_sam_hdr(const mm_idx_t *mi, const char *rg, const char *ver, int argc, char *argv[]);
|
||||
void mm_write_paf(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, const mm_reg1_t *r, void *km, int opt_flag);
|
||||
mm_seed_t *mm_collect_matches(void *km, int *_n_m, int qlen, int max_occ, int max_max_occ, int dist, const mm_idx_t *mi, const mm128_v *mv, int64_t *n_a, int *rep_len, int *n_mini_pos, uint64_t **mini_pos);
|
||||
|
||||
double mm_event_identity(const mm_reg1_t *r);
|
||||
int mm_write_sam_hdr(const mm_idx_t *mi, const char *rg, const char *ver, int argc, char *argv[]);
|
||||
void mm_write_paf(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, const mm_reg1_t *r, void *km, int64_t opt_flag);
|
||||
void mm_write_paf3(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, const mm_reg1_t *r, void *km, int64_t opt_flag, int rep_len);
|
||||
void mm_write_sam(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, const mm_reg1_t *r, int n_regs, const mm_reg1_t *regs);
|
||||
void mm_write_sam2(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, int seg_idx, int reg_idx, int n_seg, const int *n_regs, const mm_reg1_t *const* regs, void *km, int opt_flag);
|
||||
void mm_write_sam2(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, int seg_idx, int reg_idx, int n_seg, const int *n_regs, const mm_reg1_t *const* regs, void *km, int64_t opt_flag);
|
||||
void mm_write_sam3(kstring_t *s, const mm_idx_t *mi, const mm_bseq1_t *t, int seg_idx, int reg_idx, int n_seg, const int *n_regss, const mm_reg1_t *const* regss, void *km, int64_t opt_flag, int rep_len);
|
||||
|
||||
void mm_idxopt_init(mm_idxopt_t *opt);
|
||||
const uint64_t *mm_idx_get(const mm_idx_t *mi, uint64_t minier, int *n);
|
||||
int mm_idx_getseq(const mm_idx_t *mi, uint32_t rid, uint32_t st, uint32_t en, uint8_t *seq);
|
||||
int32_t mm_idx_cal_max_occ(const mm_idx_t *mi, float f);
|
||||
mm128_t *mm_chain_dp(int max_dist_x, int max_dist_y, int bw, int max_skip, int min_cnt, int min_sc, int is_cdna, int n_segs, int64_t n, mm128_t *a, int *n_u_, uint64_t **_u, void *km);
|
||||
int mm_idx_getseq2(const mm_idx_t *mi, int is_rev, uint32_t rid, uint32_t st, uint32_t en, uint8_t *seq);
|
||||
mm_reg1_t *mm_align_skeleton(void *km, const mm_mapopt_t *opt, const mm_idx_t *mi, int qlen, const char *qstr, int *n_regs_, mm_reg1_t *regs, mm128_t *a);
|
||||
mm_reg1_t *mm_gen_regs(void *km, uint32_t hash, int qlen, int n_u, uint64_t *u, mm128_t *a, int is_qstrand);
|
||||
|
||||
mm_reg1_t *mm_gen_regs(void *km, uint32_t hash, int qlen, int n_u, uint64_t *u, mm128_t *a);
|
||||
void mm_split_reg(mm_reg1_t *r, mm_reg1_t *r2, int n, int qlen, mm128_t *a);
|
||||
mm128_t *mm_chain_dp(int max_dist_x, int max_dist_y, int bw, int max_skip, int max_iter, int min_cnt, int min_sc, float gap_scale,
|
||||
int is_cdna, int n_segs, int64_t n, mm128_t *a, int *n_u_, uint64_t **_u, void *km);
|
||||
mm128_t *mg_lchain_dp(int max_dist_x, int max_dist_y, int bw, int max_skip, int max_iter, int min_cnt, int min_sc, float chn_pen_gap, float chn_pen_skip,
|
||||
int is_cdna, int n_segs, int64_t n, mm128_t *a, int *n_u_, uint64_t **_u, void *km);
|
||||
mm128_t *mg_lchain_rmq(int max_dist, int max_dist_inner, int bw, int max_chn_skip, int cap_rmq_size, int min_cnt, int min_sc, float chn_pen_gap, float chn_pen_skip,
|
||||
int64_t n, mm128_t *a, int *n_u_, uint64_t **_u, void *km);
|
||||
|
||||
void mm_mark_alt(const mm_idx_t *mi, int n, mm_reg1_t *r);
|
||||
void mm_split_reg(mm_reg1_t *r, mm_reg1_t *r2, int n, int qlen, mm128_t *a, int is_qstrand);
|
||||
void mm_sync_regs(void *km, int n_regs, mm_reg1_t *regs);
|
||||
int mm_squeeze_a(void *km, int n_regs, mm_reg1_t *regs, mm128_t *a);
|
||||
int mm_set_sam_pri(int n, mm_reg1_t *r);
|
||||
void mm_set_parent(void *km, float mask_level, int n, mm_reg1_t *r, int sub_diff);
|
||||
void mm_set_parent(void *km, float mask_level, int mask_len, int n, mm_reg1_t *r, int sub_diff, int hard_mask_level, float alt_diff_frac);
|
||||
void mm_select_sub(void *km, float pri_ratio, int min_diff, int best_n, int *n_, mm_reg1_t *r);
|
||||
void mm_select_sub_multi(void *km, float pri_ratio, float pri1, float pri2, int max_gap_ref, int min_diff, int best_n, int n_segs, const int *qlens, int *n_, mm_reg1_t *r);
|
||||
void mm_filter_regs(void *km, const mm_mapopt_t *opt, int *n_regs, mm_reg1_t *regs);
|
||||
void mm_join_long(void *km, const mm_mapopt_t *opt, int qlen, int *n_regs, mm_reg1_t *regs, mm128_t *a);
|
||||
void mm_hit_sort_by_dp(void *km, int *n_regs, mm_reg1_t *r);
|
||||
void mm_set_mapq(int n_regs, mm_reg1_t *regs, int min_chain_sc, int match_sc, int rep_len, int is_sr);
|
||||
void mm_filter_regs(const mm_mapopt_t *opt, int qlen, int *n_regs, mm_reg1_t *regs);
|
||||
void mm_hit_sort(void *km, int *n_regs, mm_reg1_t *r, float alt_diff_frac);
|
||||
void mm_set_mapq(void *km, int n_regs, mm_reg1_t *regs, int min_chain_sc, int match_sc, int rep_len, int is_sr);
|
||||
void mm_update_dp_max(int qlen, int n_regs, mm_reg1_t *regs, float frac, int a, int b);
|
||||
|
||||
void mm_est_err(const mm_idx_t *mi, int qlen, int n_regs, mm_reg1_t *regs, const mm128_t *a, int32_t n, const uint64_t *mini_pos);
|
||||
|
||||
@@ -86,6 +103,25 @@ mm_seg_t *mm_seg_gen(void *km, uint32_t hash, int n_segs, const int *qlens, int
|
||||
void mm_seg_free(void *km, int n_segs, mm_seg_t *segs);
|
||||
void mm_pair(void *km, int max_gap_ref, int dp_bonus, int sub_diff, int match_sc, const int *qlens, int *n_regs, mm_reg1_t **regs);
|
||||
|
||||
FILE *mm_split_init(const char *prefix, const mm_idx_t *mi);
|
||||
mm_idx_t *mm_split_merge_prep(const char *prefix, int n_splits, FILE **fp, uint32_t *n_seq_part);
|
||||
int mm_split_merge(int n_segs, const char **fn, const mm_mapopt_t *opt, int n_split_idx);
|
||||
void mm_split_rm_tmp(const char *prefix, int n_splits);
|
||||
|
||||
void mm_err_puts(const char *str);
|
||||
void mm_err_fwrite(const void *p, size_t size, size_t nitems, FILE *fp);
|
||||
void mm_err_fread(void *p, size_t size, size_t nitems, FILE *fp);
|
||||
|
||||
static inline float mg_log2(float x) // NB: this doesn't work when x<2
|
||||
{
|
||||
union { float f; uint32_t i; } z = { x };
|
||||
float log_2 = ((z.i >> 23) & 255) - 128;
|
||||
z.i &= ~(255 << 23);
|
||||
z.i += 127 << 23;
|
||||
log_2 += (-0.34484843f * z.f + 2.02466578f) * z.f - 0.67487759f;
|
||||
return log_2;
|
||||
}
|
||||
|
||||
#ifdef __cplusplus
|
||||
}
|
||||
#endif
|
||||
|
||||
@@ -1,35 +1,60 @@
|
||||
#include <stdio.h>
|
||||
#include <limits.h>
|
||||
#include "mmpriv.h"
|
||||
|
||||
void mm_idxopt_init(mm_idxopt_t *opt)
|
||||
{
|
||||
memset(opt, 0, sizeof(mm_idxopt_t));
|
||||
opt->k = 15, opt->w = 10, opt->flag = 0;
|
||||
opt->bucket_bits = 14;
|
||||
opt->mini_batch_size = 50000000;
|
||||
opt->batch_size = 4000000000ULL;
|
||||
}
|
||||
|
||||
void mm_mapopt_init(mm_mapopt_t *opt)
|
||||
{
|
||||
memset(opt, 0, sizeof(mm_mapopt_t));
|
||||
opt->seed = 11;
|
||||
opt->mid_occ_frac = 2e-4f;
|
||||
opt->min_mid_occ = 10;
|
||||
opt->max_mid_occ = 1000000;
|
||||
opt->sdust_thres = 0; // no SDUST masking
|
||||
|
||||
opt->min_cnt = 3;
|
||||
opt->min_chain_score = 40;
|
||||
opt->bw = 500;
|
||||
opt->bw = 500, opt->bw_long = 20000;
|
||||
opt->max_gap = 5000;
|
||||
opt->max_gap_ref = -1;
|
||||
opt->max_chain_skip = 25;
|
||||
opt->max_chain_iter = 5000;
|
||||
opt->rmq_inner_dist = 1000;
|
||||
opt->rmq_size_cap = 100000;
|
||||
opt->rmq_rescue_size = 1000;
|
||||
opt->rmq_rescue_ratio = 0.1f;
|
||||
opt->chain_gap_scale = 0.8f;
|
||||
opt->max_max_occ = 4095;
|
||||
opt->occ_dist = 500;
|
||||
|
||||
opt->mask_level = 0.5f;
|
||||
opt->mask_len = INT_MAX;
|
||||
opt->pri_ratio = 0.8f;
|
||||
opt->best_n = 5;
|
||||
|
||||
opt->max_join_long = 20000;
|
||||
opt->max_join_short = 2000;
|
||||
opt->min_join_flank_sc = 1000;
|
||||
opt->alt_drop = 0.15f;
|
||||
|
||||
opt->a = 2, opt->b = 4, opt->q = 4, opt->e = 2, opt->q2 = 24, opt->e2 = 1;
|
||||
opt->zdrop = 400;
|
||||
opt->sc_ambi = 1;
|
||||
opt->zdrop = 400, opt->zdrop_inv = 200;
|
||||
opt->end_bonus = -1;
|
||||
opt->min_dp_max = opt->min_chain_score * opt->a;
|
||||
opt->min_ksw_len = 200;
|
||||
opt->anchor_ext_len = 20, opt->anchor_ext_shift = 6;
|
||||
opt->max_clip_ratio = 1.0f;
|
||||
opt->mini_batch_size = 500000000;
|
||||
opt->max_sw_mat = 100000000;
|
||||
|
||||
opt->rank_min_len = 500;
|
||||
opt->rank_frac = 0.9f;
|
||||
|
||||
opt->pe_ori = 0; // FF
|
||||
opt->pe_bonus = 33;
|
||||
@@ -37,10 +62,15 @@ void mm_mapopt_init(mm_mapopt_t *opt)
|
||||
|
||||
void mm_mapopt_update(mm_mapopt_t *opt, const mm_idx_t *mi)
|
||||
{
|
||||
if ((opt->flag & MM_F_SPLICE_FOR) && (opt->flag & MM_F_SPLICE_REV))
|
||||
if ((opt->flag & MM_F_SPLICE_FOR) || (opt->flag & MM_F_SPLICE_REV))
|
||||
opt->flag |= MM_F_SPLICE;
|
||||
if (opt->mid_occ <= 0)
|
||||
if (opt->mid_occ <= 0) {
|
||||
opt->mid_occ = mm_idx_cal_max_occ(mi, opt->mid_occ_frac);
|
||||
if (opt->mid_occ < opt->min_mid_occ)
|
||||
opt->mid_occ = opt->min_mid_occ;
|
||||
if (opt->max_mid_occ > opt->min_mid_occ && opt->mid_occ > opt->max_mid_occ)
|
||||
opt->mid_occ = opt->max_mid_occ;
|
||||
}
|
||||
if (mm_verbose >= 3)
|
||||
fprintf(stderr, "[M::%s::%.3f*%.2f] mid_occ = %d\n", __func__, realtime() - mm_realtime0, cputime() / (realtime() - mm_realtime0), opt->mid_occ);
|
||||
}
|
||||
@@ -48,7 +78,7 @@ void mm_mapopt_update(mm_mapopt_t *opt, const mm_idx_t *mi)
|
||||
void mm_mapopt_max_intron_len(mm_mapopt_t *opt, int max_intron_len)
|
||||
{
|
||||
if ((opt->flag & MM_F_SPLICE) && max_intron_len > 0)
|
||||
opt->max_gap_ref = opt->bw = max_intron_len;
|
||||
opt->max_gap_ref = opt->bw = opt->bw_long = max_intron_len;
|
||||
}
|
||||
|
||||
int mm_set_opt(const char *preset, mm_idxopt_t *io, mm_mapopt_t *mo)
|
||||
@@ -56,38 +86,54 @@ int mm_set_opt(const char *preset, mm_idxopt_t *io, mm_mapopt_t *mo)
|
||||
if (preset == 0) {
|
||||
mm_idxopt_init(io);
|
||||
mm_mapopt_init(mo);
|
||||
} else if (strcmp(preset, "map-ont") == 0) { // this is the same as the default
|
||||
} else if (strcmp(preset, "ava-ont") == 0) {
|
||||
io->flag = 0, io->k = 15, io->w = 5;
|
||||
mo->flag |= MM_F_ALL_CHAINS | MM_F_NO_DIAG | MM_F_NO_DUAL | MM_F_NO_LJOIN;
|
||||
mo->min_chain_score = 100, mo->pri_ratio = 0.0f, mo->max_gap = 10000, mo->max_chain_skip = 25;
|
||||
mo->min_chain_score = 100, mo->pri_ratio = 0.0f, mo->max_chain_skip = 25;
|
||||
mo->bw = mo->bw_long = 2000;
|
||||
mo->occ_dist = 0;
|
||||
} else if (strcmp(preset, "map10k") == 0 || strcmp(preset, "map-pb") == 0) {
|
||||
io->flag |= MM_I_HPC, io->k = 19;
|
||||
} else if (strcmp(preset, "ava-pb") == 0) {
|
||||
io->flag |= MM_I_HPC, io->k = 19, io->w = 5;
|
||||
mo->flag |= MM_F_ALL_CHAINS | MM_F_NO_DIAG | MM_F_NO_DUAL | MM_F_NO_LJOIN;
|
||||
mo->min_chain_score = 100, mo->pri_ratio = 0.0f, mo->max_gap = 10000, mo->max_chain_skip = 25;
|
||||
} else if (strcmp(preset, "map10k") == 0 || strcmp(preset, "map-pb") == 0) {
|
||||
io->flag |= MM_I_HPC, io->k = 19;
|
||||
} else if (strcmp(preset, "map-ont") == 0) {
|
||||
io->flag = 0, io->k = 15;
|
||||
} else if (strcmp(preset, "asm5") == 0) {
|
||||
mo->min_chain_score = 100, mo->pri_ratio = 0.0f, mo->max_chain_skip = 25;
|
||||
mo->bw_long = mo->bw;
|
||||
mo->occ_dist = 0;
|
||||
} else if (strcmp(preset, "map-hifi") == 0 || strcmp(preset, "map-ccs") == 0) {
|
||||
io->flag = 0, io->k = 19, io->w = 19;
|
||||
mo->a = 1, mo->b = 19, mo->q = 39, mo->q2 = 81, mo->e = 3, mo->e2 = 1, mo->zdrop = 200;
|
||||
mo->max_gap = 10000;
|
||||
mo->a = 1, mo->b = 4, mo->q = 6, mo->q2 = 26, mo->e = 2, mo->e2 = 1;
|
||||
mo->occ_dist = 500;
|
||||
mo->min_mid_occ = 50, mo->max_mid_occ = 500;
|
||||
mo->min_dp_max = 200;
|
||||
mo->best_n = 50;
|
||||
} else if (strcmp(preset, "asm10") == 0) {
|
||||
} else if (strncmp(preset, "asm", 3) == 0) {
|
||||
io->flag = 0, io->k = 19, io->w = 19;
|
||||
mo->a = 1, mo->b = 9, mo->q = 16, mo->q2 = 41, mo->e = 2, mo->e2 = 1, mo->zdrop = 200;
|
||||
mo->bw = mo->bw_long = 100000;
|
||||
mo->max_gap = 10000;
|
||||
mo->flag |= MM_F_RMQ;
|
||||
mo->min_mid_occ = 50, mo->max_mid_occ = 500;
|
||||
mo->min_dp_max = 200;
|
||||
mo->best_n = 50;
|
||||
if (strcmp(preset, "asm5") == 0) {
|
||||
mo->a = 1, mo->b = 19, mo->q = 39, mo->q2 = 81, mo->e = 3, mo->e2 = 1, mo->zdrop = mo->zdrop_inv = 200;
|
||||
} else if (strcmp(preset, "asm10") == 0) {
|
||||
mo->a = 1, mo->b = 9, mo->q = 16, mo->q2 = 41, mo->e = 2, mo->e2 = 1, mo->zdrop = mo->zdrop_inv = 200;
|
||||
} else if (strcmp(preset, "asm20") == 0) {
|
||||
mo->a = 1, mo->b = 4, mo->q = 6, mo->q2 = 26, mo->e = 2, mo->e2 = 1, mo->zdrop = mo->zdrop_inv = 200;
|
||||
io->w = 10;
|
||||
} else return -1;
|
||||
} else if (strcmp(preset, "short") == 0 || strcmp(preset, "sr") == 0) {
|
||||
io->flag = 0, io->k = 21, io->w = 11;
|
||||
mo->flag |= MM_F_SR | MM_F_FRAG_MODE | MM_F_NO_PRINT_2ND | MM_F_2_IO_THREADS | MM_F_HEAP_SORT;
|
||||
mo->pe_ori = 0<<1|1; // FR
|
||||
mo->a = 2, mo->b = 8, mo->q = 12, mo->e = 2, mo->q2 = 24, mo->e2 = 1;
|
||||
mo->zdrop = 100;
|
||||
mo->zdrop = mo->zdrop_inv = 100;
|
||||
mo->end_bonus = 10;
|
||||
mo->max_frag_len = 800;
|
||||
mo->max_gap = 100;
|
||||
mo->bw = 100;
|
||||
mo->bw = mo->bw_long = 100;
|
||||
mo->pri_ratio = 0.5f;
|
||||
mo->min_cnt = 2;
|
||||
mo->min_chain_score = 25;
|
||||
@@ -96,19 +142,43 @@ int mm_set_opt(const char *preset, mm_idxopt_t *io, mm_mapopt_t *mo)
|
||||
mo->mid_occ = 1000;
|
||||
mo->max_occ = 5000;
|
||||
mo->mini_batch_size = 50000000;
|
||||
} else if (strcmp(preset, "splice") == 0 || strcmp(preset, "cdna") == 0) {
|
||||
} else if (strncmp(preset, "splice", 6) == 0 || strcmp(preset, "cdna") == 0) {
|
||||
io->flag = 0, io->k = 15, io->w = 5;
|
||||
mo->flag |= MM_F_SPLICE | MM_F_SPLICE_FOR | MM_F_SPLICE_REV | MM_F_SPLICE_FLANK;
|
||||
mo->max_gap = 2000, mo->max_gap_ref = mo->bw = 200000;
|
||||
mo->max_sw_mat = 0;
|
||||
mo->max_gap = 2000, mo->max_gap_ref = mo->bw = mo->bw_long = 200000;
|
||||
mo->a = 1, mo->b = 2, mo->q = 2, mo->e = 1, mo->q2 = 32, mo->e2 = 0;
|
||||
mo->noncan = 9;
|
||||
mo->zdrop = 200;
|
||||
mo->junc_bonus = 9;
|
||||
mo->zdrop = 200, mo->zdrop_inv = 100; // because mo->a is halved
|
||||
if (strcmp(preset, "splice:hq") == 0)
|
||||
mo->junc_bonus = 5, mo->b = 4, mo->q = 6, mo->q2 = 24;
|
||||
} else return -1;
|
||||
return 0;
|
||||
}
|
||||
|
||||
int mm_check_opt(const mm_idxopt_t *io, const mm_mapopt_t *mo)
|
||||
{
|
||||
if (mo->bw > mo->bw_long) {
|
||||
if (mm_verbose >= 1)
|
||||
fprintf(stderr, "[ERROR]\033[1;31m with '-rNUM1,NUM2', NUM1 (%d) can't be larger than NUM2 (%d)\033[0m\n", mo->bw, mo->bw_long);
|
||||
return -8;
|
||||
}
|
||||
if ((mo->flag & MM_F_RMQ) && (mo->flag & (MM_F_SR|MM_F_SPLICE))) {
|
||||
if (mm_verbose >= 1)
|
||||
fprintf(stderr, "[ERROR]\033[1;31m --rmq doesn't work with --sr or --splice\033[0m\n");
|
||||
return -7;
|
||||
}
|
||||
if (mo->split_prefix && (mo->flag & (MM_F_OUT_CS|MM_F_OUT_MD))) {
|
||||
if (mm_verbose >= 1)
|
||||
fprintf(stderr, "[ERROR]\033[1;31m --cs or --MD doesn't work with --split-prefix\033[0m\n");
|
||||
return -6;
|
||||
}
|
||||
if (io->k <= 0 || io->w <= 0) {
|
||||
if (mm_verbose >= 1)
|
||||
fprintf(stderr, "[ERROR]\033[1;31m -k and -w must be positive\033[0m\n");
|
||||
return -5;
|
||||
}
|
||||
if (mo->best_n < 0) {
|
||||
if (mm_verbose >= 1)
|
||||
fprintf(stderr, "[ERROR]\033[1;31m -N must be no less than 0\033[0m\n");
|
||||
@@ -126,6 +196,11 @@ int mm_check_opt(const mm_idxopt_t *io, const mm_mapopt_t *mo)
|
||||
fprintf(stderr, "[ERROR]\033[1;31m --for-only and --rev-only can't be applied at the same time\033[0m\n");
|
||||
return -3;
|
||||
}
|
||||
if (mo->e <= 0 || mo->q <= 0) {
|
||||
if (mm_verbose >= 1)
|
||||
fprintf(stderr, "[ERROR]\033[1;31m -O and -E must be positive\033[0m\n");
|
||||
return -1;
|
||||
}
|
||||
if ((mo->q != mo->q2 || mo->e != mo->e2) && !(mo->e > mo->e2 && mo->q + mo->e < mo->q2 + mo->e2)) {
|
||||
if (mm_verbose >= 1)
|
||||
fprintf(stderr, "[ERROR]\033[1;31m dual gap penalties violating E1>E2 and O1+E1<O2+E2\033[0m\n");
|
||||
@@ -136,5 +211,20 @@ int mm_check_opt(const mm_idxopt_t *io, const mm_mapopt_t *mo)
|
||||
fprintf(stderr, "[ERROR]\033[1;31m scoring system violating ({-O}+{-E})+({-O2}+{-E2}) <= 127\033[0m\n");
|
||||
return -1;
|
||||
}
|
||||
if (mo->zdrop < mo->zdrop_inv) {
|
||||
if (mm_verbose >= 1)
|
||||
fprintf(stderr, "[ERROR]\033[1;31m Z-drop should not be less than inversion-Z-drop\033[0m\n");
|
||||
return -5;
|
||||
}
|
||||
if ((mo->flag & MM_F_NO_PRINT_2ND) && (mo->flag & MM_F_ALL_CHAINS)) {
|
||||
if (mm_verbose >= 1)
|
||||
fprintf(stderr, "[ERROR]\033[1;31m -X/-P and --secondary=no can't be applied at the same time\033[0m\n");
|
||||
return -5;
|
||||
}
|
||||
if ((mo->flag & MM_F_QSTRAND) && ((mo->flag & (MM_F_OUT_SAM|MM_F_SPLICE|MM_F_FRAG_MODE)) || (io->flag & MM_I_HPC))) {
|
||||
if (mm_verbose >= 1)
|
||||
fprintf(stderr, "[ERROR]\033[1;31m --qstrand doesn't work with -a, -H, --frag or --splice\033[0m\n");
|
||||
return -5;
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
@@ -54,7 +54,7 @@ void mm_set_pe_thru(const int *qlens, int *n_regs, mm_reg1_t **regs)
|
||||
if (n_pri[0] == 1 && n_pri[1] == 1) {
|
||||
mm_reg1_t *p = ®s[0][pri[0]];
|
||||
mm_reg1_t *q = ®s[1][pri[1]];
|
||||
if (p->rid == q->rid && p->rev == q->rev && abs(p->rs - q->rs) < 3 && abs(p->re - p->re) < 3
|
||||
if (p->rid == q->rid && p->rev == q->rev && abs(p->rs - q->rs) < 3 && abs(p->re - q->re) < 3
|
||||
&& ((p->qs == 0 && qlens[1] - q->qe == 0) || (q->qs == 0 && qlens[0] - p->qe == 0)))
|
||||
{
|
||||
p->pe_thru = q->pe_thru = 1;
|
||||
@@ -105,7 +105,7 @@ void mm_pair(void *km, int max_gap_ref, int pe_bonus, int sub_diff, int match_sc
|
||||
max = -1;
|
||||
max_idx[0] = max_idx[1] = -1;
|
||||
last[0] = last[1] = -1;
|
||||
kv_resize(uint64_t, km, sc, n);
|
||||
kv_resize(uint64_t, km, sc, (size_t)n);
|
||||
for (i = 0; i < n; ++i) {
|
||||
if (a[i].key & 1) { // reverse first read or forward second read
|
||||
mm_reg1_t *q, *r;
|
||||
@@ -151,8 +151,8 @@ void mm_pair(void *km, int max_gap_ref, int pe_bonus, int sub_diff, int match_sc
|
||||
}
|
||||
}
|
||||
mapq_pe = r[0]->mapq > r[1]->mapq? r[0]->mapq : r[1]->mapq;
|
||||
for (i = 0; i < sc.n; ++i)
|
||||
if ((sc.a[i]>>32) + sub_diff >= max>>32)
|
||||
for (i = 0; i < (int)sc.n; ++i)
|
||||
if ((sc.a[i]>>32) + sub_diff >= (uint64_t)max>>32)
|
||||
++n_sub;
|
||||
if (sc.n > 1) {
|
||||
int mapq_pe_alt;
|
||||
@@ -164,7 +164,7 @@ void mm_pair(void *km, int max_gap_ref, int pe_bonus, int sub_diff, int match_sc
|
||||
if (sc.n == 1) {
|
||||
if (r[0]->mapq < 2) r[0]->mapq = 2;
|
||||
if (r[1]->mapq < 2) r[1]->mapq = 2;
|
||||
} else if (max>>32 > sc.a[sc.n - 2]>>32) {
|
||||
} else if ((uint64_t)max>>32 > sc.a[sc.n - 2]>>32) {
|
||||
if (r[0]->mapq < 1) r[0]->mapq = 1;
|
||||
if (r[1]->mapq < 1) r[1]->mapq = 1;
|
||||
}
|
||||
|
||||
+56
-11
@@ -34,6 +34,8 @@ The following Python script demonstrates the key functionality of mappy:
|
||||
import mappy as mp
|
||||
a = mp.Aligner("test/MT-human.fa") # load or build index
|
||||
if not a: raise Exception("ERROR: failed to load/build index")
|
||||
s = a.seq("MT_human", 100, 200) # retrieve a subsequence from the index
|
||||
print(mp.revcomp(s)) # reverse complement
|
||||
for name, seq, qual in mp.fastx_read("test/MT-orang.fa"): # read a fasta/q sequence
|
||||
for hit in a.map(seq): # traverse alignments
|
||||
print("{}\t{}\t{}\t{}".format(hit.ctg, hit.r_st, hit.r_en, hit.cigar_str))
|
||||
@@ -41,20 +43,24 @@ The following Python script demonstrates the key functionality of mappy:
|
||||
APIs
|
||||
----
|
||||
|
||||
Mappy implements two classes and one global function.
|
||||
Mappy implements two classes and two global function.
|
||||
|
||||
Class mappy.Aligner
|
||||
~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
.. code:: python
|
||||
|
||||
mappy.Aligner(fn_idx_in, preset=None, ...)
|
||||
mappy.Aligner(fn_idx_in=None, preset=None, ...)
|
||||
|
||||
This constructor accepts the following arguments:
|
||||
|
||||
* **fn_idx_in**: index or sequence file name. Minimap2 automatically tests the
|
||||
file type. If a sequence file is provided, minimap2 builds an index. The
|
||||
sequence file can be optionally gzip'd.
|
||||
sequence file can be optionally gzip'd. This option has no effect if **seq**
|
||||
is set.
|
||||
|
||||
* **seq**: a single sequence to index. The sequence name will be set to
|
||||
:code:`N/A`.
|
||||
|
||||
* **preset**: minimap2 preset. Currently, minimap2 supports the following
|
||||
presets: **sr** for single-end short reads; **map-pb** for PacBio
|
||||
@@ -77,17 +83,42 @@ This constructor accepts the following arguments:
|
||||
|
||||
* **n_threads**: number of indexing threads; 3 by default
|
||||
|
||||
* **fn_idx_out**: name of file to which the index is written
|
||||
* **extra_flags**: additional flags defined in minimap.h
|
||||
|
||||
* **fn_idx_out**: name of file to which the index is written. This parameter
|
||||
has no effect if **seq** is set.
|
||||
|
||||
* **scoring**: scoring system. It is a tuple/list consisting of 4, 6 or 7
|
||||
positive integers. The first 4 elements specify match scoring, mismatch
|
||||
penalty, gap open and gap extension penalty. The 5th and 6th elements, if
|
||||
present, set long-gap open and long-gap extension penalty. The 7th sets a
|
||||
mismatch penalty involving ambiguous bases.
|
||||
|
||||
.. code:: python
|
||||
|
||||
mappy.Aligner.map(seq, seq2=None)
|
||||
mappy.Aligner.map(seq, seq2=None, cs=False, MD=False)
|
||||
|
||||
This method aligns :code:`seq` against the index. It is a generator, *yielding*
|
||||
a series of :code:`mappy.Alignment` objects. If :code:`seq2` is present, mappy
|
||||
performs paired-end alignment, assuming the two ends are in the FR orientation.
|
||||
Alignments of the two ends can be distinguished by the :code:`read_num` field
|
||||
(see below).
|
||||
(see Class mappy.Alignment below). Argument :code:`cs` asks mappy to generate
|
||||
the :code:`cs` tag; :code:`MD` is similar. These two arguments might slightly
|
||||
degrade performance and are not enabled by default.
|
||||
|
||||
.. code:: python
|
||||
|
||||
mappy.Aligner.seq(name, start=0, end=0x7fffffff)
|
||||
|
||||
This method retrieves a (sub)sequence from the index and returns it as a Python
|
||||
string. :code:`None` is returned if :code:`name` is not present in the index or
|
||||
the start/end coordinates are invalid.
|
||||
|
||||
.. code:: python
|
||||
|
||||
mappy.Aligner.seq_names
|
||||
|
||||
This property gives the array of sequence names in the index.
|
||||
|
||||
Class mappy.Alignment
|
||||
~~~~~~~~~~~~~~~~~~~~~
|
||||
@@ -113,7 +144,7 @@ properties:
|
||||
* **mlen**: length of the matching bases in the alignment, excluding ambiguous
|
||||
base matches.
|
||||
|
||||
* **NM**: number of mismatches, gaps and ambiguous poistions in the alignment
|
||||
* **NM**: number of mismatches, gaps and ambiguous positions in the alignment
|
||||
|
||||
* **trans_strand**: transcript strand. +1 if on the forward strand; -1 if on the
|
||||
reverse strand; 0 if unknown
|
||||
@@ -129,6 +160,11 @@ properties:
|
||||
* **cigar**: CIGAR returned as an array of shape :code:`(n_cigar,2)`. The two
|
||||
numbers give the length and the operator of each CIGAR operation.
|
||||
|
||||
* **MD**: the :code:`MD` tag as in the SAM format. It is an empty string unless
|
||||
the :code:`MD` argument is applied when calling :code:`mappy.Aligner.map()`.
|
||||
|
||||
* **cs**: the :code:`cs` tag.
|
||||
|
||||
An :code:`Alignment` object can be converted to a string with :code:`str()` in
|
||||
the following format:
|
||||
|
||||
@@ -139,13 +175,22 @@ the following format:
|
||||
It is effectively the PAF format without the QueryName and QueryLength columns
|
||||
(the first two columns in PAF).
|
||||
|
||||
Function mappy.fastx_read
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
Miscellaneous Functions
|
||||
~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
.. code:: python
|
||||
|
||||
mappy.fastx_read(fn)
|
||||
mappy.fastx_read(fn, read_comment=False)
|
||||
|
||||
This generator function opens a FASTA/FASTQ file and *yields* a
|
||||
:code:`(name,seq,qual)` tuple for each sequence entry. The input file may be
|
||||
optionally gzip'd.
|
||||
optionally gzip'd. If :code:`read_comment` is True, this generator yields
|
||||
a :code:`(name,seq,qual,comment)` tuple instead.
|
||||
|
||||
.. code:: python
|
||||
|
||||
mappy.revcomp(seq)
|
||||
|
||||
Return the reverse complement of DNA string :code:`seq`. This function
|
||||
recognizes IUB code and preserves the letter cases. Uracil :code:`U` is
|
||||
complemented to :code:`A`.
|
||||
|
||||
+50
-2
@@ -73,13 +73,17 @@ static inline void mm_reset_timer(void)
|
||||
extern unsigned char seq_comp_table[256];
|
||||
static inline mm_reg1_t *mm_map_aux(const mm_idx_t *mi, const char *seq1, const char *seq2, int *n_regs, mm_tbuf_t *b, const mm_mapopt_t *opt)
|
||||
{
|
||||
mm_reg1_t *r;
|
||||
|
||||
Py_BEGIN_ALLOW_THREADS
|
||||
if (seq2 == 0) {
|
||||
return mm_map(mi, strlen(seq1), seq1, n_regs, b, opt, NULL);
|
||||
r = mm_map(mi, strlen(seq1), seq1, n_regs, b, opt, NULL);
|
||||
} else {
|
||||
int _n_regs[2];
|
||||
mm_reg1_t *regs[2];
|
||||
char *seq[2];
|
||||
int i, len[2];
|
||||
|
||||
len[0] = strlen(seq1);
|
||||
len[1] = strlen(seq2);
|
||||
seq[0] = (char*)seq1;
|
||||
@@ -97,8 +101,52 @@ static inline mm_reg1_t *mm_map_aux(const mm_idx_t *mi, const char *seq1, const
|
||||
regs[0] = (mm_reg1_t*)realloc(regs[0], sizeof(mm_reg1_t) * (*n_regs));
|
||||
memcpy(®s[0][_n_regs[0]], regs[1], _n_regs[1] * sizeof(mm_reg1_t));
|
||||
free(regs[1]);
|
||||
return regs[0];
|
||||
r = regs[0];
|
||||
}
|
||||
Py_END_ALLOW_THREADS
|
||||
|
||||
return r;
|
||||
}
|
||||
|
||||
static inline char *mappy_revcomp(int len, const uint8_t *seq)
|
||||
{
|
||||
int i;
|
||||
char *rev;
|
||||
rev = (char*)malloc(len + 1);
|
||||
for (i = 0; i < len; ++i)
|
||||
rev[len - i - 1] = seq_comp_table[seq[i]];
|
||||
rev[len] = 0;
|
||||
return rev;
|
||||
}
|
||||
|
||||
static char *mappy_fetch_seq(const mm_idx_t *mi, const char *name, int st, int en, int *len)
|
||||
{
|
||||
int i, rid;
|
||||
char *s;
|
||||
*len = 0;
|
||||
rid = mm_idx_name2id(mi, name);
|
||||
if (rid < 0) return 0;
|
||||
if ((uint32_t)st >= mi->seq[rid].len || st >= en) return 0;
|
||||
if (en < 0 || (uint32_t)en > mi->seq[rid].len)
|
||||
en = mi->seq[rid].len;
|
||||
s = (char*)malloc(en - st + 1);
|
||||
*len = mm_idx_getseq(mi, rid, st, en, (uint8_t*)s);
|
||||
for (i = 0; i < *len; ++i)
|
||||
s[i] = "ACGTN"[(uint8_t)s[i]];
|
||||
s[*len] = 0;
|
||||
return s;
|
||||
}
|
||||
|
||||
static mm_idx_t *mappy_idx_seq(int w, int k, int is_hpc, int bucket_bits, const char *seq, int len)
|
||||
{
|
||||
const char *fake_name = "N/A";
|
||||
char *s;
|
||||
mm_idx_t *mi;
|
||||
s = (char*)calloc(len + 1, 1);
|
||||
memcpy(s, seq, len);
|
||||
mi = mm_idx_str(w, k, is_hpc, bucket_bits, 1, (const char**)&s, (const char**)&fake_name);
|
||||
free(s);
|
||||
return mi;
|
||||
}
|
||||
|
||||
#endif
|
||||
|
||||
+36
-8
@@ -6,36 +6,55 @@ cdef extern from "minimap.h":
|
||||
#
|
||||
ctypedef struct mm_idxopt_t:
|
||||
short k, w, flag, bucket_bits
|
||||
int mini_batch_size
|
||||
int64_t mini_batch_size
|
||||
uint64_t batch_size
|
||||
|
||||
ctypedef struct mm_mapopt_t:
|
||||
int64_t flag
|
||||
int seed
|
||||
int sdust_thres
|
||||
int flag
|
||||
int bw
|
||||
|
||||
int max_qlen
|
||||
|
||||
int bw, bw_long
|
||||
int max_gap, max_gap_ref
|
||||
int max_frag_len
|
||||
int max_chain_skip
|
||||
int max_chain_skip, max_chain_iter
|
||||
int min_cnt
|
||||
int min_chain_score
|
||||
float chain_gap_scale
|
||||
int rmq_size_cap, rmq_inner_dist
|
||||
int rmq_rescue_size
|
||||
float rmq_rescue_ratio
|
||||
|
||||
float mask_level
|
||||
int mask_len
|
||||
float pri_ratio
|
||||
int best_n
|
||||
int max_join_long, max_join_short
|
||||
int min_join_flank_sc
|
||||
|
||||
float alt_drop
|
||||
|
||||
int a, b, q, e, q2, e2
|
||||
int sc_ambi
|
||||
int noncan
|
||||
int zdrop
|
||||
int junc_bonus
|
||||
int zdrop, zdrop_inv
|
||||
int end_bonus
|
||||
int min_dp_max
|
||||
int min_ksw_len
|
||||
int anchor_ext_len, anchor_ext_shift
|
||||
float max_clip_ratio
|
||||
|
||||
int pe_ori, pe_bonus
|
||||
|
||||
float mid_occ_frac
|
||||
int32_t min_mid_occ
|
||||
int32_t mid_occ
|
||||
int32_t max_occ
|
||||
int mini_batch_size
|
||||
int64_t mini_batch_size
|
||||
int64_t max_sw_mat
|
||||
|
||||
const char *split_prefix
|
||||
|
||||
int mm_set_opt(char *preset, mm_idxopt_t *io, mm_mapopt_t *mo)
|
||||
int mm_verbose
|
||||
@@ -58,6 +77,7 @@ cdef extern from "minimap.h":
|
||||
uint32_t *S
|
||||
mm_idx_bucket_t *B
|
||||
void *km
|
||||
void *h
|
||||
|
||||
ctypedef struct mm_idx_reader_t:
|
||||
pass
|
||||
@@ -68,6 +88,8 @@ cdef extern from "minimap.h":
|
||||
void mm_idx_destroy(mm_idx_t *mi)
|
||||
void mm_mapopt_update(mm_mapopt_t *opt, const mm_idx_t *mi)
|
||||
|
||||
int mm_idx_index_name(mm_idx_t *mi)
|
||||
|
||||
#
|
||||
# Mapping (key struct defined in cmappy.h below)
|
||||
#
|
||||
@@ -79,6 +101,9 @@ cdef extern from "minimap.h":
|
||||
|
||||
mm_tbuf_t *mm_tbuf_init()
|
||||
void mm_tbuf_destroy(mm_tbuf_t *b)
|
||||
void *mm_tbuf_get_km(mm_tbuf_t *b)
|
||||
int mm_gen_cs(void *km, char **buf, int *max_len, const mm_idx_t *mi, const mm_reg1_t *r, const char *seq, int no_iden)
|
||||
int mm_gen_MD(void *km, char **buf, int *max_len, const mm_idx_t *mi, const mm_reg1_t *r, const char *seq)
|
||||
|
||||
#
|
||||
# Helper header (because it is hard to expose mm_reg1_t with Cython)
|
||||
@@ -98,6 +123,8 @@ cdef extern from "cmappy.h":
|
||||
void mm_reg2hitpy(const mm_idx_t *mi, mm_reg1_t *r, mm_hitpy_t *h)
|
||||
void mm_free_reg1(mm_reg1_t *r)
|
||||
mm_reg1_t *mm_map_aux(const mm_idx_t *mi, const char *seq1, const char *seq2, int *n_regs, mm_tbuf_t *b, const mm_mapopt_t *opt)
|
||||
char *mappy_fetch_seq(const mm_idx_t *mi, const char *name, int st, int en, int *l)
|
||||
mm_idx_t *mappy_idx_seq(int w, int k, int is_hpc, int bucket_bits, const char *seq, int l)
|
||||
|
||||
ctypedef struct kstring_t:
|
||||
unsigned l, m
|
||||
@@ -115,5 +142,6 @@ cdef extern from "cmappy.h":
|
||||
void mm_fastx_close(kseq_t *ks)
|
||||
int kseq_read(kseq_t *seq)
|
||||
|
||||
char *mappy_revcomp(int l, const uint8_t *seq)
|
||||
int mm_verbose_level(int v)
|
||||
void mm_reset_timer()
|
||||
|
||||
+131
-27
@@ -1,6 +1,9 @@
|
||||
from libc.stdint cimport uint8_t, int8_t
|
||||
from libc.stdlib cimport free
|
||||
cimport cmappy
|
||||
import sys
|
||||
|
||||
__version__ = '2.21'
|
||||
|
||||
cmappy.mm_reset_timer()
|
||||
|
||||
@@ -11,9 +14,9 @@ cdef class Alignment:
|
||||
cdef int8_t _strand, _trans_strand
|
||||
cdef uint8_t _mapq, _is_primary
|
||||
cdef int _seg_id
|
||||
cdef _ctg, _cigar # these are python objects
|
||||
cdef _ctg, _cigar, _cs, _MD # these are python objects
|
||||
|
||||
def __cinit__(self, ctg, cl, cs, ce, strand, qs, qe, mapq, cigar, is_primary, mlen, blen, NM, trans_strand, seg_id):
|
||||
def __cinit__(self, ctg, cl, cs, ce, strand, qs, qe, mapq, cigar, is_primary, mlen, blen, NM, trans_strand, seg_id, cs_str, MD_str):
|
||||
self._ctg = ctg if isinstance(ctg, str) else ctg.decode()
|
||||
self._ctg_len, self._r_st, self._r_en = cl, cs, ce
|
||||
self._strand, self._q_st, self._q_en = strand, qs, qe
|
||||
@@ -23,6 +26,8 @@ cdef class Alignment:
|
||||
self._is_primary = is_primary
|
||||
self._trans_strand = trans_strand
|
||||
self._seg_id = seg_id
|
||||
self._cs = cs_str
|
||||
self._MD = MD_str
|
||||
|
||||
@property
|
||||
def ctg(self): return self._ctg
|
||||
@@ -69,9 +74,15 @@ cdef class Alignment:
|
||||
@property
|
||||
def read_num(self): return self._seg_id + 1
|
||||
|
||||
@property
|
||||
def cs(self): return self._cs
|
||||
|
||||
@property
|
||||
def MD(self): return self._MD
|
||||
|
||||
@property
|
||||
def cigar_str(self):
|
||||
return "".join(map(lambda x: str(x[0]) + 'MIDNSH'[x[1]], self._cigar))
|
||||
return "".join(map(lambda x: str(x[0]) + 'MIDNSHP=XB'[x[1]], self._cigar))
|
||||
|
||||
def __str__(self):
|
||||
if self._strand > 0: strand = '+'
|
||||
@@ -82,8 +93,10 @@ cdef class Alignment:
|
||||
if self._trans_strand > 0: ts = 'ts:A:+'
|
||||
elif self._trans_strand < 0: ts = 'ts:A:-'
|
||||
else: ts = 'ts:A:.'
|
||||
return "\t".join([str(self._q_st), str(self._q_en), strand, self._ctg, str(self._ctg_len), str(self._r_st), str(self._r_en),
|
||||
str(self._mlen), str(self._blen), str(self._mapq), tp, ts, "cg:Z:" + self.cigar_str])
|
||||
a = [str(self._q_st), str(self._q_en), strand, self._ctg, str(self._ctg_len), str(self._r_st), str(self._r_en),
|
||||
str(self._mlen), str(self._blen), str(self._mapq), tp, ts, "cg:Z:" + self.cigar_str]
|
||||
if self._cs != "": a.append("cs:Z:" + self._cs)
|
||||
return "\t".join(a)
|
||||
|
||||
cdef class ThreadBuffer:
|
||||
cdef cmappy.mm_tbuf_t *_b
|
||||
@@ -99,7 +112,8 @@ cdef class Aligner:
|
||||
cdef cmappy.mm_idxopt_t idx_opt
|
||||
cdef cmappy.mm_mapopt_t map_opt
|
||||
|
||||
def __cinit__(self, fn_idx_in, preset=None, k=None, w=None, min_cnt=None, min_chain_score=None, min_dp_score=None, bw=None, best_n=None, n_threads=3, fn_idx_out=None):
|
||||
def __cinit__(self, fn_idx_in=None, preset=None, k=None, w=None, min_cnt=None, min_chain_score=None, min_dp_score=None, bw=None, best_n=None, n_threads=3, fn_idx_out=None, max_frag_len=None, extra_flags=None, seq=None, scoring=None):
|
||||
self._idx = NULL
|
||||
cmappy.mm_set_opt(NULL, &self.idx_opt, &self.map_opt) # set the default options
|
||||
if preset is not None:
|
||||
cmappy.mm_set_opt(str.encode(preset), &self.idx_opt, &self.map_opt) # apply preset
|
||||
@@ -111,17 +125,34 @@ cdef class Aligner:
|
||||
if min_chain_score is not None: self.map_opt.min_chain_score = min_chain_score
|
||||
if min_dp_score is not None: self.map_opt.min_dp_max = min_dp_score
|
||||
if bw is not None: self.map_opt.bw = bw
|
||||
if best_n is not None: self.best_n = best_n
|
||||
if best_n is not None: self.map_opt.best_n = best_n
|
||||
if max_frag_len is not None: self.map_opt.max_frag_len = max_frag_len
|
||||
if extra_flags is not None: self.map_opt.flag |= extra_flags
|
||||
if scoring is not None and len(scoring) >= 4:
|
||||
self.map_opt.a, self.map_opt.b = scoring[0], scoring[1]
|
||||
self.map_opt.q, self.map_opt.e = scoring[2], scoring[3]
|
||||
self.map_opt.q2, self.map_opt.e2 = self.map_opt.q, self.map_opt.e
|
||||
if len(scoring) >= 6:
|
||||
self.map_opt.q2, self.map_opt.e2 = scoring[4], scoring[5]
|
||||
if len(scoring) >= 7:
|
||||
self.map_opt.sc_ambi = scoring[6]
|
||||
|
||||
cdef cmappy.mm_idx_reader_t *r;
|
||||
if fn_idx_out is None:
|
||||
r = cmappy.mm_idx_reader_open(str.encode(fn_idx_in), &self.idx_opt, NULL)
|
||||
|
||||
if seq is None:
|
||||
if fn_idx_out is None:
|
||||
r = cmappy.mm_idx_reader_open(str.encode(fn_idx_in), &self.idx_opt, NULL)
|
||||
else:
|
||||
r = cmappy.mm_idx_reader_open(str.encode(fn_idx_in), &self.idx_opt, str.encode(fn_idx_out))
|
||||
if r is not NULL:
|
||||
self._idx = cmappy.mm_idx_reader_read(r, n_threads) # NB: ONLY read the first part
|
||||
cmappy.mm_idx_reader_close(r)
|
||||
cmappy.mm_mapopt_update(&self.map_opt, self._idx)
|
||||
cmappy.mm_idx_index_name(self._idx)
|
||||
else:
|
||||
r = cmappy.mm_idx_reader_open(str.encode(fn_idx_in), &self.idx_opt, fn_idx_out)
|
||||
if r is not NULL:
|
||||
self._idx = cmappy.mm_idx_reader_read(r, n_threads) # NB: ONLY read the first part
|
||||
cmappy.mm_idx_reader_close(r)
|
||||
self._idx = cmappy.mappy_idx_seq(self.idx_opt.w, self.idx_opt.k, self.idx_opt.flag&1, self.idx_opt.bucket_bits, str.encode(seq), len(seq))
|
||||
cmappy.mm_mapopt_update(&self.map_opt, self._idx)
|
||||
self.map_opt.mid_occ = 1000 # don't filter high-occ seeds
|
||||
|
||||
def __dealloc__(self):
|
||||
if self._idx is not NULL:
|
||||
@@ -130,29 +161,89 @@ cdef class Aligner:
|
||||
def __bool__(self):
|
||||
return (self._idx != NULL)
|
||||
|
||||
def map(self, seq, seq2=None, buf=None):
|
||||
def map(self, seq, seq2=None, buf=None, cs=False, MD=False, max_frag_len=None, extra_flags=None):
|
||||
cdef cmappy.mm_reg1_t *regs
|
||||
cdef cmappy.mm_hitpy_t h
|
||||
cdef ThreadBuffer b
|
||||
cdef int n_regs
|
||||
cdef char *cs_str = NULL
|
||||
cdef int l_cs_str, m_cs_str = 0
|
||||
cdef void *km
|
||||
cdef cmappy.mm_mapopt_t map_opt
|
||||
|
||||
if self._idx == NULL: return
|
||||
map_opt = self.map_opt
|
||||
if max_frag_len is not None: map_opt.max_frag_len = max_frag_len
|
||||
if extra_flags is not None: map_opt.flag |= extra_flags
|
||||
|
||||
if self._idx is NULL: return None
|
||||
if buf is None: b = ThreadBuffer()
|
||||
else: b = buf
|
||||
if seq2 is None: regs = cmappy.mm_map_aux(self._idx, str.encode(seq), NULL, &n_regs, b._b, &self.map_opt)
|
||||
else: regs = cmappy.mm_map_aux(self._idx, str.encode(seq), str.encode(seq2), &n_regs, b._b, &self.map_opt)
|
||||
km = cmappy.mm_tbuf_get_km(b._b)
|
||||
|
||||
for i in range(n_regs):
|
||||
cmappy.mm_reg2hitpy(self._idx, ®s[i], &h)
|
||||
cigar = []
|
||||
for k in range(h.n_cigar32):
|
||||
c = h.cigar32[k]
|
||||
cigar.append([c>>4, c&0xf])
|
||||
yield Alignment(h.ctg, h.ctg_len, h.ctg_start, h.ctg_end, h.strand, h.qry_start, h.qry_end, h.mapq, cigar, h.is_primary, h.mlen, h.blen, h.NM, h.trans_strand, h.seg_id)
|
||||
cmappy.mm_free_reg1(®s[i])
|
||||
free(regs)
|
||||
_seq = seq if isinstance(seq, bytes) else seq.encode()
|
||||
if seq2 is None:
|
||||
regs = cmappy.mm_map_aux(self._idx, _seq, NULL, &n_regs, b._b, &map_opt)
|
||||
else:
|
||||
_seq2 = seq2 if isinstance(seq2, bytes) else seq2.encode()
|
||||
regs = cmappy.mm_map_aux(self._idx, _seq, _seq2, &n_regs, b._b, &map_opt)
|
||||
|
||||
def fastx_read(fn):
|
||||
try:
|
||||
i = 0
|
||||
while i < n_regs:
|
||||
cmappy.mm_reg2hitpy(self._idx, ®s[i], &h)
|
||||
cigar, _cs, _MD = [], '', ''
|
||||
for k in range(h.n_cigar32): # convert the 32-bit CIGAR encoding to Python array
|
||||
c = h.cigar32[k]
|
||||
cigar.append([c>>4, c&0xf])
|
||||
if cs or MD: # generate the cs and/or the MD tag, if requested
|
||||
if cs:
|
||||
l_cs_str = cmappy.mm_gen_cs(km, &cs_str, &m_cs_str, self._idx, ®s[i], _seq, 1)
|
||||
_cs = cs_str[:l_cs_str] if isinstance(cs_str, str) else cs_str[:l_cs_str].decode()
|
||||
if MD:
|
||||
l_cs_str = cmappy.mm_gen_MD(km, &cs_str, &m_cs_str, self._idx, ®s[i], _seq)
|
||||
_MD = cs_str[:l_cs_str] if isinstance(cs_str, str) else cs_str[:l_cs_str].decode()
|
||||
yield Alignment(h.ctg, h.ctg_len, h.ctg_start, h.ctg_end, h.strand, h.qry_start, h.qry_end, h.mapq, cigar, h.is_primary, h.mlen, h.blen, h.NM, h.trans_strand, h.seg_id, _cs, _MD)
|
||||
cmappy.mm_free_reg1(®s[i])
|
||||
i += 1
|
||||
finally:
|
||||
while i < n_regs:
|
||||
cmappy.mm_free_reg1(®s[i])
|
||||
i += 1
|
||||
free(regs)
|
||||
free(cs_str)
|
||||
|
||||
def seq(self, str name, int start=0, int end=0x7fffffff):
|
||||
cdef int l
|
||||
cdef char *s
|
||||
if self._idx == NULL: return
|
||||
s = cmappy.mappy_fetch_seq(self._idx, name.encode(), start, end, &l)
|
||||
if l == 0: return None
|
||||
r = s[:l] if isinstance(s, str) else s[:l].decode()
|
||||
free(s)
|
||||
return r
|
||||
|
||||
@property
|
||||
def k(self): return self._idx.k
|
||||
|
||||
@property
|
||||
def w(self): return self._idx.w
|
||||
|
||||
@property
|
||||
def n_seq(self): return self._idx.n_seq
|
||||
|
||||
@property
|
||||
def seq_names(self):
|
||||
cdef char *p
|
||||
if self._idx == NULL: return
|
||||
sn = []
|
||||
for i in range(self._idx.n_seq):
|
||||
p = self._idx.seq[i].name
|
||||
s = p if isinstance(p, str) else p.decode()
|
||||
sn.append(s)
|
||||
return sn
|
||||
|
||||
def fastx_read(fn, read_comment=False):
|
||||
cdef cmappy.kseq_t *ks
|
||||
ks = cmappy.mm_fastx_open(str.encode(fn))
|
||||
if ks is NULL: return None
|
||||
@@ -161,9 +252,22 @@ def fastx_read(fn):
|
||||
else: qual = None
|
||||
name = ks.name.s if isinstance(ks.name.s, str) else ks.name.s.decode()
|
||||
seq = ks.seq.s if isinstance(ks.seq.s, str) else ks.seq.s.decode()
|
||||
yield name, seq, qual
|
||||
if read_comment:
|
||||
if ks.comment.l > 0: comment = ks.comment.s if isinstance(ks.comment.s, str) else ks.comment.s.decode()
|
||||
else: comment = None
|
||||
yield name, seq, qual, comment
|
||||
else:
|
||||
yield name, seq, qual
|
||||
cmappy.mm_fastx_close(ks)
|
||||
|
||||
def revcomp(seq):
|
||||
l = len(seq)
|
||||
bseq = seq if isinstance(seq, bytes) else seq.encode()
|
||||
cdef char *s = cmappy.mappy_revcomp(l, bseq)
|
||||
r = s[:l] if isinstance(s, str) else s[:l].decode()
|
||||
free(s)
|
||||
return r
|
||||
|
||||
def verbose(v=None):
|
||||
if v is None: v = -1
|
||||
return cmappy.mm_verbose_level(v)
|
||||
|
||||
+8
-4
@@ -1,10 +1,11 @@
|
||||
#!/usr/bin/env python
|
||||
|
||||
import sys, getopt
|
||||
import sys
|
||||
import getopt
|
||||
import mappy as mp
|
||||
|
||||
def main(argv):
|
||||
opts, args = getopt.getopt(argv[1:], "x:n:m:k:w:r:")
|
||||
opts, args = getopt.getopt(argv[1:], "x:n:m:k:w:r:c")
|
||||
if len(args) < 2:
|
||||
print("Usage: minimap2.py [options] <ref.fa>|<ref.mmi> <query.fq>")
|
||||
print("Options:")
|
||||
@@ -14,9 +15,11 @@ def main(argv):
|
||||
print(" -k INT k-mer length")
|
||||
print(" -w INT minimizer window length")
|
||||
print(" -r INT band width")
|
||||
print(" -c output the cs tag")
|
||||
sys.exit(1)
|
||||
|
||||
preset, min_cnt, min_sc, k, w, bw = None, None, None, None, None, None
|
||||
preset = min_cnt = min_sc = k = w = bw = None
|
||||
out_cs = False
|
||||
for opt, arg in opts:
|
||||
if opt == '-x': preset = arg
|
||||
elif opt == '-n': min_cnt = int(arg)
|
||||
@@ -24,11 +27,12 @@ def main(argv):
|
||||
elif opt == '-r': bw = int(arg)
|
||||
elif opt == '-k': k = int(arg)
|
||||
elif opt == '-w': w = int(arg)
|
||||
elif opt == '-c': out_cs = True
|
||||
|
||||
a = mp.Aligner(args[0], preset=preset, min_cnt=min_cnt, min_chain_score=min_sc, k=k, w=w, bw=bw)
|
||||
if not a: raise Exception("ERROR: failed to load/build index file '{}'".format(args[0]))
|
||||
for name, seq, qual in mp.fastx_read(args[1]): # read one sequence
|
||||
for h in a.map(seq): # traverse hits
|
||||
for h in a.map(seq, cs=out_cs): # traverse hits
|
||||
print('{}\t{}\t{}'.format(name, len(seq), h))
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
@@ -70,10 +70,10 @@ void sdust_buf_destroy(sdust_buf_t *buf)
|
||||
static inline void shift_window(int t, kdq_t(int) *w, int T, int W, int *L, int *rw, int *rv, int *cw, int *cv)
|
||||
{
|
||||
int s;
|
||||
if (kdq_size(w) >= W - SD_WLEN + 1) { // TODO: is this right for SD_WLEN!=3?
|
||||
if ((int)kdq_size(w) >= W - SD_WLEN + 1) { // TODO: is this right for SD_WLEN!=3?
|
||||
s = *kdq_shift(int, w);
|
||||
*rw -= --cw[s];
|
||||
if (*L > kdq_size(w))
|
||||
if (*L > (int)kdq_size(w))
|
||||
--*L, *rv -= --cv[s];
|
||||
}
|
||||
kdq_push(int, w, t);
|
||||
@@ -114,7 +114,7 @@ static void find_perfect(void *km, perf_intv_v *P, const kdq_t(int) *w, int T, i
|
||||
r += c[t]++;
|
||||
new_r = r, new_l = kdq_size(w) - i - 1;
|
||||
if (new_r * 10 > T * new_l) {
|
||||
for (j = 0; j < P->n && P->a[j].start >= i + start; ++j) { // find insertion position
|
||||
for (j = 0; j < (int)P->n && P->a[j].start >= i + start; ++j) { // find insertion position
|
||||
perf_intv_t *p = &P->a[j];
|
||||
if (max_r == 0 || p->r * max_l > max_r * p->l)
|
||||
max_r = p->r, max_l = p->l;
|
||||
@@ -177,7 +177,7 @@ uint64_t *sdust(void *km, const uint8_t *seq, int l_seq, int T, int W, int *n)
|
||||
#ifdef _SDUST_MAIN
|
||||
#include <zlib.h>
|
||||
#include <stdio.h>
|
||||
#include "getopt.h"
|
||||
#include "ketopt.h"
|
||||
#include "kseq.h"
|
||||
KSEQ_INIT(gzFile, gzread)
|
||||
|
||||
@@ -186,16 +186,17 @@ int main(int argc, char *argv[])
|
||||
gzFile fp;
|
||||
kseq_t *ks;
|
||||
int W = 64, T = 20, c;
|
||||
ketopt_t o = KETOPT_INIT;
|
||||
|
||||
while ((c = getopt(argc, argv, "w:t:")) >= 0) {
|
||||
if (c == 'w') W = atoi(optarg);
|
||||
else if (c == 't') T = atoi(optarg);
|
||||
while ((c = ketopt(&o, argc, argv, 1, "w:t:", 0)) >= 0) {
|
||||
if (c == 'w') W = atoi(o.arg);
|
||||
else if (c == 't') T = atoi(o.arg);
|
||||
}
|
||||
if (optind == argc) {
|
||||
if (o.ind == argc) {
|
||||
fprintf(stderr, "Usage: sdust [-w %d] [-t %d] <in.fa>\n", W, T);
|
||||
return 1;
|
||||
}
|
||||
fp = strcmp(argv[optind], "-")? gzopen(argv[optind], "r") : gzdopen(fileno(stdin), "r");
|
||||
fp = strcmp(argv[o.ind], "-")? gzopen(argv[o.ind], "r") : gzdopen(fileno(stdin), "r");
|
||||
ks = kseq_init(fp);
|
||||
while (kseq_read(ks) >= 0) {
|
||||
uint64_t *r;
|
||||
|
||||
@@ -0,0 +1,106 @@
|
||||
#include "mmpriv.h"
|
||||
#include "kalloc.h"
|
||||
#include "ksort.h"
|
||||
|
||||
mm_seed_t *mm_seed_collect_all(void *km, const mm_idx_t *mi, const mm128_v *mv, int32_t *n_m_)
|
||||
{
|
||||
mm_seed_t *m;
|
||||
size_t i;
|
||||
int32_t k;
|
||||
m = (mm_seed_t*)kmalloc(km, mv->n * sizeof(mm_seed_t));
|
||||
for (i = k = 0; i < mv->n; ++i) {
|
||||
const uint64_t *cr;
|
||||
mm_seed_t *q;
|
||||
mm128_t *p = &mv->a[i];
|
||||
uint32_t q_pos = (uint32_t)p->y, q_span = p->x & 0xff;
|
||||
int t;
|
||||
cr = mm_idx_get(mi, p->x>>8, &t);
|
||||
if (t == 0) continue;
|
||||
q = &m[k++];
|
||||
q->q_pos = q_pos, q->q_span = q_span, q->cr = cr, q->n = t, q->seg_id = p->y >> 32;
|
||||
q->is_tandem = q->flt = 0;
|
||||
if (i > 0 && p->x>>8 == mv->a[i - 1].x>>8) q->is_tandem = 1;
|
||||
if (i < mv->n - 1 && p->x>>8 == mv->a[i + 1].x>>8) q->is_tandem = 1;
|
||||
}
|
||||
*n_m_ = k;
|
||||
return m;
|
||||
}
|
||||
|
||||
#define MAX_MAX_HIGH_OCC 128
|
||||
|
||||
void mm_seed_select(int32_t n, mm_seed_t *a, int len, int max_occ, int max_max_occ, int dist)
|
||||
{ // for high-occ minimizers, choose up to max_high_occ in each high-occ streak
|
||||
extern void ks_heapdown_uint64_t(size_t i, size_t n, uint64_t*);
|
||||
extern void ks_heapmake_uint64_t(size_t n, uint64_t*);
|
||||
int32_t i, last0, m;
|
||||
uint64_t b[MAX_MAX_HIGH_OCC]; // this is to avoid a heap allocation
|
||||
|
||||
if (n == 0 || n == 1) return;
|
||||
for (i = m = 0; i < n; ++i)
|
||||
if (a[i].n > max_occ) ++m;
|
||||
if (m == 0) return; // no high-frequency k-mers; do nothing
|
||||
for (i = 0, last0 = -1; i <= n; ++i) {
|
||||
if (i == n || a[i].n <= max_occ) {
|
||||
if (i - last0 > 1) {
|
||||
int32_t ps = last0 < 0? 0 : (uint32_t)a[last0].q_pos>>1;
|
||||
int32_t pe = i == n? len : (uint32_t)a[i].q_pos>>1;
|
||||
int32_t j, k, st = last0 + 1, en = i;
|
||||
int32_t max_high_occ = (int32_t)((double)(pe - ps) / dist + .499);
|
||||
if (max_high_occ > 0) {
|
||||
if (max_high_occ > MAX_MAX_HIGH_OCC)
|
||||
max_high_occ = MAX_MAX_HIGH_OCC;
|
||||
for (j = st, k = 0; j < en && k < max_high_occ; ++j, ++k)
|
||||
b[k] = (uint64_t)a[j].n<<32 | j;
|
||||
ks_heapmake_uint64_t(k, b); // initialize the binomial heap
|
||||
for (; j < en; ++j) { // if there are more, choose top max_high_occ
|
||||
if (a[j].n < (int32_t)(b[0]>>32)) { // then update the heap
|
||||
b[0] = (uint64_t)a[j].n<<32 | j;
|
||||
ks_heapdown_uint64_t(0, k, b);
|
||||
}
|
||||
}
|
||||
for (j = 0; j < k; ++j) a[(uint32_t)b[j]].flt = 1;
|
||||
}
|
||||
for (j = st; j < en; ++j) a[j].flt ^= 1;
|
||||
for (j = st; j < en; ++j)
|
||||
if (a[j].n > max_max_occ)
|
||||
a[j].flt = 1;
|
||||
}
|
||||
last0 = i;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
mm_seed_t *mm_collect_matches(void *km, int *_n_m, int qlen, int max_occ, int max_max_occ, int dist, const mm_idx_t *mi, const mm128_v *mv, int64_t *n_a, int *rep_len, int *n_mini_pos, uint64_t **mini_pos)
|
||||
{
|
||||
int rep_st = 0, rep_en = 0, n_m, n_m0;
|
||||
size_t i;
|
||||
mm_seed_t *m;
|
||||
*n_mini_pos = 0;
|
||||
*mini_pos = (uint64_t*)kmalloc(km, mv->n * sizeof(uint64_t));
|
||||
m = mm_seed_collect_all(km, mi, mv, &n_m0);
|
||||
if (dist > 0 && max_max_occ > max_occ) {
|
||||
mm_seed_select(n_m0, m, qlen, max_occ, max_max_occ, dist);
|
||||
} else {
|
||||
for (i = 0; i < n_m0; ++i)
|
||||
if (m[i].n > max_occ)
|
||||
m[i].flt = 1;
|
||||
}
|
||||
for (i = 0, n_m = 0, *rep_len = 0, *n_a = 0; i < n_m0; ++i) {
|
||||
mm_seed_t *q = &m[i];
|
||||
//fprintf(stderr, "X\t%d\t%d\t%d\n", q->q_pos>>1, q->n, q->flt);
|
||||
if (q->flt) {
|
||||
int en = (q->q_pos >> 1) + 1, st = en - q->q_span;
|
||||
if (st > rep_en) {
|
||||
*rep_len += rep_en - rep_st;
|
||||
rep_st = st, rep_en = en;
|
||||
} else rep_en = en;
|
||||
} else {
|
||||
*n_a += q->n;
|
||||
(*mini_pos)[(*n_mini_pos)++] = (uint64_t)q->q_span<<32 | q->q_pos>>1;
|
||||
m[n_m++] = *q;
|
||||
}
|
||||
}
|
||||
*rep_len += rep_en - rep_st;
|
||||
*_n_m = n_m;
|
||||
return m;
|
||||
}
|
||||
@@ -4,26 +4,26 @@ except ImportError:
|
||||
from distutils.core import setup
|
||||
from distutils.extension import Extension
|
||||
|
||||
cmdclass = {}
|
||||
import sys, platform
|
||||
|
||||
try:
|
||||
from Cython.Build import build_ext
|
||||
except ImportError: # without Cython
|
||||
module_src = 'python/mappy.c'
|
||||
else: # with Cython
|
||||
module_src = 'python/mappy.pyx'
|
||||
cmdclass['build_ext'] = build_ext
|
||||
|
||||
import sys
|
||||
sys.path.append('python')
|
||||
|
||||
extra_compile_args = ['-DHAVE_KALLOC']
|
||||
include_dirs = ["."]
|
||||
|
||||
if platform.machine() in ["aarch64", "arm64"]:
|
||||
include_dirs.append("sse2neon/")
|
||||
extra_compile_args.extend(['-ftree-vectorize', '-DKSW_SSE2_ONLY', '-D__SSE2__'])
|
||||
else:
|
||||
extra_compile_args.append('-msse4.1') # WARNING: ancient x86_64 CPUs don't have SSE4
|
||||
|
||||
def readme():
|
||||
with open('python/README.rst') as f:
|
||||
return f.read()
|
||||
with open('python/README.rst') as f:
|
||||
return f.read()
|
||||
|
||||
setup(
|
||||
name = 'mappy',
|
||||
version = '2.8',
|
||||
version = '2.21',
|
||||
url = 'https://github.com/lh3/minimap2',
|
||||
description = 'Minimap2 python binding',
|
||||
long_description = readme(),
|
||||
@@ -32,18 +32,18 @@ setup(
|
||||
license = 'MIT',
|
||||
keywords = 'sequence-alignment',
|
||||
scripts = ['python/minimap2.py'],
|
||||
ext_modules = [Extension('mappy',
|
||||
sources = [module_src, 'align.c', 'bseq.c', 'chain.c', 'format.c', 'hit.c', 'index.c', 'pe.c', 'options.c',
|
||||
ext_modules = [Extension('mappy',
|
||||
sources = ['python/mappy.pyx', 'align.c', 'bseq.c', 'lchain.c', 'seed.c', 'format.c', 'hit.c', 'index.c', 'pe.c', 'options.c',
|
||||
'ksw2_extd2_sse.c', 'ksw2_exts2_sse.c', 'ksw2_extz2_sse.c', 'ksw2_ll_sse.c',
|
||||
'kalloc.c', 'kthread.c', 'map.c', 'misc.c', 'sdust.c', 'sketch.c', 'esterr.c'],
|
||||
'kalloc.c', 'kthread.c', 'map.c', 'misc.c', 'sdust.c', 'sketch.c', 'esterr.c', 'splitidx.c'],
|
||||
depends = ['minimap.h', 'bseq.h', 'kalloc.h', 'kdq.h', 'khash.h', 'kseq.h', 'ksort.h',
|
||||
'ksw2.h', 'kthread.h', 'kvec.h', 'mmpriv.h', 'sdust.h',
|
||||
'python/cmappy.h', 'python/cmappy.pxd'],
|
||||
extra_compile_args = ['-DHAVE_KALLOC', '-msse4'], # WARNING: ancient x86_64 CPUs don't have SSE4
|
||||
include_dirs = ['.'],
|
||||
extra_compile_args = extra_compile_args,
|
||||
include_dirs = include_dirs,
|
||||
libraries = ['z', 'm', 'pthread'])],
|
||||
classifiers = [
|
||||
'Development Status :: 4 - Beta',
|
||||
'Development Status :: 5 - Production/Stable',
|
||||
'License :: OSI Approved :: MIT License',
|
||||
'Operating System :: POSIX',
|
||||
'Programming Language :: C',
|
||||
@@ -52,4 +52,4 @@ setup(
|
||||
'Programming Language :: Python :: 3',
|
||||
'Intended Audience :: Science/Research',
|
||||
'Topic :: Scientific/Engineering :: Bio-Informatics'],
|
||||
cmdclass = cmdclass)
|
||||
setup_requires=["cython"])
|
||||
|
||||
+84
@@ -0,0 +1,84 @@
|
||||
#include <string.h>
|
||||
#include <assert.h>
|
||||
#include <stdlib.h>
|
||||
#include <stdio.h>
|
||||
#include <errno.h>
|
||||
#include "mmpriv.h"
|
||||
|
||||
FILE *mm_split_init(const char *prefix, const mm_idx_t *mi)
|
||||
{
|
||||
char *fn;
|
||||
FILE *fp;
|
||||
uint32_t i, k = mi->k;
|
||||
fn = (char*)calloc(strlen(prefix) + 10, 1);
|
||||
sprintf(fn, "%s.%.4d.tmp", prefix, mi->index);
|
||||
if ((fp = fopen(fn, "wb")) == NULL) {
|
||||
if (mm_verbose >= 1)
|
||||
fprintf(stderr, "[ERROR]\033[1;31m failed to write to temporary file '%s'\033[0m: %s\n", fn, strerror(errno));
|
||||
exit(1);
|
||||
}
|
||||
mm_err_fwrite(&k, 4, 1, fp);
|
||||
mm_err_fwrite(&mi->n_seq, 4, 1, fp);
|
||||
for (i = 0; i < mi->n_seq; ++i) {
|
||||
uint32_t l;
|
||||
l = strlen(mi->seq[i].name);
|
||||
mm_err_fwrite(&l, 1, 4, fp);
|
||||
mm_err_fwrite(mi->seq[i].name, 1, l, fp);
|
||||
mm_err_fwrite(&mi->seq[i].len, 4, 1, fp);
|
||||
}
|
||||
free(fn);
|
||||
return fp;
|
||||
}
|
||||
|
||||
mm_idx_t *mm_split_merge_prep(const char *prefix, int n_splits, FILE **fp, uint32_t *n_seq_part)
|
||||
{
|
||||
mm_idx_t *mi = 0;
|
||||
char *fn;
|
||||
int i, j;
|
||||
|
||||
if (n_splits < 1) return 0;
|
||||
fn = CALLOC(char, strlen(prefix) + 10);
|
||||
for (i = 0; i < n_splits; ++i) {
|
||||
sprintf(fn, "%s.%.4d.tmp", prefix, i);
|
||||
if ((fp[i] = fopen(fn, "rb")) == 0) {
|
||||
if (mm_verbose >= 1)
|
||||
fprintf(stderr, "ERROR: failed to open temporary file '%s': %s\n", fn, strerror(errno));
|
||||
for (j = 0; j < i; ++j)
|
||||
fclose(fp[j]);
|
||||
free(fn);
|
||||
return 0;
|
||||
}
|
||||
}
|
||||
free(fn);
|
||||
|
||||
mi = CALLOC(mm_idx_t, 1);
|
||||
for (i = 0; i < n_splits; ++i) {
|
||||
mm_err_fread(&mi->k, 4, 1, fp[i]); // TODO: check if k is all the same
|
||||
mm_err_fread(&n_seq_part[i], 4, 1, fp[i]);
|
||||
mi->n_seq += n_seq_part[i];
|
||||
}
|
||||
mi->seq = CALLOC(mm_idx_seq_t, mi->n_seq);
|
||||
for (i = j = 0; i < n_splits; ++i) {
|
||||
uint32_t k;
|
||||
for (k = 0; k < n_seq_part[i]; ++k, ++j) {
|
||||
uint32_t l;
|
||||
mm_err_fread(&l, 1, 4, fp[i]);
|
||||
mi->seq[j].name = (char*)calloc(l + 1, 1);
|
||||
mm_err_fread(mi->seq[j].name, 1, l, fp[i]);
|
||||
mm_err_fread(&mi->seq[j].len, 4, 1, fp[i]);
|
||||
}
|
||||
}
|
||||
return mi;
|
||||
}
|
||||
|
||||
void mm_split_rm_tmp(const char *prefix, int n_splits)
|
||||
{
|
||||
int i;
|
||||
char *fn;
|
||||
fn = CALLOC(char, strlen(prefix) + 10);
|
||||
for (i = 0; i < n_splits; ++i) {
|
||||
sprintf(fn, "%s.%.4d.tmp", prefix, i);
|
||||
remove(fn);
|
||||
}
|
||||
free(fn);
|
||||
}
|
||||
+1
-1
@@ -1,4 +1,4 @@
|
||||
>MT_orang
|
||||
>MT_orang co:Z:comment
|
||||
GTTTATGTAGCTTATTCTATCCAAAGCAATGCACTGAAAATGTCTCGACGGGCCCACACG
|
||||
CCCCATAAACAAATAGGTTTGGTCCTAGCCTTTCTATTAGCTCTTAGTGAGGTTACACAT
|
||||
GCAAGCATCCCCGCCCCAGTGAGTCGCCCTCCAAGTCACTCTGACTAAGAGGAGCAAGCA
|
||||
|
||||
File diff suppressed because one or more lines are too long
+127
@@ -0,0 +1,127 @@
|
||||
>ref
|
||||
TGCGGAGGCTGAAGCAACTCCATCTTGGAAGCTAATCTACCATGTTGGCTTCTGATTAAC
|
||||
ATCAGTTCTGGGAAGGCTTGTAAGATTTCCTGTTTGTCTATTATTTCCTAGGTAAGAGCA
|
||||
GATACTTACTGTAAATCCTGCCCCTAGATTAAACAACCTTGGTGTTATCGTACTTCCATT
|
||||
GTCCTATACATCCCTTCGGAATCCCCCTTTCCCTATGGTCCTCAAGCCCTTGGTCTGGGG
|
||||
AGTAACAGCATAGGGATCAACCATCTCGTCTTGCCACTGCCCGAAATACAGACATGGCTT
|
||||
CTGTTCCTAAGTCCCTATTCAACTTTTCTTTCTAAGAAACTGGATTTGTCAGCCTCTTTC
|
||||
TTCACCTCTCAGCTTCCTTGGACTTTGGGGGTAGGTTTGCGTAGACATGCTCACCACAGA
|
||||
CACAATATCAGCTTCATTCTACAGATGAGGAAGGCAAGCCTTGGGGAGCTTAACCAACTT
|
||||
GTCGAGACTCATGTATATACCAACACTGAAAAGCAGATATTCCAGACTCCCAGTCATGCC
|
||||
ACAGGCACACCCCTCAGTGAGAGGTGGGGTTTGTAGTTGAGGCTATTTCCTGCCCAGGGA
|
||||
GCAGGGAGGCACTCTAGCTTCCCTGAGCTAACGTGGTTCTGCTTGTGTCTGACTTCCAGG
|
||||
TCTCTGCCCTTTCCAAGCTCACTAGGATGGGCTTCGGGTGTGTCAAATGCCTCAGACAGT
|
||||
ACAGATCCACACAGAATGGGCATATGCAACCAATCAGTGTCATAAAAAAGAAGGAAATGA
|
||||
CTCGGGCCCCCTGTGTGTTCAACATGTCGAAGGTATCTGTGCAGCAGAAGAAAGAGGGGC
|
||||
AAAAGCCCCCAGTGCCACAGGCCAGAGGCAGCAGCTTGGGCCCATGTGGGAGGGTTTGCT
|
||||
TTCCCCTGCCAAAGTGATGGGCTGCTGCAGCCTGGGGCTTGTGGGAATCCTTCCTGGGCC
|
||||
TGTGTGGGAAGTGTAGGCAGGGAGAGTGCTGCTTTCCCAAGCTCATCCCAGCTACAGCTA
|
||||
CCTTTGTGCTCTGGGATTCAGGACCCCCGAGGGGGCTGGCAGGAGAGTCTCTGTTCTCGG
|
||||
ATGGGTTGTCACCAGGGCATACATGGGAAGTGGGCTCTCTGGAGTCACCCTCCAGGGGAC
|
||||
AATGCCAATTCCAGACACATTTACTGGAACCCCTACACTGATGACCTTTTGTTGAGGGTT
|
||||
GAATTATGTCCCCAAAAAAGATACATTGAAGTCCAAACCTCTGGTGTCTATAAATGTGAT
|
||||
TTTATTTGAAAATGAGGTTTCTATGGACTAAATTGTGTCCCTCCCAAATTCATATTTTGA
|
||||
AGCCCTAGCCCCCAGTGTGACTATACCTAGAGACAGAGATCTTTAGGAGGTAATTAAGGT
|
||||
TCAATGAGGTCAGGTGGGTGGGGCCCTAAACCAACAGGAAGGACTGTGGCCTTACTAGAA
|
||||
AAGGAAGAAAAAGCATTTCCTCTCTTCTAGTATAAAAGGACACAGAAAGAAGGCAGATAT
|
||||
CTACAAGCCACGAAGAGAGACGTCACTGAGAACTGAATTTGTGTACATTGATCTGGAACT
|
||||
TCCAGCCTCCAGAACTTGAGAAATACATTTCTGTTGTTTATTTTTTTTTCATGTAATCAA
|
||||
TTCATTTATCATATATTTATTGAGTGCCTACTATGTGCCAGAGGATACAGCAGTAACAAA
|
||||
ACTAGGCAAAAATTGTGCCTAAAAGAGGGAAGATGACTTTTCTTAAAGTGTGGAATAAAG
|
||||
AAAAGTAAGATAGCGGATAGAAGCTTGAAGTGAAAGCAGGTTCACAGGAAGTTTCTTTGG
|
||||
TCATTTGTTTTGTTTTTAAATAGTGGAAAGATGTATATGTTTATGGAGAAAGATTGCCTT
|
||||
GAAGATGCAAGAGGAAGAGATGATCAAAATTCAAGAAGAAGCAGAAAGTGATAGAATAAA
|
||||
GAGCACAAGTGGAGAATTAGTGTTAATGAAAAGAAGGATGCTTCCTTTGATATGAAGTGA
|
||||
AGGAAGAGAGAATGAGTAAAGACCAAGACTTGAAGTCCCTAGTTTAATAGAGGGAGATTT
|
||||
CTTCTTTTGATAGCAACAATGGTATTCTGAATTATTTGAAGACATGTCATATTTCTCTTG
|
||||
TGCCATTTTCCTCCCAGTTTAAACATTCTCATAACCTCTATTCCTCACATGATGTTTTTC
|
||||
CAGGTCCTTTATTCTTTGGCACTCTCTTCTCTGGACACATTGTATTCTGTCATTGGTCCT
|
||||
AAAATTTAGATACCCACAATTGAACATACTCCTCTAGATATGGTCTAGCTAATGCAAAAG
|
||||
AACTGCTGCCTTCCAACTTGTTCAGACATCATATGTTTGTTGTCAAACGCTAAGTTGAGT
|
||||
TGTTATCTTTTAAGTTTTGTTTTTGTTTTTTTTTTTTTTTTTTAATTCCAAGAGGTGCCC
|
||||
ACGTTGGCTAAGTACCAAACAGGGTACTAGGGAATTTTACTTCTGAGTTAAATGCCATTC
|
||||
TAGTTGTTTTTTCTTCATCTCCAGTAAGGTTATCTTTATTCACCAGTTGTTACAATAGCT
|
||||
GTGGGTCTTGCTTCTCACAGTTTTATGCTGTCTGTGCTATTTTCTCTACTGATCATCACC
|
||||
ACAATCATTATTGCTTATCATAATTGTTATCTTTATTTTCTCCTTTAATCAAGAATCAGT
|
||||
CTTCCTTTATCTCATTATTCTCTTTTGCAGGCTTCAGGATAATTATGGTTGGAGTGCACT
|
||||
GGGGGAACCAGTGCAGCTAAGCTCTGACATCTTTGCATCCCTTTTCCATCTGCTGTTTTG
|
||||
GCACTCTGGTAGAATAGATAACCTAAAAACGACTTTAAAACATCTAGAAATTTTGGATAA
|
||||
AATATAACAAACATCCCTTTAAATGCACAACTGATCTTCCATGGAAGTCACAGAAATATA
|
||||
TAACGCCAAAAAGAAGGGAAGCTGAAACCCAGGGCTGTAAACATGAACATCATCTTCTCT
|
||||
CCCTTTTTCTTGTGACTTATCTTGTTTTTCTCAGCTTTGGTGCTACCAAGGCTTGACTTT
|
||||
AATAGGCATTTCCAATCAATGAGAGAATTTCTTTTGCTTTCATCAACAATTCAGTTATTG
|
||||
ATGTTAACATATATATCATTTGAGTACTTTTCTTTTTTTTATTATTATTATACTTTAAGT
|
||||
TTTAGGGTCCATGTGCACAATGTGCAGGTTAGTTACGTATGTATACATGTGCCATGCTGG
|
||||
TGTGCTGCACCCATTAACTCATCATTTAGCATTAGGTATATCTCCTAATGCTATCCCTTC
|
||||
CCCCTCTCCCCACCCCACAACAGTCCCCAGAGTGTTCCCCTTCCTGTGTCCATGTGTTCT
|
||||
CATTGTTCAATCCCCATCTATGAGTGAGAACATGCGGTGTTTGGTTTTTTGTCCTTGCAA
|
||||
TAGTTTACTGAGAATGATGATTTCTAATTTCATCCATGTCCCTAAAGAGCTTCTGCACAG
|
||||
CAAAAGAAACTACCATCAGAGTGAACAGGCAACCTACAAAATGGGAGAAAATTTTCACAA
|
||||
CCTGCTCATCTGACAAAGGGCTAATATCCAGAATCTACAATGAACTCAAACAAATTTACA
|
||||
AGAAAAAAACAAACAACCCCATCAAAAAGTGGGCAAAGGATATGAACAGACACTTCTCAA
|
||||
AAGAAGACATTTATGCAGCCAAAAGACACATGAAAAAATGCTCATCATCACTGGCCATCA
|
||||
GAGAAATGCAAACCAAAACCACAATGAGATACCATCTCACACCAGTTAAAATGGCAATCA
|
||||
TTAAAAAGTCAGGAAACAACAGGTGCTGGAGAGGATGTGGAGAAACAGGAACACTTTTAC
|
||||
ACTGTTGGTGGGACTGTAAACTAGTTCAACCATTGTGGAAGTCAGTGTGCTGATTCCTCA
|
||||
GGGATCTAGAACTAGAAATACCATTTGACCCAGCCATCCCATTACTGGGTATATACCCAA
|
||||
AGGACTATAAATCATGCTGCTATAAAGACACATGCACACGTATGTTTATTGCGGCACTAT
|
||||
TCACAATAGCAAAGACTTGGAACCAACCCAAATGTCCAACAATGATAGACTGGATTAAGA
|
||||
AAATGTGGCACATATACACCACGGAATACTGTGCAGCCATAAAAAATGATGAGTTCATGT
|
||||
CCTTTGTAGGGACACGGATGAAATTGGAAATCATTTCTGTTGTTTAAACCACGAAGTCTA
|
||||
TGGTATCTGGTTATGACAACCTGAGAATACTAACTCAAGGGTCTTTCGCAGATGTCATTA
|
||||
AGTTGTTAAAGTGAGGTCATTATGGTGGGTCCTAATCCAAGAGAAGAGATGCATGGACAG
|
||||
ACGTGCACAACGGGAGGACCAAGCCAAGACACACAGGGAGAATGGCCATGGGAAGATGGA
|
||||
GGCAGAGATCAAAGTGAGGCACCCACAAGCCAAGAAATGGCAGGAGCTACCAGCAGCTGG
|
||||
AAGATGCAGAGAAGCATTCCTTCTTAGAGGTTTCAGAGAGAGTATGGTGCTACTGACACC
|
||||
TTGATTTTGAACTTCTAGTCTCCAGAACTATGAGAGAATAAATTTCTGTTGGTTAAGCCA
|
||||
TCGAGTTTGTGTAAGTTTGTTATAAGAGCCCTAGGAAATAAACATATCCATTTATTCAGG
|
||||
AAAGCCTGCTAGAGTGCAAATATTTGGAAAAGATACTACTATGCAAATGTTTGAAAAAGA
|
||||
TATTGCTCTTGATTCTGCCTTATGGGTTTTTCATTTCTGTAAGCTATTCTCAAAGTTTTG
|
||||
TTCTTGGACTACTATTGGTAATTAAGACTGCAACATGTTTGGCAACATCAGTTGAGAACT
|
||||
GTTGCTCTGGGAACGTTTTCGGCAAGCCTCAGCCCTTCTTTTCCCTTGGCTTGCATTGAG
|
||||
GAGTTAGGTGATACTCTGCTGCTCAGGCCCAGCACCTTTATGGACCGTATTCCCCTGGTG
|
||||
GAATGACCATCTCTGCTTGCTCTGATTGGCTGTTGGGGTTTTCTAGCATGCCCTATTTAA
|
||||
TATGTATGATTTATCTCTTACTTCAGTTGGAAGGTACAGTTGCTCTGTAGTTGGCATGCA
|
||||
GTCATGGTGACTATGAAAATATAAAATAATGTTTTGGTTTACAGACACTTAGAAATAAGT
|
||||
TGTGTCTCAAAATTGGGTGACTATTCTAGTTATCTGCTACTCAATATCCTTGTGCGAGCC
|
||||
CTCTTTACCCAGAATCAAACTAAACCATGAGGGGCACTATAGAATGTCACCCCTGGGTCC
|
||||
AGGATACTATGGGGACTCAGAAGCCAAGCTCCCACTGGGGGATCTAGGGCATGCCCCCAA
|
||||
GGTAAGATTCCCACCTCTTTGTTCAGCAGGAAGCACCCATCACACAAGGAGGTAGGAATA
|
||||
AACAAGCATTCGTCAAGAACAAAAGATACAGATGTTCTGCTGGAGCTTGGATACATAGCA
|
||||
TAAGAGGGAACAGTTCTCACAGGTAAGAGTAAGTTTTCCTCTGGTGGTGACAGTGGGACC
|
||||
TGTGGGGGAGAGAATTGGGAGTACTGACAGGAAGGCAGAGTGGCTGTCCAAATGAACGGA
|
||||
TTGTTTGCACATGGCCTTTAGGGCACGTTGTGTTAGCCTTCCATTGCTGCTTATATTAGT
|
||||
CTGTTTTCACACTGCCCATAAATGCATACCTGAGACTGGATAATTTATAAAGAAAAAGAG
|
||||
CCTTAATGTACTCATAGTTGCATGTGGCTGGGGAGGCCTCACAATCATGGCAGAAGGTGA
|
||||
AAGGCACATCTTACATGGAAGCAGACAAGAGAGAATTGAGGACCAAGTGAAAGGGGTTTC
|
||||
CCCTTATAAAACCATCAGATCACATGAGACTTTTTCACCACCATGAGAACAGTAAGGGGA
|
||||
AAACTATGCTCATGATTCAATTGTCTCCCACTGGATTCCTCCCACAACACATAGGAATTA
|
||||
TGGGAGCTAAAATTCAAGATGAGATTTGGGTGAGGACACAGCCAAACCCTATCACTGCTG
|
||||
TAATCAATTCCCACCAACTTAGTGGCTCGAAACATCACAGATTTATGATCTTATGACGGT
|
||||
GGAGGTCCCCAAATGGATCTTCTAGGTCTAGAATCAAGGTATCAGCAGACCACTTCTTTT
|
||||
GGAGGCTCTGGTGGAGAAACCATTTCCTCGCCTTTTCCAGCTTCTAGAGGCTGCCCTTCT
|
||||
CATTCCTTGGTTCACGGCCACACTCATTTCCATCTCTGCTTCCACTGTGACAACTTCTCT
|
||||
GCCTCAGACCCTCCTGCTTTGCCTTTGTAAGGACCCTTGTGATGAGATCAGGCCCATCCA
|
||||
GGATTATCCCTCATCTCAAGACCTTTACCTTAATCACATTTGCAAGGTCTCTTCCACTGT
|
||||
GTCAGGTAACATTTTCACAGGTTCCAGGGATTAGGGTGTGGACATCTTGGGGAGCTGGAG
|
||||
GATATTATTTCATCTACCACACACATCTCTACCTTGTACAGGCAAGCACTTGCAAAGTGC
|
||||
AATGTGATCCTCTGGAGCCACTGTCCTCCCAGAGCTTATATATACTCTGAAAGTCAACTC
|
||||
TCAGACCACAGCCTCCTGTCCATGCACCACTCTCATCAACACCCCCACCCGAAACACTTT
|
||||
CACTCCACCCTCTTTGTCCCCTAACTCATGGAGAAGAAAATCTAATTAGTAGGAGTGGAA
|
||||
TTTGGCTTTCATCTTTACCAGTACTAGAAATATGGTGTGTGTCTTTTTGTAAAAATTCTC
|
||||
TCAACTAAATTGTTTTTATTAATTTCTGCAAAATGTGAACATCAACTCCCTTCATGTGAA
|
||||
TGTCAATAAGATTAAATGAGCTGTCTCAGCTCCTAGCCTGTGCAAGCTAACAGCTCAGGA
|
||||
GATGTTTATTTCTTTCCCTCTTCTTTCCTTAATGAAGCCCTCTCCTTTGACATCTTCAAT
|
||||
TCTGGAGCGCTTCTTTTCTGAGGCCTTGGCTCCCCCACATTGCCCACCCTTTTCCTGCTC
|
||||
GTCCACATTTCTGGCTTCTATTCTCTTGTCTTTACCATCTCCCTGAACAATGTTATCCGT
|
||||
TCCAATGACTTCAACAGTCTCTCCGCTTACATATGATGCCTCTCAAACTCTGATCTCCAA
|
||||
CTCTTCCAAAGAGCTCTGGACCTTTGTTCCAATTACCTGAAAAACATCTTCTTGGATGTC
|
||||
CCATTAGCACTGTTAAATCAAACAAGAATTTCCCTCCCTCCTGCCTTGCTGTAGTTCCCC
|
||||
TAGGGATTCGGTTGTGTGGGAAGATGTGTGGAGAGCTCTTAGTTGACTCCCTTCTCTGCA
|
||||
GTTCTACCTCTCTAGAGACTTGGAGGACCCACTGTTTCCGCCTCGCTTTTTCAGGCCTAG
|
||||
AGATTGCTCGCTCCTGGGCTGGCTGCTTCATAATTCCTTATTAGTAGTTTCCCAAGCTTA
|
||||
CATATCTGTAAATATTTACTTTAGTTAAATTCTCCCCAATTTCCACAATATGTTGGCTGC
|
||||
ACATGCTTTCTACTAGGAGTCACACAACTATGATAAGAACCAAGAAATATTAGTAAACGT
|
||||
TTTTTACCATTATTGGCCTATACCCTGGAATAGCCAACAATAACCTAGAACCTATGCAAC
|
||||
AAGAATATCCAACAAGAACCTAGAGACCTGTCAGTCTATAGGTGGGAACTACAGGATGAG
|
||||
A
|
||||
+151
-15
@@ -61,13 +61,6 @@
|
||||
Volume = {32},
|
||||
Year = {2016}}
|
||||
|
||||
@misc{Suzuki:2016,
|
||||
title = {Fast and accurate alignment tool for PacBio and Nanopore long reads},
|
||||
author = {Hajime Suzuki},
|
||||
journal = {Unpublished},
|
||||
howpublished = {\href{https://github.com/ocxtal/minialign}{https://github.com/ocxtal/minialign}},
|
||||
year = {2016}}
|
||||
|
||||
@misc{Ruan:2016,
|
||||
title = {Ultra-fast de novo assembler using long noisy reads},
|
||||
author = {Jue Ruan},
|
||||
@@ -172,14 +165,6 @@
|
||||
Volume = {29},
|
||||
Year = {2011}}
|
||||
|
||||
@article {Suzuki130633,
|
||||
author = {Suzuki, Hajime and Kasahara, Masahiro},
|
||||
title = {Acceleration Of Nucleotide Semi-Global Alignment With Adaptive Banded Dynamic Programming},
|
||||
year = {2017},
|
||||
note = {doi:10.1101/130633},
|
||||
publisher = {Cold Spring Harbor Labs Journals},
|
||||
journal = {bioRxiv}}
|
||||
|
||||
@article{Gotoh:1982aa,
|
||||
Author = {Gotoh, O},
|
||||
Journal = {J Mol Biol},
|
||||
@@ -313,3 +298,154 @@
|
||||
Title = {Assembling large genomes with single-molecule sequencing and locality-sensitive hashing},
|
||||
Volume = {33},
|
||||
Year = {2015}}
|
||||
|
||||
@article{Gurevich:2013aa,
|
||||
Author = {Gurevich, Alexey and others},
|
||||
Journal = {Bioinformatics},
|
||||
Pages = {1072-5},
|
||||
Title = {{QUAST}: quality assessment tool for genome assemblies},
|
||||
Volume = {29},
|
||||
Year = {2013}}
|
||||
|
||||
@article{Li:2010fk,
|
||||
Author = {Li, Heng and Durbin, Richard},
|
||||
Journal = {Bioinformatics},
|
||||
Pages = {589-95},
|
||||
Title = {Fast and accurate long-read alignment with {Burrows-Wheeler} transform},
|
||||
Volume = {26},
|
||||
Year = {2010}}
|
||||
|
||||
@article{Marcais:2018aa,
|
||||
Author = {Mar{\c c}ais, Guillaume and others},
|
||||
Journal = {PLoS Comput Biol},
|
||||
Pages = {e1005944},
|
||||
Title = {{MUMmer4}: A fast and versatile genome alignment system},
|
||||
Volume = {14},
|
||||
Year = {2018}}
|
||||
|
||||
@article{Li:2009ys,
|
||||
Author = {Li, Heng and others},
|
||||
Journal = {Bioinformatics},
|
||||
Pages = {2078-9},
|
||||
Title = {The {Sequence Alignment/Map format and SAMtools}},
|
||||
Volume = {25},
|
||||
Year = {2009}}
|
||||
|
||||
@article{Suzuki:2018aa,
|
||||
Author = {Suzuki, Hajime and Kasahara, Masahiro},
|
||||
Journal = {BMC Bioinformatics},
|
||||
Pages = {45},
|
||||
Title = {Introducing difference recurrence relations for faster semi-global alignment of long sequences},
|
||||
Volume = {19},
|
||||
Year = {2018}}
|
||||
|
||||
@article{Li:2018ab,
|
||||
Author = {Li, Heng},
|
||||
Journal = {Bioinformatics},
|
||||
Pages = {3094-3100},
|
||||
Title = {Minimap2: pairwise alignment for nucleotide sequences},
|
||||
Volume = {34},
|
||||
Year = {2018}}
|
||||
|
||||
@article{Jain:2020aa,
|
||||
Author = {Jain, Chirag and others},
|
||||
Journal = {Bioinformatics},
|
||||
Pages = {i111-i118},
|
||||
Title = {Weighted minimizer sampling improves long read mapping},
|
||||
Volume = {36},
|
||||
Year = {2020}}
|
||||
|
||||
@article{Miga:2020aa,
|
||||
Author = {Miga, Karen H and others},
|
||||
Journal = {Nature},
|
||||
Pages = {79-84},
|
||||
Title = {Telomere-to-telomere assembly of a complete human {X} chromosome},
|
||||
Volume = {585},
|
||||
Year = {2020}}
|
||||
|
||||
@article {Jain2020.11.01.363887,
|
||||
author = {Jain, Chirag and others},
|
||||
title = {A long read mapping method for highly repetitive reference sequences},
|
||||
elocation-id = {2020.11.01.363887},
|
||||
year = {2020},
|
||||
doi = {10.1101/2020.11.01.363887},
|
||||
publisher = {Cold Spring Harbor Laboratory},
|
||||
abstract = {About 5-10\% of the human genome remains inaccessible for functional analysis due to the presence of repetitive sequences such as segmental duplications and tandem repeat arrays. To enable high-quality resequencing of personal genomes, it is crucial to support end-to-end genome variant discovery using repeat-aware read mapping methods. In this study, we highlight the fact that existing long read mappers often yield incorrect alignments and variant calls within long, near-identical repeats, as they remain vulnerable to allelic bias. In the presence of a non-reference allele within a repeat, a read sampled from that region could be mapped to an incorrect repeat copy because the standard pairwise sequence alignment scoring system penalizes true variants.To address the above problem, we propose a novel, long read mapping method that addresses allelic bias by making use of minimal confidently alignable substrings (MCASs). MCASs are formulated as minimal length substrings of a read that have unique alignments to a reference locus with sufficient mapping confidence (i.e., a mapping quality score above a user-specified threshold). This approach treats each read mapping as a collection of confident sub-alignments, which is more tolerant of structural variation and more sensitive to paralog-specific variants (PSVs) within repeats. We mathematically define MCASs and discuss an exact algorithm as well as a practical heuristic to compute them. The proposed method, referred to as Winnowmap2, is evaluated using simulated as well as real long read benchmarks using the recently completed gapless assemblies of human chromosomes X and 8 as a reference. We show that Winnowmap2 successfully addresses the issue of allelic bias, enabling more accurate downstream variant calls in repetitive sequences. As an example, using simulated PacBio HiFi reads and structural variants in chromosome 8, Winnowmap2 alignments achieved the lowest false-negative and false-positive rates (1.89\%, 1.89\%) for calling structural variants within near-identical repeats compared to minimap2 (39.62\%, 5.88\%) and NGMLR (56.60\%, 36.11\%) respectively.Winnowmap2 code is accessible at https://github.com/marbl/WinnowmapCompeting Interest StatementThe authors have declared no competing interest.},
|
||||
URL = {https://www.biorxiv.org/content/early/2020/11/02/2020.11.01.363887},
|
||||
eprint = {https://www.biorxiv.org/content/early/2020/11/02/2020.11.01.363887.full.pdf},
|
||||
journal = {bioRxiv}
|
||||
}
|
||||
|
||||
@article{Li:2020aa,
|
||||
Author = {Li, Heng and others},
|
||||
Journal = {Genome Biol},
|
||||
Pages = {265},
|
||||
Title = {The design and construction of reference pangenome graphs with minigraph},
|
||||
Volume = {21},
|
||||
Year = {2020}}
|
||||
|
||||
@article{Ren:2021aa,
|
||||
Author = {Ren, Jingwen and Chaisson, Mark J P},
|
||||
Journal = {PLoS Comput Biol},
|
||||
Pages = {e1009078},
|
||||
Title = {lra: A long read aligner for sequences and contigs},
|
||||
Volume = {17},
|
||||
Year = {2021}}
|
||||
|
||||
@inproceedings{DBLP:conf/wabi/AbouelhodaO03,
|
||||
Author = {Mohamed Ibrahim Abouelhoda and Enno Ohlebusch},
|
||||
Booktitle = {Algorithms in Bioinformatics, Third International Workshop, {WABI} 2003, Budapest, Hungary, September 15-20, 2003, Proceedings},
|
||||
Crossref = {DBLP:conf/wabi/2003},
|
||||
Pages = {1--16},
|
||||
Title = {A Local Chaining Algorithm and Its Applications in Comparative Genomics},
|
||||
Year = {2003}}
|
||||
|
||||
@article{Ono:2021aa,
|
||||
Author = {Ono, Yukiteru and others},
|
||||
Journal = {Bioinformatics},
|
||||
Pages = {589-595},
|
||||
Title = {{PBSIM2}: a simulator for long-read sequencers with a novel generative model of quality scores},
|
||||
Volume = {37},
|
||||
Year = {2021}}
|
||||
|
||||
@article{Sedlazeck:2018ab,
|
||||
Author = {Sedlazeck, Fritz J and others},
|
||||
Journal = {Nat Methods},
|
||||
Pages = {461-468},
|
||||
Title = {Accurate detection of complex structural variations using single-molecule sequencing},
|
||||
Volume = {15},
|
||||
Year = {2018}}
|
||||
|
||||
@article{Jeffares:2017aa,
|
||||
Author = {Jeffares, Daniel C and others},
|
||||
Journal = {Nat Commun},
|
||||
Pages = {14061},
|
||||
Title = {Transient structural variations have strong effects on quantitative traits and reproductive isolation in fission yeast},
|
||||
Volume = {8},
|
||||
Year = {2017}}
|
||||
|
||||
@article{Zook:2020aa,
|
||||
Author = {Zook, Justin M and others},
|
||||
Journal = {Nat Biotechnol},
|
||||
Pages = {1347-1355},
|
||||
Title = {A robust benchmark for detection of germline large deletions and insertions},
|
||||
Volume = {38},
|
||||
Year = {2020}}
|
||||
|
||||
@article{Harpak:2017aa,
|
||||
Author = {Harpak, Arbel and others},
|
||||
Journal = {Proc Natl Acad Sci U S A},
|
||||
Pages = {12779-12784},
|
||||
Title = {Frequent nonallelic gene conversion on the human lineage and its effect on the divergence of gene duplicates},
|
||||
Volume = {114},
|
||||
Year = {2017}}
|
||||
|
||||
@article{Li:2018aa,
|
||||
Author = {Li, Heng and others},
|
||||
Journal = {Nat Methods},
|
||||
Month = {Aug},
|
||||
Number = {8},
|
||||
Pages = {595-597},
|
||||
Title = {A synthetic-diploid benchmark for accurate variant-calling evaluation},
|
||||
Volume = {15},
|
||||
Year = {2018}}
|
||||
|
||||
+123
-61
@@ -1,6 +1,6 @@
|
||||
\documentclass{bioinfo}
|
||||
\copyrightyear{2017}
|
||||
\pubyear{2017}
|
||||
\copyrightyear{2018}
|
||||
\pubyear{2018}
|
||||
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
@@ -19,7 +19,7 @@
|
||||
\begin{document}
|
||||
\firstpage{1}
|
||||
|
||||
\title[Aligning nucleotide sequences with minimap2]{Minimap2: versatile pairwise alignment for nucleotide sequences}
|
||||
\title[Aligning nucleotide sequences with minimap2]{Minimap2: pairwise alignment for nucleotide sequences}
|
||||
\author[Li]{Heng Li}
|
||||
\address{Broad Institute, 415 Main Street, Cambridge, MA 02142, USA}
|
||||
|
||||
@@ -40,9 +40,10 @@ full-length noisy Direct RNA or cDNA reads, and assembly contigs or closely
|
||||
related full chromosomes of hundreds of megabases in length. Minimap2 does
|
||||
split-read alignment, employs concave gap cost for long insertions and
|
||||
deletions (INDELs) and introduces new heuristics to reduce spurious alignments.
|
||||
It is 3--4 times faster than mainstream short-read mappers at comparable
|
||||
accuracy and $\ge$30 times faster at higher accuracy for both genomic and mRNA
|
||||
reads, surpassing most aligners specialized in one type of alignment.
|
||||
It is 3--4 times as fast as mainstream short-read mappers at comparable
|
||||
accuracy, and is $\ge$30 times faster than long-read genomic or cDNA
|
||||
mappers at higher accuracy, surpassing most aligners specialized in one type of
|
||||
alignment.
|
||||
|
||||
\section{Availability and implementation:}
|
||||
\href{https://github.com/lh3/minimap2}{https://github.com/lh3/minimap2}
|
||||
@@ -63,7 +64,7 @@ the thought that 10kb long sequences should be easier to map than 100bp reads
|
||||
because we can more effectively skip repetitive regions, which are often the
|
||||
bottleneck of short-read alignment. We confirmed our speculation by achieving
|
||||
approximate mapping 50 times faster than BWA-MEM~\citep{Li:2016aa}.
|
||||
\citet{Suzuki130633} extended our work with a fast and novel algorithm on
|
||||
\citet{Suzuki:2018aa} extended our work with a fast and novel algorithm on
|
||||
generating base-level alignment, which in turn inspired us to develop minimap2
|
||||
with added functionality.
|
||||
|
||||
@@ -87,12 +88,14 @@ the versatility of minimap2.
|
||||
|
||||
Minimap2 follows a typical seed-chain-align procedure as is used by most
|
||||
full-genome aligners. It collects minimizers~\citep{Roberts:2004fv} of the
|
||||
reference sequences and indexes them in a hash table. Then for each query
|
||||
sequence, minimap2 takes query minimizers as \emph{seeds}, finds matches to the
|
||||
reference, and identifies sets of colinear seeds, which are called
|
||||
reference sequences and indexes them in a hash table, with the key being the
|
||||
hash of a minimizer and the value being a list of locations of the minimizer
|
||||
copies. Then for each query
|
||||
sequence, minimap2 takes query minimizers as \emph{seeds}, finds exact matches
|
||||
(i.e. \emph{anchors}) to the reference, and identifies sets of colinear anchors as
|
||||
\emph{chains}. If base-level alignment is requested, minimap2 applies dynamic
|
||||
programming (DP) to extend from the ends of chains and to close unseeded
|
||||
regions between adjacent seeds in chains.
|
||||
programming (DP) to extend from the ends of chains and to close
|
||||
regions between adjacent anchors in chains.
|
||||
|
||||
Minimap2 uses indexing and seeding algorithms similar to
|
||||
minimap~\citep{Li:2016aa}, and furthers the predecessor with more accurate
|
||||
@@ -119,10 +122,13 @@ distance between two anchors is too large); otherwise
|
||||
\end{equation}
|
||||
In implementation, a gap of length $l$ costs
|
||||
\[
|
||||
\gamma_c(l)=0.01\cdot \bar{w}\cdot|l|+0.5\log_2|l|
|
||||
\gamma_c(l)=\left\{\begin{array}{ll}
|
||||
0.01\cdot \bar{w}\cdot|l|+0.5\log_2|l| & (l\not=0) \\
|
||||
0 & (l=0)
|
||||
\end{array}\right.
|
||||
\]
|
||||
where $\bar{w}$ is the average seed length. For $m$ anchors, directly computing all $f(\cdot)$ with
|
||||
Eq.~(\ref{eq:chain}) takes $O(m^2)$ time. Although theoretically faster
|
||||
where $\bar{w}$ is the average seed length. For $N$ anchors, directly computing all $f(\cdot)$ with
|
||||
Eq.~(\ref{eq:chain}) takes $O(N^2)$ time. Although theoretically faster
|
||||
chaining algorithms exist~\citep{Abouelhoda:2005aa}, they
|
||||
are inapplicable to generic gap cost, complex to implement and usually
|
||||
associated with a large constant. We introduced a simple heuristic to
|
||||
@@ -132,7 +138,7 @@ We note that if anchor $i$ is chained to $j$, chaining $i$ to a predecessor
|
||||
of $j$ is likely to yield a lower score. When evaluating Eq.~(\ref{eq:chain}),
|
||||
we start from anchor $i-1$ and stop the process if we cannot find a better
|
||||
score after up to $h$ iterations. This approach reduces the average time to
|
||||
$O(h\cdot m)$. In practice, we can almost always find the optimal chain with
|
||||
$O(hN)$. In practice, we can almost always find the optimal chain with
|
||||
$h=50$; even if the heuristic fails, the optimal chain is often close.
|
||||
|
||||
\subsubsection{Backtracking}
|
||||
@@ -146,9 +152,11 @@ in more than one chains.
|
||||
\subsubsection{Identifying primary chains}\label{sec:primary}
|
||||
In the absence of copy number changes, each query segment should not be mapped
|
||||
to two places in the reference. However, chains found at the previous step may
|
||||
have significant or complete overlaps due to repeats in the reference.
|
||||
have significant or complete overlaps due to repeats in the reference~\citep{Li:2010fk}.
|
||||
Minimap2 used the following procedure to identify \emph{primary chains} that do
|
||||
not greatly overlap on the query. Let $Q$ be an empty set initially. For each
|
||||
not greatly overlap on the query.
|
||||
|
||||
Let $Q$ be an empty set initially. For each
|
||||
chain from the best to the worst according to their chaining scores: if on the
|
||||
query, the chain overlaps with a chain in $Q$ by 50\% or higher percentage of
|
||||
the shorter chain, mark the chain as secondary to the chain in $Q$; otherwise,
|
||||
@@ -156,6 +164,16 @@ add the chain to $Q$. In the end, $Q$ contains all the primary chains. We did
|
||||
not choose a more sophisticated data structure (e.g. range tree or k-d tree)
|
||||
because this step is not the performance bottleneck.
|
||||
|
||||
For each primary chain, minimap2 estimates its mapping quality with an
|
||||
empirical formula:
|
||||
\[
|
||||
{\rm mapQ}=40\cdot (1-f_2/f_1)\cdot\min\{1,m/10\}\cdot\log f_1
|
||||
\]
|
||||
where $\log$ denotes natural logarithm, $m$ is the number of anchors on the primary chain, $f_1$ is the chaining
|
||||
score, and $f_2\le f_1$ is the score of the best chain that is secondary to the
|
||||
primary chain. Intuitively, a chain is assigned to a higher mapping quality if
|
||||
it is long and its best secondary chain is weak.
|
||||
|
||||
\subsubsection{Estimating per-base sequence divergence}
|
||||
Suppose a query sequence harbors $n$ seeds of length $k$, $m$ of which are
|
||||
present in a chain. We want to estimate the sequence divergence $\epsilon$
|
||||
@@ -186,7 +204,7 @@ $0.9$.
|
||||
\subsubsection{Indexing with homopolymer compressed $k$-mers}
|
||||
SmartDenovo
|
||||
(\href{https://github.com/ruanjue/smartdenovo}{https://github.com/ruanjue/smartdenovo};
|
||||
J Ruan, personal communication) indexes reads with homopolymer-compressed (HPC)
|
||||
J. Ruan, personal communication) indexes reads with homopolymer-compressed (HPC)
|
||||
$k$-mers and finds the strategy improves overlap sensitivity for SMRT reads.
|
||||
Minimap2 adopts the same heuristic.
|
||||
|
||||
@@ -199,9 +217,9 @@ To demonstrate the effectiveness of HPC $k$-mers, we performed read overlapping
|
||||
for the example {\it E. coli} SMRT reads from PBcR~\citep{Berlin:2015xy}, using
|
||||
different types of $k$-mers. With normal 15bp minimizers per 5bp window,
|
||||
minimap2 finds 90.9\% of $\ge$2kb overlaps inferred from the read-to-reference
|
||||
alignment. With HPC 19-mers, minimap2 finds 97.4\% of overlaps. It achieves this
|
||||
alignment. With HPC 19-mers per 5bp window, minimap2 finds 97.4\% of overlaps. It achieves this
|
||||
higher sensitivity by indexing 1/3 fewer minimizers, which further helps
|
||||
performance. HPC-based indexing reduces the sensitivity for ONT reads, though.
|
||||
performance. HPC-based indexing reduces the sensitivity for current ONT reads, though.
|
||||
|
||||
\subsection{Aligning genomic DNA}\label{sec:genomic}
|
||||
|
||||
@@ -240,14 +258,14 @@ performance of minimap2. Traditional SSE implementations~\citep{Farrar:2007hs}
|
||||
based on Eq.~(\ref{eq:ae86}) can achieve 16-way parallelization for short
|
||||
sequences, but only 4-way parallelization when the peak alignment score reaches
|
||||
32767. Long sequence alignment may exceed this threshold. Inspired by
|
||||
\citet{Wu:1996aa} and the following work, \citet{Suzuki130633} proposed a
|
||||
\citet{Wu:1996aa} and the following work, \citet{Suzuki:2018aa} proposed a
|
||||
difference-based formulation that lifted this limitation.
|
||||
In case of 2-piece gap cost, define
|
||||
\[
|
||||
\left\{\begin{array}{ll}
|
||||
u_{ij}\triangleq H_{ij}-H_{i-1,j} & v_{ij}\triangleq H_{ij}-H_{i,j-1} \\
|
||||
x_{ij}\triangleq E_{i+1,j}-H_{ij} & \tilde{x}_{ij}\triangleq \tilde{E}_{i+1,j}-\tilde{H}_{ij} \\
|
||||
y_{ij}\triangleq F_{i,j+1}-H_{ij} & \tilde{y}_{ij}\triangleq \tilde{F}_{i,j+1}-\tilde{H}_{ij}
|
||||
x_{ij}\triangleq E_{i+1,j}-H_{ij} & \tilde{x}_{ij}\triangleq \tilde{E}_{i+1,j}-H_{ij} \\
|
||||
y_{ij}\triangleq F_{i,j+1}-H_{ij} & \tilde{y}_{ij}\triangleq \tilde{F}_{i,j+1}-H_{ij}
|
||||
\end{array}\right.
|
||||
\]
|
||||
We can transform Eq.~(\ref{eq:ae86}) to
|
||||
@@ -307,11 +325,11 @@ y_{rt}&=&\max\{0,y_{r-1,t}+u_{r-1,t}-z_{rt}+q\}-q-e\\
|
||||
\end{equation*}
|
||||
In this formulation, cells with the same diagonal index $r$ are independent of
|
||||
each other. This allows us to fully vectorize the computation of all cells on
|
||||
the same anti-diagonal in one inner loop. It also simplifies banded alignment,
|
||||
the same anti-diagonal in one inner loop. It also simplifies banded alignment (500bp band width by default),
|
||||
which would be difficult with striped vectorization~\citep{Farrar:2007hs}.
|
||||
|
||||
On the condition that $q+e<\tilde{q}+\tilde{e}$ and $e>\tilde{e}$, the initial
|
||||
values in the diagonal-antidiagonal formuation is
|
||||
values in the diagonal-antidiagonal formuation are
|
||||
\[
|
||||
\left\{\begin{array}{l}
|
||||
x_{r-1,-1}=y_{r-1,r}=-q-e\\
|
||||
@@ -330,12 +348,19 @@ r\cdot(e-\tilde{e})-(\tilde{q}-q)-\tilde{e} & (r=\lceil\frac{\tilde{q}-q}{e-\til
|
||||
\]
|
||||
These can be derived from the initial values for Eq.~(\ref{eq:ae86}).
|
||||
|
||||
When performing global alignment, we do not need to compute $H_{rt}$ in each cell.
|
||||
We use 16-way vectorization throughout the alignment process. When extending
|
||||
alignments from ends of chains, we need to find the cell $(r,t)$ where $H_{rt}$
|
||||
reaches the maximum. We resort to 4-way vectorization to compute
|
||||
$H_{rt}=H_{r-1,t}+u_{rt}$. Because this computation is simple,
|
||||
Eq.~(\ref{eq:suzuki}) is still the dominant performance bottleneck.
|
||||
|
||||
In practice, our 16-way vectorized implementation of global alignment is three
|
||||
times as fast as Parasail's 4-way vectorization~\citep{Daily:2016aa}. Without
|
||||
banding, our implementation is slower than Edlib~\citep{Sosic:2017aa}, but with
|
||||
a 1000bp band, it is considerably faster. When performing global alignment
|
||||
between anchors, we expect the alignment to stay close to the diagonal of the
|
||||
DP matrix. Banding is applicable most of time.
|
||||
DP matrix. Banding is applicable most of the time.
|
||||
|
||||
\subsubsection{The Z-drop heuristic}
|
||||
|
||||
@@ -359,6 +384,16 @@ alignment between the two subsequences involved in the global alignment, but
|
||||
this time with the one subsequence reverse complemented. This additional
|
||||
alignment step may identify short inversions that are missed during chaining.
|
||||
|
||||
\subsubsection{Filtering out misplaced anchors}
|
||||
Due to sequencing errors and local homology, some anchors in a chain may be
|
||||
wrong. If we blindly align regions between two misplaced anchors, we will
|
||||
produce a suboptimal alignment. To reduce this artifact, we filter out
|
||||
anchors that lead to a $>$10bp insertion and a $>$10bp deletion at the same
|
||||
time, and filter out terminal anchors that lead to a long gap towards the ends
|
||||
of a chain. These heuristics greatly alleviate the issues with misplaced
|
||||
anchors, but they are unable to fix all such errors. Local misalignment is a
|
||||
limitation of minimap2 which we hope to address in future.
|
||||
|
||||
\subsection{Aligning spliced sequences}
|
||||
|
||||
The algorithm described above can be adapted to spliced alignment. In this
|
||||
@@ -397,7 +432,7 @@ p/2 & \mbox{if $T[i+1,i+3]$ is ${\tt GTC}$ or ${\tt GTT}$} \\
|
||||
p & \mbox{otherwise}
|
||||
\end{array}\right.\]
|
||||
where $T[i,j]$ extracts a substring of $T$ between $i$ and $j$ inclusively.
|
||||
$d(i)$ penalizes non-canonical donor sites with $p$ and less frequent Eukayotic
|
||||
$d(i)$ penalizes non-canonical donor sites with $p$ and less frequent Eukaryotic
|
||||
splicing signal ${\tt GT[C/T]}$ with $p/2$~\citep{Irimia:2008aa}. Similarly,
|
||||
\[a(i)=\left\{\begin{array}{ll}
|
||||
0 & \mbox{if $T[i-2,i]$ is ${\tt CAG}$ or ${\tt TAG}$} \\
|
||||
@@ -424,13 +459,13 @@ alignment.
|
||||
|
||||
\subsection{Aligning short paired-end reads}
|
||||
|
||||
During chainging, minimap2 takes a pair of reads as one fragment with a gap of
|
||||
During chaining, minimap2 takes a pair of reads as one fragment with a gap of
|
||||
unknown length in the middle. It applies a normal gap cost between seeds on the
|
||||
same read but is a more permissive gap cost between seeds on different reads.
|
||||
More precisely, the gap cost during chaining is:
|
||||
More precisely, the gap cost during chaining is ($l\not=0$):
|
||||
\[
|
||||
\gamma_c(l)=\left\{\begin{array}{ll}
|
||||
0.01\cdot\bar{w}\cdot l+0.5\log_2 l & \mbox{if two seeds on the same read} \\
|
||||
0.01\cdot\bar{w}\cdot |l|+0.5\log_2 |l| & \mbox{if two seeds on the same read} \\
|
||||
\min\{0.01\cdot\bar{w}\cdot|l|,\log_2|l|\} & \mbox{otherwise}
|
||||
\end{array}\right.
|
||||
\]
|
||||
@@ -443,6 +478,17 @@ consistent paired-end alignments.
|
||||
|
||||
\section{Results}
|
||||
|
||||
Minimap2 is implemented in the C programming language and comes with APIs in
|
||||
both C and Python. It is distributed under the MIT license, free to both
|
||||
commercial and academic uses. Minimap2 uses the same base algorithm for all
|
||||
applications, but it has to apply different sets of parameters depending on
|
||||
input data types. Similar to BWA-MEM, minimap2 introduces `presets' that
|
||||
modify multiple parameters with a simple invocation. Detailed settings
|
||||
and command-line options can be found in the minimap2 manpage. In addition to
|
||||
the applications evaluated in the following sections, minimap2 also retains
|
||||
minimap's functionality to find overlaps between long reads and to search
|
||||
against large multi-species databases such as \emph{nt} from NCBI.
|
||||
|
||||
\subsection{Aligning long genomic reads}\label{sec:long-genomic}
|
||||
|
||||
\begin{figure}[!tb]
|
||||
@@ -450,19 +496,22 @@ consistent paired-end alignments.
|
||||
\includegraphics[width=.5\textwidth]{roc-color.pdf}
|
||||
\caption{Evaluation on aligning simulated reads. Simulated reads were mapped
|
||||
to the primary assembly of human genome GRCh38. A read is considered correctly
|
||||
mapped if the true position overlaps with the best mapping position by 10\% of
|
||||
the read length. Read alignments are sorted by mapping quality in the
|
||||
descending order. For each mapping quality threshold, the fraction of
|
||||
alignments with mapping quality above the threshold and their error rate are
|
||||
mapped if its longest alignment overlaps with the true interval, and the
|
||||
overlap length is $\ge$10\% of the true interval length. Read alignments are
|
||||
sorted by mapping quality in the descending order. For each mapping quality
|
||||
threshold, the fraction of alignments (out of the number of input reads) with
|
||||
mapping quality above the threshold and their error rate are
|
||||
plotted along the curve. (a) long-read alignment evaluation. 33,088 $\ge$1000bp
|
||||
reads were simulated using pbsim~\citep{Ono:2013aa} with error profile sampled
|
||||
from file `m131017\_060208\_42213\_*.1.*' downloaded at
|
||||
\href{http://bit.ly/chm1p5c3}{http://bit.ly/chm1p5c3}. The N50 read length is
|
||||
11,628. Aligners were run under the default setting for SMRT reads.
|
||||
(b) short-read alignment evaluation. 10 million pairs of 150bp reads were
|
||||
simulated using mason2~\citep{Holtgrewe:2010aa} with option
|
||||
`\mbox{--illumina-prob-mismatch-scale 2.5}'. Short-read aligners were run under the
|
||||
default setting except for changing the maximum fragment length to
|
||||
Kart outputted all alignments at mapping quality 60, so is not shown in the
|
||||
figure. It mapped nearly all reads with 4.1\% of alignments being wrong, less
|
||||
accurate than others. (b) short-read alignment evaluation. 10 million pairs of
|
||||
150bp reads were simulated using mason2~\citep{Holtgrewe:2010aa} with option
|
||||
`\mbox{--illumina-prob-mismatch-scale 2.5}'. Short-read aligners were run under
|
||||
the default setting except for changing the maximum fragment length to
|
||||
800bp.}\label{fig:eval}
|
||||
\end{figure}
|
||||
|
||||
@@ -471,7 +520,7 @@ BLASR~(v1.MC.rc64; \citealp{Chaisson:2012aa}),
|
||||
BWA-MEM~(v0.7.15; \citealp{Li:2013aa}),
|
||||
GraphMap~(v0.5.2; \citealp{Sovic:2016aa}),
|
||||
Kart~(v2.2.5; \citealp{Lin:2017aa}),
|
||||
minialign~(v0.5.3; \citealp{Suzuki:2016}) and
|
||||
minialign~(v0.5.3; \href{https://github.com/ocxtal/minialign}{https://github.com/ocxtal/minialign}) and
|
||||
NGMLR~(v0.2.5; \citealp{Sedlazeck169557}). We excluded rHAT~\citep{Liu:2016ab}
|
||||
and LAMSA~\citep{Liu:2017aa} because they either
|
||||
crashed or produced malformatted output. In this evaluation, minimap2 has
|
||||
@@ -480,11 +529,11 @@ higher mapping accuracy (Fig.~\ref{fig:eval}a). Minimap2 and
|
||||
NGMLR provide better mapping quality estimate: they rarely give repetitive hits
|
||||
high mapping quality. Apparently, other aligners may
|
||||
occasionally miss close suboptimal hits and be overconfident in wrong mappings.
|
||||
On run time, minialign is slightly faster than minimap2 and Kart. They are over
|
||||
30 times faster than the rest. Minimap2 consumed 6.1GB memory at the peak,
|
||||
more than BWA-MEM but less than others.
|
||||
On run time, minimap2 took 200 CPU seconds, comparable to minialign and Kart, and is over
|
||||
30 times faster than the rest. Minimap2 consumed 6.8GB memory at the peak,
|
||||
more than BWA-MEM (5.4GB), similar to NGMLR and less than others.
|
||||
|
||||
On real human SMRT reads, the relative performance and sensitivity of
|
||||
On real human SMRT reads, the relative performance and fraction of mapped reads reported by
|
||||
these aligners are broadly similar to the metrics on simulated data. We are
|
||||
unable to provide a good estimate of mapping error rate due to the lack of the
|
||||
truth. On ONT $\sim$100kb human reads~\citep{Jain128835}, BWA-MEM failed.
|
||||
@@ -526,7 +575,7 @@ Peak RAM (GByte) & 8.9 & 14.5 & 3.2 & 29.2\vspace{1em}\\
|
||||
\% approx. introns & 91.8\% & 96.9\% & 92.5\% & 82.4\% \\
|
||||
\botrule
|
||||
\end{tabular}
|
||||
}{Mouse reads (AC:SRR5286960) were mapped to the primary assembly of mouse
|
||||
}{Mouse cDNA reads (AC:SRR5286960; R9.4 chemistry) were mapped to the primary assembly of mouse
|
||||
genome GRCm38 with the following tools and command options: minimap2 (`-ax
|
||||
splice'); GMAP (`-n 0 --min-intronlength 30 --cross-species'); SpAln (`-Q7 -LS
|
||||
-S3'); STARlong (according to
|
||||
@@ -535,7 +584,7 @@ compared to the EnsEMBL gene annotation, release 89. A predicted intron
|
||||
is \emph{novel} if it has no overlaps with any annotated introns. An intron
|
||||
is \emph{exact} if it is identical to an annotated intron. An intron is
|
||||
\emph{approximate} if both its 5'- and 3'-end are within 10bp around the ends
|
||||
of an annotated intron.}
|
||||
of an annotated intron. Chimeric alignments are defined in the SAM spec~\citep{Li:2009ys}.}
|
||||
\end{table}
|
||||
|
||||
We next aligned real mouse reads~\citep{Byrne:2017aa} with GMAP~(v2017-06-20;
|
||||
@@ -592,7 +641,7 @@ simulated data set than Bowtie2 and SNAP but less accurate than BWA-MEM
|
||||
(Fig.~\ref{fig:eval}b). Closer investigation reveals that BWA-MEM achieves
|
||||
a higher accuracy partly because it tries to locally align a read in a small
|
||||
region close to its mate. If we disable this feature, BWA-MEM becomes slightly
|
||||
less accurate than minimap2. We might consider to implement a similar heuristic
|
||||
less accurate than minimap2. We might implement a similar heuristic
|
||||
in minimap2 in future.
|
||||
|
||||
To evaluate the accuracy of minimap2 on real data, we aligned human reads
|
||||
@@ -603,19 +652,29 @@ across the whole genome and have been \emph{de novo} assembled with SMRT reads
|
||||
to high quality. This allowed us to construct an independent truth variant
|
||||
dataset~\citep{Li223297} for
|
||||
ERR1341796. In this evaluation, minimap2 has higher SNP false negative rate
|
||||
(FNR; 2.5\% of minimap2 vs 2.2\% of BWA-MEM), but fewer false positive SNPs per
|
||||
million bases (FPPM; 3.0 vs 3.9), lower 2--50bp INDEL FNR (7.3\% vs 7.5\%) and
|
||||
similar INDEL FPPM (both 1.0). Minimap2 is broadly similar to BWA-MEM in the
|
||||
(FNR; 2.6\% of minimap2 vs 2.3\% of BWA-MEM), but fewer false positive SNPs per
|
||||
million bases (FPPM; 7.0 vs 8.8), similar INDEL FNR (11.2\% vs 11.3\%) and
|
||||
similar INDEL FPPM (6.4 vs 6.5). Minimap2 is broadly comparable to BWA-MEM in the
|
||||
context of small variant calling.
|
||||
|
||||
\subsection{Other applications}
|
||||
\subsection{Aligning long-read assemblies}
|
||||
|
||||
Minimap2 retains minimap's functionality to find overlaps between long reads
|
||||
and to search against large multi-species databases such as \emph{nt} from
|
||||
NCBI. Minimap2 can also align similar genomes or different assemblies of the
|
||||
same species. It took 7 wall-clock minutes over 8 CPU cores to align a human
|
||||
SMRT assembly (AC:GCA\_001297185.1) to GRCh38, over 20 times faster
|
||||
MUMmer4~\citep{Kurtz:2004zr}.
|
||||
Minimap2 can align a SMRT assembly (AC:GCA\_001297185.1) against GRCh38 in 7
|
||||
minutes using 8 CPU cores, over 20 times faster than nucmer from
|
||||
MUMmer4~\citep{Marcais:2018aa}. With the paftools.js script from the minimap2
|
||||
package, we called 2.67 million single-base substitutions out of 2.78Gbp
|
||||
genomic regions. The transition-to-transversion ratio (ts/tv) is 2.01. In
|
||||
comparison, using MUMmer4's dnadiff pipeline, we called 2.86 million
|
||||
substitutions in 2.83Gbp at ts/tv=1.87. Given that ts/tv averaged across the
|
||||
human genome is about 2 but ts/tv averaged over random errors is 0.5, the
|
||||
minimap2 callset arguably has higher precision at lower sensitivity.
|
||||
|
||||
The sample being assembled is a female. Minimap2 still called 201 substitutions
|
||||
on the Y chromosome. These substitutions all come from one contig aligned at
|
||||
96.8\% sequence identity. The contig could be a segmental duplication
|
||||
absent from GRCh38. In constrast, dnadiff called 9070 substitutions on the Y
|
||||
chromosome across 73 SMRT contigs. This again implies our minimap2-based
|
||||
pipeline has higher precision.
|
||||
|
||||
\section{Discussions}
|
||||
|
||||
@@ -633,9 +692,9 @@ involving $>$100kb introns, which was impractically slow ten years ago. The
|
||||
minimap2 chaining algorithm is fast and highly accurate by itself. In fact,
|
||||
chaining alone is more accurate than all the other long-read mappers in
|
||||
Fig.~\ref{fig:eval}a (data not shown). This accuracy helps to reduce downstream
|
||||
base-level alignment of candidate chains, which is still times slower than
|
||||
base-level alignment of candidate chains, which is still several times slower than
|
||||
chaining even with the Suzuki-Kasahara improvement. In addition, taking a
|
||||
general form, minimap2 chaining can be adapted to non-typical data types such
|
||||
general form, minimap2 chaining can be adapted to non-typical data types such as
|
||||
spliced reads and multiple reads per fragment. This gives us the opportunity to
|
||||
extend the same base algorithm to a variety of use cases.
|
||||
|
||||
@@ -647,8 +706,9 @@ k-mers with a hash table instead. Such fixed-length seeds are inferior to
|
||||
variable-length seeds in theory, but can be computed much more efficiently in
|
||||
practice. When a query sequence has multiple seed hits, we can afford to skip
|
||||
highly repetitive seeds without affecting the final accuracy. This further
|
||||
alleviates the concern with the uniqueness of seeds. Hash table is the ideal
|
||||
data structure for mapping long query sequences.
|
||||
alleviates the concern with the seeding uniqueness. At the same time, at low
|
||||
sequence identity, it is rare to see long seeds anyway. Hash table is the ideal
|
||||
data structure for mapping long noisy sequences.
|
||||
|
||||
\section*{Acknowledgements}
|
||||
We owe a debt of gratitude to H. Suzuki and M. Kasahara for releasing their
|
||||
@@ -657,6 +717,8 @@ Schatz, P. Rescheneder and F. Sedlazeck for pointing out the limitation of
|
||||
BWA-MEM. We are also grateful to minimap2 users who have greatly helped to
|
||||
suggest features and to fix various issues.
|
||||
|
||||
\paragraph{Funding\textcolon} NHGRI 1R01HG010040-01
|
||||
|
||||
\bibliography{minimap2}
|
||||
|
||||
\end{document}
|
||||
|
||||
@@ -0,0 +1,225 @@
|
||||
\documentclass{bioinfo}
|
||||
\copyrightyear{2021}
|
||||
\pubyear{2021}
|
||||
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{url}
|
||||
\usepackage{amsmath}
|
||||
\usepackage[ruled,vlined]{algorithm2e}
|
||||
\newcommand\mycommfont[1]{\footnotesize\rmfamily{\it #1}}
|
||||
\SetCommentSty{mycommfont}
|
||||
\SetKwComment{Comment}{$\triangleright$\ }{}
|
||||
|
||||
\usepackage{natbib}
|
||||
\bibliographystyle{apalike}
|
||||
|
||||
\DeclareMathOperator*{\argmax}{argmax}
|
||||
|
||||
\begin{document}
|
||||
\firstpage{1}
|
||||
|
||||
\title[Improvements to minimap2]{New strategies to improve minimap2 alignment accuracy}
|
||||
\author[Li]{Heng Li$^{1,2}$}
|
||||
\address{$^1$Dana-Farber Cancer Institute, 450 Brookline Ave, Boston, MA 02215, USA,
|
||||
$^2$Harvard Medical School, 10 Shattuck St, Boston, MA 02215, USA}
|
||||
|
||||
\maketitle
|
||||
|
||||
\begin{abstract}
|
||||
|
||||
\section{Summary:} We present several recent improvements to minimap2, a
|
||||
versatile pairwise aligner for nucleotide sequences. Now minimap2 v2.22 can
|
||||
more accurately map long reads to highly repetitive regions and align through
|
||||
insertions or deletions up to 100kb by default, addressing major weakness in
|
||||
minimap2 v2.18 or earlier.
|
||||
|
||||
\section{Availability and implementation:}
|
||||
\href{https://github.com/lh3/minimap2}{https://github.com/lh3/minimap2}
|
||||
|
||||
\section{Contact:} hli@ds.dfci.harvard.edu
|
||||
\end{abstract}
|
||||
|
||||
\section{Introduction}
|
||||
Minimap2~\citep{Li:2018ab} is widely used for maping long sequence
|
||||
reads and assembly contigs. \citet{Jain:2020aa} found minimap2 v2.18 or earlier occasionally
|
||||
misaligned reads from highly repetitive regions as minimap2 ignored seeds of
|
||||
high occurrence. They also noticed minimap2 may misplace reads with structural
|
||||
variations (SVs) in such regions~\citep{Jain2020.11.01.363887}. These
|
||||
misalignments have become a pressing issue in the advent of
|
||||
temolere-to-telomore human assembly~\citep{Miga:2020aa}. Meanwhile, old minimap2
|
||||
was unable to efficiently align long insertions/deletions (INDELs) and often
|
||||
breaks an alignment around variable-number tandem repeats (VNTRs). This has
|
||||
inspired new chaining algorithms~\citep{Li:2020aa,Ren:2021aa} which are not
|
||||
integrated into minimap2. Here we will describe recent efforts implemented
|
||||
in v2.19 through v2.22 to improve mapping results.
|
||||
|
||||
\begin{methods}
|
||||
\section{Methods}
|
||||
|
||||
\subsection{Rescuing high-occurrence $k$-mers}
|
||||
Minimap2 keeps all $k$-mer minimizers during indexing. Its original
|
||||
implementation only selected low-occurrence minimizers during mapping. The
|
||||
cutoff is a few hundred for mapping long reads against a human genome. If a
|
||||
read habors only a few or even no low-occurrence minimizers, it will fail
|
||||
chaining due to insufficient anchors.
|
||||
|
||||
To resolve this issue, we implemented a new heuristic to add additional
|
||||
minimizers. Suppose we are looking at two adjacent low-occurence $k$-mers
|
||||
located at position $x_1$ and $x_2$, respectively. If $|x_1-x_2|\ge500$,
|
||||
minimap2 v2.22 additionally selects $\lfloor|x_1-x_2|/500\rfloor$ minimizers
|
||||
of the lowest occurrence among minimizers between $x_1$ and $x_2$.
|
||||
We use a binary heap data
|
||||
structure to select minimizers of the lowest occurrence in this interval.
|
||||
This strategy adds necessary anchors at the cost of increasing total alignment
|
||||
time by a few percent on real data.
|
||||
|
||||
\subsection{Aligning through longer INDELs}
|
||||
The original minimap2 may fail to align long INDELs due to its chaining
|
||||
heuristics. Briefly, minimap2 applies dynamic programming (DP) to chain
|
||||
minimizer anchors. This is a quadratic algorithm, which is slow for chaining
|
||||
contigs. For acceptable performance, the original minimap2 uses a 500bp band by
|
||||
default. If there is an INDEL longer than 500bp and the two chains around the INDEL
|
||||
have no overlaps on either the query or the reference sequence, minimap2 may
|
||||
join the two short chains later at a later step. We call it the
|
||||
long-join heuristic. This heuristic may fail around VNTRs because short chains
|
||||
often have overlaps in VNTRs. More subtly, minimap2 may escape the inner DP
|
||||
loop early, again for performance, if the chaining result is not improved for
|
||||
50 iterations. When there is a copy number change in a long segmental
|
||||
duplication, the early escape may break around the event even if users
|
||||
specify a large band.
|
||||
|
||||
In minigraph~\citep{Li:2020aa}, we developed a new chaining algorithm that
|
||||
finds short INDELs with DP-based chaining and goes through long INDELs with a
|
||||
subquadratic algorithm~\citep{DBLP:conf/wabi/AbouelhodaO03}. We ported the same
|
||||
algorithm to minimap2 for contig mapping. For long-read mapping, the minigraph
|
||||
algorithm is slower. Minimap2 v2.22 now still uses the DP-based algorithm to
|
||||
find short chains and then invokes the minigraph algorithm to rechain anchors in
|
||||
these short chains. The rechaining step achieves the same goal as long-join
|
||||
but is more reliable as it can resolve overlaps between short chains. The old
|
||||
long-join heuristic has since been removed.
|
||||
|
||||
\subsection{Properly mapping long reads with SVs}
|
||||
The original minimap2 ranks an alignment by its Smith-Waterman score and
|
||||
outputs the best scoring alignment. However, when there are SVs on the read,
|
||||
the best scoring alignment is sometimes not the correct alignment.
|
||||
\citet{Jain2020.11.01.363887} resolved this dilemma by altering the mapping
|
||||
algorithm.
|
||||
|
||||
In our view, this problem is rooted in impropriate scoring: affine-gap penalty
|
||||
over-penalizes a long INDEL that was often evolutionarily created in one event.
|
||||
We should not penalize a SV linearly in its length. Minimap2 v2.22 rescores
|
||||
an alignment with the following scoring function. Suppose an alignment consists
|
||||
of $M$ matching bases, $N$ substitutions and $G$ gap opens, we empirically
|
||||
score the alignment with
|
||||
$$
|
||||
M-\frac{N+G}{2d}-\sum_{i=1}^G\log_2(1+g_i)
|
||||
$$
|
||||
where $g_i\ge1$ is the length of the $i$-th gap and
|
||||
$$
|
||||
d=\max\left\{\frac{N+G}{M+N+G},0.02\right\}
|
||||
$$
|
||||
Here $d$ approximates per-base sequence divergence with the smallest value set
|
||||
to 2\%. As an analogy to affine-gap scoring, the matching score in our scheme
|
||||
is 1, the mismatch and gap open penalties are both $1/2d$ and the gap extension
|
||||
penalty is a logarithm function of the gap length. Our scoring gives a long SV
|
||||
a much milder penalty. In terms of time complexity, scoring an alignment is
|
||||
linear in the length of the alignment. Time spent on rescoring is negligible in
|
||||
practice.
|
||||
|
||||
%If we assume sequences evolve under a duplication-mutation model, we may have a
|
||||
%better way to choose the best alignment. If a long read can be mapped to $n$
|
||||
%loci, we can take the read as the template and build a
|
||||
%pseudo-multi-sequence-alignment (pMSA) of $n+1$ sequences. In this pMSA, we say
|
||||
%a site on the read is informative if the $n$ reference subsequences differ at
|
||||
%the position.
|
||||
|
||||
\end{methods}
|
||||
|
||||
\section{Results}
|
||||
|
||||
\begin{table}
|
||||
\processtable{Evaluation of minimap2 v2.22}
|
||||
{\footnotesize\label{tab:1}\begin{tabular}{p{4.2cm}rrrr}
|
||||
\toprule
|
||||
$[$Benchmark$]$ Metric & v2.22 & v2.18 & Winno & lra \\
|
||||
\midrule
|
||||
$[$sim-map$]$ \% mapped reads at Q10 & 97.9 & 97.6 & {\bf 99.0} & 97.3 \\
|
||||
$[$sim-map$]$ err. rate at Q10 (phredQ) & {\bf 52} & {\bf 52} & 38 & 24 \\
|
||||
$[$winno-cmp$]$ rate of diff. (phredQ) & {\bf 41} & 37 & N/A & 18 \\
|
||||
$[$sim-sv$]$ \% false negative rate & {\bf 0.5} & 2.0 & {\bf 0.5} & 1.4 \\
|
||||
$[$sim-sv$]$ \% false discovery rate & {\bf 0.0} & 0.1 & {\bf 0.0} & 0.1 \\
|
||||
$[$real-sv-1k$]$ \% false negative rate & {\bf 7.3} & 20.0 & 13.0 & N/A \\
|
||||
$[$real-sv-1k$]$ \% false discovery rate & 2.7 & {\bf 2.4} & 2.7 & N/A \\
|
||||
\botrule
|
||||
\end{tabular}}
|
||||
{In $[$sim-map$]$, 152,713 reads were simulated from the CHM13 telomere-to-telomere assembly v1.1
|
||||
(AC: GCA\_009914755.3) with pbsim2~\citep{Ono:2021aa}: ``pbsim2 -{}-hmm\_model R94.model -{}-length-min
|
||||
5000 -{}-length-mean 20000 -{}-accuracy-mean 0.95''. Alignments of mapping quality
|
||||
10 or higher were evaluated by ``paftools.js mapeval''. The mapping error rate
|
||||
is measured in the phred scale: if the error rate is $e$, $-10\log_{10}e$ is
|
||||
reported in the table. In $[$winno-cmp$]$, 1.39 million CHM13 HiFi reads from
|
||||
SRR11292121 were mapped against CHM13. 99.3\% of them were mapped by Winnowmap2
|
||||
at mapping quality 10 or higher and were taken as ground truth to evaluate
|
||||
minimap2 and lra with ``paftools.js pafcmp''. $[$sim-sv$]$ simulated 1,000
|
||||
50bp to 1000bp INDELs from chr8 in CHM13 using SURVIVOR~\citep{Jeffares:2017aa} and simulated Nanopore
|
||||
reads at 30 folds with the same pbsim2 command line. SVs were called with
|
||||
``sniffles -q 10''~\citep{Sedlazeck:2018ab} and compared to the simulated truth with ``SURVIVOR eval
|
||||
call.vcf truth.bed 50''. In $[$real-sv-1k$]$, small and long variants were
|
||||
called by dipcall-0.3~\citep{Li:2018aa} for HG002 assemblies (AC: GCA\_018852605.1 and
|
||||
GCA\_018852615.1) and compared to the GIAB truth~\citep{Zook:2020aa} using ``truvari -r 2000 -s
|
||||
1000 -S 400 -{}-multimatch -{}-passonly'' which sets the minimum INDEL size to 1kb in evaluation. }
|
||||
\end{table}
|
||||
|
||||
We evaluated minimap2 v2.22 along with v2.18, Winnowmap2 v2.03 and lra v1.3.2
|
||||
(Table~\ref{tab:1}). Both versions of minimap2 achieved high mapping accuracy on
|
||||
simulated Nanopore reads (sim-map). Winnowmap2 aligned more reads at mapping
|
||||
quality 10 or higher (mapQ10). However, it may occasionally assign a high mapping
|
||||
quality to a read with multiple identical best alignments. This reduced its
|
||||
mapping accuracy.
|
||||
|
||||
In lack of groud truth for real data, so we took Winnowmap2 mapping as ground
|
||||
truth to evaluate other mappers (winno-cmp). Out of 1,378,092 reads with mapQ10
|
||||
alignments by Winnowmap2, minimap2 v2.22 could map all of them. 118 reads, less
|
||||
than 0.01\% of all reads, were mapped differently by v2.22. 51 of them have
|
||||
multiple identical best alignments. We believe these are more likely to be
|
||||
Winnowmap2 errors. Most of the remaining 67 (=118-51) reads have multiple
|
||||
highly similar but not identical alignments. We are not sure what are real
|
||||
mapping errors.
|
||||
|
||||
The two benchmarks above only evaluate read mappings without variations.
|
||||
To measure the mapping accuracy in the presence of SVs (sim-sv), we reproduced
|
||||
the results by~\citep{Jain2020.11.01.363887}. Minimap2 v2.22 is as good as
|
||||
Winnowmap2 now. Note that we were setting the Sniffles mapping quality
|
||||
threshold to 10 in consistent with the benchmarks above. If we used the
|
||||
default threshold 20, v2.22 would miss additional 0.5\% SVs, suggesting
|
||||
minimap2 v2.22 could map variant reads correctly but with conservative mapping
|
||||
quality. This observation is more about the interaction between mappers and
|
||||
callers. Furthermore, the simulation here only considers a simple scenario in
|
||||
evolution. Non-allelic gene conversions, which happen often in segmental
|
||||
duplications~\citep{Harpak:2017aa}, would obscure the optimal mapping
|
||||
strategies. How much such simple SV simulation informs real-world SV calling
|
||||
remains a question.
|
||||
|
||||
To see if minimap2 v2.22 could improve long INDEL alignment, we ran dipcall on
|
||||
contig-to-reference alignments and focused on INDELs longer than 1kb
|
||||
(real-sv-1k). v2.22 is more sensitive at comparable specificity, confirming its
|
||||
advantage in more contiguous alignment. lra is supposed to handle long INDELs
|
||||
better, too. However, we could not get lra to work well with dipcall, so did
|
||||
not report the numbers.
|
||||
|
||||
Minimap2 spends most computing time on base alignment. As recent improvements
|
||||
in v2.22 incur little additional computing and do not change the base alignment
|
||||
algorithm, the new version has similar performance to older verions. It is
|
||||
consistently faster than Winnowmap2 by several times. Sometimes simple
|
||||
heuristics can be as effective as more sophisticated yet slower solutions.
|
||||
|
||||
\section*{Acknowledgements}
|
||||
We thank Arang Rhie and Chirag Jain for providing motivating examples where
|
||||
older minimap2 underperforms.
|
||||
|
||||
\paragraph{Funding\textcolon} This work is funded by NHGRI grant R01HG010040.
|
||||
|
||||
\bibliography{minimap2}
|
||||
|
||||
\end{document}
|
||||
Reference in New Issue
Block a user