{"id":114,"date":"2015-09-03T18:02:07","date_gmt":"2015-09-03T16:02:07","guid":{"rendered":"https:\/\/blogs.fu-berlin.de\/gruenstaeudl\/?p=114"},"modified":"2015-09-03T20:55:35","modified_gmt":"2015-09-03T18:55:35","slug":"blastng-away-part-1","status":"publish","type":"post","link":"https:\/\/blogs.fu-berlin.de\/gruenstaeudl\/2015\/09\/03\/blastng-away-part-1\/","title":{"rendered":"BlastNg Away &#8211; Part 1"},"content":{"rendered":"<p><strong>A post without spello in the title<\/strong><\/p>\n<p>Today, I attempted to blast entire chloroplast genomes against NCBI&#8217;s <a href=\"https:\/\/www.ncbi.nlm.nih.gov\/nuccore\" target=\"_blank\">nucleotide database<\/a> via the <em>BLASTn<\/em> command-line tool. Since typical plastomes are between 150,000 and 160,000 bp in length, <em>BLASTn<\/em> searches that are conducted remotely take approximately 20 min. on average.<\/p>\n<p style=\"padding-left: 30px\"><code>time blastn -db nt -query myinputseq.fasta -remote -out results.txt<\/code><br \/>\n<samp><br \/>\nreal 21m25.189s<br \/>\nuser 0m0.070s<br \/>\nsys 0m0.010s<br \/>\n<\/samp><\/p>\n<p>Can we speed up such searches by splitting the input query sequence into equally-sized, smaller pieces and blasting each piece separately?<\/p>\n<p style=\"padding-left: 30px\"><samp># Splitting input query sequence into ten equally-sized, smaller pieces<\/samp><br \/>\n<code><br \/>\nINF=myinputseq.fasta<br \/>\nsplit -d -b $(bc &lt;&lt;&lt; $(tail -n1 $INF | wc -c)\/10) $INF prt<br \/>\n<\/code><br \/>\n<samp># Blasting each region against NCBI&#8217;s nucleotide database<\/samp><br \/>\n<code><br \/>\nfor i in $(ls prt*); do<br \/>\necho $i &gt;&gt; results.txt;<br \/>\ntime blastn -db nt -query $i -remote -outfmt '7 length pident sscinames' -max_target_seqs 10 -out $i.result &gt;&gt; time.txt;<br \/>\nrm $i;<br \/>\ndone<br \/>\n<\/code><\/p>\n<p style=\"padding-left: 30px\"><samp><br \/>\nreal 0m15.161s<br \/>\nuser 0m0.060s<br \/>\nsys 0m0.010s<\/samp><\/p>\n<p style=\"padding-left: 30px\"><samp><br \/>\nreal 0m28.901s<br \/>\nuser 0m0.060s<br \/>\nsys 0m0.013s<\/samp><\/p>\n<p style=\"padding-left: 30px\"><samp><br \/>\nreal 0m27.328s<br \/>\nuser 0m0.057s<br \/>\nsys 0m0.013s<\/samp><\/p>\n<p style=\"padding-left: 30px\"><samp><br \/>\nreal 0m14.909s<br \/>\nuser 0m0.067s<br \/>\nsys 0m0.010s<\/samp><\/p>\n<p style=\"padding-left: 30px\"><samp><br \/>\nreal 0m13.689s<br \/>\nuser 0m0.043s<br \/>\nsys 0m0.023s<br \/>\n<em>etc.<\/em><\/samp><\/p>\n<p>&nbsp;<\/p>\n<p>Yes, we can! (However, I bet that the above scenario was a lucky incidence, and that the time improvement caused by reducing the query sequence size is not always that pronounced).<\/p>\n<p>&nbsp;<\/p>\n","protected":false},"excerpt":{"rendered":"<p>A post without spello in the title Today, I attempted to blast entire chloroplast genomes against NCBI&#8217;s nucleotide database via the BLASTn command-line tool. Since typical plastomes are between 150,000 and 160,000 bp in length, BLASTn searches that are conducted remotely take approximately 20 min. on average. time blastn -db nt -query myinputseq.fasta -remote -out [&hellip;]<\/p>\n","protected":false},"author":2306,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[57598],"tags":[],"class_list":["post-114","post","type-post","status-publish","format-standard","hentry","category-bioinformatics"],"_links":{"self":[{"href":"https:\/\/blogs.fu-berlin.de\/gruenstaeudl\/wp-json\/wp\/v2\/posts\/114","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.fu-berlin.de\/gruenstaeudl\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.fu-berlin.de\/gruenstaeudl\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.fu-berlin.de\/gruenstaeudl\/wp-json\/wp\/v2\/users\/2306"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.fu-berlin.de\/gruenstaeudl\/wp-json\/wp\/v2\/comments?post=114"}],"version-history":[{"count":20,"href":"https:\/\/blogs.fu-berlin.de\/gruenstaeudl\/wp-json\/wp\/v2\/posts\/114\/revisions"}],"predecessor-version":[{"id":134,"href":"https:\/\/blogs.fu-berlin.de\/gruenstaeudl\/wp-json\/wp\/v2\/posts\/114\/revisions\/134"}],"wp:attachment":[{"href":"https:\/\/blogs.fu-berlin.de\/gruenstaeudl\/wp-json\/wp\/v2\/media?parent=114"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.fu-berlin.de\/gruenstaeudl\/wp-json\/wp\/v2\/categories?post=114"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.fu-berlin.de\/gruenstaeudl\/wp-json\/wp\/v2\/tags?post=114"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}