Featured Research

from universities, journals, and other organizations

An error-eliminating fix overcomes big problem in '3rd-gen' genome sequencing

Date:
July 1, 2012
Source:
Cold Spring Harbor Laboratory
Summary:
A team has developed a software package that fixes a serious problem inherent in "3rd-gen" single-molecule genome sequencing: the fact that every fifth or sixth DNA "letter" it generates is incorrect. The high error rate is the flip side of the new method's chief virtue: it generates much longer genome "reads," providing a much more complete picture of genomes.

Error correction provided by the team's algorithm provides more accurate mapping of RNA transcripts in addition to genome sequences. A genome browser view of cDNA alignments using uncorrected (dark blue) and 2nd-gen Illumina-corrected (green) PacBio reads generated from cDNAs of a maize variant. The corrected sequences ("After") match the reference annotations end to end and include two isoforms.
Credit: Image courtesy of Cold Spring Harbor Laboratory

A quantitative biologist at Cold Spring Harbor Laboratory (CSHL) and collaborators have published results of experiments that demonstrate the power of so-called single-molecule sequencing, which was recently introduced but whose use has so far been limited by technical issues.

Related Articles


The team, led by CSHL Assistant Professor Michael Schatz and Adam Phillippy and Sergey Koren of the National Biodefense Analysis and Countermeasures Center and the University of Maryland (UMD), has developed a software package that corrects a serious problem inherent in the newest sequencing technology: the fact that every fifth or sixth DNA "letter" it generates is incorrect. The high error rate is the flip side of the new method's chief virtue: it generates much longer genome "reads" than other technologies currently used, up to 100 times longer, and thus can provide a much more complete picture of genome structure than can be obtained with current, "2nd-gen" sequencing technology.

Using mathematical algorithms, Schatz and the team have preserved the great advantage of the "3rd-gen" method while all but eliminating its chief flaw. They have reduced the error rate from about 15% or greater to less than one-tenth of one percent. This mathematical "fix" -- which has been published in open-source code to the World Wide Web -- greatly increases the practical utility of 3rd-gen sequencing for the entire biomedical research community. The team demonstrates the breadth of potential applications of single-molecule sequencing by applying their fix to sequencing tasks ranging from the tiny bacteriophage virus at one end of the difficulty scale to the large and vastly more complex genome of the parrot, at the other. The parrot genome is more than a third the size of the human genome and is published online July 1 with the team's paper in Nature Biotechnology. The parrot sequence is "far superior to that of any previously sequenced bird genome," Schatz says.

To understand why it is better is to appreciate the advantages of 3rd-gen sequencing. The main advantage has to do with the average length of each "read" (i.e., genome segments read by a sequencer). The individual sequences are assembled into "contigs" -- shorthand for contiguous sequences -- much the way pieces in a jigsaw puzzle are assembled. In currently used 2nd-gen technology, the contigs are very small, and are massively redundant. A "consensus" version of each segment, representing the results of many layered reads, tends to be extremely accurate. But the small size of puzzle pieces prevents accurate assembly of certain genome portions, like those containing long repetitive sequences.

Obtaining superior versions of complete genomes was the objective that motivated Schatz and his collaborators, who also include HHMI Investigator Erich D. Jarvis of Duke University and CSHL Professor W. Richard McCombie, a sequencing pioneer, among others.

Combining the best of both generations

With single-molecule sequencing, the assembled contigs are much longer -- affording a much better picture of relatively larger genome segments, including those occupied by lengthy repeats. This is what Schatz and his team wanted to preserve, while at the same time boosting the error-free rate. They did so by effectively taking the best of both 2nd- and 3rd-gen technologies.

"We call our approach 'hybrid error correction,'" Schatz explains.

The team's major insight was to take advantage of the long-read data offered by a 3rd-gen machine like that used in their experiments, a Pacific Biosciences RS sequencer, and mixing in highly accurate short reads obtained from a separate 2nd-gen sequencer. The two data types were run through an open-source genome assembly program called Celera Assembler to generate a clean final assembly that has proven 99.9% error-free and composed of contigs whose median size is at least double that obtainable with 2nd-gen "short-read" sequencers. Contig sizes are expected to increase appreciably in subsequent iterations of the hybrid approach as single molecule long-read sequencing improves.

High-quality genome assemblies are especially important for genome annotation and comparative genome analyses. Many microbial genome analyses depend on finished genomes, but their cost is prohibitive using older technologies. High-quality analysis of the genomes of higher organisms depends upon continuous sequences that capture long stretches of DNA that spell out genes. Discoveries in recent years of spontaneously occurring structural changes in genomes called copy number variations -- such as those made by CSHL Professor Mike Wigler and his team in their research on schizophrenia and autism -- make clear the importance of being able to obtain clean and accurate pictures of the entire genomes of affected individuals.

With hybrid error correction, Schatz and his colleagues have "demonstrated that high error rates associated with long reads need not be a barrier to genome assembly," he summarizes. "High-error long reads can be efficiently assembled in combination with complementary short reads to produce assemblies not previously possible."


Story Source:

The above story is based on materials provided by Cold Spring Harbor Laboratory. Note: Materials may be edited for content and length.


Journal Reference:

  1. Sergey Koren, Michael C Schatz, Brian P Walenz, Jeffrey Martin, Jason T Howard, Ganeshkumar Ganapathy, Zhong Wang, David A Rasko, W Richard McCombie, Erich D Jarvis, Adam M Phillippy. Hybrid error correction and de novo assembly of single-molecule sequencing reads. Nature Biotechnology, 2012; DOI: 10.1038/nbt.2280

Cite This Page:

Cold Spring Harbor Laboratory. "An error-eliminating fix overcomes big problem in '3rd-gen' genome sequencing." ScienceDaily. ScienceDaily, 1 July 2012. <www.sciencedaily.com/releases/2012/07/120701191609.htm>.
Cold Spring Harbor Laboratory. (2012, July 1). An error-eliminating fix overcomes big problem in '3rd-gen' genome sequencing. ScienceDaily. Retrieved November 28, 2014 from www.sciencedaily.com/releases/2012/07/120701191609.htm
Cold Spring Harbor Laboratory. "An error-eliminating fix overcomes big problem in '3rd-gen' genome sequencing." ScienceDaily. www.sciencedaily.com/releases/2012/07/120701191609.htm (accessed November 28, 2014).

Share This


More From ScienceDaily



More Plants & Animals News

Friday, November 28, 2014

Featured Research

from universities, journals, and other organizations


Featured Videos

from AP, Reuters, AFP, and other news services

Research on Bats Could Help Develop Drugs Against Ebola

Research on Bats Could Help Develop Drugs Against Ebola

AFP (Nov. 28, 2014) In Africa's only biosafety level 4 laboratory, scientists have been carrying out experiments on bats to understand how virus like Ebola are being transmitted, and how some of them resist to it. Duration: 01:18 Video provided by AFP
Powered by NewsLook.com
New Dinosaur Species Found in Museum Collection

New Dinosaur Species Found in Museum Collection

Reuters - Innovations Video Online (Nov. 27, 2014) A British palaeontologist has discovered a new species of dinosaur while studying fossils in a Canadian museum. Pentaceratops aquilonius was related to Triceratops and lived at the end of the Cretaceous Period, around 75 million years ago. Jim Drury has more. Video provided by Reuters
Powered by NewsLook.com
Tryptophan Isn't Making You Sleepy On Thanksgiving

Tryptophan Isn't Making You Sleepy On Thanksgiving

Newsy (Nov. 27, 2014) Tryptophan, a chemical found naturally in turkey meat, gets blamed for sleepiness after Thanksgiving meals. But science points to other culprits. Video provided by Newsy
Powered by NewsLook.com
Classic Hollywood Memorabilia Goes Under the Hammer

Classic Hollywood Memorabilia Goes Under the Hammer

Reuters - Entertainment Video Online (Nov. 26, 2014) The iconic piano from "Casablanca" and the Cowardly Lion suit from "The Wizard of Oz" fetch millions at auction. Sara Hemrajani reports. Video provided by Reuters
Powered by NewsLook.com

Search ScienceDaily

Number of stories in archives: 140,361

Find with keyword(s):
Enter a keyword or phrase to search ScienceDaily for related topics and research stories.

Save/Print:
Share:

Breaking News:

Strange & Offbeat Stories


Plants & Animals

Earth & Climate

Fossils & Ruins

In Other News

... from NewsDaily.com

Science News

Health News

Environment News

Technology News



Save/Print:
Share:

Free Subscriptions


Get the latest science news with ScienceDaily's free email newsletters, updated daily and weekly. Or view hourly updated newsfeeds in your RSS reader:

Get Social & Mobile


Keep up to date with the latest news from ScienceDaily via social networks and mobile apps:

Have Feedback?


Tell us what you think of ScienceDaily -- we welcome both positive and negative comments. Have any problems using the site? Questions?
Mobile: iPhone Android Web
Follow: Facebook Twitter Google+
Subscribe: RSS Feeds Email Newsletters
Latest Headlines Health & Medicine Mind & Brain Space & Time Matter & Energy Computers & Math Plants & Animals Earth & Climate Fossils & Ruins