Skip to main page content
Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2011 Jul;11(4):743-8.
doi: 10.1111/j.1755-0998.2011.03005.x. Epub 2011 Mar 24.

PRGmatic: An Efficient Pipeline for Collating Genome-Enriched Second-Generation Sequencing Data Using a 'Provisional-Reference Genome'

Affiliations

PRGmatic: An Efficient Pipeline for Collating Genome-Enriched Second-Generation Sequencing Data Using a 'Provisional-Reference Genome'

Sarah M Hird et al. Mol Ecol Resour. .

Abstract

Second-generation sequencing is increasingly being used in combination with genome-enrichment techniques to amplify a large number of loci in many individuals for the purpose of population genetic and phylogeographic analysis. Compiling all the necessary tools to analyse these data is complex and time-consuming. Here, we assemble a set of programs and pipe them together with Perl, enabling research laboratories without a dedicated bioinformatician to utilize second-generation sequencing. User input is a folder of the second-generation sequencing reads sorted by individual (in FASTA format) and pipeline output is a folder of multi-FASTA files that correspond to loci (with 2 alleles called per individual). Additional output includes a summary file of the number of individuals per locus, observed and expected heterozygosity for each locus, distribution of multiple hits and summary statistics (θ, Tajima's D, etc.). This user-friendly, open source pipeline, which requires no a priori reference genome because it constructs its own, allows the user to set various parameters (e.g. minimum coverage) in the dependent programs (CAP3, BWA, SAMtools and VarScan) and facilitates evaluation of the nature and quality of data collected prior to analysis in software packages.

Similar articles

See all similar articles

Cited by 5 articles

Publication types

MeSH terms

LinkOut - more resources

Feedback