Dataset ID
JGAD000625
- Type of data
- bam/gvcf data of NGS (WGS)
- Access criteria
- Controlled-access (Type II)
- Total data volume
- 17 B
- File formats
- TXT
- Research
- hum0184
- Date published
- 2022-02-17
- Date modified
- 2022-02-17
- DDBJ Search
- JGAD000625 (opens in a new tab)
- JGA Study
- JGAS000239 (opens in a new tab)
Analysis method
WGS
- Materials and participants
- 4,566 Japanese general residents
- Subject count4566 (Individual)
- PopulationJapanese
- Sample description
- fastq files of JGAD000338
- Tumor / normalNormal
- Experimental method
- WGS
- Platform
- Illumina HiSeq 2500
Illumina NovaSeq 6000 - Reference genome
- GRCh37
- Mapping
- BWA mem 0.7.12
- Read deduplication
- Picard 2.10.6
- Realignment and base quality recalibration
- GATK 3.7
- Mapping quality
- Reads with MAPQ< 20 were excluded at variant calling with GATK 3.7 HaplotypeCaller
- QC and filtering
- Data with bad base quality and high %GC content were removed.
Aligment:
Data matched for the following condition were removed.
- Low mapping rate
- Different insert size
- Gender information mismatch between meta-data and genotype data
- Suspected sex chromosome aberration
Genotyping:
GATK's best practices includes a variant filtering step following Variant Quality Score Recalibration (VQSR)
- DP/GP (DP < 5, GQ < 20, DP > 60, GQ < 95)
- Heterozygosity (F>=0.05)
- Hardy-Weinberg equilibrium (p < 10^-6)
- Repeat & Low Complexity
Principal Component Analysis (PCA):
PCA was performed with individuals included in the 1000 genomes project and outliers from Japanese cluster were removed.
After these filtering steps, variants located in the regions listed as the HighConfidenceRegion (Genome-In-A-Bottle project) were flagged. - Analysis method
- GATK 3.7 HaplotypeCaller
- Coverage (depth)
- HiSeq 2500: 31.8x
NovaSeq 6000: 28x - Variant count
- Autosomes: 10,202,908
X chromosome: 410,435
Autosomes: 76,768,387
X chromosome: 2,898,518 - Data summary
- Whole genome sequencing analyzed data included in the JGAD000338 were mapped to the GRCh37 reference genome sequence, and variant detection was carried out using the GATK (Genome Analysis Toolkit) standards. This project is an initiative of the GEnome Medical alliance Japan (GEM Japan, GEM-J). Learn more
- Data use policy
- JGAP000011 Policy
NBDC data sharing policy (JGAP000001)